Appendix A — Floating-point arithmetic

Published

October 9, 2026

Let’s start off with a deliberate incitement

print(0.1 + 0.2 == 0.3)
False

What the heck? Did we just discovered a major bug in one of the most widely used programming languages on the face of the Earth or, rather, we are missing something obvious?

I hate to break it to you, but Python floating-point arithmetic is just fine. The thing is, floating point arithmetic per se is interesting, and we take this chance to review a few facts that are relevant to all of us who happen to work with numbers.

A.1 Base 2 numeral system

You might have heard that virtually all digital computers work internally in base 2—the flip-flop being the basic building block underlying our modern digital ecosystem, everything inside the computer is a long stream of carefully placed zeros and ones. If we want to understand floating-point arithmetic in a digital computer, we have to learn how to count in base 2, first.

When we write the number \(2.25\) in the ordinary decimal positional notation, what we actually mean is \[\begin{align*} 2.25 = 2 + 2 \times 10^{-1} + 5 \times 10^{-2}. \end{align*}\] We can generalize our positional representation to an arbitrary base \(b\) by saying that \[\begin{align} (c_m\ldots c_2c_1c_0.c_{-1}c_{-2}\ldots)_b = c_m \times b^m + \cdots + c_2 \times b^2 + c_1 \times n + c_0 + c_{-1} b^{-1} + c_{-2} b^{-2} \ldots \end{align}\]

How would we go about writing \(2.25\) in base \(2\)? Well, it is not too difficult to convince ourselves that \[\begin{align*} (10.01)_2 \triangleq 0b10.01 = 1 \times 2 + 0 \times 1 + 0 \times 2^{-1} + 1 \times 2^{-2} = 2 + \frac{1}{4} = 2.25. \end{align*}\]

That was easy. How about \(0.1\)? We can try and approximate our target from above and from below, and see if we can get exactly to the point. Unfortunately that is not the case

0 0/2 0b0.0 0.1 0b0.1 1/2 0.5
0 0/4 0b0.00 0.1 0b0.01 1/4 0.25
0 0/8 0b0.000 0.1 0b0.001 1/8 0.125
0.0625 1/16 0b0.0001 0.1 0b0.0010 2/16 0.125
0.09375 3/32 0b0.00011 0.1 0b0.00100 4/32 0.125
0.09375 6/64 0b0.000110 0.1 0b0.000111 7/64 0.109375
0.09375 12/128 0b0.0001100 0.1 0b0.0001101 13/128 0.101562
0.0976562 25/256 0b0.00011001 0.1 0b0.00011010 26/256 0.101562
0.0996094 51/512 0b0.000110011 0.1 0b0.000110100 52/512 0.101562

We can continue as much as we want, but \(0.1\) does not have a finite binary representation—it is actually periodic in base \(2\) \[ 0.1 = 0b0.00011001100110011001100110011001100110011001100110011 \ldots = 0b0.0\overline{0}\overline{0}\overline{1}\overline{1} \]

That is hardly surprising: many rational numbers do not have a finite decimal representation—take for example \[ \frac{1}{3} = 0.\overline{3} \] In a sense, we could argue that in any base \(b\) most rational numbers lack a finite representation—exactly which one clearly depends on \(b\).

A.2 Floating-point representation

The representation of floating-point numbers in a binary digital computer, nowadays, is pretty much everywhere conforming to the IEEE 754 standard (Cowlishaw 2008), coming in two different flavors at 32 (single precision) and 64 bits (double precision). Python internally uses the latter by default.

More precisely, a floating point number is represented in the computer in the form \[ x = (-1)^s \times m \times 2^{e - b} \quad \text{with} \quad 1 \leq m < 2. \tag{A.1}\] Moving from the most to the least significant bit:

  • \(s\) is the sign—the most significant bit is 0 if the number is positive, and 1 if it’s negative;
  • \(e\) is the exponent, for which the standard reserves 8 of the 32 bits in single precision and 11 of the 64 bits in double precision; \(b\), the bias, is an additive, positive constant that amounts to 127 in single precision and 1023 in double precision—its purpose is to be able to represent numbers smaller than 1 without the need of a sign for the exponent;
  • \(m\) is the mantissa, and it is represented as the sequence of binary digits after the separator (assuming that the integer part is 1), taking up 23 bits in single precision and 52 in double precision.

A couple of technical details. The values \(e = 0\) and \(e = 2^8 - 1 = 255\) (in single precision) or \(e = 2^{11} - 1 = 2047\) (in double precision) are reserved (more about this in a second). Therefore for normal numbers \(e\) ranges between 1 and 254 (in single precision) or 2046 (in double precision). Conversely, \(e - b\) goes between -126 and 127 in single precision and -1022 a 1023 in double precision.

For the mantissa, we have effectively one extra bit (that is, 24 in single precision and 53 in double precision) coming from the fact that a leading one is assumed for normal numbers.

And now the two special cases:

  • when \(e = 0\) the number is treated as a subnormal number, where the exponent is by definition \(1 - b\) (-126 in single precision and -1022 in double precision) but the leading bit of the mantissa is assumed to be 0 instead of 1 (which is to say that \(m < 1\));
  • \(e = 255\) (in single precision) or \(e = 2047\) signal an infinity.

A.2.1 A simple example

That was a fairly terse introduction to the IEEE 754 standard. You don’t have to commit to memory the whole thing, but you should at least remember the general form Equation A.1

Now, how about we fully work out a simple example, say \(x = 2.25\)? First thing first, we need to write it in the form of the product of a mantissa \(1 < m < 2\) times a power of 2. That would be \[ x = 2.25 = 1.125 \times 2 \] If we assume double precision, we have \[\begin{align*} s & = 0\\ e & - b = 1 \implies e = 1024 = 0b10000000000\\ m & = 1.125 = 0b(1).0010000000000000000000000000000000000000000000000000 \end{align*}\]

Note we have indicated all the 11 bits for the exponent and all the 52 bits for the mantissa, including the trailing zeroes, and the MSB of the mantissa in parenthesis. If you wonder why the mantissa reads like that, keep in mind that \[ 1.125 = 1 + \frac{1}{8}, \] if that makes sense.

And we are officially ready to assemble the full, glorious sequence of 64 bits that would represent \(x = 2.25\) in the memory of our computer. That would be \[ x \rightarrow 0b0\overbrace{0010000000}^{e-b} \overbrace{0010000000000000000000000000000000000000000000000000}^{m~\text{(MSB omitted)}} \]

A.3 Back to the problem

What is this offering, exactly, to the solution of our initial dilemma? Well, the point is: many of the numbers that we can represent exactly in base \(10\) cannot be represented with a finite number of digits in base \(2\), and this is exactly where things get potentially confusing.

The three numbers that we have started with, \(0.1\), \(0.2\) and \(0.3\), all look harmless because they are easy to represent in base \(10\). Unfortunately, none of the fractions \[ \frac{1}{10},~\frac{1}{5}~\text{e}~\frac{3}{10} \] can be represented exactly in base \(2\), because all of them have the prime factor \(5\) at the denominator. In other words, within a binary digital computer, we can only store an approximation, in the form of a fraction where the denominator is a power of \(2\).

Which fraction exactly? Well, fortunately the Python method as_integer_ratio() is designed to tell us exactly that

for x in (0.1, 0.2, 0.3):
    numerator, denominator = x.as_integer_ratio()
    print(f"{x} -> {numerator} / {denominator} = {x:.20f}")
0.1 -> 3602879701896397 / 36028797018963968 = 0.10000000000000000555
0.2 -> 3602879701896397 / 18014398509481984 = 0.20000000000000001110
0.3 -> 5404319552844595 / 18014398509481984 = 0.29999999999999998890

And now we start to see the solution: when we write \(0.1\) in our text editor, what we are committing to the memory is in fact a slightly different number—very close, but not quite the same. No wonder things do not add up as we naively expected

x = (0.1 + 0.2)
numerator, denominator = x.as_integer_ratio()
print(f"(0.1 + 0.2) -> {numerator} / {denominator} = {x:.20f}")
(0.1 + 0.2) -> 1351079888211149 / 4503599627370496 = 0.30000000000000004441

See? It’s not like we are off by much. We start seeing a discrepancy around the seventeenth decimal place, which in most cases we would hardly care about. And yet the two things, when compared bit by bit, are not identical.

A.4 Properties of floating-point arithmetic

Good: we have learned a very fundamental thing: floating-point arithmetic on a digital computer is, by its very nature, intrinsically inexact. There is really nothing we can do about its—it is one of the facts of life that we have to accept and learn to live with.

A.4.1 Range and resolution

The available range for a floating-point number is primarily dictated by the number of bits allocated for the exponents. In the double-precision IEEE 754 standard

  • the mantissa ranges from \(1\) to \(2\) (excluded);
  • the exponent ranges from \(-1022\) to \(1023\)

and normal numbers span the range between \[ 1 \times 2^{-1022} \approx 2.2 \times 10^{-308} \quad\text{and}\quad (2 - 2^{-52}) \times 2^{1023} \approx 2 \times 2^{1023} \approx 1.8 \times 10^{308} \]

Note

For completeness: when the exponent takes its smallest possible value, which would be technically \(-1023\), the IEEE standard switches to the sub-normal representation, where the exponent is still \(-1022\), and the most significant bit of the mantissa is assumed to be \(0\) instead of \(1\). This is a completely different regime, where the relative precision is not constant anymore, and the range is extended toward the zero, down to \[ 2^{-52} \times 2^{-1022} = 2^{-1074} \approx 4.9 \times 10^{-324} \]

Back to the Python prompt, this reads

print(1.e-324)
print(5.e-324)
print(1.5e308)
print(2.e308)
0.0
5e-324
1.5e+308
inf

(Watch out: you can trigger an overflow quite easily, if you are not careful. \(171!\) is too big to be represented, and so is \(144^{144}\) and \(e^{710}\).)

Clearly, floating-point arithmetic is designed to operate at constant relative precision, determined by the number of bits allocated for the mantissa. With \(53\) effective bits we have \[ 2^{53} \approx 9 \times 10^{15} \] possible values in the interval \([1, 2[\), or about \(16\) significant digits. The maximum (relative) rounding error is \(2^{-53}\) which is sometimes called half spacing. Things work out pretty good when we multiply and divide, not necessarily so much when we subtract.

Since the relative precision is constant, the physical spacing between consecutive floating point number is not. More precisely, the latter is approximately proportional to the value we are representing. For numbers between \(1\) e \(2\) the spacing is \(2^{-52}\); between \(2\) e \(4\) it becomes \(2^{-51}\) and so on and so forth: it doubles every time you cross a power of \(2\). The math module of the standard library provides the ulp (unit in the last place) function which, for a given float \(x\) returns the absolute distance to the one immediately above that can be represented

import math

for x in (1., 1.e10, 1.e20, 1.e50, 1.e100):
    print(math.ulp(x))
2.220446049250313e-16
1.9073486328125e-06
16384.0
2.076918743413931e+34
1.942668892225729e+84

By the time we reach \(2^{52} \approx 4.5 \times 10^{15}\) the unit in the last place becomes greater than \(1\). Not a big deal, per se, but if for a moment we wanted to treat the Avogadro number as an integer, we would have a hard time adding a single molecule to a mole of gas, as the ulp if of the order of \(70\) millions:

import math

N = 6.02e23

print(math.ulp(N))
print(N == N + 1)
67108864.0
True

Let’s wrap this up. With \(11\) bits for the exponent and \(52 + 1\) bits for the mantissa, the double-precision IEEE standard provides a dynamic range of about \(616\) orders of magnitude and a precision of \(16\) digits. This is not terrible, considering that there are about \(42\) orders of magnitude between the proton radius and that of the visible universe, and that \(16\) significant digits are enough for basically all the precision measurements available.

A.4.2 More on arithmetics

Fun fact: some of the properties of elementary arithmetics that we are taught in elementary school, such as the commutative, associative and distributive properties of some operations, cannot be taken for granted when dealing with floating-point numbers in a digital computer. At this point, this should be hardly a surprise, but just for the reference

a = 0.2
b = 0.1
c = 0.3

print((a + b) - c == a + (b - c))
print(a * (b + c) == a * b + a * c)
False
False

(Think about: if is true that we are making approximations at every step in the way, then it is all but unexpected that the final answer might depend on the exact sequence of operations.)

Most of the times we are dealing with small differences, at the level of the machine precision, but there are situations that can be engineered to make everything blow up:

a = 1.e16
b = -1.e16
c = 1.

print((a + b) + c)
print(a + (b + c))
1.0
0.0

(This brings us back to the example of the Avogadro number: summing very large numbers to much smaller numbers is bad!)

A.4.3 Floating-point comparisons

By now we know that the very idea of comparing floating-point number is ill-defined, and yet you see all the time programs like

a = 0.1
b = 0.2
c = 0.3

if a + b <= c:
    print("win!")
else:
    print("lose...")
lose...

Now, if you think about this in terms of the left-hand side of the inequality being a generic operation on random numbers reasonably distributed over an interval much larger than the machine precision, you will get the expected answer most of the time. (And yet you cannot rely on that being true.)

Luckily, both the Python standard library and numpy offer facilities to handle comparisons between floats in a more rigorous manner.

import math
import numpy as np

a = 0.1
b = 0.2
c = 0.3

print(math.isclose(a + b, c))
print(np.allclose(a + b, c))
True
True

A.5 What did we learn?

Well, that the floating point arithmetic is intrinsically inexact, and you do have to pay attention. The IEEE 754 standard is also very well designed, and things usually work out fine—just not all the time.

Pay extra attention to:

  • addition and subtractions between numbers differing by many orders of magnitude;
  • subtractions between numbers very close to each other;
  • long chains of multiplications and divisions, with potential for underflow or overflow at some intermediate point;
  • coherent accumulation of small errors within loops (if you have 10 minutes to spare, take a look at this report).

Consider yourself advised!

Cowlishaw, Mike, ed. 2008. “IEEE Standard for Floating-Point Arithmetic.” Standard IEEE Std 754-2008. New York, NY, USA: IEEE Computer Society. https://web.archive.org/web/20160806053349/http://www.csee.umbc.edu/~tsimo1/CMSC455/IEEE-754-2008.pdf.