Skip to content

Floating-Point Arithmetic and Loss of Significance

When we perform arithmetic on floating-point numbers, each individual operation can introduce a small rounding error. Most of the time these errors stay small. But when subtracting two nearly equal numbers, the errors grow drastically. This is called Loss of Significance (or catastrophic cancellation).


Rounding in Arithmetic Operations

Setup

Use Convention 1 with β=2\beta = 2, m=3m = 3, emin⁡=−1e_{\min}=-1, emax⁡=2e_{\max}=2.

Let: x=58=(0.101)2×20y=78=(0.111)2×20x = \frac{5}{8} = (0.101)_2 \times 2^0 \qquad y = \frac{7}{8} = (0.111)_2 \times 2^0

Both are already exact floating-point numbers, so fl(x)=xfl(x)=x and fl(y)=yfl(y)=y.

Computing x×yx \times y

The exact product:

x×y=58×78=3564x \times y = \frac{5}{8} \times \frac{7}{8} = \frac{35}{64}

Converting to binary:

3564=(0.100011)2×20\frac{35}{64} = (0.100011)_2 \times 2^0

This has 6 mantissa digits, but our system only holds m=3m = 3. We must round.

The 4th bit (the bit immediately past d3d_3) is 0, so we truncate:

fl(x×y)=(0.100)2×20=3264=12fl(x \times y) = (0.100)_2 \times 2^0 = \frac{32}{64} = \frac{1}{2}

Rounding error:

∣3564−3264∣=364≈0.047\left|\frac{35}{64} - \frac{32}{64}\right| = \frac{3}{64} \approx 0.047

This is a normal, bounded rounding error, therefore δ≤ξM\delta \le \xi_M holds.


Where Loss of Significance Comes From

Suppose xx and yy are not exact: fl(x)=x(1+δ1)fl(x) = x(1+\delta_1) and fl(y)=y(1+δ2)fl(y) = y(1+\delta_2).

Computing fl(x)−fl(y)fl(x) - fl(y):

fl(x)−fl(y)=x(1+δ1)−y(1+δ2)=(x−y)+xδ1−yδ2=(x−y) ⁣(1+xδ1−yδ2x−y)fl(x) - fl(y) = x(1+\delta_1) - y(1+\delta_2) = (x-y) + x\delta_1 - y\delta_2 = (x-y)\!\left(1 + \frac{x\delta_1 - y\delta_2}{x - y}\right)

The relative error in the result is:

xδ1−yδ2x−y\frac{x\delta_1 - y\delta_2}{x - y}


Classic Example: Quadratic Roots

Solve x2−56x+1=0x^2 - 56x + 1 = 0 on a toy computer that rounds to 4 significant figures.

Exact roots (7 s.f.)

x=56±562−42=28±783x = \frac{56 \pm \sqrt{56^2 - 4}}{2} = 28 \pm \sqrt{783}

x1=28+783≈55.9822x2=28−783≈0.017862x_1 = 28 + \sqrt{783} \approx 55.9822 \qquad x_2 = 28 - \sqrt{783} \approx 0.017862

With 4 significant figures

The toy computer computes 783≈27.98\sqrt{783} \approx 27.98 (4 s.f.).

x1=28+27.98=55.98✓ accuratex_1 = 28 + 27.98 = \mathbf{55.98} \qquad \text{✓ accurate}

x2=28−27.98=0.02000✗ wrong! (actual: 0.017862)x_2 = 28 - 27.98 = \mathbf{0.02000} \qquad \text{✗ wrong! (actual: 0.017862)}

For x2x_2, the two nearly equal numbers 2828 and 27.9827.98 cancel, leaving only 1 significant figure in the result, which isa severe Loss of Significance.


Avoiding Loss of Significance

The Workaround: Use Vieta’s Formulas

For x2−56x+1=0x^2 - 56x + 1 = 0, Vieta’s formulas give:

x1⋅x2=ca=11=1x1+x2=ba=56x_1 \cdot x_2 = \frac{c}{a} = \frac{1}{1} = 1 \qquad x_1 + x_2 = \frac{b}{a} = 56

Strategy: compute the root that is far from zero (no cancellation) using the standard formula, then obtain the other root by division.

Step 1: compute x1x_1 (addition, no cancellation):

x1=28+27.98=55.98x_1 = 28 + 27.98 = 55.98

Step 2: compute x2x_2 from the product relation:

x1⋅x2=1  ⟹  x2=1x1=155.98=0.01786x_1 \cdot x_2 = 1 \;\Longrightarrow\; x_2 = \frac{1}{x_1} = \frac{1}{55.98} = \mathbf{0.01786}

This perfectly recovers the accurate value of x2≈0.017862x_2 \approx 0.017862 to 4 significant figures.


Summary of Conditions That Cause Loss of Significance

OperationWhen dangerousWhy
x−yx - yx≈yx \approx yLeading significant digits cancel
x+yx + y with opposite signs∥x∥≈∥y∥\|x\| \approx \|y\|Same as subtraction
Quadratic formula −b−b2−4ac2a\dfrac{-b - \sqrt{b^2-4ac}}{2a}b2≫4acb^2 \gg 4acb2−4ac≈∥b∥\sqrt{b^2-4ac} \approx \|b\|