When we perform arithmetic on floating-point numbers, each individual operation can introduce a small rounding error. Most of the time these errors stay small. But when subtracting two nearly equal numbers , the errors grow drastically. This is called Loss of Significance (or catastrophic cancellation ).
Rounding in Arithmetic Operations
Setup
Use Convention 1 with β = 2 \beta = 2 β = 2 , m = 3 m = 3 m = 3 , e min = − 1 e_{\min}=-1 e m i n = − 1 , e max = 2 e_{\max}=2 e m a x = 2 .
Let:
x = 5 8 = ( 0.101 ) 2 × 2 0 y = 7 8 = ( 0.111 ) 2 × 2 0 x = \frac{5}{8} = (0.101)_2 \times 2^0 \qquad y = \frac{7}{8} = (0.111)_2 \times 2^0 x = 8 5 = ( 0.101 ) 2 × 2 0 y = 8 7 = ( 0.111 ) 2 × 2 0
Both are already exact floating-point numbers, so f l ( x ) = x fl(x)=x f l ( x ) = x and f l ( y ) = y fl(y)=y f l ( y ) = y .
Computing x × y x \times y x × y
The exact product:
x × y = 5 8 × 7 8 = 35 64 x \times y = \frac{5}{8} \times \frac{7}{8} = \frac{35}{64} x × y = 8 5 × 8 7 = 64 35
Converting to binary:
35 64 = ( 0.100011 ) 2 × 2 0 \frac{35}{64} = (0.100011)_2 \times 2^0 64 35 = ( 0.100011 ) 2 × 2 0
This has 6 mantissa digits, but our system only holds m = 3 m = 3 m = 3 . We must round.
The 4th bit (the bit immediately past d 3 d_3 d 3 ) is 0, so we truncate :
f l ( x × y ) = ( 0.100 ) 2 × 2 0 = 32 64 = 1 2 fl(x \times y) = (0.100)_2 \times 2^0 = \frac{32}{64} = \frac{1}{2} f l ( x × y ) = ( 0.100 ) 2 × 2 0 = 64 32 = 2 1
Rounding error:
∣ 35 64 − 32 64 ∣ = 3 64 ≈ 0.047 \left|\frac{35}{64} - \frac{32}{64}\right| = \frac{3}{64} \approx 0.047 64 35 − 64 32 = 64 3 ≈ 0.047
This is a normal, bounded rounding error, therefore δ ≤ ξ M \delta \le \xi_M δ ≤ ξ M holds.
Where Loss of Significance Comes From
Suppose x x x and y y y are not exact: f l ( x ) = x ( 1 + δ 1 ) fl(x) = x(1+\delta_1) f l ( x ) = x ( 1 + δ 1 ) and f l ( y ) = y ( 1 + δ 2 ) fl(y) = y(1+\delta_2) f l ( y ) = y ( 1 + δ 2 ) .
Computing f l ( x ) − f l ( y ) fl(x) - fl(y) f l ( x ) − f l ( y ) :
f l ( x ) − f l ( y ) = x ( 1 + δ 1 ) − y ( 1 + δ 2 ) = ( x − y ) + x δ 1 − y δ 2 = ( x − y ) ( 1 + x δ 1 − y δ 2 x − y ) fl(x) - fl(y)
= x(1+\delta_1) - y(1+\delta_2)
= (x-y) + x\delta_1 - y\delta_2
= (x-y)\!\left(1 + \frac{x\delta_1 - y\delta_2}{x - y}\right) f l ( x ) − f l ( y ) = x ( 1 + δ 1 ) − y ( 1 + δ 2 ) = ( x − y ) + x δ 1 − y δ 2 = ( x − y ) ( 1 + x − y x δ 1 − y δ 2 )
The relative error in the result is:
x δ 1 − y δ 2 x − y \frac{x\delta_1 - y\delta_2}{x - y} x − y x δ 1 − y δ 2
Danger
If x ≈ y x \approx y x ≈ y , the denominator x − y x - y x − y is tiny. Even though x δ 1 − y δ 2 x\delta_1 - y\delta_2 x δ 1 − y δ 2 is small, dividing by a tiny denominator makes the relative error enormous . This is Loss of Significance.
Classic Example: Quadratic Roots
Solve x 2 − 56 x + 1 = 0 x^2 - 56x + 1 = 0 x 2 − 56 x + 1 = 0 on a toy computer that rounds to 4 significant figures .
Exact roots (7 s.f.)
x = 56 ± 56 2 − 4 2 = 28 ± 783 x = \frac{56 \pm \sqrt{56^2 - 4}}{2} = 28 \pm \sqrt{783} x = 2 56 ± 5 6 2 − 4 = 28 ± 783
x 1 = 28 + 783 ≈ 55.9822 x 2 = 28 − 783 ≈ 0.017862 x_1 = 28 + \sqrt{783} \approx 55.9822 \qquad x_2 = 28 - \sqrt{783} \approx 0.017862 x 1 = 28 + 783 ≈ 55.9822 x 2 = 28 − 783 ≈ 0.017862
The toy computer computes 783 ≈ 27.98 \sqrt{783} \approx 27.98 783 ≈ 27.98 (4 s.f.).
x 1 = 28 + 27.98 = 55.98 ✓ accurate x_1 = 28 + 27.98 = \mathbf{55.98} \qquad \text{✓ accurate} x 1 = 28 + 27.98 = 55.98 ✓ accurate
x 2 = 28 − 27.98 = 0.02000 ✗ wrong! (actual: 0.017862) x_2 = 28 - 27.98 = \mathbf{0.02000} \qquad \text{✗ wrong! (actual: 0.017862)} x 2 = 28 − 27.98 = 0.02000 ✗ wrong! (actual: 0.017862)
For x 2 x_2 x 2 , the two nearly equal numbers 28 28 28 and 27.98 27.98 27.98 cancel, leaving only 1 significant figure in the result, which isa severe Loss of Significance.
Avoiding Loss of Significance
For x 2 − 56 x + 1 = 0 x^2 - 56x + 1 = 0 x 2 − 56 x + 1 = 0 , Vieta’s formulas give:
x 1 ⋅ x 2 = c a = 1 1 = 1 x 1 + x 2 = b a = 56 x_1 \cdot x_2 = \frac{c}{a} = \frac{1}{1} = 1 \qquad x_1 + x_2 = \frac{b}{a} = 56 x 1 ⋅ x 2 = a c = 1 1 = 1 x 1 + x 2 = a b = 56
Strategy: compute the root that is far from zero (no cancellation) using the standard formula, then obtain the other root by division.
Step 1: compute x 1 x_1 x 1 (addition, no cancellation):
x 1 = 28 + 27.98 = 55.98 x_1 = 28 + 27.98 = 55.98 x 1 = 28 + 27.98 = 55.98
Step 2: compute x 2 x_2 x 2 from the product relation:
x 1 ⋅ x 2 = 1 ⟹ x 2 = 1 x 1 = 1 55.98 = 0.01786 x_1 \cdot x_2 = 1 \;\Longrightarrow\; x_2 = \frac{1}{x_1} = \frac{1}{55.98} = \mathbf{0.01786} x 1 ⋅ x 2 = 1 ⟹ x 2 = x 1 1 = 55.98 1 = 0.01786
This perfectly recovers the accurate value of x 2 ≈ 0.017862 x_2 \approx 0.017862 x 2 ≈ 0.017862 to 4 significant figures.
Caution
General Rule: If two roots x 1 , x 2 x_1, x_2 x 1 , x 2 of x 2 − b x + c = 0 x^2 - bx + c = 0 x 2 − b x + c = 0 satisfy x 1 x 2 = c x_1 x_2 = c x 1 x 2 = c , compute the numerically stable root first (the one that does not subtract nearly equal quantities), then find the other as c / x 1 c / x_1 c / x 1 . This completely avoids the cancellation.
Summary of Conditions That Cause Loss of Significance
Operation When dangerous Why x − y x - y x − y x ≈ y x \approx y x ≈ y Leading significant digits cancel x + y x + y x + y with opposite signs∥ x ∥ ≈ ∥ y ∥ \|x\| \approx \|y\| ∥ x ∥ ≈ ∥ y ∥ Same as subtraction Quadratic formula − b − b 2 − 4 a c 2 a \dfrac{-b - \sqrt{b^2-4ac}}{2a} 2 a − b − b 2 − 4 a c b 2 ≫ 4 a c b^2 \gg 4ac b 2 ≫ 4 a c b 2 − 4 a c ≈ ∥ b ∥ \sqrt{b^2-4ac} \approx \|b\| b 2 − 4 a c ≈ ∥ b ∥