Skip to content

Rounding and Machine Epsilon

Rounding and Machine Epsilon

The floating-point number line is a discrete set of points. Any real number xx that does not fall exactly on one of those points must be rounded to the nearest one. This section formalises that process and introduces machine epsilon ξM\xi_M which represents the worst-case relative rounding error.


The Floating-Point Representation fl(x)fl(x)

Given a real number xx, its floating-point representation fl(x)fl(x) is the nearest number in the system FF:

fl(x)=the closest element of F to xfl(x) = \text{the closest element of } F \text{ to } x

Because the floating-point number line is discrete, values that fall between representable numbers are rounded to the closest available point. For example, a number closer to the left bound will round down, while a number closer to the right bound will round up.

Number line showing standard rounding to the nearest floating-point number

As shown above:

  • 0.1000010.100001 is closer to 0.1000.100 than to 0.1010.101. It rounds down to 0.1000.100, which we define as truncation or rounding down.
  • 0.1001010.100101 is closer to 0.1010.101 than to 0.1000.100. It rounds up to 0.1010.101, which we define as rounding up.

Tie-Breaking Rule

When xx falls exactly halfway between two adjacent floating-point numbers, standard proximity rounding can’t decide. To prevent statistical bias from always rounding up or always rounding down, IEEE 754 uses the round-to-nearest-even (or banker’s rounding) rule.

The number is rounded to the adjacent floating-point value whose last digit is even (which means it ends in 0 in binary).

Number line demonstrating the round-to-nearest-even tie-breaking rule for midpoints

As shown above:

  • 0.10010.1001 is the exact midpoint between 0.1000.100 and 0.1010.101. It rounds down to 0.1000.100 because it ends in 0 (even).
  • 0.10110.1011 is the exact midpoint between 0.1010.101 and 0.1100.110. It rounds up to 0.1100.110 because it ends in 0 (even).

Quick Tricks for Binary Rounding

When rounding a binary fraction to mm mantissa bits, inspect the bits immediately following the mm-th position:

  • Round Down (Truncate)
    If the immediate next bit (the (m+1)(m+1)-th bit) is 0. The value is closer to the lower representable number, regardless of any bits that follow.

  • Round Up
    If the immediate next bit is 1 AND there are additional non-zero bits after it. This means the value is strictly greater than the halfway point.

  • Round to Nearest Even (Tie-Breaker)
    If the immediate next bit is 1 AND it is the only remaining bit (or all subsequent bits are 0). The value sits exactly at the midpoint, so you round up or down to ensure the final mm-th bit is 0 (even).

Scale-Invariant (Relative) Error

δ=∣fl(x)−x∣∣x∣\delta = \frac{|fl(x) - x|}{|x|}

Rearranging:

fl(x)−x=δ⋅x⟹fl(x)=x (1+δ)fl(x) - x = \delta \cdot x \qquad\Longrightarrow\qquad \boxed{fl(x) = x\,(1 + \delta)}

This compact form says: fl(x)fl(x) is xx multiplied by a factor that deviates from 1 by at most δ\delta.


Machine Epsilon ξM\xi_M

Machine epsilon is the maximum possible relative rounding error in a floating-point system:

ξM=max⁡x  δ=max⁡x∣fl(x)−x∣∣x∣\xi_M = \max_x \; \delta = \max_x \frac{|fl(x) - x|}{|x|}

To find the worst-case (maximum) relative error, our goal is to maximize the numerator (the absolute rounding error) and minimize the denominator (the true value xx).

The absolute rounding error is maximized when a number falls exactly halfway between two adjacent representable numbers. The minimum value depends on the floating-point convention being used.

Deriving ξM\xi_M: Standard Form

In the Standard Form representation, the mantissa is written as 0.d1d2…dm0.d_1 d_2 \ldots d_m where d1≠0d_1 \neq 0. Let’s assume m=3m=3 and β=2\beta=2.

  • Minimize Denominator: The smallest normalized magnitude is ∣x∣min⁡=0.1002=β−1|x|_{\min} = 0.100_2 = \beta^{-1}.
  • Maximize Numerator: The maximum absolute error (R.E.)max⁡(\text{R.E.})_{\max} is half the distance between 0.1000.100 and 0.1010.101. (R.E.)max⁡=12×(0.101−0.100)=12×β−3=12β−m(\text{R.E.})_{\max} = \frac{1}{2} \times (0.101 - 0.100) = \frac{1}{2} \times \beta^{-3} = \frac{1}{2}\beta^{-m}
Number line showing maximum rounding error for Standard Form

Dividing the maximum error by the minimum value gives:

ξM=12 β−mβ−1=12 β1−m\xi_M = \frac{\frac{1}{2}\,\beta^{-m}}{\beta^{-1}} = \frac{1}{2}\,\beta^{1-m}

Deriving ξM\xi_M: Normalized Form

In this Normalized Form, let’s look at an m=3m=3 setup where the mantissa is padded (e.g., 0.1d2d3d40.1 d_2 d_3 d_4).

  • Minimize Denominator: The smallest magnitude is ∣x∣min⁡=0.10002=β−1|x|_{\min} = 0.1000_2 = \beta^{-1}.
  • Maximize Numerator: The maximum error is half the distance between 0.10000.1000 and 0.10010.1001. (R.E.)max⁡=12×(0.1001−0.1000)=12×β−4=12β−(m+1)(\text{R.E.})_{\max} = \frac{1}{2} \times (0.1001 - 0.1000) = \frac{1}{2} \times \beta^{-4} = \frac{1}{2}\beta^{-(m+1)}
Number line showing maximum rounding error for Normalized Form

Dividing the two yields:

ξM=12 β−(m+1)β−1=12 β−m\xi_M = \frac{\frac{1}{2}\,\beta^{-(m+1)}}{\beta^{-1}} = \frac{1}{2}\,\beta^{-m}

Deriving ξM\xi_M: De-normalized Form

In the De-normalized (or implicit leading bit) form, numbers are represented as 1.d1d2d31.d_1 d_2 d_3. Let’s again use m=3m=3.

  • Minimize Denominator: The smallest magnitude is ∣x∣min⁡=1.0002=β0=1|x|_{\min} = 1.000_2 = \beta^0 = 1.
  • Maximize Numerator: The maximum error is half the distance between 1.0001.000 and 1.0011.001. (R.E.)max⁡=12×(1.001−1.000)=12×β−3=12β−m(\text{R.E.})_{\max} = \frac{1}{2} \times (1.001 - 1.000) = \frac{1}{2} \times \beta^{-3} = \frac{1}{2}\beta^{-m}
Number line showing maximum rounding error for De-normalized Form

Dividing them provides:

ξM=12 β−mβ0=12 β−m\xi_M = \frac{\frac{1}{2}\,\beta^{-m}}{\beta^0} = \frac{1}{2}\,\beta^{-m}

Comparison Summary

Convention∣x∣min⁡\lvert x\rvert_{\min}(R.E.)max⁡(\text{R.E.})_{\max}ξM\xi_M
Standard Formβ−1\beta^{-1}12β−m\frac{1}{2}\beta^{-m}12β1−m\frac{1}{2}\beta^{1-m}
Normalized Formβ−1\beta^{-1}12β−(m+1)\frac{1}{2}\beta^{-(m+1)}12β−m\frac{1}{2}\beta^{-m}
De-normalized Formβ0\beta^012β−m\frac{1}{2}\beta^{-m}12β−m\frac{1}{2}\beta^{-m}

Notice how both the Normalized and De-normalized forms achieve a tighter (smaller) machine epsilon, granting higher precision by effectively utilizing the available bits.


Key Property

For any representable number xx:

δ≤ξM\delta \le \xi_M

That is, every individual rounding satisfies the machine epsilon bound.


Quick Example

β=2\beta = 2, m=3m = 3, emin⁡=−1e_{\min}=-1, emax⁡=2e_{\max}=2, Convention 1.

Between (0.100)2×20=0.5(0.100)_2 \times 2^0 = 0.5 and (0.101)2×20=0.625(0.101)_2 \times 2^0 = 0.625:

  • Gap =0.125= 0.125
  • Midpoint =0.5625= 0.5625
  • A real value of x=0.5625x = 0.5625 would round to the nearest even-last-digit number.
    • 0.50.5 ends in 0 (even), 0.6250.625 ends in 1 (odd) → rounds to fl(0.5625)=0.5fl(0.5625) = 0.5
  • Relative error =∣0.5−0.5625∣/0.5625=0.0625/0.5625≈0.111= |0.5 - 0.5625| / 0.5625 = 0.0625 / 0.5625 \approx 0.111
  • ξM=12×21−3=18=0.125≥0.111\xi_M = \frac{1}{2} \times 2^{1-3} = \frac{1}{8} = 0.125 \ge 0.111 ✓