IEEE 754 Double-Precision Standard
The theoretical floating-point system we have been studying becomes real hardware through the IEEE 754 standard. The 64-bit double-precision format is the workhorse of scientific computing.
64-Bit Layout
A double-precision number occupies exactly 64 bits, split into three fields:
| Field | Bits | Purpose |
|---|---|---|
| Sign | 1 | 0 = positive, 1 = negative |
| Exponent | 11 | Stores the (biased) exponent |
| Fraction | 52 | Stores of the mantissa |
The Denormalized Form
IEEE 754 uses Convention 3 (Denormalized Form). Every normal number is interpreted as:
The leading 1 before the decimal point is always present for normal numbers and is implicit. It is not stored in the 52 fraction bits, giving one extra bit of precision for free.
Why Exponent Biasing Is Necessary
The 11 exponent bits can hold unsigned integers from 0 to .
Without biasing, the smallest representable number would be:
That is far too large since it cannot represent numbers like . Biasing fixes this by interpreting the stored value as .
Exponent Biasing (Bias = 1023)
We subtract 1023 from the stored exponent to get the effective exponent. So, in denormalized form, we can write it as:
Using the equivalent 0.1 form:
or,
Naive Minimum and Maximum
Naively, if we used every possible stored exponent value, our minimum and maximum would be:
- Naive Minimum: Storing yields an effective exponent of
- Naive Maximum: Storing yields an effective exponent of
The Problem: How Do We Represent Zero?
There is a major flaw with the naive approach above: we cannot represent exactly zero.
Because denormalized numbers rely on an implicit leading 1 (), the fraction part is always at least . No matter how small the exponent gets, the final calculated value never truly reaches .
The Solution: Reserved Exponents
To fix the zero problem (and handle math errors), IEEE 754 sacrifices the two extreme stored exponents ( and ) and reserves them for special cases:
| Stored exponent | Fraction bits | Represents |
|---|---|---|
| 0 | all zeros | |
| 0 | any non-zero | very small numbers near 0, consider as 0 |
| 2047 | all zeros | (overflow result) |
| 2047 | any non-zero | Not a Number (e.g. ) |
Overflow and Underflow
With the extremes reserved, pushing outside the boundaries of normal numbers triggers specific behaviors:
| Condition | When | Result |
|---|---|---|
| Overflow | Computed exponent | Stored as |
| Underflow | Computed exponent | Stored as 0 |
Actual Practical Limits
By reserving and , the usable stored exponents for are strictly bounded between 1 and 2046. This gives us an effective exponent range of to .
Here are the actual limits for double-precision normal numbers:
These are the limits you see when a language reports Double.MAX_VALUE or DBL_MAX.
Where Do These Come From?
Smallest normal: Using the minimum usable exponent () and all zeros in the fraction:
Largest normal: Using the maximum usable exponent () and all ones in the fraction: