Skip to content

IEEE 754 Double-Precision Standard

The theoretical floating-point system we have been studying becomes real hardware through the IEEE 754 standard. The 64-bit double-precision format is the workhorse of scientific computing.


64-Bit Layout

A double-precision number occupies exactly 64 bits, split into three fields:

64-bit IEEE 754 layout: Sign (1 bit, bit 63), Exponent (11 bits, bits 62-52), Fraction/Mantissa (52 bits, bits 51-0)
FieldBitsPurpose
Sign10 = positive, 1 = negative
Exponent11Stores the (biased) exponent
Fraction52Stores d1d2…d52d_1 d_2 \ldots d_{52} of the mantissa

The Denormalized Form

IEEE 754 uses Convention 3 (Denormalized Form). Every normal number is interpreted as:

± (1.d1 d2 … d52)2  ×  2 e\pm\,(1.d_1\,d_2\,\ldots\,d_{52})_2 \;\times\; 2^{\,e}

The leading 1 before the decimal point is always present for normal numbers and is implicit. It is not stored in the 52 fraction bits, giving one extra bit of precision for free.


Why Exponent Biasing Is Necessary

The 11 exponent bits can hold unsigned integers from 0 to 211−1=20472^{11}-1 = 2047.
Without biasing, the smallest representable number would be:

(1.00…0)2×20=1(1.00\ldots 0)_2 \times 2^0 = 1

That is far too large since it cannot represent numbers like 0.0010.001. Biasing fixes this by interpreting the stored value as eeffective=estored−biase_{\text{effective}} = e_{\text{stored}} - \text{bias}.


Exponent Biasing (Bias = 1023)

We subtract 1023 from the stored exponent to get the effective exponent. So, in denormalized form, we can write it as:

± (1.d1 d2 … d52)2  ×  2 e−1023\pm\,(1.d_1\,d_2\,\ldots\,d_{52})_2 \;\times\; 2^{\,e - 1023}

Using the equivalent 0.1 form:

± (0.1 d1 d2 … d52)2  ×  2 e−1023×21\pm\,(0.1\,d_1\,d_2\,\ldots\,d_{52})_2 \;\times\; 2^{\,e - 1023}\times 2^1 or, ± (0.1 d1 d2 … d52)2  ×  2 e−1022\pm\,(0.1\,d_1\,d_2\,\ldots\,d_{52})_2 \;\times\; 2^{\,e - 1022}

Naive Minimum and Maximum

Naively, if we used every possible stored exponent value, our minimum and maximum would be:

  • Naive Minimum: Storing 00 yields an effective exponent of −1022-1022 (0.100…0)2×2−1022(0.100\ldots 0)_2 \times 2^{-1022}
  • Naive Maximum: Storing 20472047 yields an effective exponent of 10251025 (0.111…1)2×21025(0.111\ldots 1)_2 \times 2^{1025}

The Problem: How Do We Represent Zero?

There is a major flaw with the naive approach above: we cannot represent exactly zero.

Because denormalized numbers rely on an implicit leading 1 (1.d1d2…1.d_1d_2\dots), the fraction part is always at least 1.01.0. No matter how small the exponent gets, the final calculated value never truly reaches 00.


The Solution: Reserved Exponents

To fix the zero problem (and handle math errors), IEEE 754 sacrifices the two extreme stored exponents (00 and 20472047) and reserves them for special cases:

Stored exponentFraction bitsRepresents
0all zeros±0\pm 0
0any non-zerovery small numbers near 0, consider as 0
2047all zeros±∞\pm\infty (overflow result)
2047any non-zeroNot a Number (e.g. 0/00/0)

Overflow and Underflow

With the extremes reserved, pushing outside the boundaries of normal numbers triggers specific behaviors:

ConditionWhenResult
OverflowComputed exponent >1024> 1024Stored as ±∞\pm\infty
UnderflowComputed exponent <−1021< -1021Stored as 0

Actual Practical Limits

By reserving 00 and 20472047, the usable stored exponents for are strictly bounded between 1 and 2046. This gives us an effective exponent range of −1021-1021 to 10241024.

Here are the actual limits for double-precision normal numbers:

Smallest normal≈2.225×10−308\text{Smallest normal} \approx 2.225 \times 10^{-308}

Largest normal≈1.798×10308\text{Largest normal} \approx 1.798 \times 10^{308}

These are the limits you see when a language reports Double.MAX_VALUE or DBL_MAX.

Where Do These Come From?

Smallest normal: Using the minimum usable exponent (1−1022=−10211 - 1022 = -1021) and all zeros in the fraction:

(0.100…0)2×2−1021=2−1021≈2.225×10−308(0.100\ldots 0)_2 \times 2^{-1021} = 2^{-1021} \approx 2.225 \times 10^{-308}

Largest normal: Using the maximum usable exponent (2046−1022=10242046 - 1022 = 1024) and all ones in the fraction:

(0.111…1)2×21024≈(2−2−52)×21024≈1.798×10308(0.111\ldots 1)_2 \times 2^{1024} \approx (2 - 2^{-52}) \times 2^{1024} \approx 1.798 \times 10^{308}