Skip to content

Floating Point Representation

Fixed-point representation breaks down whenever a problem involves both very large and very small quantities simultaneously. Floating-point solves this by separating a number into a significant digits part and a scale part, much like scientific notation.


The Floating-Point System F

The complete set of floating-point numbers for a given set of parameters is:

F={  ± (0.d1 d2 … dn)β  ×  β e  }F = \left\{\; \pm\,(0.d_1\,d_2\,\ldots\,d_n)_\beta \;\times\; \beta^{\,e} \;\right\}

Components

SymbolNameDescription
±\pmSignPositive or negative
0.d1d2…dn0.d_1 d_2 \ldots d_nMantissaThe fractional part carries all significant digits
β\betaBase (radix)The number system base (2 for computers, 10 for decimal)
eeExponentAn integer that scales the mantissa by a power of β\beta

Constraints

All parameters are integers (β, di, e∈Z\beta,\, d_i,\, e \in \mathbb{Z}), subject to:

0  ≤  di  ≤  β−1andemin⁡  ≤  e  ≤  emax⁡0 \;\le\; d_i \;\le\; \beta - 1 \qquad \text{and} \qquad e_{\min} \;\le\; e \;\le\; e_{\max}


Converting a Number to F-Format

The goal is to rewrite a number so that its integer part becomes 0 and the significant digits appear after the radix point, while ee accounts for the shift.

Example 1: Base 10

Convert 123.45123.45 to floating-point:

123.45  =  12.345×101  =  1.2345×102  =  (0.12345)10×103123.45 \;=\; 12.345 \times 10^1 \;=\; 1.2345 \times 10^2 \;=\; \mathbf{(0.12345)_{10} \times 10^3}

So d1=1,  d2=2,  d3=3,  d4=4,  d5=5d_1=1,\; d_2=2,\; d_3=3,\; d_4=4,\; d_5=5 and e=3e=3.

Example 2: Base 2

Convert (1001.11)2(1001.11)_2 to floating-point:

Shift the radix point 4 places to the left (each left-shift multiplies ee by one factor of 2):

1001.112=(0.100111)2×241001.11_2 = \mathbf{(0.100111)_2 \times 2^4}

So d1=1,  d2=0,  d3=0,  d4=1,  d5=1,  d6=1d_1=1,\; d_2=0,\; d_3=0,\; d_4=1,\; d_5=1,\; d_6=1 and e=4e=4.


Precision vs. Range

Every floating-point system makes two independent choices:

ParameterControlsEffect of increasing
nn (mantissa digits)PrecisionFiner granularity: smaller gaps between adjacent numbers
emax⁡−emin⁡e_{\max} - e_{\min} (exponent span)RangeWider range of representable magnitudes

In hardware, both are fixed by the standard. The most important standard is IEEE 754, covered in Section 5.