Floating Point Representation
Fixed-point representation breaks down whenever a problem involves both very large and very small quantities simultaneously. Floating-point solves this by separating a number into a significant digits part and a scale part, much like scientific notation.
The Floating-Point System F
The complete set of floating-point numbers for a given set of parameters is:
Components
| Symbol | Name | Description |
|---|---|---|
| Sign | Positive or negative | |
| Mantissa | The fractional part carries all significant digits | |
| Base (radix) | The number system base (2 for computers, 10 for decimal) | |
| Exponent | An integer that scales the mantissa by a power of |
Constraints
All parameters are integers (), subject to:
Converting a Number to F-Format
The goal is to rewrite a number so that its integer part becomes 0 and the significant digits appear after the radix point, while accounts for the shift.
Example 1: Base 10
Convert to floating-point:
So and .
Example 2: Base 2
Convert to floating-point:
Shift the radix point 4 places to the left (each left-shift multiplies by one factor of 2):
So and .
Precision vs. Range
Every floating-point system makes two independent choices:
| Parameter | Controls | Effect of increasing |
|---|---|---|
| (mantissa digits) | Precision | Finer granularity: smaller gaps between adjacent numbers |
| (exponent span) | Range | Wider range of representable magnitudes |
In hardware, both are fixed by the standard. The most important standard is IEEE 754, covered in Section 5.