Skip to content

Floating Point Arithmetic

Chapter 1 · Floating Point Arithmetic

This chapter covers how computers represent and operate on real numbers. Because only a finite, discrete set of values can be stored in a fixed number of bits, almost every real-number computation involves some approximation. Understanding the nature and magnitude of that approximation is the foundation of all numerical analysis.


Learning Objectives

By the end of this chapter you should be able to:

  1. Write a number in fixed-point and floating-point binary notation and convert it to base 10.
  2. Distinguish between Standard Form (Convention 1) and Normalized Form (Convention 2).
  3. Determine the count, smallest value, largest value, and spacing of representable numbers for a given system.
  4. Explain the structure of IEEE 754 double-precision numbers and the role of exponent biasing.
  5. Compute the floating-point representation fl(x)fl(x) and its rounding error.
  6. Define machine epsilon ξM\xi_M and derive its value for a given convention.
  7. Identify Loss of Significance in an arithmetic expression and apply a workaround.

Chapter Sections

#SectionKey Concepts
1Fixed Point RepresentationNotation, digit range, base conversion
2Floating Point RepresentationF-format, mantissa, exponent
3Number ConventionsStandard Form, Normalized Form
4Range and SpacingCount of representable numbers, spacing
5IEEE 754 Standard64-bit layout, exponent biasing, overflow
6Rounding and Machine Epsilonfl(x)fl(x), δ\delta, ξM\xi_M
7Loss of SignificanceCancellation error, avoidance strategies
8Practice ProblemsCSE 330 exercises with full solutions