Floating Point Arithmetic
Chapter 1 · Floating Point Arithmetic
This chapter covers how computers represent and operate on real numbers. Because only a finite, discrete set of values can be stored in a fixed number of bits, almost every real-number computation involves some approximation. Understanding the nature and magnitude of that approximation is the foundation of all numerical analysis.
Learning Objectives
By the end of this chapter you should be able to:
- Write a number in fixed-point and floating-point binary notation and convert it to base 10.
- Distinguish between Standard Form (Convention 1) and Normalized Form (Convention 2).
- Determine the count, smallest value, largest value, and spacing of representable numbers for a given system.
- Explain the structure of IEEE 754 double-precision numbers and the role of exponent biasing.
- Compute the floating-point representation and its rounding error.
- Define machine epsilon and derive its value for a given convention.
- Identify Loss of Significance in an arithmetic expression and apply a workaround.
Chapter Sections
| # | Section | Key Concepts |
|---|---|---|
| 1 | Fixed Point Representation | Notation, digit range, base conversion |
| 2 | Floating Point Representation | F-format, mantissa, exponent |
| 3 | Number Conventions | Standard Form, Normalized Form |
| 4 | Range and Spacing | Count of representable numbers, spacing |
| 5 | IEEE 754 Standard | 64-bit layout, exponent biasing, overflow |
| 6 | Rounding and Machine Epsilon | , , |
| 7 | Loss of Significance | Cancellation error, avoidance strategies |
| 8 | Practice Problems | CSE 330 exercises with full solutions |