Skip to content

Practice Problems

The following problems are drawn from CSE 330 coursework. Each problem includes a complete worked solution. Try solving each one yourself before expanding the solution.


Problem 1: Form Types and Representable Ranges

Given β=2\beta = 2, m=3m = 3, and −2≤e≤1-2 \le e \le 1, find the minimum, maximum, and total count of representable positive numbers for each of the three forms below.


(a) Standard Form

x=± (0.d1 d2 d3)2×2e,d1≠0  ⇒  d1=1x = \pm\,(0.d_1\,d_2\,d_3)_2 \times 2^e, \qquad d_1 \ne 0 \;\Rightarrow\; d_1 = 1

  • Free digits: d2,d3∈{0,1}d_2, d_3 \in \{0,1\} → 22=42^2 = 4 mantissas
  • Exponents: e∈{−2,−1,0,1}e \in \{-2,-1,0,1\} → 4 values
Minimum(0.100)2×2−2=12×14=18(0.100)_2 \times 2^{-2} = \frac{1}{2} \times \frac{1}{4} = \dfrac{1}{8}
Maximum(0.111)2×21=78×2=74(0.111)_2 \times 2^{1} = \frac{7}{8} \times 2 = \dfrac{7}{4}
Total (positive)22×4=162^2 \times 4 = 16
Total (with sign)16×2=3216 \times 2 = 32

(b) Normalized Form

In this form the leading digit after the radix point is 1 and no other digit is fixed. All three digits 0.1d1 d2 d30.1d_1\,d_2\,d_3 vary, giving a wider set of representable mantissas.

x=± (0.1d1 d2 d3)2×2ex = \pm\,(0.1d_1\,d_2\,d_3)_2 \times 2^e

Minimum(0.1000)2×2−2=12×14=18(0.1000)_2 \times 2^{-2} = \frac{1}{2} \times \frac{1}{4} = \dfrac{1}{8}
Maximum(0.1111)2×21=1516×2=158(0.1111)_2 \times 2^{1} = \frac{15}{16} \times 2 = \dfrac{15}{8}
Total (positive)23×4=322^3 \times 4 = 32
Total (with sign)32×2=6432 \times 2 = 64

(c) Denormalized Form

x=± (1.d1 d2 d3)2×2ex = \pm\,(1.d_1\,d_2\,d_3)_2 \times 2^e

The leading 1 is explicit; all three digits d1,d2,d3∈{0,1}d_1, d_2, d_3 \in \{0,1\} are free → 23=82^3 = 8 mantissas.

Equivalently in the 0. style: (0.1 d1 d2 d3)2×2e+1(0.1\,d_1\,d_2\,d_3)_2 \times 2^{e+1}.

Minimum(1.000)2×2−2=1×14=14(1.000)_2 \times 2^{-2} = 1 \times \frac{1}{4} = \dfrac{1}{4}
Maximum(1.111)2×21=158×2=154(1.111)_2 \times 2^{1} = \frac{15}{8} \times 2 = \dfrac{15}{4}
Total (positive)8×4=328 \times 4 = 32
Total (with sign)32×2=6432 \times 2 = 64

Problem 2: Converting Decimal to Standard Floating-Point

Convert (6.25)10(6.25)_{10} to Standard floating-point representation with β=2\beta = 2, m=3m = 3, and −1≤e≤3-1 \le e \le 3.

Step 1: Convert 6.256.25 to binary:

6.25=4+2+0.25=(110.01)26.25 = 4 + 2 + 0.25 = (110.01)_2

Step 2: Write in F-format (shift radix point left so the integer part becomes 0):

110.012=0.11001×23110.01_2 = 0.11001 \times 2^3

Step 3: Round to m=3m = 3 mantissa digits.

The mantissa so far: 0.110010.11001. We keep d1d2d3=110d_1 d_2 d_3 = 110. The next digit is d4=0d_4 = 0, so we truncate:

fl(6.25)=(0.110)2×23fl(6.25) = (0.110)_2 \times 2^3

Step 4: Verify:

(0.110)2×23=(12+14)×8=34×8=6.0(0.110)_2 \times 2^3 = \left(\frac{1}{2}+\frac{1}{4}\right) \times 8 = \frac{3}{4} \times 8 = 6.0

Rounding Error:

R.E.=∣6.25−6.0∣=0.25\text{R.E.} = |6.25 - 6.0| = 0.25

Relative Error:

δ=∣6.25−6.0∣6.25=0.04\delta = \frac{|6.25 - 6.0|}{6.25} = 0.04


Problem 3: Floating-Point Multiplication

Given x=38x = \dfrac{3}{8} and y=58y = \dfrac{5}{8}, find fl(x×y)fl(x \times y) using m=4m = 4 (Convention 1, β=2\beta = 2), and compute the rounding error.

Step 1: Convert to Standard Form:

x=38=(0.375)10=(0.011)2=(0.11)2×2−1x = \frac{3}{8} = (0.375)_{10} = (0.011)_2 = (0.11)_2 \times 2^{-1}

y=58=(0.625)10=(0.101)2×20y = \frac{5}{8} = (0.625)_{10} = (0.101)_2 \times 2^0

As m=4m = 4, so there is no need to round. Therefore, fl(x)=x=(0.11)2×2−1=38fl(x) = x = (0.11)_2 \times 2^{-1} = \frac{3}{8} and fl(y)=y=(0.101)2×20=58fl(y) = y = (0.101)_2 \times 2^0 = \frac{5}{8}.

Step 2: Compute the exact product:

x×y=fl(x)×fl(y)=38×58=1564=0.234375x \times y = fl(x) \times fl(y) = \frac{3}{8} \times \frac{5}{8} = \frac{15}{64} = 0.234375

Step 3: Convert to binary F-format:

1564=(0.001111)2=(0.1111)2×2−2\frac{15}{64} = (0.001111)_2 = (0.1111)_2 \times 2^{-2}

The mantissa 0.11110.1111 has exactly m=4m = 4 digits. Therefore, no rounding needed.

fl(x×y)=(0.1111)2×2−2=1516×14=1564fl(x \times y) = (0.1111)_2 \times 2^{-2} = \frac{15}{16} \times \frac{1}{4} = \frac{15}{64}

Rounding Error:

R.E.=∣1564−1564∣=0\text{R.E.} = \left|\frac{15}{64} - \frac{15}{64}\right| = 0

No rounding error in this case since the exact result fits within m=4m = 4 digits.


Problem 4: Finding Machine Epsilon

Given β=2\beta = 2, m=4m = 4, −100≤e≤100-100 \le e \le 100, find the machine epsilon ξM\xi_M using the IEEE Normalized Form.

For the Normalized Form:

ξM=12 β−m\xi_M = \frac{1}{2}\,\beta^{-m}

Substituting β=2\beta = 2, m=4m = 4:

ξM=12×2−4=12×116=132\xi_M = \frac{1}{2} \times 2^{-4} = \frac{1}{2} \times \frac{1}{16} = \boxed{\frac{1}{32}}


Problem 5 — Quadratic Roots and Loss of Significance

Compute the roots of x2−12x+5=0x^2 - 12x + 5 = 0 keeping four significant figures throughout. Show the Loss of Significance and apply the workaround.

Exact Roots (high precision)

x=12±144−202=12±1242=6±31x = \frac{12 \pm \sqrt{144 - 20}}{2} = \frac{12 \pm \sqrt{124}}{2} = 6 \pm \sqrt{31}

x1=6+31≈11.5678x2=6−31≈0.43218x_1 = 6 + \sqrt{31} \approx 11.5678 \qquad x_2 = 6 - \sqrt{31} \approx 0.43218

With 4 Significant Figures

A toy computer approximates 31≈5.568\sqrt{31} \approx 5.568 (4 s.f., rounding 5.5677...5.5677...).

x1=6+5.568=11.568  →  11.57✓ close to 11.5678x_1 = 6 + 5.568 = 11.568 \;\to\; \mathbf{11.57} \qquad\text{✓ close to } 11.5678

x2=6−5.568=0.4320  →  0.4320✗ actual: 0.43218x_2 = 6 - 5.568 = 0.4320 \;\to\; \mathbf{0.4320} \qquad\text{✗ actual: } 0.43218

The subtraction 6−5.5686 - 5.568 causes cancellation (1 leading digit lost), giving 3–4 digits of accuracy instead of 4.
(The error grows worse with more extreme cases(see Problem 2 analogue with 56x).)

Workaround Using Vieta’s Formulas

For x2−12x+5=0x^2 - 12x + 5 = 0:

x1+x2=12,x1⋅x2=5x_1 + x_2 = 12, \qquad x_1 \cdot x_2 = 5

Step 1 — compute the stable root first (addition, no cancellation):

x1=6+5.568=11.57x_1 = 6 + 5.568 = 11.57

Step 2 — recover x2x_2 from the product:

x1⋅x2=5  ⟹  x2=5x1=511.57=0.4322x_1 \cdot x_2 = 5 \;\Longrightarrow\; x_2 = \frac{5}{x_1} = \frac{5}{11.57} = \mathbf{0.4322}

This matches the accurate value 0.432180.43218 to 4 significant figures. The division completely avoids the catastrophic cancellation.