Files
ai-engineering-from-scratch/phases/01-math-foundations/15-statistics-for-ml
Divya aa9adcdd45 fix(phase-01/15): chi-squared p-value was 7.5x too large (#376)
_lower_incomplete_gamma_ratio approximated the regularized lower
incomplete gamma with a 500-step uniform midpoint Riemann sum. At a = 0.5
(df = 1) the integrand has an integrable pole at t = 0 that a uniform grid
cannot resolve, so the sum under-integrated and 1 - ratio came out too
large: chi_squared_test([120, 80], [100, 100]) returned p = 0.0352 where
the exact value is 0.004678, contradicting docs/en.md:247 ("With 1 degree
of freedom, chi^2 = 8 gives p < 0.005").

The same function also overflowed for df >= 344, because it divided by
math.exp(math.lgamma(a)) instead of subtracting lgamma in the exponent.

Replaced with the standard split: the power series for x < a + 1 and a
Lentz continued fraction otherwise, both scaled by
exp(-x + a*ln(x) - lgamma(a)).

Verified against scipy.special.gammainc over 130 (a, x) pairs spanning
a in [0.5, 100000] and x in [0.01, 50000]: max absolute error 7.8e-16.
chi_squared_p_value(450, 400) now returns 0.04249935069791977 instead of
raising OverflowError. The shipped demo uses df = 3, where the old error
was only ~0.2%, so its printed output is byte-identical before and after.
2026-08-02 11:42:43 +01:00
..