_lower_incomplete_gamma_ratio approximated the regularized lower
incomplete gamma with a 500-step uniform midpoint Riemann sum. At a = 0.5
(df = 1) the integrand has an integrable pole at t = 0 that a uniform grid
cannot resolve, so the sum under-integrated and 1 - ratio came out too
large: chi_squared_test([120, 80], [100, 100]) returned p = 0.0352 where
the exact value is 0.004678, contradicting docs/en.md:247 ("With 1 degree
of freedom, chi^2 = 8 gives p < 0.005").
The same function also overflowed for df >= 344, because it divided by
math.exp(math.lgamma(a)) instead of subtracting lgamma in the exponent.
Replaced with the standard split: the power series for x < a + 1 and a
Lentz continued fraction otherwise, both scaled by
exp(-x + a*ln(x) - lgamma(a)).
Verified against scipy.special.gammainc over 130 (a, x) pairs spanning
a in [0.5, 100000] and x in [0.01, 50000]: max absolute error 7.8e-16.
chi_squared_p_value(450, 400) now returns 0.04249935069791977 instead of
raising OverflowError. The shipped demo uses df = 3, where the old error
was only ~0.2%, so its printed output is byte-identical before and after.
_regularized_beta approximated the regularized incomplete beta with a
200-step uniform midpoint Riemann sum. t_cdf_approx calls it with
x = df / (df + t^2) and a = df / 2, so as df grows the integrand becomes a
needle at the right endpoint and 200 uniform nodes step over it. At
df = 19992, t = 1.855 the code returned 1.33e-11 where the true p-value is
0.0636 — not significant.
That made the lesson's headline demo an artifact of the bug: the
STATISTICAL VS PRACTICAL SIGNIFICANCE section printed "p-value: 0.0000 /
Significant: True" only because the quadrature had collapsed to zero. The
n = 200 A/B path was also ~13% off (0.4082 vs 0.4710).
Replaced with the Lentz continued fraction for the regularized incomplete
beta, using the x < (a+1)/(a+b+2) symmetry switch and computing the front
factor in log space.
With correct p-values, effect = 0.1 and sigma = 10 at n = 10000 per group
is no longer significant, so the demo would have printed
"Significant: False" directly above its own line "Lesson: large n can make
a negligible effect 'significant'." Raised large_n to 100000, where the
documented point holds on the merits: p = 0.0019 with Cohen's d = -0.0139,
still negligible.
Verified against scipy.special.betainc over 216 (x, a, b) cases spanning
a in [0.5, 100000]: max absolute error 2.5e-10; and against scipy.stats.t
over 63 (df, t) pairs: max absolute error 2.0e-11.