Files
ai-engineering-from-scratch/phases/19-capstone-projects/29-end-to-end-coding-task-demo
Rohit Ghumare 8050434e5c fix: repair lessons, harden security, and gate lesson tests and quiz bias (#480)
* fix(phase-13/06): bootstrap sys.path so the documented test command runs

The lesson doc says to run unittest discovery over code/tests, but the test
imported main with no path setup, so discovery from the lesson root failed
with ImportError. Insert the code directory on sys.path the way the later
protocol lessons already do. Passes from the lesson root and from code/.

* fix(phase-13/07): bootstrap sys.path so the documented test command runs

Discovery over code/tests failed with ImportError because the test imported
main with no path setup. Insert the code directory on sys.path to match the
later protocol lessons.

* fix(phase-13/08): bootstrap sys.path so the documented test command runs

The documented discovery command raised ImportError on import main. Add the
sys.path bootstrap used by the sibling lessons so the test runs from the
lesson root and from code/.

* fix(phase-13/09): bootstrap sys.path so the documented test command runs

The documented discovery command raised ImportError on import main. Add the
sys.path bootstrap used by the sibling lessons so the test runs from the
lesson root and from code/.

* fix(phase-19/29): make the fixture tests directory an importable package

The demo fixture ships a namespace-style tests directory with no __init__.py,
so a regular tests package elsewhere on sys.path could shadow it and the
in-repo test runner failed to import the fixture module. Add an empty
__init__.py so the fixture tests resolve regardless of what else is installed.

* fix(phase-19/39): bound generation by chars and gate the demo on loss

decode_response replaces invalid UTF-8 bytes, so a 4-byte generation can
re-encode to more than 4 bytes; the test now bounds the decoded character
count, which is the real invariant that generate enforces. The demo's success
gate keyed on exact-match improvement, which the documented padding mask makes
unreachable, so it now checks that training loss decreased.

* fix(phase-19/59): test patch-projection grad via a nonzero reduction

Summing the CLS token straight out of the final LayerNorm is an algebraic
zero for any input, so the gradient reaching the patch projection was always
zero and the test failed for a reason unrelated to the encoder. Reduce with a
sum of squares, which the demo already uses, so the test exercises a real
gradient path.

* fix(phase-19/47): load checkpoints weights-only, jail shard paths

torch.load ran with weights_only=False, so restoring an untrusted checkpoint
executes whatever the pickle names. Store RNG state as primitives so it
survives the weights-only loader, load every payload with weights_only=True,
reject shard paths that resolve outside the checkpoint directory, and raise
ValueError on integrity failures instead of assert. Add regression tests for a
pickled object and a shard path escape, and update the docs, skill, and quiz.

* fix(phase-11/09): guard run_code with an AST check, not a blocklist

The substring blocklist read the snippet as text, so an attribute chain such
as the object-subclass walk reached the real interpreter through __globals__.
Parse the snippet and walk the tree, rejecting import statements, dunder
attribute access, and unsafe builtin names by structure. Add regression tests
for the escape, and make the tool description and docs state plainly that this
is a teaching filter, not real isolation. Align the TypeScript port's wording
and block the constructor and globalThis gadgets in its blocklist.

* fix(site): serve markdown at / before the static file wins

The homepage markdown-negotiation rewrite lost to the static index.html, so
requesting text/markdown at / returned cached HTML. Move it into a legacy
route, which is evaluated before the filesystem. Add a readiness-test guard
that fails if any negotiation rewrite is shadowed by a static file, and assert
the root route.

* fix(phase-19/86): parse nested sequences in the stdlib YAML fallback

The PyYAML-free fallback stopped gathering a rule's lines at the first nested
sequence item, so a rule with an any_of block lost that block and every field
after it, and the engine rejected the rule for a missing explanation. The
gather now ends on indentation alone, so a deeper sequence stays part of its
rule. Verified the fallback matches PyYAML byte for byte on rules.yml, which
restores this lesson and the end-to-end safety gate that composes it on a
standard Python install.

* test(ci): run each lesson's own tests on push and pull request

CI executed the certification labs but never the 523 lessons, so a lesson that
could not import its own module reached main unnoticed. Adds
scripts/run_lesson_tests.py, which discovers each lesson's tests across the four
layouts in the repo and runs them the way the lesson docs say, and a
lesson-tests job that runs it. The runner installs no scientific dependencies:
a lesson that needs one is skipped by scanning its source, so the job stays
green while the stdlib lessons run for real. Documents the command in
CONTRIBUTING.

* test(ci): gate quiz answer-length bias and ratchet it down

The quizzes de-bias answer position but not answer length: on 84% of questions
the correct option is the longest by a wide margin, so a reader can guess it
without knowing the material. Adds scripts/check_quiz_bias.py, which measures
the share of length-biased questions and, in --check, fails when a change pushes
the rate above a baseline set to today's level. New or edited quizzes cannot add
bias, and the baseline ratchets down as quizzes are rebalanced. Wires it into
the curriculum workflow next to the position gate and documents the rule in
AGENTS.md. Reducing the existing rate is a separate content pass; this stops it
getting worse and makes it measurable.

* fix(phase-11/09): add module docstring, honest TS wording, ast import

Address review: give function_calling.py a module docstring like the sibling
lessons, align the TypeScript run_code tool description with the honest wording
already used on the Python side (a teaching filter, not real isolation), and add
the missing ast import to the docs example's import block so the example runs.

* fix(phase-19/39): fail the demo on a non-finite final loss

The success gate compared the final loss to the first with >=, so a NaN final
loss made the comparison false and the demo exited 0. Require the final loss to
be finite before comparing, so a diverged run is reported as a failure.

* docs(phase-19/47): note the torch 2.6 requirement for weights_only

Before torch 2.6 the weights_only loader had a known bypass (CVE-2025-32434),
so the security guarantee this lesson relies on holds only from 2.6 on. Say so
next to the weights_only explanation.

* fix(site): match the root markdown Accept header case-insensitively

Media types are case-insensitive, so a client sending Text/Markdown should
still reach the markdown route. Add the (?i) flag to the root route's Accept
matcher.

* fix(ci): fail a lesson when a real test fails next to a missing dep

The lesson-test runner skipped a suite whenever its output mentioned a missing
optional module, which could hide a genuine assertion failure in the same run.
Only skip when the output shows no test failure alongside the missing module.

* fix(ci): fail the quiz-bias check on an unreadable quiz file

The scan skipped a quiz.json it could not parse and carried on, so a malformed
file would leave the gate green on an incomplete scan. Collect read and parse
errors, report each file, and exit non-zero when any are found.
2026-09-24 17:27:49 +05:30
..