5 Commits
Author SHA1 Message Date
Rohit Ghumare dda194f840 fix(quiz): correct answer is always in the same position (slot B) (#381)
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.

scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.

Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.

The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.

Fixes #368
2026-08-01 14:24:15 +01:00
Rohit Ghumare d1cb9d1933 feat: book edition pipeline, fundamentals-first headlines, animated figures, verified bug fixes (#348)
Book pipeline: six-volume EPUB/PDF compilation built by CI from lesson
sources (book/, scripts/build_book.py, themed title pages with edition
stamps, site-matching print theme), attached to every GitHub release.
Homepage Books section and README section link the latest release.

Fundamentals-first headline policy across the course: 16 lesson titles
and 30+ taglines/section headings now lead with the concept (agent
state machines, actor model, role-based teams, memory paging, serving
engine internals, permission modes); framework and product names are
demoted to attributed in-body examples. README, ROADMAP, quizzes, and
prerequisite references synced. New agent-memory taxonomy section maps
memory types to representative implementations.

Vendor-neutral model policy: runnable defaults read the LLM_MODEL env
var with undated aliases; dated snapshot ids removed; multi-provider
phrasing in the setup lesson.

Lessons deepened with original material: prediction-game origins of
perplexity (05/16), scripted-era chatbot lineage 1950-2001 (05/17),
causal-triangle derivation from prefix averaging plus GPT-5 date fix
(07/07). Three new animated site figures back them (figures-history.js).

llms.txt now carries per-lesson raw markdown links so agents can fetch
full lesson text directly.

Bug fixes verified with executed repros: capstone solved flag keyed to
test results, 405B cost estimator overflow, f-string crash on
Python <3.12, no-torch demo path, negative stable BCE, all-zero
stationary distribution, inverted Cohens d, per-lesson quiz panel,
lesson-fetch retry with honest errors, decision-trees doc completed,
editor shortcuts, rustc run command, Docker python3.12 build with doc
sync, git lesson fork flow, FIPA receiver field, fnm under Rosetta,
15 curl-verified link fixes, remaining imdb dataset id spot.
2026-07-25 20:24:56 +01:00
Rohit Ghumare cb55ea0cc6 feat(site): curriculum-wide interactive figure system (134 widgets, 13 modules) (#279)
* feat(site): interactive training-foundations figures in 5 lessons

Add five theme-aware interactive widgets to lesson-figures.js, embedded
via the existing ```figure fence:

- gradient-descent (P1.08 optimization): drag learning rate, watch the
  descent path converge or diverge past lr > 1
- softmax-temperature (P3.04 activations): divide logits by T, reshape
  the distribution from argmax to uniform
- bias-variance (P2.10): slide model complexity across the U-shaped
  test-error curve, see the sweet spot move
- l2-regularization (P3.07): raise lambda, watch every weight shrink
- lr-schedule (P3.09): compare warmup, cosine, step, exponential decay

Validated headless: all five mount with no console errors, sliders and
selects drive re-render, both light and dark themes render correctly.

* feat(site): interactive LLM-internals figures in 5 lessons

Batch 2, building on the same widget system:

- sampling-decoder (P10.04 mini-gpt): temperature then top-k then top-p
  filtering over the logits, survivors renormalized
- scaling-laws (P7.13): Chinchilla loss from params and tokens, with the
  20-tokens-per-parameter compute-optimal rule
- quantization (P10.11): bits per weight against model size and the
  precision lost at fp16/int8/int4/int2
- rope-explorer (P7.04): rotary frequencies across position and dimension,
  base controls wavelength and usable context
- lora-params (P11.08): rank against the 2r/d trainable fraction

Validated headless: all five mount with no console errors, sliders and
selects drive re-render, both light and dark render correctly.

* feat(site): interactive evaluation and representation figures in 5 lessons

Batch 3, same widget system:

- precision-recall-threshold (P2.09 model-evaluation): slide the cutoff
  across two class distributions, watch precision/recall/F1 trade
- cross-entropy-loss (P3.05 loss-functions): -log(p_true), the price of
  being confident and wrong
- cosine-similarity (P11.04 embeddings): the angle between two vectors is
  the similarity, magnitude drops out
- tokenizer-tradeoff (P10.01 tokenizers): vocab size against tokens-per-word
  and the embedding table cost
- rag-chunking (P11.06 rag): chunk size, overlap, and top-k against chunk
  count and context tokens per query

Validated headless: all five mount with no console errors, math checks out
(thr 0.8 -> P 1.00/R 0.11, -ln(0.05)=2.996, cos 90 deg = 0, 224 chunks),
sliders drive re-render, both light and dark render correctly.

* feat(site): interactive figure system — 74 new widgets across 11 phases

Expand the lesson-figure system from a handful of widgets into a curriculum-wide
library. Refactor lesson-figures.js to expose a shared LF toolkit (el, svgEl,
slider, select, fmtInt, clamp, lerp, raf, register) and split widgets into eight
per-phase module files that plug in via LF.register.

New module files (3,682 LOC) and the concepts they make draggable:
- figures-math.js (P1, 11): vector projection, matrix transform + determinant,
  eigenvectors, derivative tangent, chain rule, gaussian, bayes update,
  entropy/KL, PCA axes, fourier synthesis, convex vs nonconvex
- figures-ml.js (P2, 10): regression fit/MSE, logistic boundary, SVM margin,
  kNN smoothness, k-means steps, tree depth, feature scaling, naive bayes,
  class imbalance, k-fold CV
- figures-dl.js (P3, 9): perceptron boundary, MLP forward pass, vanishing
  gradients, optimizer trajectories, weight-init variance, dropout, batchnorm,
  learning curves, gradient clipping
- figures-vision-speech.js (P4/P6, 8): convolution kernel, pooling, receptive
  field, conv output size, CNN params, spectrogram window, mel scale, aliasing
- figures-transformers.js (P5/P7, 9): attention heatmap, multihead split, causal
  mask, sqrt(d_k) scaling, word2vec arithmetic, BPE merges, GQA sharing,
  residual stream, flash-attention memory
- figures-genai-rl.js (P8/P9, 9): diffusion denoise, noise schedule, VAE latent,
  GAN minimax, Q-learning gridworld, value iteration, epsilon-greedy, discount
  horizon, policy-gradient ascent
- figures-llms-systems.js (P10/P12/P13, 9): beam search, speculative decoding,
  MoE routing, context window, perplexity, continuous batching, ViT patches,
  multimodal fusion, MCP round trip
- figures-agents-alignment.js (P11/P14/P16/P18, 9): agent loop, ReAct trace,
  tool routing, swarm message scaling, supervisor tree, RLHF reward-KL,
  DPO margin, context budget, guardrail gates

Each widget embedded in its lesson via the figure fence (74 lessons). All
theme-aware through CSS vars, vanilla ES5, no dependencies.

Validated headless: all 90 registered figures (16 prior + 74) mount with zero
console errors in a master harness; rich SVG visualizations (attention heatmap,
gridworld policy, convolution feature map, swarm graphs) render correctly in
both light and dark.

* feat(site): 44 more interactive figures — NLP, LLM internals, infra, autonomy

Wave 2 extends the figure system into the phases that were still bare,
plus deeper coverage of the large NLP and LLM phases. Five new module
files (2,219 LOC), each plugging into the shared LF toolkit:

- figures-math2.js (P1, 9): SVD low-rank reconstruction, tensor broadcasting,
  log-sum-exp stability, Lp unit balls, monte-carlo pi, system conditioning,
  random-walk diffusion, roots of unity, graph degree
- figures-nlp2.js (P5, 8): BoW/TF-IDF, RNN unroll, LSTM gates, seq2seq
  alignment, edit distance, n-gram backoff, BIO tagging, sentiment logits
- figures-llms2.js (P10, 9): RMSNorm vs LayerNorm, SwiGLU, RLHF pipeline,
  DPO loss, paged KV cache, expert capacity, sliding-window attention,
  differential attention, weight tying
- figures-infra.js (P17, 9): data/tensor/pipeline parallelism, ZeRO sharding,
  GPU memory breakdown, throughput-latency, autoscaling, cost-per-token,
  roofline
- figures-frontier.js (P15/P19, 9): task decomposition, reflection loop,
  memory consolidation, world-model rollout, autonomy oversight, pass@k,
  eval-harness matrix, canary rollout, trace spans

Embedded in 44 lessons via the figure fence. Validated headless: all 134
registered figures (16 core + 118 module) mount with zero console errors in
a full harness; pipeline-bubble, SVD energy, and trace-span visualizations
render correctly in light and dark.

* fix(site): address review findings on figure widgets

- sampling-decoder: formula now reads 'cumulative >= p' (nucleus keeps the
  smallest set covering p, matching the implementation)
- supervisor-hierarchy: drop the dead capped-total accumulator; show the exact
  geometric total and note when the diagram caps a level at 64 so the number
  and the drawn nodes stay consistent; handle b=1 (total = depth + 1) instead
  of the closed form that is undefined at b=1
- image-patch-tokens: use ceil(size/patch) so non-divisible sizes count the
  partial patch row; formula shows the ceil and meta notes the padded size
- debugging-neural-networks: normalize the one-off Type 'Practice' to 'Build'

Verified in browser: all three widgets render with the corrected text/math,
no console errors.

Skipped: the 'figure fence is not an approved language tag' findings. lesson.html
keys on codeLang === 'figure' to emit the widget mount point; the fence body is
the figure id. Renaming the fence to the figure id would stop it rendering.
There is no fence-language allowlist for these lesson docs.
2026-06-10 19:35:56 +01:00
Rohit Ghumare e2c27325aa feat(phase-17): quiz backfill, 0/28 -> 28/28 (#151)
* feat(phase-17/01): add quiz.json

* feat(phase-17/02): add quiz.json

* feat(phase-17/03): add quiz.json

* feat(phase-17/04): add quiz.json

* feat(phase-17/05): add quiz.json

* feat(phase-17/06): add quiz.json

* feat(phase-17/07): add quiz.json

* feat(phase-17/08): add quiz.json

* feat(phase-17/09): add quiz.json

* feat(phase-17/10): add quiz.json

* feat(phase-17/11): add quiz.json

* feat(phase-17/12): add quiz.json

* feat(phase-17/13): add quiz.json

* feat(phase-17/14): add quiz.json

* feat(phase-17/15): add quiz.json

* feat(phase-17/16): add quiz.json

* feat(phase-17/17): add quiz.json

* feat(phase-17/18): add quiz.json

* feat(phase-17/19): add quiz.json

* feat(phase-17/20): add quiz.json

* feat(phase-17/21): add quiz.json

* feat(phase-17/22): add quiz.json

* feat(phase-17/23): add quiz.json

* feat(phase-17/24): add quiz.json

* feat(phase-17/25): add quiz.json

* feat(phase-17/26): add quiz.json

* feat(phase-17/27): add quiz.json

* feat(phase-17/28): add quiz.json

* chore(catalog): rebuild after phase 17 quiz backfill

* fix(phase-17): randomize correct-answer positions across quizzes
2026-05-23 01:12:25 +01:00
Rohit Ghumare daef29e8ae feat(phase-17/07): TensorRT-LLM on Blackwell with FP8 and NVFP4 2026-04-23 18:45:40 +01:00