7 Commits
Author SHA1 Message Date
Rohit Ghumare 62cdefbe4c feat(site): full curriculum figure coverage (503 lessons) + lang-picker CSS cachebust (#380)
* feat(site): 39 more interactive figures — wave 3 across thin phases

Extends the figure system into the phases that were still sparse: LLM
engineering, multimodal, agents depth, alignment, plus vision/speech/genai
remainders. Five new module files (1,959 LOC) on the shared LF toolkit:

- figures-llmeng.js (P11/P13, 8): few-shot curve, chain-of-thought, constrained
  decoding, prompt-cache hit, semantic cache, function-call args, LLM-judge
  rubric, lost-in-the-middle
- figures-multimodal.js (P12, 7): contrastive matrix, cross-attention fusion,
  modality projection, CFG guidance scale, VQ codebook, video patches, CTC align
- figures-agents2.js (P14/P16, 8): ReWOO plan, tree-of-thoughts, self-refine,
  memory blocks, Voyager skills, LangGraph state, orchestration patterns, debate
- figures-alignment2.js (P9/P18, 8): PPO clip, reward model, constitutional AI,
  actor-critic, interpretability probe, SAE features, jailbreak defense,
  scalable oversight
- figures-foundations2.js (P4/P6/P8, 8): augmentation, transfer learning, BN
  train/eval, CTC collapse, MFCC pipeline, autoencoder bottleneck, normalizing
  flow, score matching

Embedded in 39 figure-free lessons. Validated headless: all 173 registered
figures (16 core + 157 module) mount with zero console errors; PPO clip,
CLIP contrastive matrix, and tree-of-thoughts verified in light and dark.

* feat(site): 81 animated SVG figures across capstone, agents, CV, NLP, infra

Wave 4a. Shift from slider widgets to unique, concept-specific SMIL-animated
SVG illustrations (no JS loops, no real compute, light DOM). 10 new module
files, each figure a distinct visual matched to its lesson:

- figures-capstone-a/b (P19, 16): tokenizer merges, sliding window, training
  loop, DPO, RAG flow, eval grid sweep, sandbox runner, safety checkpoints
- figures-agents3 (P14, 8): HTN tree, workflow chain, actor mailbox, debate
  convergence, computer-use cursor, voice pipeline, injection hijack, cascade
- figures-nlp3 (P5, 9): POS tags, dependency arcs, QA span, summarize collapse,
  topic drift, coref links, NLI router, relation triples, constrained decode
- figures-cv2 (P4, 8): detection NMS, segmentation flood, GAN, diffusion denoise,
  NeRF rays, CLIP matrix, metric embedding, depth sweep
- figures-llms3 (P7/P10, 8): MoE routing, encoder-decoder, RNN vs parallel,
  speculative draft-verify, multi-token predict, self-critique, loss masking,
  activation recompute
- figures-autonomous2 (P15, 8): AlphaEvolve loop, Darwin-Godel archive, bounded
  gates, circuit breaker, checkpoint replay, cost governor, injection boundary
- figures-swarms2 (P16, 8): consensus wave, auction, stigmergy, hierarchy token,
  message bus, roles, blackboard, speaker election
- figures-infra2 (P17, 8): cache-aware router, cold start, model cascade,
  prefill/decode split, batch lanes, semantic cache, edge bandwidth, load waves
- figures-systems3 (P11/12/13/8/6, 8): masked diffusion, any-to-any stream,
  video diffusion, inpaint, agentic RAG, MCP NxM, A2A lifecycle, RVQ codec

Embedded in 80 figure-free lessons. Validated headless: all 254 registered
figures mount with zero console errors, 1062 SMIL animation nodes present,
samples (NMS, stigmergy, RAG, detection) verified rendering.

* feat(site): 80 more animated SVG figures across capstone, CV, speech, tools

Wave 4b. Ten more SMIL-animated module files, each figure a unique
concept-specific illustration (no JS loops, no real compute, light DOM):

- figures-capstone-c/d (P19, 16): embedding lookup, transformer block, GPT
  assembly, weight remap, grad accumulation, atomic checkpoint, HyDE, BLEU,
  reliability diagram, ZeRO shard, pipeline bubble, constitution loop
- figures-cv3 (P4, 8): RoIAlign, latent compression, CTC, pose heatmap,
  gaussian splat, rectified flow, open-vocab, track association
- figures-speech2 (P6, 8): ASR attention, EER crossover, TTS stack, codec
  tokens, VAD cascade, full-duplex, WER alignment, voice factorize
- figures-multimodal2 (P12, 8): patch-n-pack, LLaVA projector, M-RoPE axes,
  video token budget, action tokens, doc layout, MaxSim, agent loop
- figures-tools2 (P13, 8): tool loop, parallel fanout, schema routing, client
  merge, transport handshake, task lifecycle, tool poisoning, router failover
- figures-agents4 (P14, 8): memory fusion, crew-vs-flow, handoff, subagent
  isolation, SWE-bench gate, agent-human gap, span tree, eval layers
- figures-swarms3 (P16, 8): contract-net, work-stealing, handoff routing,
  agent-card discovery, debate topology, theory-of-mind, CTDE, checkpoint
- figures-genai3 (P8/P5, 8): VAR next-scale, FID, PatchGAN, StyleGAN mapping,
  hybrid retrieval, Matryoshka, entity linking, needle-in-haystack
- figures-misc2 (P15/P17/P11, 8): propose-then-commit, priority tiers, research
  loop, speculative tree, gateway fallback, sequential test, schema funnel

Embedded in 80 figure-free lessons. Validated headless: all 334 registered
figures mount with zero console errors, 2061 SMIL animation nodes; samples
(contract-net auction, TTS stack) verified rendering.

* fix(site): bump asset versions so the language picker CSS refreshes

The picker markup and CSS shipped, but the style.css link kept the old
?v=20260525a query, so returning visitors' browsers served cached CSS
without the .lang-panel rules and the picker rendered unstyled and
always-open. Bump every asset version (and version the new langs.js /
lang-picker.js) to force a fresh fetch.

* feat(site): full curriculum figure coverage — 169 animated figures, all 503 lessons

Wave 5 completes the interactive figure system: every lesson in every phase
now carries a concept-specific animated SVG. Sixteen new module files
(7,688 LOC), each figure a unique SMIL illustration with motion craft applied
throughout (spline ease-out entries from opacity 0 at 95 percent scale,
staggered cascades, exits faster than entries, calm 2.5-6s loops, no JS
animation loops, no real compute):

- figures-capstone-e/f/g/h/i (P19, 49 figures)
- figures-alignment3/4 (P18, 23)
- figures-workbench (P14, 15): the agent workbench mini-track animated
- figures-tools3 (P13, 11)
- figures-setup (P00, 12): commit DAG, GPU dispatch, secret injection, venv
  isolation, docker layers, LSP round trip, flame graph, more
- figures-foundations3 (P01/02/09, 11)
- figures-visaudio4 (P04/06/08, 9)
- figures-nlp5 (P05/07, 8)
- figures-llmstack5 (P10/11/12, 11)
- figures-autoswarm5 (P15/16, 12)
- figures-infra4 (P17, 8)

Coverage: 0 figure-free lessons remain; all 503 lesson docs carry a figure.
Validated headless: 506 registered figures mount with zero console errors,
4,505 SMIL animation nodes.
2026-08-01 14:43:01 +01:00
Rohit Ghumare dda194f840 fix(quiz): correct answer is always in the same position (slot B) (#381)
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.

scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.

Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.

The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.

Fixes #368
2026-08-01 14:24:15 +01:00
Rohit Ghumare 4026f961a1 fix(phase-05): randomize correct positions + fix ambiguous MCQ 2026-05-23 01:10:18 +01:00
Rohit Ghumare c7bb439018 feat(phase-05/04): add quiz.json 2026-05-23 01:04:58 +01:00
Rohit Ghumare 42471796f1 fix(outputs): rename skills to remove name collisions (#141)
* fix(outputs): rename skills to remove name collisions

* fix(outputs): restore skill- prefix in renamed frontmatter names
2026-05-22 14:40:27 +01:00
Rohit Ghumare da76702a05 fix(phase-05): remove dead asset links and stray markdown link match
The 10 audit findings in phase 05 all pointed at SVG assets that were
never created. Nine were broken image embeds (./assets/<name>.svg) and
one was a code-fence false positive where '[tool_name](**args)' inside a
Python snippet looked like a Markdown link to the audit's regex.

This commit:
- Removes the nine broken figure embeds across lessons 01-09. The
  surrounding prose stands on its own; no caption text needed rewriting.
- Splits the offending Python expression in lesson 17 onto two lines
  (fn = tools[tool_name]; result = fn(**args)) so '](**args)' no longer
  appears as adjacent characters.
2026-05-20 18:10:25 +01:00
Rohit Ghumare fd53ad179e feat(phase-05/04): GloVe, FastText, and subword embeddings
Three approaches side by side. GloVe as co-occurrence matrix
factorization with weighted log-loss. FastText as sum of character
n-grams (with OOV-tolerant lookup demonstrated). BPE as iterative
pair-merging that produces the subword vocabulary every modern LLM
tokenizer uses.

The teaching thread: Word2Vec had no OOV story. FastText fixed it for
words. BPE solved it for the whole vocabulary at a different level and
became the transformer-era default.

Ship artifact: tokenizer-picker skill with vocab-size targets (32k
English, 64-100k multilingual) and the reproducibility pitfall
(tokenizer-model mismatch is the single most common silent production
bug).

~45 minutes. Prerequisites lesson 05/03.
2026-04-22 15:45:57 +01:00