mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Extends the figure system into the phases that were still sparse: LLM engineering, multimodal, agents depth, alignment, plus vision/speech/genai remainders. Five new module files (1,959 LOC) on the shared LF toolkit: - figures-llmeng.js (P11/P13, 8): few-shot curve, chain-of-thought, constrained decoding, prompt-cache hit, semantic cache, function-call args, LLM-judge rubric, lost-in-the-middle - figures-multimodal.js (P12, 7): contrastive matrix, cross-attention fusion, modality projection, CFG guidance scale, VQ codebook, video patches, CTC align - figures-agents2.js (P14/P16, 8): ReWOO plan, tree-of-thoughts, self-refine, memory blocks, Voyager skills, LangGraph state, orchestration patterns, debate - figures-alignment2.js (P9/P18, 8): PPO clip, reward model, constitutional AI, actor-critic, interpretability probe, SAE features, jailbreak defense, scalable oversight - figures-foundations2.js (P4/P6/P8, 8): augmentation, transfer learning, BN train/eval, CTC collapse, MFCC pipeline, autoencoder bottleneck, normalizing flow, score matching Embedded in 39 figure-free lessons. Validated headless: all 173 registered figures (16 core + 157 module) mount with zero console errors; PPO clip, CLIP contrastive matrix, and tree-of-thoughts verified in light and dark.
Phase 8: Generative AI
Create images, video, audio, 3D, and more.
14 lessons, ~14 hours total. Each lesson ships: a 180-230 line doc, a runnable stdlib Python demo, a diagram, and a named skill for your agent.
| # | Lesson | Time |
|---|---|---|
| 01 | Generative Models — Taxonomy & History | ~45 min |
| 02 | Autoencoders & VAE | ~75 min |
| 03 | GANs — Generator vs Discriminator | ~75 min |
| 04 | Conditional GANs & Pix2Pix | ~75 min |
| 05 | StyleGAN | ~45 min |
| 06 | Diffusion Models — DDPM from Scratch | ~75 min |
| 07 | Latent Diffusion & Stable Diffusion | ~75 min |
| 08 | ControlNet, LoRA & Conditioning | ~75 min |
| 09 | Inpainting, Outpainting & Editing | ~75 min |
| 10 | Video Generation | ~45 min |
| 11 | Audio Generation | ~45 min |
| 12 | 3D Generation | ~45 min |
| 13 | Flow Matching & Rectified Flows | ~45 min |
| 14 | Evaluation — FID, CLIP Score, Human Preference | ~45 min |
See ROADMAP.md for the full cross-phase plan.