mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
main
1810
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1bafaa88bb | chore(site): rebuild data.js | ||
|
|
29c480a6f9 |
fix(site): refresh cached assets and align project commands (#513)
* fix(site): revalidate pages and version deployed assets * fix(projects): align setup cards and add copy feedback * fix(site): keep desktop navigation compact and controls usable |
||
|
|
61bc244ced | chore(site): rebuild data.js | ||
|
|
3ecf630b0f |
feat(projects): add 48 builds and a 52-project roadmap (#497)
* feat(projects): workflow-hooks stage 1 * feat(projects): workflow-hooks stage 2 * feat(projects): workflow-hooks stage 3 * feat(projects): workflow-hooks stage 4 * fix(projects): preserve sponsor navigation in project footer * feat(projects): agent-budget-planner stage 1 * feat(projects): agent-budget-planner stage 2 * feat(projects): agent-budget-planner stage 3 * feat(projects): agent-budget-planner stage 4 * feat(projects): agent-trace-debugger stage 1 * feat(projects): agent-trace-debugger stage 2 * feat(projects): agent-trace-debugger stage 3 * feat(projects): agent-trace-debugger stage 4 * feat(projects): browser-agent stage 1 * feat(projects): browser-agent stage 2 * feat(projects): browser-agent stage 3 * feat(projects): browser-agent stage 4 * feat(projects): calendar-focus-planner stage 1 * feat(projects): calendar-focus-planner stage 2 * feat(projects): calendar-focus-planner stage 3 * feat(projects): calendar-focus-planner stage 4 * feat(projects): changelog-writer-from-git stage 1 * feat(projects): changelog-writer-from-git stage 2 * feat(projects): changelog-writer-from-git stage 3 * feat(projects): changelog-writer-from-git stage 4 * feat(projects): cloud-agent-with-aws-strands stage 1 * feat(projects): cloud-agent-with-aws-strands stage 2 * feat(projects): cloud-agent-with-aws-strands stage 3 * feat(projects): cloud-agent-with-aws-strands stage 4 * feat(projects): csv-sql-question-workbench stage 1 * feat(projects): csv-sql-question-workbench stage 2 * feat(projects): csv-sql-question-workbench stage 3 * feat(projects): csv-sql-question-workbench stage 4 * feat(projects): dataset-split-auditor stage 1 * feat(projects): dataset-split-auditor stage 2 * feat(projects): dataset-split-auditor stage 3 * feat(projects): dataset-split-auditor stage 4 * feat(projects): desktop-control stage 1 * feat(projects): desktop-control stage 2 * feat(projects): desktop-control stage 3 * feat(projects): desktop-control stage 4 * feat(projects): distributed-eval-farm stage 1 * feat(projects): distributed-eval-farm stage 2 * feat(projects): distributed-eval-farm stage 3 * feat(projects): distributed-eval-farm stage 4 * feat(projects): doc-qa-with-citations stage 1 * feat(projects): doc-qa-with-citations stage 2 * feat(projects): doc-qa-with-citations stage 3 * feat(projects): doc-qa-with-citations stage 4 * feat(projects): document-extraction-desk stage 1 * feat(projects): document-extraction-desk stage 2 * feat(projects): document-extraction-desk stage 3 * feat(projects): document-extraction-desk stage 4 * feat(projects): durable-agent-jobs stage 1 * feat(projects): durable-agent-jobs stage 2 * feat(projects): durable-agent-jobs stage 3 * feat(projects): durable-agent-jobs stage 4 * feat(projects): feedback-theme-board stage 1 * feat(projects): feedback-theme-board stage 2 * feat(projects): feedback-theme-board stage 3 * feat(projects): feedback-theme-board stage 4 * feat(projects): harness-bench stage 1 * feat(projects): harness-bench stage 2 * feat(projects): harness-bench stage 3 * feat(projects): harness-bench stage 4 * feat(projects): inbox-triage-desk stage 1 * feat(projects): inbox-triage-desk stage 2 * feat(projects): inbox-triage-desk stage 3 * feat(projects): inbox-triage-desk stage 4 * feat(projects): json-schema-output-guard stage 1 * feat(projects): json-schema-output-guard stage 2 * feat(projects): json-schema-output-guard stage 3 * feat(projects): json-schema-output-guard stage 4 * feat(projects): llm-gateway-with-fallbacks stage 1 * feat(projects): llm-gateway-with-fallbacks stage 2 * feat(projects): llm-gateway-with-fallbacks stage 3 * feat(projects): llm-gateway-with-fallbacks stage 4 * feat(projects): local-model-eval-harness stage 1 * feat(projects): local-model-eval-harness stage 2 * feat(projects): local-model-eval-harness stage 3 * feat(projects): local-model-eval-harness stage 4 * feat(projects): mcp-at-scale stage 1 * feat(projects): mcp-at-scale stage 2 * feat(projects): mcp-at-scale stage 3 * feat(projects): mcp-at-scale stage 4 * feat(projects): mcp-at-scale stage 5 * feat(projects): meeting-notes-to-actions stage 1 * feat(projects): meeting-notes-to-actions stage 2 * feat(projects): meeting-notes-to-actions stage 3 * feat(projects): meeting-notes-to-actions stage 4 * feat(projects): memory-server stage 1 * feat(projects): memory-server stage 2 * feat(projects): memory-server stage 3 * feat(projects): memory-server stage 4 * feat(projects): multi-agent-code-review-panel stage 1 * feat(projects): multi-agent-code-review-panel stage 2 * feat(projects): multi-agent-code-review-panel stage 3 * feat(projects): multi-agent-code-review-panel stage 4 * feat(projects): postmortem-writer stage 1 * feat(projects): postmortem-writer stage 2 * feat(projects): postmortem-writer stage 3 * feat(projects): postmortem-writer stage 4 * feat(projects): pr-review-reporter stage 1 * feat(projects): pr-review-reporter stage 2 * feat(projects): pr-review-reporter stage 3 * feat(projects): pr-review-reporter stage 4 * feat(projects): prompt-regression-tester stage 1 * feat(projects): prompt-regression-tester stage 2 * feat(projects): prompt-regression-tester stage 3 * feat(projects): prompt-regression-tester stage 4 * feat(projects): rag-freshness-pipeline stage 1 * feat(projects): rag-freshness-pipeline stage 2 * feat(projects): rag-freshness-pipeline stage 3 * feat(projects): rag-freshness-pipeline stage 4 * feat(projects): report-judge stage 1 * feat(projects): report-judge stage 2 * feat(projects): report-judge stage 3 * feat(projects): report-judge stage 4 * feat(projects): research-report-agent stage 1 * feat(projects): research-report-agent stage 2 * feat(projects): research-report-agent stage 3 * feat(projects): research-report-agent stage 4 * feat(projects): research-report-agent stage 5 * feat(projects): research-report-agent stage 6 * feat(projects): research-report-agent stage 7 * feat(projects): retrieval-evaluation-lab stage 1 * feat(projects): retrieval-evaluation-lab stage 2 * feat(projects): retrieval-evaluation-lab stage 3 * feat(projects): retrieval-evaluation-lab stage 4 * feat(projects): rust-agent-shell stage 1 * feat(projects): rust-agent-shell stage 2 * feat(projects): rust-agent-shell stage 3 * feat(projects): rust-agent-shell stage 4 * feat(projects): sandbox-ladder stage 1 * feat(projects): sandbox-ladder stage 2 * feat(projects): sandbox-ladder stage 3 * feat(projects): sandbox-ladder stage 4 * feat(projects): self-improving-skill-loop stage 1 * feat(projects): self-improving-skill-loop stage 2 * feat(projects): self-improving-skill-loop stage 3 * feat(projects): self-improving-skill-loop stage 4 * feat(projects): semantic-notes-search stage 1 * feat(projects): semantic-notes-search stage 2 * feat(projects): semantic-notes-search stage 3 * feat(projects): semantic-notes-search stage 4 * feat(projects): skill-installer stage 1 * feat(projects): skill-installer stage 2 * feat(projects): skill-installer stage 3 * feat(projects): skill-installer stage 4 * feat(projects): skill-router stage 1 * feat(projects): skill-router stage 2 * feat(projects): skill-router stage 3 * feat(projects): skill-router stage 4 * feat(projects): skill-scanner stage 1 * feat(projects): skill-scanner stage 2 * feat(projects): skill-scanner stage 3 * feat(projects): skill-scanner stage 4 * feat(projects): skill-validator stage 1 * feat(projects): skill-validator stage 2 * feat(projects): skill-validator stage 3 * feat(projects): skill-validator stage 4 * feat(projects): source-grounded-study-coach stage 1 * feat(projects): source-grounded-study-coach stage 2 * feat(projects): source-grounded-study-coach stage 3 * feat(projects): source-grounded-study-coach stage 4 * feat(projects): support-agent-with-google-adk stage 1 * feat(projects): support-agent-with-google-adk stage 2 * feat(projects): support-agent-with-google-adk stage 3 * feat(projects): support-agent-with-google-adk stage 4 * feat(projects): tiny-coding-agent stage 1 * feat(projects): tiny-coding-agent stage 2 * feat(projects): tiny-coding-agent stage 3 * feat(projects): tiny-coding-agent stage 4 * feat(projects): token-counter-and-cost-meter stage 1 * feat(projects): token-counter-and-cost-meter stage 2 * feat(projects): token-counter-and-cost-meter stage 3 * feat(projects): token-counter-and-cost-meter stage 4 * feat(projects): tool-call-firewall stage 1 * feat(projects): tool-call-firewall stage 2 * feat(projects): tool-call-firewall stage 3 * feat(projects): tool-call-firewall stage 4 * feat(projects): typed-workflow-agent-with-mastra stage 1 * feat(projects): typed-workflow-agent-with-mastra stage 2 * feat(projects): typed-workflow-agent-with-mastra stage 3 * feat(projects): typed-workflow-agent-with-mastra stage 4 * feat(projects): visual-evidence-library stage 1 * feat(projects): visual-evidence-library stage 2 * feat(projects): visual-evidence-library stage 3 * feat(projects): visual-evidence-library stage 4 * feat(projects): voice-note-transcriber-pipeline stage 1 * feat(projects): voice-note-transcriber-pipeline stage 2 * feat(projects): voice-note-transcriber-pipeline stage 3 * feat(projects): voice-note-transcriber-pipeline stage 4 * feat(projects): web-change-brief stage 1 * feat(projects): web-change-brief stage 2 * feat(projects): web-change-brief stage 3 * feat(projects): web-change-brief stage 4 * feat(projects): workflow-hooks stage 1 * feat(projects): workflow-hooks stage 2 * feat(projects): workflow-hooks stage 3 * feat(projects): workflow-hooks stage 4 * feat(projects): publish the practical application catalog * feat(projects): add 52 planned builds to the roadmap * fix(projects): make lesson animations explain computed changes Stage mechanisms should expose the decisions learners are implementing. Preserve record identity as state changes and keep project context below the selected lesson. * fix(projects): agent-budget-planner stage 01 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): agent-budget-planner stage 02 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): agent-budget-planner stage 03 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): agent-budget-planner stage 04 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): agent-trace-debugger stage 03 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): browser-agent stage 02 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): browser-agent stage 03 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): browser-agent stage 04 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): calendar-focus-planner stage 01 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): cloud-agent-with-aws-strands stage 01 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): cloud-agent-with-aws-strands stage 02 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): cloud-agent-with-aws-strands stage 04 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): desktop-control stage 04 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): distributed-eval-farm stage 01 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): distributed-eval-farm stage 02 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): distributed-eval-farm stage 04 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): document-extraction-desk stage 03 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): durable-agent-jobs stage 01 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): durable-agent-jobs stage 02 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): durable-agent-jobs stage 03 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): durable-agent-jobs stage 04 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): feedback-theme-board stage 01 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): inbox-triage-desk stage 03 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): llm-gateway-with-fallbacks stage 01 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): llm-gateway-with-fallbacks stage 04 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): mcp-at-scale review corrections Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): meeting-notes-to-actions stage 01 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): meeting-notes-to-actions stage 04 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): memory-server stage 04 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): pr-review-reporter stage 01 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): pr-review-reporter stage 03 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): pr-review-reporter stage 04 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): research-report-agent stage 01 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): retrieval-evaluation-lab review corrections Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): rust-agent-shell stage 01 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): rust-agent-shell stage 04 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): semantic-notes-search stage 03 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): skill-installer review corrections Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): skill-router stage 02 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): skill-router stage 04 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): tiny-coding-agent stage 02 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): voice-note-transcriber-pipeline stage 03 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): web-change-brief stage 03 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): workflow-hooks stage 01 contracts Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * fix(projects): site review corrections Keep the stage contract, learner scaffold and verification evidence consistent at the reviewed failure boundary. * docs(projects): record verified review corrections |
||
|
|
7589cf88e1 | chore(site): rebuild data.js | ||
|
|
d864ffe8d8 | fix: read readiness inputs as UTF-8 (#501) | ||
|
|
bf7791e140 | chore(site): rebuild data.js | ||
|
|
dbd08c5387 |
feat(site): add complete lesson button at end of lesson (#482)
* feat(site): add complete lesson button at end of lesson Learners navigating lesson pages need an explicit action at the end of each lesson to mark it as completed without having to rely solely on the top checkpoint loop. This adds a completion section after lesson exercises, updates local progress immediately, and provides clear visual feedback and unmark capability. Fixes #479. * fix(site): disable complete button when completed and remove aria-pressed Avoid exposing the completed action as a toggle control on screen readers while keeping the dedicated unmark button available. * test(site): remove comments in test per contributing guidelines * fix(site): clear stale completion status and translate the panel labels The status line under the completion buttons was only set by the panel's own clicks, so completing or reopening the lesson from the Continue control at the top left a message that contradicted the button. The status now records the state it describes, and a sync from any source clears it once the lesson's state no longer matches. The panel's labels join site/ui-strings.json so the interface translation layer covers them. The check mark is now a literal character, so the "Completed ✓" key matches the page source. * fix(site): pin translations for the lesson completion panel Every interface string in site/ui-strings.json carries a curated translation for each CI language. A key without one falls back to the machine translator, which renders short labels poorly (the unpinned "Sponsor us" key came out as "He is our sponsor" in Spanish). Pin the eight completion-panel labels for zh, hi, es, ar, fr, pt, tr, and vi, using the terms the existing lesson strings use, and make the panel test require a translation of every label in every language. --------- Co-authored-by: Rohit Ghumare <ghumare64@gmail.com> |
||
|
|
968da0791b | chore(site): rebuild data.js v2026.10 | ||
|
|
a05d392590 |
fix: repair broken lesson references and guard homepage storage (#495)
* fix(site): guard theme storage access on the homepage and about page
Reading or writing localStorage throws a SecurityError when storage is
blocked (strict privacy settings, some embedded webviews). site/app.js
read the saved theme outside any try/catch, before it registered its
DOMContentLoaded handler, so the throw stopped the whole homepage from
initializing. The about page's inline theme script had the same
unguarded read and write.
Wrap all four accesses the way the other pages already do and fall
back to the system theme. Bump the app.js cache key so browsers pick
up the fix, and give app.js its own release constant in the cache-key
test.
Fixes #490
* fix(lessons): repair dataset, model, and tool references that no longer resolve
Lesson snippets pointed at resources that are gone or never existed, so
they fail when a student runs them:
- wikimedia/wikipedia only ships 20231101.* configs; 20220301.en is gone
- MMAU-Pro lives at gamma-lab-umd/MMAU-Pro, with audio_path and answer
fields; open-ended rows are filtered out of the exact-match score
- meta-llama/Llama-3-70B-Instruct is meta-llama/Meta-Llama-3-70B-Instruct
- Depth Anything V2 for transformers is
depth-anything/Depth-Anything-V2-Large-hf
- microsoft/deberta-v3-large-mnli and microsoft/BEATs-base are not on
the Hub; BEATs checkpoints ship with microsoft/unilm
- sayakpaul/sd-lora-ghibli does not exist; use a published SD 1.5
Ghibli LoRA with its trigger phrase
- Qwen/Qwen3-0.6B-spec and meta-llama/Llama-3.2-1B-Instruct-spec are
not real draft models
- gemini-3-pro is not a Gemini model code, gemini-1.5-pro is shut down,
and the google.generativeai SDK reached end of life on 2025-11-30
- vLLM replaced --speculative-model and --num-speculative-tokens with
--speculative-config, and --dtype has no float8_e4m3fn choice
- convert_hf_to_gguf.py cannot write q4_k_m; llama-quantize does
- AutoGPTQ and AutoAWQ are archived; GPTQModel and LLM Compressor are
the maintained successors, and vLLM reads their quantization method
from the checkpoint config
Fixes #493
* fix(lessons): replace dead reference links
A link check over every lesson found 46 references that return 404 or
410 or no longer resolve. Each replacement was fetched and matched
against the cited title or content:
- pages that moved on the same site: vLLM, librosa, Letta, Apollo
Research, Stability AI, NVIDIA, NeurIPS, the MCP spec, the Claude
docs, the A2A spec, the Julia docs, W&B, Baseten, and Arena (formerly
LMSYS Chatbot Arena, whose old domain no longer resolves)
- papers and books pointed at DOIs or publisher pages: Friedman on
gradient boosting, Golub and Van Loan, Kuttruff, Littman's thesis,
Milne and Witten, and Zave and Jackson (whose DOI was wrong)
- Wayback snapshots where the source only survives in the archive:
Poynton's color space tour and a model-routing article
- citations of pages that never existed replaced with the real source:
the DINOv2 paper, the MMAU-Pro project page, Anthropic's Contextual
Retrieval post, the February 2026 risk report, and the correct
Anthropic alignment post
- removed where no source exists: a Spinning Up DQN page, the
andrewgarst/agentic_harness repo, and an Akira blog post
URLs in code-file reference headers get the same treatment.
* chore(site): rebuild data.js
* Revert "chore(site): rebuild data.js"
The main-only CI job rebuilds site/data.js after merge, so the PR
should not carry it; a stale "Last built" line would conflict.
This reverts commit
|
||
|
|
c257687012 | chore(site): rebuild data.js | ||
|
|
744520ac08 | fix(site): refresh cached header for newsletter signup (#494) | ||
|
|
c0cac771e8 | chore(site): rebuild data.js | ||
|
|
6c0c06dce2 |
feat(site): add Substack newsletter signup (#492)
* feat(site): add Substack newsletter signup * fix(site): restore compact newsletter signup design |
||
|
|
8bc378c2e0 | chore(site): rebuild data.js | ||
|
|
4b9c61beb2 |
feat(site): animate the MCPA and Claude certification figures (#489)
* feat(site): animate the Claude certification labs Each Claude certification lab now opens its output with an animated strip. The inputs flow into the formula and then into the result: a gauge for threshold, equation, and readiness labs, one bar per option for decision labs, and a stage rail for pipeline labs. The strip replaces the HTML meter bars and stage cards that showed the same values. The labs now share one builder and use the course lf card markup, so the runtime names each SVG from the card title and caption. Bars grow through stroke-dash rather than transforms, because the runtime marks nodes it reconciles after a slider drag with transform-origin center. The redundant control labelling is gone, since LF.slider and LF.select already bind their labels. * feat(site): animate the MCPA certification figures The 34 MCPA lesson figures were static SVG strings. Each one now builds its diagram with a small shared SMIL kit and plays the mechanism in causal order: messages draw along their arrows with a packet at the tip, gates light up when a request reaches them, and outcomes grow in when the packet arrives. Each loop holds the finished diagram, fades, and restarts. Only opacity, transform, stroke-dashoffset, and motion are animated, so every animation loops and no replay button appears, and each figure sets its reduced-motion and print frame to the complete diagram. Text, facts, and layout carry over, apart from fixes where the original overflowed a box or the viewBox, put a label on top of a line, or left an empty table column. Cards now use the course lf markup, so the runtime names each SVG from the card title and caption, and the original aria-label text becomes the SVG description instead of being overridden by aria-labelledby. Below 640px wide, the diagram scrolls sideways inside its card at a readable size, where it used to shrink labels to about 5px. |
||
|
|
319c898d87 | chore(site): rebuild data.js | ||
|
|
38ed3c95b0 |
fix(site): show published exam facts on certification cards (#487)
Catalog cards show a track's first three exam facts: questions, time limit, and passing score. The official MCPA page publishes only the 90-minute duration, so the MCPA card read "Not published" twice out of three and looked broken. Cards now skip facts a track marks unpublished and fill the row with the next published ones, so the MCPA card shows its time limit, exam fee, and format. Claude cards are unchanged, and the track page still lists every fact, including the ones the provider does not publish. |
||
|
|
48c565b9f1 | chore(site): rebuild data.js | ||
|
|
352b2eac89 |
fix(readme): rebuild the banner curriculum stack so every phase appears once (#486)
The banner's curriculum stack skipped Phase 11 (LLM Engineering), grouped Generative AI and Multimodal under "LLMs · Transformers", drew translucent layers whose edges showed through each other, and ran its last label into the right edge past the 56px margin the header keeps. The stack is now eight solid slabs with thickness, foundations at the bottom and capstones on top, covering phases 00 to 19 exactly once and in order: setup and math, ML and deep learning, vision, NLP, and speech, transformers and generative AI, RL and LLMs, multimodal and tools, agents and swarms, then production, safety, and capstones. Each label sits on its slab's corner, and every label stays inside the margin. The left column gains a certification line for the Claude and MCPA preparation tracks, and the separators drop em and en dashes. The count line is unchanged and still pinned by the build test. |
||
|
|
4b134e9c52 | chore(site): rebuild data.js | ||
|
|
2341733d60 |
fix(readme): redraw the FIG_001 artifact icons so no label collides with geometry (#485)
The four "every lesson ships something" icons had collisions at README size: the agent loop's arc ran through the TOOL box and its label, OBS sat on its box edge with a detached arrowhead, the MCP labels hung outside the rack at about 4px, and the skill icon ended in a stray plug shape. The README also drew the 120-unit artwork at 96px, shrinking every label by a fifth. Each icon is redrawn on the same grid in the blueprint style: a prompt page with a >_ line sent to the model, a SKILL.md file dropping into an agent's skills tray, the agent loop as LLM to TOOL to OBS and back, and an MCP server whose tools, resources, and prompts are labeled inside its units with a two-way client link. Every label sits inside its container with clearance, checked by bounding-box tests at 96, 120, and 360px, and the README renders the icons at their native 120px. The translated READMEs are regenerated. |
||
|
|
3f628a302b | chore(site): rebuild data.js | ||
|
|
9f06b4b740 |
feat(certifications): MCPA exam prep on the MCP 2026-07-28 protocol (#484)
* feat(mcpa): add MCPA program, track blueprint, and source ledger
Start the Model Context Protocol Associate certification prep, mirroring the
Claude certification system. Adds the program manifest, the mcpa-f track with
the five published domains and weights (Fundamentals 16, Architecture 14,
Interactions 26, Security 24, Use Cases 20) and exam mechanics, and a source
verification ledger that maps every exam fact to the official certification
page, the launch announcement, and the MCP 2026-07-28 specification with the
retrieval date. Objectives are original, derived from the published
sub-competency names and the specification.
* feat(mcpa): add the MCP fundamentals lesson and figure runtime
First MCPA cert lesson, mirroring the Claude certification lesson contract: a
concept-first explainer of the integration problem MCP solves, the four roles,
and the handshake, discovery, and invocation flow, with a stdlib JSON-RPC mock
(no SDK, no network), five-plus tests, a six-question quiz, a field-brief
artifact, and a registered host-client-server figure. Aligned to MCP 2026-07-28.
* feat(certifications): make cert audit + debias provider-aware
Discover providers via certifications/*/program.json and run the same
contract per provider instead of hardcoding certifications/claude.
- audit_certifications.py: PROVIDERS registry + use_provider(slug);
check_public_pages runs once, run_provider_audit(audit, slug) per
provider; LESSON_REF_PREFIX / PROVIDER_ENTITY / PROVIDER_AI_NATIVE_SURFACE
replace the former claude-only literals; AI-native-surface check gated
per provider.
- debias_certification_questions.py: glob certifications/*/lessons and
certifications/*/assessments.
- claude verdict unchanged: 33 lessons, 8 assessments, 295 questions, 0 issues.
- balance mcpa lesson 01 quiz answer positions and option lengths.
* feat(certifications): add MCPA cert README and generalize route-count regex
- certifications/mcpa/README.md: route table (MCPA, 10 lessons), blueprint
weights, AI-tutor onboarding, GitHub lesson index, non-affiliation notice;
states the official item count and passing score are not published and the
90-vs-120 minute source discrepancy.
- audit_certifications.py: widen the README route-count regex from CC-only to
any uppercase exam code so MCPA is parsed; claude codes still match and the
claude verdict stays 0 issues.
* feat(certifications): build MCPA track lessons, figures, and wiring
- Nine new lessons (00, 02-09) covering all five MCPA domains. Each ships a
concept-first doc, a standard-library MCP mock in code/main.py, a Python
test suite (88 lesson tests total, all passing), a six-question quiz, a
reusable output artifact, and a mechanism figure.
- site/figures-mcpa-certifications.js: register all ten lesson figures
(mcpa-00 through mcpa-09), each a themed inline-SVG diagram; render-checked.
- tracks/mcpa-f.json: wire the ten lessons to their domains and roles.
- prerequisites.json: linear lesson dependency chain over all ten lessons.
- audit_certifications.py: enforce the mcpa expectedFigures map.
* feat(certifications): add MCPA assessments and AI-native tutor surface
Completes the MCPA certification track. Full audit is green: 43 lessons,
10 assessments, 380 questions, 0 issues; the Claude verdict is unchanged
(33 lessons, 8 assessments, 295 questions).
- assessments/mcpa-f/diagnostic.json (25 questions) and mock-01.json (60
questions) with the blueprint domain mix (10/8/16/14/12); every question
maps to a declared domain objective and cites an existing lesson for
remediation. Declared in tracks/mcpa-f.json.
- debias pass balances answer positions across the ten lesson quizzes and
the two assessments.
- skills/mcpa-certification (SKILL.md + agents/openai.yaml) and its
.claude/skills mirror: an AI-native tutor for the MCPA route, mirroring
the Claude certification tutor.
- certifications/mcpa/GETTING_STARTED.md: the GitHub learner guide.
- curriculum.yml: run certification lab tests and demos for every provider,
not just Claude.
- README.md: MCPA onboarding section and quick-start row.
* feat(certifications): add MCP 2026-07-28 protocol brief and wire-shape checker
The MCPA exam is aligned to the 2026-07-28 revision, which made MCP
stateless (SEP-2575, SEP-2567): no initialize handshake, no sessions,
per-request _meta version and capabilities, mandatory server/discover,
Multi Round-Trip Requests instead of server-initiated requests (SEP-2322),
a required resultType, subscriptions/listen, cacheable list results
(SEP-2549), and a new error-code allocation policy.
- certifications/mcpa/research/mcp-2026-07-28-brief.md: the protocol
source of truth for every MCPA lesson and question, read from the
primary specification pages, schema.ts, the extension pages, and the
SEPs, with a table of exam traps where legacy-era beliefs are now wrong.
- scripts/check_mcpa_wire.py: imports each MCPA lesson's transcript() and
enforces the 2026-07-28 wire invariants (required _meta fields,
resultType, cache hints on the six cacheable operations, MRTR retry
rules, subscription ids on listen streams, header and body agreement,
the error-code allocation policy) and rejects legacy methods such as
initialize unless a lesson marks them as compatibility examples.
- scripts/test_check_mcpa_wire.py: 16 tests for the checker's rules.
* refactor(certifications): restructure the MCPA track for the stateless 2026-07-28 protocol
The first MCPA build taught the legacy initialize handshake as current,
answered unknown tools with -32601 and schema-invalid arguments with
-32602, and allocated custom errors in the legacy -32000..-32019 range.
MCP 2026-07-28 removed the handshake and sessions (SEP-2575, SEP-2567),
made input validation a tool execution error (SEP-1303), and partitioned
the server-error range into a legacy block and an MCP-reserved block.
- remove the nine lessons and two assessments built on the legacy model
- rewrite the 18 domain objectives against 2026-07-28: the stateless
core, discovery and capability negotiation, Multi Round-Trip Requests,
subscriptions, caching, extensions, and the deprecation lifecycle
- lay out a 34-lesson route across the five domains with a prerequisite
chain and the expected figure for each lesson
- link all 17 phase 13 MCP lessons as deep dives
* feat(mcpa): integration-problem lesson on the stateless core
- one client discovers and calls two unrelated servers using per-request
_meta (protocolVersion, clientCapabilities, clientInfo) and
server/discover with cache hints, with no handshake or session
- N times M versus N plus M integration arithmetic
- the two error channels: an unknown tool is a -32602 protocol error, a
missing argument is an isError tool execution error the model can fix
(SEP-1303)
- UnsupportedProtocolVersion -32022 carrying data.supported and
data.requested
- transcript() passes scripts/check_mcpa_wire.py; this lesson is the
structural exemplar for the rest of the track
* feat(mcpa): discovery and capability negotiation lesson
- server/discover as the only discovery call a server must implement:
DiscoverResult with supportedVersions, capabilities, instructions,
serverInfo, and CacheableResult hints (ttlMs, cacheScope)
- server capabilities are cached per server; client capabilities are
declared fresh on every request in _meta
- MissingRequiredClientCapability (-32021) with data.requiredCapabilities
when a tool needs elicitation the request did not declare
- UnsupportedProtocolVersion (-32022) then a retry with a supported
version and a new request id; era is message shape, not version string
- unknown tool (-32602) versus unknown method (-32601)
* feat(mcpa): stateless core lesson
- statelessness as a protocol invariant (SEP-2575): every request carries
its own version and capabilities, no initialize, no session
- two replicas behind a round-robin router serve interleaved requests
from one shared store, so any replica answers any request
- cross-request state through opaque server-minted handles (SEP-2567):
authorized per call, bounded lifetime, expiry and foreign-principal
use returned as isError tool execution errors
- list results identical across connections; a request without the
required _meta fields rejected with -32602
* feat(mcpa): hosts, clients, and servers topology lesson
- one client per server inside the host, with local stdio and remote
Streamable HTTP servers and trust following process and network
boundaries
- server features (tools, resources, prompts, completion) versus client
features (elicitation; sampling and roots deprecated) and the control
model: model-controlled tools, application-driven resources,
user-controlled prompts
- host aggregation across servers: tool-name collisions resolved by
server-id prefixing, serverInfo.name never used as a routing key
because it is self-reported and not unique
- capability-gated discovery: no tools/list against a server that did
not declare the tools capability
* feat(mcpa): exam strategy lesson for the 34-lesson route
- blueprint-weighted study allocation and weighted readiness (a heavy
domain moves the estimate more than a plain average)
- exam mechanics from the source ledger, including the 90-minute page
value versus the 120-minute launch announcement
- a reading strategy for legacy-era distractors: the initialize
handshake, sessions, and -32601 for an unknown tool
- the full 34-lesson route with its domain tags, validated so every
lesson and every domain is covered; the capstone spans all five
* feat(mcpa): tool schema and structured content lesson
- tool definition fields and naming rules (1 to 128 characters of
A-Za-z0-9_-., unique per server, aggregator prefixing)
- JSON Schema 2020-12 as the default dialect (SEP-1613), any 2020-12
keyword allowed (SEP-2106), and the recommended no-parameter schema
- network $ref never auto-dereferenced; the registration path refuses it
- outputSchema with structuredContent plus a serialized text mirror
- schema-invalid arguments returned as isError tool execution errors
(SEP-1303); -32602 kept for unknown tools and missing _meta, with the
legacy behavior shown only as a labeled counterexample
* feat(mcpa): JSON-RPC envelope and _meta lesson
- the four message shapes under MCP: requests with a non-null unique id,
results that must carry resultType, errors, and id-less notifications
that are never answered; no batching
- resultType semantics: complete, input_required, unknown values
invalid, absent treated as complete for earlier servers
- _meta key grammar (optional reverse-DNS prefix plus name) and the
reserved-prefix rule on the second label: io.modelcontextprotocol and
dev.mcp reserved, com.example.mcp not
- the reserved keys, including the OpenTelemetry traceparent, tracestate,
and baggage exception (SEP-414)
- a request missing protocolVersion or clientCapabilities rejected with
-32602, shown next to three labeled malformed messages
* feat(mcpa): reading server manifests lesson
- review a server/discover result, a tools/list page, and a registry
server.json before any tool is called
- tool annotation defaults applied when omitted (readOnlyHint false,
destructiveHint true and only meaningful when not read-only,
idempotentHint false, openWorldHint true), treated as untrusted hints
- x-mcp-header rules: HTTP token syntax, case-insensitive uniqueness, no
number types, never on secrets
- cacheScope as a sharing hint, not access control; instructions as
self-reported text that can try to steer the model
- reverse-DNS registry namespaces tied to verified owners
- a manifest linter that flags six defects in a careless server and none
in a clean one
* feat(mcpa): reading the specification lesson
- how the 2026-07-28 specification is organized and the floor every
implementation must support (base protocol, versioning, message
patterns) versus optional components
- RFC 2119 and RFC 8174 keyword strength, including the capitals rule
- schema.ts as the source of truth and schema.json as generated output
- revision states (Draft, Current, Final) versus feature states (Active,
Deprecated, Removed) under the lifecycle policy (SEP-2596): a 12-month
minimum window measured from the deprecating revision's release, and
the earliest removal on or after it
- JSON-RPC batching (added 2025-03-26, removed 2025-06-18) as the
reversal the lifecycle policy was written to prevent
- tracing changelog entries back to their SEPs
* feat(mcpa): resources primitive lesson
- resources/list, resources/read, and resources/templates/list with an
RFC 6570 template, text and base64 blob contents, and a directory read
returning several contents
- resource not found as -32602 with data.uri (SEP-2164), never -32002 and
never an empty contents array
- URI scheme choice: https only when the client can fetch directly;
file, git, or a custom scheme otherwise
- traversal-safe resolution of file URIs so a path containing ".." can
never escape the served root
- cache hints on every cacheable result, with a private cacheScope for
user-specific content
* feat(mcpa): prompts and argument completion lesson
- prompts as user-controlled templates: prompts/list with pagination and
cache hints, prompts/get substituting arguments into messages that can
carry resource links
- unknown prompt, missing required argument, and invalid cursor all
reported as -32602
- completion/complete for ref/prompt and ref/resource references, with
context.arguments narrowing later suggestions
- the 100-value cap with total and hasMore, exercised against a real
catalog larger than the cap
* feat(mcpa): error handling lesson
- protocol errors versus tool execution errors (SEP-1303): an unknown
tool stays -32602, invalid or missing arguments return isError results
the model can correct
- MCP-reserved codes with their data shapes: HeaderMismatch -32020,
MissingRequiredClientCapability -32021 (data.requiredCapabilities),
UnsupportedProtocolVersion -32022 (data.supported, data.requested)
- the 2026-07-28 allocation policy enforced in code: a guard refuses the
legacy -32000..-32019 block, retired -32002 and -32042, and undefined
reserved codes before any response is built
- HTTP status mapping only where the specification states one (400,
404, 202, 401, 403, 405); statuses the spec leaves open are marked so
* feat(mcpa): client registration and identity lesson
- registration priority: pre-registered credentials, Client ID Metadata
Documents when the authorization server advertises support,
deprecated Dynamic Client Registration, then asking the user
- CIMD validation: an HTTPS client_id with a path, exact client_id match,
required fields, and redirect URI checks (SEP-991)
- application_type native versus web for DCR clients
- credentials keyed by issuer and never reused across authorization
servers; re-registration when the authorization server changes
- per-client consent at a proxy to prevent the confused-deputy attack
- the OAuth Client Credentials and Enterprise-Managed Authorization
extensions for machine-to-machine and IdP-governed access
* feat(mcpa): deprecated client features lesson
- roots, sampling, and logging deprecated by SEP-2577 but still valid in
2026-07-28: roots/list and sampling/createMessage travel as MRTR
inputRequests, gated by declared client capabilities (-32021 when
missing)
- per-request io.modelcontextprotocol/logLevel with notifications/message
only on that request's own stream, and -32602 for an unknown level
- removed versus deprecated: logging/setLevel and
notifications/roots/list_changed are gone (SEP-2575), shown only as a
labeled legacy contrast
- earliest removal computed from the 12-month window, with migration
paths for each feature, DCR, includeContext, and HTTP+SSE
* feat(mcpa): multi round-trip requests and elicitation lesson
- MRTR replacing server-initiated requests (SEP-2322): an input_required
result with inputRequests and requestState, then a retry with a new id,
inputResponses, and the state echoed exactly
- requestState protected with an HMAC that binds the principal, a short
expiry, a digest of the originating request, and a single-use nonce
- tampering, expiry, cross-principal replay, and a retargeted retry all
rejected before any side effect
- form-mode elicitation with accept, decline, and cancel; URL mode for
out-of-band sensitive steps (SEP-1036); a client without the
elicitation capability refused with -32021
* feat(mcpa): risk and safety controls lesson
- a threat model mapped to the clause that mitigates each threat: tool
poisoning, prompt injection through results, rug pulls, tool
shadowing, confused deputy, token passthrough, requestState tampering,
SSRF through CIMD fetches and network $ref, DNS rebinding, malicious
icons, and supply-chain drift
- a gateway that pins tool definitions by hash and holds a changed
definition for review, quarantines descriptions carrying injected
instructions, refuses network $ref, blocks token passthrough, and
enforces a per-tool rate limit
- policy refusals returned as isError tool results, never as invented
codes in the reserved -32000..-32099 range
* feat(mcpa): tools primitive lesson
- tools/list with opaque cursors (an empty-string cursor still means more
pages), cache hints, and results identical across connections
- CallToolResult: content, structuredContent, and isError defaulting to
false
- every content block type (text, image, audio, resource_link, embedded
resource) with audience, priority, and lastModified annotations; an
embedded resource's annotations sit beside resource, per schema.ts
- tool annotation defaults (readOnlyHint false, destructiveHint true,
idempotentHint false, openWorldHint true) treated as untrusted hints
- listChanged through subscriptions/listen, from acknowledgment to
notifications/tools/list_changed to a fresh tools/list
* feat(mcpa): notifications, subscriptions, and cancellation lesson
- subscriptions/listen replacing resources/subscribe and the HTTP GET
stream: the notification filter, the acknowledgment as the first
message, and subscriptionId equal to the listen request id so several
subscriptions can be demultiplexed
- stream notifications (list_changed, resources/updated) versus
request-scoped progress and message notifications that never travel on
a listen stream
- progress tokens with strictly increasing progress
- cancellation per transport: closing the SSE stream on HTTP,
notifications/cancelled on stdio; the server sends that notification
only to tear down a listen stream; graceful closure and late-message
races
- a fresh subscriptions/listen after a stdio reconnect
* feat(mcpa): transports and HTTP header contract lesson
- stdio: newline-delimited framing with embedded newlines rejected,
stdout reserved for MCP messages, stderr for logs, cancellation by
notification, shutdown by closing stdin
- Streamable HTTP without sessions: one POST endpoint, JSON or
per-request SSE responses, 202 for notifications, 405 for GET and
DELETE, no resumability
- Origin validation with 403 and localhost binding against DNS rebinding
- the header contract (SEP-2243): MCP-Protocol-Version, Mcp-Method, and
Mcp-Name mirrored from the body, x-mcp-header parameters as
Mcp-Param-{Name}, and base64 sentinel encoding checked against the
specification's own worked examples
- a header-body mismatch rejected with 400 and HeaderMismatch -32020
* feat(mcpa): trust zones lesson
- five trust zones in one exchange: user and host, client, server,
upstream systems, and the model
- every untrusted input that reaches model context labeled and
quarantined: tool descriptions, annotations, icons, results, and
resource contents
- clientInfo and serverInfo treated as self-reported display data
- multi-server isolation: an instruction embedded in one server's result
cannot trigger a call on another server
- annotations from an untrusted server fall back to safe defaults;
icon URIs other than https or data rejected
- local server launch commands accepted only from the host's own
configuration (SEP-1024)
* docs(mcpa): record primary-source conflicts resolved during the lesson build
Seven places where specification pages, schema.ts, and SEP texts
disagree, each with the normative resolution: removed versus deprecated
roots and logging methods (SEP-2575 over SEP-2577's feature grouping),
embedded-resource annotation placement (schema.ts over the rendered
example), notification-POST headers, requestState rejection channel,
HTTP statuses for three JSON-RPC codes, server-side subscription
teardown, and tool-name format (published page over SEP-986's draft).
Points the specification leaves open are marked so assessments never
test them as a single rule.
* feat(mcpa): model interaction flow lesson
- the host loop from user request through model selection, tools/call,
and back into model context, driven by a deterministic scripted model
- tool execution errors (isError true) fed back to the model and
corrected on a new request id; protocol errors such as an unknown tool
(-32602) never retried blindly
- an input_required result gathered from a human through
elicitation/create and retried with a new id, inputResponses, and the
requestState echoed exactly
- a tool with no annotations treated as destructive under the
specification defaults: one call confirmed and sent, one denied and
never put on the wire
- deterministic tools/list ordering across identical calls
* feat(mcpa): cache freshness and cursor pagination lesson
- ttlMs and cacheScope on every cacheable result (SEP-2549): absent or
negative ttlMs clamps to 0, cacheScope has no default
- public entries shared across tokens behind a shared cache, private
entries never crossing identities, proven by server call counts
- input_required results and MRTR completions never cached
- opaque cursors, including an empty-string cursor that still means
another page, and invalid cursors rejected with -32602
- notifications/resources/list_changed on a subscriptions/listen stream
invalidating a cached list before its ttlMs runs out
- stable list ordering so client caches and model prompt caches see a
byte-identical catalog
* feat(mcpa): audit trail and trace propagation lesson
- W3C traceparent in _meta (SEP-414): version, trace id, parent id, and
flags validated as lowercase hex of exact length, all-zero ids
rejected, tracestate and baggage passed through untouched
- a three-hop call (client, server, upstream server) where each hop
mints a child span under the same trace id
- a hash-chained audit log recording the authenticated principal (never
self-reported clientInfo), method, tool, redacted arguments, result
channel, request id, and trace id
- tamper detection for both a naive edit and an edit that rewrites its
own hash, caught one entry later through prev_hash
- denied attempts (unknown tool, bad bearer token) logged as protocol
errors, not dropped
- deprecated logging replaced by stderr and OpenTelemetry, with the
per-request logLevel still producing notifications/message
* feat(mcpa): long-running work and the tasks extension lesson
- the tasks extension (SEP-2663) negotiated per request: the client
declares io.modelcontextprotocol/tasks in clientCapabilities and the
server advertises it in server/discover
- a tool that runs synchronously without the extension and returns
resultType task (CreateTaskResult) with it
- tasks/get polling through working, input_required, and completed,
with tasks/get itself always a complete result
- mid-flight input supplied with tasks/update and inputResponses;
cooperative cancellation with tasks/cancel
- -32602 for an unknown or expired taskId and -32021 naming the
extension for a client that never declared it
- what changed from the 2025-11-25 experimental tasks: no tasks/result,
no tasks/list, and no task parameter on tools/call
- choosing among a plain call, an MRTR round trip, and a task
* feat(mcpa): OAuth authorization lesson
- the three roles: the MCP server as resource server, the MCP client,
and the authorization server
- Protected Resource Metadata discovery (RFC 9728): the
WWW-Authenticate resource_metadata URL first, then the path-specific
and root well-known fallbacks
- authorization server metadata discovery order for path and root
issuers, with the issuer required to match
- PKCE S256 generated with the standard library, and the client
refusing to proceed when code_challenge_methods_supported is absent
- the RFC 8707 resource parameter carrying the canonical server URI on
both the authorization and token requests
- the RFC 9207 iss decision table, compared without normalization and
applied to error responses too
- bearer tokens in the Authorization header on every request, audience
validation, no token passthrough, and 401 versus 403 versus 400
* feat(mcpa): operational use-case selection lesson
- the control model as the first design question: model-controlled
tools, application-driven resources, user-controlled prompts, and the
skills extension for cataloged multi-step procedures
- local systems on stdio with credentials from the environment; remote
systems on Streamable HTTP with the interactive OAuth flow, the client
credentials extension for unattended callers, or enterprise-managed
authorization for a central identity provider
- the tasks extension for long jobs and MCP Apps for interactive views,
with a text fallback for hosts that do not declare the ui extension
- cacheScope chosen from data sensitivity, and a no-external-system case
where MCP is not the right fit
- a decision engine that turns a use-case profile into a primitive,
transport, auth path, and extension recommendation
* feat(mcpa): deployment roles and adoption lesson
- six deployment roles on top of the host, client, and server
architecture: server author, host and client developer, platform or
gateway operator, security and governance owner, registry publisher,
and end user
- a responsibility matrix built from a catalog of exact MUST and SHOULD
requirements with page citations, where the same requirement (Origin
validation) moves from the server author to the gateway operator as
the deployment shape changes
- gap reporting for a requirement no role owns
- stdio credentials supplied by the operator's environment and never
carried on the wire
- governance: Agentic AI Foundation stewardship, the maintainer
hierarchy and contributor ladder minimums, Working Groups versus
Interest Groups, the SEP status workflow, and the feature lifecycle
- SDK tiers (SEP-1730) as an adoption risk decision: conformance,
feature, triage, and critical-bug commitments per tier
* feat(mcpa): tool invocation lifecycle lesson
- eight checkpoints from discover and list through select, confirm,
call, validate, execute, and result, with select and confirm as
host-only steps that never touch the wire
- validate ends only in a protocol error (-32602 for an unknown tool);
schema violations, upstream failures, and business-rule failures in
execute become tool execution errors with isError true
- a genuine internal fault during execute surfaced as -32603
- notifications/progress bound to the request's progressToken with a
strictly increasing progress value, and a hard timeout that progress
does not extend
- transport-specific cancellation: closing the stream on Streamable HTTP,
notifications/cancelled on stdio, and no cancelled or timed-out
resultType
- reissuing after a broken stream with a new request id, guided by
idempotentHint as an untrusted hint and a server-minted handle
* feat(mcpa): protocol eras and compatibility lesson
- the revision timeline from 2024-11-05 to 2026-07-28 and the legacy,
modern, and dual-era terms
- the stdio probe for a dual-era client: server/discover first, a
DiscoverResult or a recognized modern error such as -32022 means
modern (retry with a supported version, never fall back), any other
error or a timeout means legacy
- the fallback never keyed on one specific error code
- the Streamable HTTP probe: reading a 400 body for a recognized modern
JSON-RPC error before assuming legacy
- era cached per server process on stdio or per origin on HTTP, with a
re-probe when the cached assumption fails
- a modern-only server naming its supported versions when it rejects a
legacy initialize, and the legacy opening sequence shown only inside
wrapped legacy examples
* feat(mcpa): consent and least privilege lesson
- consent gathered through a multi round-trip request: a tools/call
returns input_required with an elicitation/create request, and the
retry carries a new id, inputResponses, and the signed requestState
- decline, cancel, and a rejected retry returned as tool execution
errors with isError true, never an invented consent error code
- consent scoped per tool and bound to the approved arguments, with
requestState signed, single-use, and rejected when a retry changes
the arguments
- annotation defaults (destructiveHint and openWorldHint true,
readOnlyHint false) deciding when to prompt, treated as untrusted
hints rather than enforcement
- step-up authorization: a 403 insufficient_scope challenge, a new token
requested for the union of granted and challenged scopes, and a retry
cap when the scope is never granted
- tools/list filtered by granted scopes and cached with cacheScope
private because it varies by authorization
* feat(mcpa): extensions framework lesson
- extension identifiers as {vendor-prefix}/{name}: the
io.modelcontextprotocol/ prefix for official extensions and a reversed
owned domain for third parties
- negotiation on every request: the client declares extensions in
clientCapabilities and the server advertises its own in server/discover,
with the settings object as the value
- an optional extension falling back to core behavior and a mandatory
one rejected with -32021 naming it in data.requiredCapabilities
- extensions disabled by default and opt-in; a breaking change needs a
new identifier
- the SEP-2133 process, experimental-ext- repositories owned by a
Working Group or Interest Group, and the official roster: tasks, MCP
Apps, skills, and the two authorization extensions
- what changed from initialize-time declaration to per-request
negotiation
* feat(mcpa): registry, gateways, and SDK tiers lesson
- the MCP Registry as a preview metadata index for public servers:
server.json with packages (npm, pypi, nuget, cargo, oci, mcpb),
remotes (streamable-http, deprecated sse), or both
- reverse-DNS namespaces admitted only after GitHub, DNS, or HTTP
verification; spoofed namespaces and private servers rejected
- exact version strings (ranges prohibited), and the server.json schema
version kept separate from the protocol version
- aggregators and subregistries built on top of the registry
- gateways that route on the Mcp-Method and Mcp-Name headers, reject a
header and body mismatch with -32020 before any backend is touched,
partition private cache entries by caller, and never pass the
caller's token through to a backend
- SDK tiers (SEP-1730): conformance, feature, triage, and critical-bug
commitments, and relegation after four weeks of continuous failures
* docs(mcpa): pin MCP Apps metadata shapes in the protocol brief
The extension summary named _meta.ui.csp and _meta.ui.permissions
without their shape or location. Checked against the MCP Apps
specification (ext-apps 2026-01-26, draft agrees):
- a tool's _meta.ui carries resourceUri and visibility, which defaults
to model and app and gates what the agent lists and what an app may
call
- the UI resource's _meta.ui carries csp, an object of connectDomains,
resourceDomains, frameDomains, and baseUriDomains origin lists, and
permissions, an object of camera, microphone, geolocation, and
clipboardWrite flags
- the host builds CSP from declared domains only, may restrict further,
applies a restrictive default when csp is omitted, and may honor
permissions that apps must not assume
- the app's ui/initialize is unrelated to the removed core initialize
* feat(mcpa): capstone that reads one exchange end to end
One incident-console server driven through a single transcript that
exercises every domain:
- server/discover with cache hints after an unsupported version is
corrected from the -32022 data.supported list
- protocol version and client capabilities in _meta on every request
- a schema-invalid call returned with isError true and corrected
- a consent round trip with an HMAC-signed, principal-bound
requestState, a retry on a new id, and a tampered state rejected
- -32021 when the client never declared elicitation or the tasks
extension
- a long diagnostic run as a task, polled to completion, and a second
task cancelled
- request-scoped notifications/progress and the Mcp-Method and Mcp-Name
header contract
- a wrong-audience token rejected with 401 before any JSON-RPC body is
read
- one traceparent trace id carried through every hop, and a hash-chained
audit log that pinpoints a tampered entry
- a readiness checklist mapping every objective to what a candidate must
be able to do
* docs(mcpa): teach the tutor and guides the stateless 34-lesson route
- tutor skill (and its Claude Code mirror): the 2026-07-28 revision is
taught as current, with no initialize handshake, no sessions, and
per-request _meta plus server/discover; older revisions appear only as
what changed, and deprecated features as still working until removal
- the protocol brief and the wire-shape checker join the tutor's source
list, and the tutor runs the checker when a learner edits a transcript
- three full mocks with distinct emphasis, and a fresh mock for every
retake so a second score measures readiness rather than recall
- the capstone description now matches lesson 33's integrated exchange
- learner guide: 34-lesson route, 30-question diagnostic, three mocks,
the multi round-trip lesson as the worked example, and the wire
checker in the local verification suite
- root README: the MCPA summary describes the new route, and the skills
table and install list include mcpa-certification
* ci(certifications): gate MCPA transcripts on the 2026-07-28 wire shape
Runs the wire-shape checker's own tests and then the checker over every
MCPA lesson transcript, so a lesson that reintroduces the legacy
handshake, drops the per-request _meta fields, omits resultType, or
emits an undefined error code fails CI. Both scripts are added to the
workflow's path filters.
* feat(mcpa): MCP Apps interactive interfaces lesson
- the io.modelcontextprotocol/ui extension negotiated per request, with
a text-only fallback for hosts that never declare it
- a tool's _meta.ui carrying resourceUri and visibility: the agent's tool
list excludes tools without "model", and an app's tools/call is refused
for tools without "app", before any consent prompt
- the ui:// resource fetched with an ordinary resources/read and
recognized only by the text/html;profile=mcp-app mime type
- the resource's _meta.ui csp object (connectDomains, resourceDomains,
frameDomains, baseUriDomains) turned into a Content Security Policy
that starts from default-src 'none' and admits only declared domains,
with the specification's restrictive default when csp is omitted
- permissions flags (camera, microphone, geolocation, clipboardWrite)
honored at the host's discretion and never assumed by the app
- the sandboxed iframe, a sandbox proxy on its own origin for web hosts,
and the app's ui/ JSON-RPC dialect over postMessage, including a
ui/initialize unrelated to the removed core initialize
* feat(mcpa): register all 34 lesson figures and index the route
- site/figures-mcpa-certifications.js rebuilt from each lesson's figure
snippet: 34 mechanism figures (blueprint weights through the capstone
flow), each with its own CSS prefix and marker ids, all rendering an
SVG under a stub DOM; the ten figures of the legacy route are gone
- certifications/mcpa/README.md: the lesson index now lists all 34
lessons by title, the route diagram follows the new domains, and the
overview names the protocol brief, the wire-shape checker, the
30-question diagnostic, and the three mocks
* docs(mcpa): give each deprecated feature its own source and removal floor
The brief grouped all six deprecated features under SEP-2577 with one
2027-07-28 removal floor. The specification's deprecated registry
records them separately:
- Roots, Sampling, and Logging: SEP-2577, deprecated in 2026-07-28,
earliest removal on or after 2027-07-28
- Dynamic Client Registration: PR #2858, deprecated in 2026-07-28,
earliest removal on or after 2027-07-28
- includeContext "thisServer" and "allServers": SEP-2596, deprecated in
2025-11-25, removal follows Sampling
- HTTP+SSE: SEP-2596, deprecated in 2025-03-26, earliest removal three
months after SEP-2596 reaches Final
Lesson 15 already teaches the registry's values; this brings the brief
in line so later questions do not inherit the grouped floor.
* docs(mcpa): state the x-mcp-header limits with their RFC 2119 strength
The brief said x-mcp-header was "never for secrets" and that clients
drop invalid tools. The tools page is more precise:
- only integer, string, and boolean parameters that are statically
reachable from the schema root can be mirrored; number is excluded
- server developers SHOULD NOT mark sensitive parameters (passwords,
API keys, tokens, PII), because header values are visible to
intermediaries; this is a SHOULD NOT, not a prohibition
- HTTP clients MUST reject a tool whose x-mcp-header values violate the
constraints by excluding it from their tools/list result
* docs(mcpa): teach what a revision date means and how hosts load skills
Two facts the assessments test were not taught by any lesson:
- lesson 01 and the brief: a revision identifier is a YYYY-MM-DD date
marking the last backwards incompatible change, so a Current revision
can take compatible fixes without being renamed (Draft, Current,
Final)
- lesson 30 and the brief: reading a skill's SKILL.md through
resources/read only retrieves text; skill content is untrusted input
tagged with its origin, explicit user policy decides whether a skill
is loaded, hosts should let users inspect a skill first, may not let
it trigger host-side execution without per-skill approval, and ignore
permission-widening frontmatter such as allowed-tools unless the user
approved it (SEP-2640)
* docs(mcpa): teach the x-mcp-header mirroring limits in the transports lesson
The transports lesson introduced x-mcp-header mirroring and the base64
sentinel but not its limits, which the assessments test:
- only integer, string, and boolean parameters statically reachable
from the schema root can be mirrored, never number
- a Streamable HTTP client must exclude a tool whose x-mcp-header
values break those constraints from its tools/list result
- server developers should not mark sensitive parameters such as API
keys or tokens, because header values are visible to intermediaries
* fix(site): refresh the figure manifest cache key for the rebuilt MCPA figures
lesson.html pins figure-manifest.js with a content hash. Rebuilding
site/figures-mcpa-certifications.js with the 34 new figures changed the
generated manifest, so the pinned key moves from 0980822a99ac to
339ef88a8714; test_build_artifacts.js checks that the committed page
matches the manifest the build produces.
* docs(i18n): regenerate the translated READMEs with the MCPA sections
The English README gained the MCPA route row, the MCPA section, the
mcpa-certification skill row, and the updated install list, but the
twelve i18n/<lang>/README.md files were never regenerated, so
build_readme_i18n.py --check failed. Regenerated with the script; the
new blocks have no hand-authored translations yet and fall back to
English as the generator documents, and no existing translation was
dropped.
* feat(mcpa): diagnostic and three full-length mocks on the 2026-07-28 protocol
Four original assessments declared on the mcpa-f track, 210 questions,
every key checked against the protocol brief and the specification
pages it cites:
- diagnostic (30 questions, 30 minutes): one concept per item across
all 18 objectives, for placing a learner by domain
- mock 1 (60 questions, 90 minutes): operational scenarios, such as
scaling a stateful tool behind a load balancer, one-click local
server consent, registry preview planning, and DCR application_type
- mock 2 (60 questions, 90 minutes): wire-level messages, such as
missing _meta fields, -32021 and -32022 error data, progress
monotonicity, HTTP cancellation by closing the stream, all-zero
traceparent ids, and MCP Apps mime types
- mock 3 (60 questions, 90 minutes): design and security trade-offs,
such as where state lives, requestState protection, token audience,
skill loading consent, and SDK tier risk
Each mock follows the blueprint split (10/8/16/14/12 across the five
domains), uses every objective, and references all 34 lessons. Items
that repeated another file's scenario were rewritten to test a
different fact, and answer positions are balanced by
debias_certification_questions.py.
* chore(mcpa): balance answer positions across the 34 lesson quizzes
Applies debias_certification_questions.py to the MCPA lesson quizzes so
correct answers follow the per-file balanced position cycle the CI
check enforces. Only option order changes; every question, option, and
explanation is unchanged, and no Claude certification file is touched.
* fix(mcpa): answer unknown methods with -32601 in four lesson servers
The dispatchers in lessons 04, 11, 12, and 14 fell through to -32602
(invalid params) for a JSON-RPC method they do not implement. An
unknown method is -32601 (method not found); -32602 is reserved for an
unknown tool, missing _meta, and other invalid parameters. Every other
lesson dispatcher already used -32601, so a learner reading these four
servers top to bottom would have learned the wrong code.
Also adds a test for lesson 03's validate_error_shape, the one shape
validator its tests never exercised.
* fix(mcpa): align three labs with what their lessons say they show
- lesson 30: the text pointed learners at the sixth exchange for the
malformed no-slash-here identifier, but it is the seventh
- lesson 32: the demo printed "prefer remote" while calling
resolve_install_target with prefer="package"
- lesson 33: acknowledge_incident is listed by tools/list but a plain
tools/call answered "Unknown tool" (-32602); it now returns a tool
execution error saying the tool is served only over the authorized
HTTP endpoint, with a test
* fix(scripts): flag legacy and violation wrappers that do not wrap a message
check_mcpa_wire.py skipped every entry marked legacy or violation, and
when the wrapped message was not an object (a typo such as None or a
bare method string) the entry produced no finding at all, hiding an
authoring mistake. Such wrappers are now reported, matching how every
other malformed entry is handled, with a regression test for both
wrapper kinds.
* docs(mcpa): give every lab the source header the lesson contract requires
AGENTS.md asks each lesson's code/main.py to open with a 4 to 6 line
header citing its docs/en.md path and its spec or RFC sources. The MCPA
labs had a one-line docstring. Each header now names the lesson's full
docs/en.md path, what the lab does, and the specification pages, SEPs,
RFCs, or W3C documents it implements.
* fix(mcpa): require an explicit approval before consent-gated work runs
Two labs treated any elicitation answer with action "accept" as consent,
even when the form content said no:
- lesson 25: accepting the delete_file confirmation with approved false
still deleted the file; consent now needs action accept and a content
object whose approved field is exactly true, and a non-object answer
is handled safely
- lesson 20: accepting the vault read with proceed false still returned
the private note; the vault now needs proceed exactly true
Each fix has a test showing the refused path leaves no side effect.
* fix(mcpa): turn malformed wire input into refusals instead of exceptions
Six labs raised on attacker-controlled or malformed input instead of
answering it:
- lesson 05: a -32022 error with an empty supported list raised
IndexError; the probe now records a modern era with no confirmed
version and only caches a retry version after a successful discover
- lesson 09: an x-mcp-header value that is not a string crashed the
linter, and object or array properties were not flagged; only string,
integer, and boolean parameters may be mirrored
- lesson 14: a non-ASCII requestState raised during signature checks; it
is now rejected as malformed
- lesson 19: a base64 sentinel header with invalid base64 or UTF-8
raised; it now decodes to a mismatch and gets the HeaderMismatch reply
- lessons 27 and 33: a malformed traceparent raised ValueError; lesson
27 restarts the trace with a new root and lesson 33 continues without
a trace id
Each case has a regression test.
* fix(mcpa): close gaps in the labs' own security controls
- lesson 22: only the first CALL server.tool instruction in server
content was checked, so a harmless first instruction could hide a
cross-server one; every match is now quarantined and checked on relay
- lesson 26: the pinned tool descriptor left out annotations, so a
silent destructiveHint flip was not caught as a rug pull, although the
lesson's own threat matrix pins annotations; they are now part of the
hashed descriptor
- lesson 32: the gateway log stored the caller's raw bearer token as the
principal and printed it in the transcript; it now records a
non-reversible principal reference, and the backend answers an
unimplemented method with -32601 instead of -32602
Each fix has a test.
* fix(mcpa): expire idle baskets from their last activity
Lesson 04 tells learners a basket expires after five idle ticks, but
is_expired measured from creation, so an actively used basket still
expired. Baskets now record their last activity, adding an item
refreshes it, and a test shows an item added inside the window keeps
the basket alive.
* docs(mcpa): match run locations and rejection rules to the labs and spec
- lessons 09, 10, and 12 told learners to run python3 code/main.py from
the repository root, where that relative path does not exist; they
now say to run it from the lesson directory like the other labs
- the lesson 14 checklist said a protocol error is the wrong channel for
a tampered or expired requestState; the spec requires rejection but
does not prescribe the channel, so the checklist now names the lab's
choice, a tool execution error, and a fresh input_required result as
valid options
- the lesson 21 patterns sheet said a tools/call without the tasks
extension always gets -32021; per SEP-2663 the server returns an
ordinary result when it can finish within the request and -32021 only
when it cannot, and tasks/get, tasks/update, and tasks/cancel without
the declaration get -32021
* docs(mcpa): say the MCPA track is on GitHub only for now
The website build renders only the first certification program, so no
MCPA page is published there yet. GETTING_STARTED, both copies of the
tutor skill, and the README certification row pointed learners at
website routes that do not exist. They now send learners to the GitHub
lesson and assessment paths and say the website does not carry MCPA
yet. The translated READMEs are regenerated from the English one.
* fix(audit): stop treating the MCPA practice mock size as an exam fact
The official MCPA page does not publish an item count, and the fact
ledger records it as not published. The audit's MCPA-F exam facts listed
items 60, so check_track verified the practice mock size as if it were
official. The item count is removed from the verified facts; the track
keeps 60 as the curriculum's chosen mock length.
* fix(mcpa): refuse differently cased cross-server instructions in lesson 22
The embedded-instruction pattern matches CALL server.tool without regard
to letter case, but the relay check compared the captured names exactly,
so CALL Tickets.delete_all_tickets in untrusted server content still let
a relay to tickets.delete_all_tickets through as allowed. Quarantine and
relay checks now compare server and tool names case-insensitively, so a
case variant is refused, and a test covers the case-shifted instruction
and the same-server case.
* feat(site): render every certification program, not only the first
parseCertifications read only the first folder under certifications/, so
the website showed Claude alone: the catalog, track, assessment, and
lesson pages, the homepage spotlight, search, sitemap, and llms.txt never
saw MCPA. The build now loads every program, tags each track, lesson, and
assessment with its programId, derives each program's learner guide and
tutor skill paths, and fails on duplicate program, track, or assessment
ids across programs.
- the catalog renders one section per program with its own access
notice, GitHub tutor links, and track grid, and the no-JavaScript
discovery block follows the same structure
- track, assessment, and lesson pages show the disclaimer and scoring
notice of the track's own program instead of a hardcoded Anthropic one
- lesson data loading, the language picker, and the api lesson and
certification routes accept any certifications/<program>/lessons path
- the homepage spotlight shows both badges and names both providers, and
build.js keeps its track, lesson, and question counts in sync
- program.json gains shortName and accessNoticeTitle, and the MCPA badge
is marked square so it is not clipped to the Claude badge outline
- exam facts read "Not published" when a track marks its item count or
passing score unpublished, so the MCPA card no longer presents the
practice mock size as the official question count
- a track with four assessments lays them out two by two
- tests derive track counts from the data and cover a second program in
the lesson and certification routes
* docs(mcpa): send learners to the MCPA track now that the site renders it
The GitHub-only wording existed because the website could not show a
second program. With the multi-program build, GETTING_STARTED, both
copies of the tutor skill, and the README point to the MCPA track page
again, the certifications index lists both programs with their onboarding
guides, and the translated READMEs are regenerated.
* docs(i18n): translate the MCPA goal row in the Portuguese and Russian READMEs
The Portuguese and Russian READMEs translate every row of the "choose
what you want to build" table, but the new MCPA row had no entry in
readme_translations.py, so it rendered in English between translated
rows. Both languages now carry the row, with the onboarding guide and
the MCPA track links unchanged, and the translated READMEs are
regenerated.
* feat(site): add a Sponsor us page rendered from SPONSORS.md
SPONSORS.md promises a sponsor page on the curriculum site, but the site
had no sponsor page and no route to one. build.js now renders SPONSORS.md
into site/sponsors.html on every build, so the page cannot drift from the
file the maintainer edits.
- the renderer covers what SPONSORS.md uses: headings with GitHub-style
anchor ids, paragraphs, lists with wrapped items, aligned tables, bold,
inline code, and links, where repository files resolve to GitHub and
javascript: or parent-directory links render as plain labels
- raw HTML is limited to a, picture, source, and img with https-only
URLs; any other tag is escaped, and a closing tag is kept only when its
opening tag was kept
- the SerpApi logo follows the site's theme toggle instead of the
operating system color scheme
- "Sponsor us" joins the hamburger menu through header.js and the footer
of every page, and is added to the interface strings for translation
- /sponsors is served from sponsors.html with the same markdown
negotiation as the other public pages, and the page is listed in the
sitemap and llms.txt
- tests pin the rendered page to SPONSORS.md, the HTML allowlist, and
the menu and footer links
* fix(site): give the privacy, contact, and developer pages the shared header
The three trust pages shipped a stripped header with no header.js, no
search or theme controls, no site fonts, and an old stylesheet key. At
1400px and below the stylesheet renders the header nav as a dropdown
panel that header.js normally hides behind the menu button, so on these
pages the panel stayed open over the content.
They now use the same header markup as the other pages, load the site
fonts, the current stylesheet, and the shared theme, progress, header,
and search scripts, gain a skip link to main, and style their eyebrow
line. The shared asset test now covers all three pages so their cache
keys cannot fall behind again.
|
||
|
|
f6721a0167 | chore(site): rebuild data.js | ||
|
|
8050434e5c |
fix: repair lessons, harden security, and gate lesson tests and quiz bias (#480)
* fix(phase-13/06): bootstrap sys.path so the documented test command runs The lesson doc says to run unittest discovery over code/tests, but the test imported main with no path setup, so discovery from the lesson root failed with ImportError. Insert the code directory on sys.path the way the later protocol lessons already do. Passes from the lesson root and from code/. * fix(phase-13/07): bootstrap sys.path so the documented test command runs Discovery over code/tests failed with ImportError because the test imported main with no path setup. Insert the code directory on sys.path to match the later protocol lessons. * fix(phase-13/08): bootstrap sys.path so the documented test command runs The documented discovery command raised ImportError on import main. Add the sys.path bootstrap used by the sibling lessons so the test runs from the lesson root and from code/. * fix(phase-13/09): bootstrap sys.path so the documented test command runs The documented discovery command raised ImportError on import main. Add the sys.path bootstrap used by the sibling lessons so the test runs from the lesson root and from code/. * fix(phase-19/29): make the fixture tests directory an importable package The demo fixture ships a namespace-style tests directory with no __init__.py, so a regular tests package elsewhere on sys.path could shadow it and the in-repo test runner failed to import the fixture module. Add an empty __init__.py so the fixture tests resolve regardless of what else is installed. * fix(phase-19/39): bound generation by chars and gate the demo on loss decode_response replaces invalid UTF-8 bytes, so a 4-byte generation can re-encode to more than 4 bytes; the test now bounds the decoded character count, which is the real invariant that generate enforces. The demo's success gate keyed on exact-match improvement, which the documented padding mask makes unreachable, so it now checks that training loss decreased. * fix(phase-19/59): test patch-projection grad via a nonzero reduction Summing the CLS token straight out of the final LayerNorm is an algebraic zero for any input, so the gradient reaching the patch projection was always zero and the test failed for a reason unrelated to the encoder. Reduce with a sum of squares, which the demo already uses, so the test exercises a real gradient path. * fix(phase-19/47): load checkpoints weights-only, jail shard paths torch.load ran with weights_only=False, so restoring an untrusted checkpoint executes whatever the pickle names. Store RNG state as primitives so it survives the weights-only loader, load every payload with weights_only=True, reject shard paths that resolve outside the checkpoint directory, and raise ValueError on integrity failures instead of assert. Add regression tests for a pickled object and a shard path escape, and update the docs, skill, and quiz. * fix(phase-11/09): guard run_code with an AST check, not a blocklist The substring blocklist read the snippet as text, so an attribute chain such as the object-subclass walk reached the real interpreter through __globals__. Parse the snippet and walk the tree, rejecting import statements, dunder attribute access, and unsafe builtin names by structure. Add regression tests for the escape, and make the tool description and docs state plainly that this is a teaching filter, not real isolation. Align the TypeScript port's wording and block the constructor and globalThis gadgets in its blocklist. * fix(site): serve markdown at / before the static file wins The homepage markdown-negotiation rewrite lost to the static index.html, so requesting text/markdown at / returned cached HTML. Move it into a legacy route, which is evaluated before the filesystem. Add a readiness-test guard that fails if any negotiation rewrite is shadowed by a static file, and assert the root route. * fix(phase-19/86): parse nested sequences in the stdlib YAML fallback The PyYAML-free fallback stopped gathering a rule's lines at the first nested sequence item, so a rule with an any_of block lost that block and every field after it, and the engine rejected the rule for a missing explanation. The gather now ends on indentation alone, so a deeper sequence stays part of its rule. Verified the fallback matches PyYAML byte for byte on rules.yml, which restores this lesson and the end-to-end safety gate that composes it on a standard Python install. * test(ci): run each lesson's own tests on push and pull request CI executed the certification labs but never the 523 lessons, so a lesson that could not import its own module reached main unnoticed. Adds scripts/run_lesson_tests.py, which discovers each lesson's tests across the four layouts in the repo and runs them the way the lesson docs say, and a lesson-tests job that runs it. The runner installs no scientific dependencies: a lesson that needs one is skipped by scanning its source, so the job stays green while the stdlib lessons run for real. Documents the command in CONTRIBUTING. * test(ci): gate quiz answer-length bias and ratchet it down The quizzes de-bias answer position but not answer length: on 84% of questions the correct option is the longest by a wide margin, so a reader can guess it without knowing the material. Adds scripts/check_quiz_bias.py, which measures the share of length-biased questions and, in --check, fails when a change pushes the rate above a baseline set to today's level. New or edited quizzes cannot add bias, and the baseline ratchets down as quizzes are rebalanced. Wires it into the curriculum workflow next to the position gate and documents the rule in AGENTS.md. Reducing the existing rate is a separate content pass; this stops it getting worse and makes it measurable. * fix(phase-11/09): add module docstring, honest TS wording, ast import Address review: give function_calling.py a module docstring like the sibling lessons, align the TypeScript run_code tool description with the honest wording already used on the Python side (a teaching filter, not real isolation), and add the missing ast import to the docs example's import block so the example runs. * fix(phase-19/39): fail the demo on a non-finite final loss The success gate compared the final loss to the first with >=, so a NaN final loss made the comparison false and the demo exited 0. Require the final loss to be finite before comparing, so a diverged run is reported as a failure. * docs(phase-19/47): note the torch 2.6 requirement for weights_only Before torch 2.6 the weights_only loader had a known bypass (CVE-2025-32434), so the security guarantee this lesson relies on holds only from 2.6 on. Say so next to the weights_only explanation. * fix(site): match the root markdown Accept header case-insensitively Media types are case-insensitive, so a client sending Text/Markdown should still reach the markdown route. Add the (?i) flag to the root route's Accept matcher. * fix(ci): fail a lesson when a real test fails next to a missing dep The lesson-test runner skipped a suite whenever its output mentioned a missing optional module, which could hide a genuine assertion failure in the same run. Only skip when the output shows no test failure alongside the missing module. * fix(ci): fail the quiz-bias check on an unreadable quiz file The scan skipped a quiz.json it could not parse and carried on, so a malformed file would leave the gate green on an incomplete scan. Collect read and parse errors, report each file, and exit non-zero when any are found. |
||
|
|
0285d9bd92 | chore(site): rebuild data.js | ||
|
|
6e2a5868a8 |
fix(scripts): seed quiz de-bias from a POSIX-normalized path (#466)
* fix(scripts): seed quiz de-bias from a POSIX-normalized path
`seed_for` seeded the option shuffle with the raw path from
`glob.glob("phases/*/*/quiz.json")`, which carries OS-native separators.
Windows therefore computed a different shuffle than Linux, so `--check`
reported 2098 questions as not de-biased and a bare `--fix` run rewrote
373 files into an arrangement the Linux CI gate then rejects.
Normalize the separator before hashing, mirroring
`debias_certification_questions.py`, which already seeds off
`path.relative_to(ROOT).as_posix()`. POSIX paths are unchanged, so the
committed arrangement and the CI gate are unaffected.
Verified on Windows: `--check` now reports 0 rewrites and the same
positional distribution as Linux (A 539, B 582, C 528, D 588), and
`seed_for` maps "a/b" and "a\b" to one seed. An assertion in `main()`
pins both forms so the gate fails if the normalization is ever lost.
* fix(scripts): keep quiz seeding change comment-free
|
||
|
|
37ae4f22c2 | chore(site): rebuild data.js | ||
|
|
ddfbb4cd91 |
fix(phase-10/10): perplexity was different on every run (#358)
The simulated log-probs are seeded from the text so the same text scores the same way, and STEP 3 compares Strong/Medium/Weak on that basis. hash() of a str is salted per interpreter process, so the seed changed every run - three consecutive runs gave Strong-model perplexity 1.20, 1.16, 1.13. Seed from hashlib.sha256 instead. Three runs now produce byte-identical output, and the Strong < Medium < Weak ordering the lesson relies on is preserved (1.13 / 1.44 / 2.52). Co-authored-by: thejesh23 <thejesh23@users.noreply.github.com> |
||
|
|
805bdef902 | chore(site): rebuild data.js | ||
|
|
3851215fa5 |
fix(phase-16/03): ACP audit trail was always empty (#357)
docs/en.md says every agent execution produces an audit entry with the complete
trajectory of tool calls in between, and the gateway wraps the execution in that
trail. The demo printed 'Trajectory steps: 0' and 'Artifacts: 0'.
Two causes, both needed: sendMessage fired processTask without awaiting, so it
returned before the handler ran; and delegateTask ran auditRunner.run before
dispatching, so structuredClone(result.trajectory) snapshotted an array the
handler had not filled yet.
Await the handler, and audit after dispatch.
Task state: completed
Artifacts: 1
Trajectory steps: 2
- Searching for React 19 documentation Tool: web_search
- Extracting key findings from search results Tool: doc_analysis
Co-authored-by: thejesh23 <thejesh23@users.noreply.github.com>
|
||
|
|
1a17799065 | chore(site): rebuild data.js | ||
|
|
c3dfda93ef | chore(phase-00/03): remove unused sys import in gpu_check (#286) | ||
|
|
56d3877d98 |
fix: sandbox path jail bypassed by a slashless symlink (escape) (#249)
* fix: sandbox path jail bypassed by a slashless symlink (escape) 26-sandbox-runner-denylist: the path jail (_check_path_jail) only resolves and prefix-checks arguments for which _looks_like_path() is true, which requires a path separator (or exactly ./..). A symlink whose name has no slash — e.g. `link.txt` in the project root pointing at /etc/passwd — is therefore never jail-checked, so `cat link.txt` escapes the jail and reads files outside the project root, exactly what the module's advertised "symlink-safe path jail" is meant to prevent. (`sub/link.txt`, having a slash, is correctly denied; the asymmetry is the tell.) Also jail-check arguments that exist on disk under the root (os.path.lexists), so a slashless symlink is resolved and refused. lexists catches the symlink even when its target is missing. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test: skip slashless-symlink jail test where symlinks are unsupported os.symlink raises OSError on Windows without Developer Mode/admin, which would fail this test spuriously in such CI. Skip instead of failing. Addresses CodeRabbit review feedback on the PR. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c2e9899b20 | chore(site): rebuild data.js | ||
|
|
f6472e3cd9 |
imp(phase-00/01): centralize Check defaults in const fn constructor (#361)
Replace repeated struct literals with Check::new(name, program, optional). `args: &["--version"]` now lives in one place instead of eight. |
||
|
|
a371adb375 | chore(site): rebuild data.js | ||
|
|
ff34db8a6b | fix: align code with course statement (#436) | ||
|
|
6331da9d91 | chore(site): rebuild data.js | ||
|
|
d86fedaa51 |
fix: dark-mode mermaid fills, Apple Silicon Docker build, fnm unzip, ROADMAP rows, site chrome translations (#477)
* fix: dark-mode mermaid fills, Apple Silicon Docker build, fnm unzip, ROADMAP rows, site chrome translations - site/lesson.html: mermaidPreprocess handles 3-digit hex fills so fill:#dfd stays legible in dark mode (#433) - 07-docker-for-ai: pin FROM to linux/amd64 and document why the cu124 layer fails on Apple Silicon (#476) - 01-dev-environment: add unzip to the apt line, fnm's installer needs it (#473) - ROADMAP.md: add Phase 11 lessons 16 and 17, 521 rows to 523 - site: translate interface strings for the eight CI languages via ui-i18n.js + ui-strings.js, picker dispatches aifs:lang, drift guarded by test_ui_i18n.js in CI (#469, #404) * docs: describe fnm's unzip check and scope the CUDA image to x86_64 hosts - 01-dev-environment: the installer checks for unzip up front and exits with its own message; macOS goes through Homebrew - 07-docker-for-ai: the linux/amd64 pin means the image is for x86_64 Linux hosts with an NVIDIA GPU * feat(i18n): generate site chrome translations in the translate pipeline The hand dictionary becomes site/ui-strings.json: an English key list plus per-language overrides. scripts/translate_ui_strings.py fills every other key through the lesson translator's provider layer, reuses what is already published, protects file names, paths, and commands behind the same placeholders, and writes i18n/<lang>/ui.json. A new ui-strings job in translate.yml runs it per CI language and publishes to the translations branch with the same race-safe worktree loop, so a new label or a new language needs no translation written by hand. site/ui-i18n.js now fetches i18n/<lang>/ui.json from the translations branch at runtime, caches it per language, and falls back to English until a language is published. Tests cover source consistency, precedence, placeholder protection, the page drift guard, and the fetch path. * fix(i18n): retranslate keys whose override was removed and keep failed dictionary loads retryable - ui.json now carries {strings, pinned}; a key that was pinned in the last publication but has no override now is translated again instead of reusing the old pin - the runtime no longer caches a failed fetch, so a language published later is picked up on the next switch without a reload * fix(i18n): refresh flat ui.json values once when pin provenance is missing A flat published file predates the {strings, pinned} shape and cannot say which values came from overrides, so every value in it counts as pinned and is retranslated on the first run instead of being reused blindly. |
||
|
|
d18b8fe5a9 |
fix(book): wrap inline code and fail incomplete PDF builds (#460)
* fix(book): keep inline table code inside PDF margins * fix(book): preserve Unicode and fail incomplete PDF builds * fix(book): wrap inline code in PDF prose without extra symbols * fix(book): wrap long plain-text identifiers in PDF tables * fix(book): preserve Unicode sequences in table wrappingv2026.09 |
||
|
|
f068c1e63f | fix(book): wrap long code lines and reference URLs in PDFs (#459) | ||
|
|
bcd09d9102 | fix(book): preserve literal tokens in EPUB and PDF (#458) | ||
|
|
a63ead4002 | chore(site): rebuild data.js | ||
|
|
2cde5bc42b | fix(site): sync counts, repair actions, and harden public resources (#456) | ||
|
|
a56b4b8ad4 | chore(site): rebuild data.js | ||
|
|
6cdb135c5c | feat: add AI Engineering Learning Paths (#444) | ||
|
|
39ea8a1c6d | chore(site): rebuild data.js | ||
|
|
6571f5430a | feat(site): improve motion, responsive UX, and agent readiness (#426) |