Rohit Ghumare 9f06b4b740 feat(certifications): MCPA exam prep on the MCP 2026-07-28 protocol (#484)
* feat(mcpa): add MCPA program, track blueprint, and source ledger

Start the Model Context Protocol Associate certification prep, mirroring the
Claude certification system. Adds the program manifest, the mcpa-f track with
the five published domains and weights (Fundamentals 16, Architecture 14,
Interactions 26, Security 24, Use Cases 20) and exam mechanics, and a source
verification ledger that maps every exam fact to the official certification
page, the launch announcement, and the MCP 2026-07-28 specification with the
retrieval date. Objectives are original, derived from the published
sub-competency names and the specification.

* feat(mcpa): add the MCP fundamentals lesson and figure runtime

First MCPA cert lesson, mirroring the Claude certification lesson contract: a
concept-first explainer of the integration problem MCP solves, the four roles,
and the handshake, discovery, and invocation flow, with a stdlib JSON-RPC mock
(no SDK, no network), five-plus tests, a six-question quiz, a field-brief
artifact, and a registered host-client-server figure. Aligned to MCP 2026-07-28.

* feat(certifications): make cert audit + debias provider-aware

Discover providers via certifications/*/program.json and run the same
contract per provider instead of hardcoding certifications/claude.

- audit_certifications.py: PROVIDERS registry + use_provider(slug);
  check_public_pages runs once, run_provider_audit(audit, slug) per
  provider; LESSON_REF_PREFIX / PROVIDER_ENTITY / PROVIDER_AI_NATIVE_SURFACE
  replace the former claude-only literals; AI-native-surface check gated
  per provider.
- debias_certification_questions.py: glob certifications/*/lessons and
  certifications/*/assessments.
- claude verdict unchanged: 33 lessons, 8 assessments, 295 questions, 0 issues.
- balance mcpa lesson 01 quiz answer positions and option lengths.

* feat(certifications): add MCPA cert README and generalize route-count regex

- certifications/mcpa/README.md: route table (MCPA, 10 lessons), blueprint
  weights, AI-tutor onboarding, GitHub lesson index, non-affiliation notice;
  states the official item count and passing score are not published and the
  90-vs-120 minute source discrepancy.
- audit_certifications.py: widen the README route-count regex from CC-only to
  any uppercase exam code so MCPA is parsed; claude codes still match and the
  claude verdict stays 0 issues.

* feat(certifications): build MCPA track lessons, figures, and wiring

- Nine new lessons (00, 02-09) covering all five MCPA domains. Each ships a
  concept-first doc, a standard-library MCP mock in code/main.py, a Python
  test suite (88 lesson tests total, all passing), a six-question quiz, a
  reusable output artifact, and a mechanism figure.
- site/figures-mcpa-certifications.js: register all ten lesson figures
  (mcpa-00 through mcpa-09), each a themed inline-SVG diagram; render-checked.
- tracks/mcpa-f.json: wire the ten lessons to their domains and roles.
- prerequisites.json: linear lesson dependency chain over all ten lessons.
- audit_certifications.py: enforce the mcpa expectedFigures map.

* feat(certifications): add MCPA assessments and AI-native tutor surface

Completes the MCPA certification track. Full audit is green: 43 lessons,
10 assessments, 380 questions, 0 issues; the Claude verdict is unchanged
(33 lessons, 8 assessments, 295 questions).

- assessments/mcpa-f/diagnostic.json (25 questions) and mock-01.json (60
  questions) with the blueprint domain mix (10/8/16/14/12); every question
  maps to a declared domain objective and cites an existing lesson for
  remediation. Declared in tracks/mcpa-f.json.
- debias pass balances answer positions across the ten lesson quizzes and
  the two assessments.
- skills/mcpa-certification (SKILL.md + agents/openai.yaml) and its
  .claude/skills mirror: an AI-native tutor for the MCPA route, mirroring
  the Claude certification tutor.
- certifications/mcpa/GETTING_STARTED.md: the GitHub learner guide.
- curriculum.yml: run certification lab tests and demos for every provider,
  not just Claude.
- README.md: MCPA onboarding section and quick-start row.

* feat(certifications): add MCP 2026-07-28 protocol brief and wire-shape checker

The MCPA exam is aligned to the 2026-07-28 revision, which made MCP
stateless (SEP-2575, SEP-2567): no initialize handshake, no sessions,
per-request _meta version and capabilities, mandatory server/discover,
Multi Round-Trip Requests instead of server-initiated requests (SEP-2322),
a required resultType, subscriptions/listen, cacheable list results
(SEP-2549), and a new error-code allocation policy.

- certifications/mcpa/research/mcp-2026-07-28-brief.md: the protocol
  source of truth for every MCPA lesson and question, read from the
  primary specification pages, schema.ts, the extension pages, and the
  SEPs, with a table of exam traps where legacy-era beliefs are now wrong.
- scripts/check_mcpa_wire.py: imports each MCPA lesson's transcript() and
  enforces the 2026-07-28 wire invariants (required _meta fields,
  resultType, cache hints on the six cacheable operations, MRTR retry
  rules, subscription ids on listen streams, header and body agreement,
  the error-code allocation policy) and rejects legacy methods such as
  initialize unless a lesson marks them as compatibility examples.
- scripts/test_check_mcpa_wire.py: 16 tests for the checker's rules.

* refactor(certifications): restructure the MCPA track for the stateless 2026-07-28 protocol

The first MCPA build taught the legacy initialize handshake as current,
answered unknown tools with -32601 and schema-invalid arguments with
-32602, and allocated custom errors in the legacy -32000..-32019 range.
MCP 2026-07-28 removed the handshake and sessions (SEP-2575, SEP-2567),
made input validation a tool execution error (SEP-1303), and partitioned
the server-error range into a legacy block and an MCP-reserved block.

- remove the nine lessons and two assessments built on the legacy model
- rewrite the 18 domain objectives against 2026-07-28: the stateless
  core, discovery and capability negotiation, Multi Round-Trip Requests,
  subscriptions, caching, extensions, and the deprecation lifecycle
- lay out a 34-lesson route across the five domains with a prerequisite
  chain and the expected figure for each lesson
- link all 17 phase 13 MCP lessons as deep dives

* feat(mcpa): integration-problem lesson on the stateless core

- one client discovers and calls two unrelated servers using per-request
  _meta (protocolVersion, clientCapabilities, clientInfo) and
  server/discover with cache hints, with no handshake or session
- N times M versus N plus M integration arithmetic
- the two error channels: an unknown tool is a -32602 protocol error, a
  missing argument is an isError tool execution error the model can fix
  (SEP-1303)
- UnsupportedProtocolVersion -32022 carrying data.supported and
  data.requested
- transcript() passes scripts/check_mcpa_wire.py; this lesson is the
  structural exemplar for the rest of the track

* feat(mcpa): discovery and capability negotiation lesson

- server/discover as the only discovery call a server must implement:
  DiscoverResult with supportedVersions, capabilities, instructions,
  serverInfo, and CacheableResult hints (ttlMs, cacheScope)
- server capabilities are cached per server; client capabilities are
  declared fresh on every request in _meta
- MissingRequiredClientCapability (-32021) with data.requiredCapabilities
  when a tool needs elicitation the request did not declare
- UnsupportedProtocolVersion (-32022) then a retry with a supported
  version and a new request id; era is message shape, not version string
- unknown tool (-32602) versus unknown method (-32601)

* feat(mcpa): stateless core lesson

- statelessness as a protocol invariant (SEP-2575): every request carries
  its own version and capabilities, no initialize, no session
- two replicas behind a round-robin router serve interleaved requests
  from one shared store, so any replica answers any request
- cross-request state through opaque server-minted handles (SEP-2567):
  authorized per call, bounded lifetime, expiry and foreign-principal
  use returned as isError tool execution errors
- list results identical across connections; a request without the
  required _meta fields rejected with -32602

* feat(mcpa): hosts, clients, and servers topology lesson

- one client per server inside the host, with local stdio and remote
  Streamable HTTP servers and trust following process and network
  boundaries
- server features (tools, resources, prompts, completion) versus client
  features (elicitation; sampling and roots deprecated) and the control
  model: model-controlled tools, application-driven resources,
  user-controlled prompts
- host aggregation across servers: tool-name collisions resolved by
  server-id prefixing, serverInfo.name never used as a routing key
  because it is self-reported and not unique
- capability-gated discovery: no tools/list against a server that did
  not declare the tools capability

* feat(mcpa): exam strategy lesson for the 34-lesson route

- blueprint-weighted study allocation and weighted readiness (a heavy
  domain moves the estimate more than a plain average)
- exam mechanics from the source ledger, including the 90-minute page
  value versus the 120-minute launch announcement
- a reading strategy for legacy-era distractors: the initialize
  handshake, sessions, and -32601 for an unknown tool
- the full 34-lesson route with its domain tags, validated so every
  lesson and every domain is covered; the capstone spans all five

* feat(mcpa): tool schema and structured content lesson

- tool definition fields and naming rules (1 to 128 characters of
  A-Za-z0-9_-., unique per server, aggregator prefixing)
- JSON Schema 2020-12 as the default dialect (SEP-1613), any 2020-12
  keyword allowed (SEP-2106), and the recommended no-parameter schema
- network $ref never auto-dereferenced; the registration path refuses it
- outputSchema with structuredContent plus a serialized text mirror
- schema-invalid arguments returned as isError tool execution errors
  (SEP-1303); -32602 kept for unknown tools and missing _meta, with the
  legacy behavior shown only as a labeled counterexample

* feat(mcpa): JSON-RPC envelope and _meta lesson

- the four message shapes under MCP: requests with a non-null unique id,
  results that must carry resultType, errors, and id-less notifications
  that are never answered; no batching
- resultType semantics: complete, input_required, unknown values
  invalid, absent treated as complete for earlier servers
- _meta key grammar (optional reverse-DNS prefix plus name) and the
  reserved-prefix rule on the second label: io.modelcontextprotocol and
  dev.mcp reserved, com.example.mcp not
- the reserved keys, including the OpenTelemetry traceparent, tracestate,
  and baggage exception (SEP-414)
- a request missing protocolVersion or clientCapabilities rejected with
  -32602, shown next to three labeled malformed messages

* feat(mcpa): reading server manifests lesson

- review a server/discover result, a tools/list page, and a registry
  server.json before any tool is called
- tool annotation defaults applied when omitted (readOnlyHint false,
  destructiveHint true and only meaningful when not read-only,
  idempotentHint false, openWorldHint true), treated as untrusted hints
- x-mcp-header rules: HTTP token syntax, case-insensitive uniqueness, no
  number types, never on secrets
- cacheScope as a sharing hint, not access control; instructions as
  self-reported text that can try to steer the model
- reverse-DNS registry namespaces tied to verified owners
- a manifest linter that flags six defects in a careless server and none
  in a clean one

* feat(mcpa): reading the specification lesson

- how the 2026-07-28 specification is organized and the floor every
  implementation must support (base protocol, versioning, message
  patterns) versus optional components
- RFC 2119 and RFC 8174 keyword strength, including the capitals rule
- schema.ts as the source of truth and schema.json as generated output
- revision states (Draft, Current, Final) versus feature states (Active,
  Deprecated, Removed) under the lifecycle policy (SEP-2596): a 12-month
  minimum window measured from the deprecating revision's release, and
  the earliest removal on or after it
- JSON-RPC batching (added 2025-03-26, removed 2025-06-18) as the
  reversal the lifecycle policy was written to prevent
- tracing changelog entries back to their SEPs

* feat(mcpa): resources primitive lesson

- resources/list, resources/read, and resources/templates/list with an
  RFC 6570 template, text and base64 blob contents, and a directory read
  returning several contents
- resource not found as -32602 with data.uri (SEP-2164), never -32002 and
  never an empty contents array
- URI scheme choice: https only when the client can fetch directly;
  file, git, or a custom scheme otherwise
- traversal-safe resolution of file URIs so a path containing ".." can
  never escape the served root
- cache hints on every cacheable result, with a private cacheScope for
  user-specific content

* feat(mcpa): prompts and argument completion lesson

- prompts as user-controlled templates: prompts/list with pagination and
  cache hints, prompts/get substituting arguments into messages that can
  carry resource links
- unknown prompt, missing required argument, and invalid cursor all
  reported as -32602
- completion/complete for ref/prompt and ref/resource references, with
  context.arguments narrowing later suggestions
- the 100-value cap with total and hasMore, exercised against a real
  catalog larger than the cap

* feat(mcpa): error handling lesson

- protocol errors versus tool execution errors (SEP-1303): an unknown
  tool stays -32602, invalid or missing arguments return isError results
  the model can correct
- MCP-reserved codes with their data shapes: HeaderMismatch -32020,
  MissingRequiredClientCapability -32021 (data.requiredCapabilities),
  UnsupportedProtocolVersion -32022 (data.supported, data.requested)
- the 2026-07-28 allocation policy enforced in code: a guard refuses the
  legacy -32000..-32019 block, retired -32002 and -32042, and undefined
  reserved codes before any response is built
- HTTP status mapping only where the specification states one (400,
  404, 202, 401, 403, 405); statuses the spec leaves open are marked so

* feat(mcpa): client registration and identity lesson

- registration priority: pre-registered credentials, Client ID Metadata
  Documents when the authorization server advertises support,
  deprecated Dynamic Client Registration, then asking the user
- CIMD validation: an HTTPS client_id with a path, exact client_id match,
  required fields, and redirect URI checks (SEP-991)
- application_type native versus web for DCR clients
- credentials keyed by issuer and never reused across authorization
  servers; re-registration when the authorization server changes
- per-client consent at a proxy to prevent the confused-deputy attack
- the OAuth Client Credentials and Enterprise-Managed Authorization
  extensions for machine-to-machine and IdP-governed access

* feat(mcpa): deprecated client features lesson

- roots, sampling, and logging deprecated by SEP-2577 but still valid in
  2026-07-28: roots/list and sampling/createMessage travel as MRTR
  inputRequests, gated by declared client capabilities (-32021 when
  missing)
- per-request io.modelcontextprotocol/logLevel with notifications/message
  only on that request's own stream, and -32602 for an unknown level
- removed versus deprecated: logging/setLevel and
  notifications/roots/list_changed are gone (SEP-2575), shown only as a
  labeled legacy contrast
- earliest removal computed from the 12-month window, with migration
  paths for each feature, DCR, includeContext, and HTTP+SSE

* feat(mcpa): multi round-trip requests and elicitation lesson

- MRTR replacing server-initiated requests (SEP-2322): an input_required
  result with inputRequests and requestState, then a retry with a new id,
  inputResponses, and the state echoed exactly
- requestState protected with an HMAC that binds the principal, a short
  expiry, a digest of the originating request, and a single-use nonce
- tampering, expiry, cross-principal replay, and a retargeted retry all
  rejected before any side effect
- form-mode elicitation with accept, decline, and cancel; URL mode for
  out-of-band sensitive steps (SEP-1036); a client without the
  elicitation capability refused with -32021

* feat(mcpa): risk and safety controls lesson

- a threat model mapped to the clause that mitigates each threat: tool
  poisoning, prompt injection through results, rug pulls, tool
  shadowing, confused deputy, token passthrough, requestState tampering,
  SSRF through CIMD fetches and network $ref, DNS rebinding, malicious
  icons, and supply-chain drift
- a gateway that pins tool definitions by hash and holds a changed
  definition for review, quarantines descriptions carrying injected
  instructions, refuses network $ref, blocks token passthrough, and
  enforces a per-tool rate limit
- policy refusals returned as isError tool results, never as invented
  codes in the reserved -32000..-32099 range

* feat(mcpa): tools primitive lesson

- tools/list with opaque cursors (an empty-string cursor still means more
  pages), cache hints, and results identical across connections
- CallToolResult: content, structuredContent, and isError defaulting to
  false
- every content block type (text, image, audio, resource_link, embedded
  resource) with audience, priority, and lastModified annotations; an
  embedded resource's annotations sit beside resource, per schema.ts
- tool annotation defaults (readOnlyHint false, destructiveHint true,
  idempotentHint false, openWorldHint true) treated as untrusted hints
- listChanged through subscriptions/listen, from acknowledgment to
  notifications/tools/list_changed to a fresh tools/list

* feat(mcpa): notifications, subscriptions, and cancellation lesson

- subscriptions/listen replacing resources/subscribe and the HTTP GET
  stream: the notification filter, the acknowledgment as the first
  message, and subscriptionId equal to the listen request id so several
  subscriptions can be demultiplexed
- stream notifications (list_changed, resources/updated) versus
  request-scoped progress and message notifications that never travel on
  a listen stream
- progress tokens with strictly increasing progress
- cancellation per transport: closing the SSE stream on HTTP,
  notifications/cancelled on stdio; the server sends that notification
  only to tear down a listen stream; graceful closure and late-message
  races
- a fresh subscriptions/listen after a stdio reconnect

* feat(mcpa): transports and HTTP header contract lesson

- stdio: newline-delimited framing with embedded newlines rejected,
  stdout reserved for MCP messages, stderr for logs, cancellation by
  notification, shutdown by closing stdin
- Streamable HTTP without sessions: one POST endpoint, JSON or
  per-request SSE responses, 202 for notifications, 405 for GET and
  DELETE, no resumability
- Origin validation with 403 and localhost binding against DNS rebinding
- the header contract (SEP-2243): MCP-Protocol-Version, Mcp-Method, and
  Mcp-Name mirrored from the body, x-mcp-header parameters as
  Mcp-Param-{Name}, and base64 sentinel encoding checked against the
  specification's own worked examples
- a header-body mismatch rejected with 400 and HeaderMismatch -32020

* feat(mcpa): trust zones lesson

- five trust zones in one exchange: user and host, client, server,
  upstream systems, and the model
- every untrusted input that reaches model context labeled and
  quarantined: tool descriptions, annotations, icons, results, and
  resource contents
- clientInfo and serverInfo treated as self-reported display data
- multi-server isolation: an instruction embedded in one server's result
  cannot trigger a call on another server
- annotations from an untrusted server fall back to safe defaults;
  icon URIs other than https or data rejected
- local server launch commands accepted only from the host's own
  configuration (SEP-1024)

* docs(mcpa): record primary-source conflicts resolved during the lesson build

Seven places where specification pages, schema.ts, and SEP texts
disagree, each with the normative resolution: removed versus deprecated
roots and logging methods (SEP-2575 over SEP-2577's feature grouping),
embedded-resource annotation placement (schema.ts over the rendered
example), notification-POST headers, requestState rejection channel,
HTTP statuses for three JSON-RPC codes, server-side subscription
teardown, and tool-name format (published page over SEP-986's draft).
Points the specification leaves open are marked so assessments never
test them as a single rule.

* feat(mcpa): model interaction flow lesson

- the host loop from user request through model selection, tools/call,
  and back into model context, driven by a deterministic scripted model
- tool execution errors (isError true) fed back to the model and
  corrected on a new request id; protocol errors such as an unknown tool
  (-32602) never retried blindly
- an input_required result gathered from a human through
  elicitation/create and retried with a new id, inputResponses, and the
  requestState echoed exactly
- a tool with no annotations treated as destructive under the
  specification defaults: one call confirmed and sent, one denied and
  never put on the wire
- deterministic tools/list ordering across identical calls

* feat(mcpa): cache freshness and cursor pagination lesson

- ttlMs and cacheScope on every cacheable result (SEP-2549): absent or
  negative ttlMs clamps to 0, cacheScope has no default
- public entries shared across tokens behind a shared cache, private
  entries never crossing identities, proven by server call counts
- input_required results and MRTR completions never cached
- opaque cursors, including an empty-string cursor that still means
  another page, and invalid cursors rejected with -32602
- notifications/resources/list_changed on a subscriptions/listen stream
  invalidating a cached list before its ttlMs runs out
- stable list ordering so client caches and model prompt caches see a
  byte-identical catalog

* feat(mcpa): audit trail and trace propagation lesson

- W3C traceparent in _meta (SEP-414): version, trace id, parent id, and
  flags validated as lowercase hex of exact length, all-zero ids
  rejected, tracestate and baggage passed through untouched
- a three-hop call (client, server, upstream server) where each hop
  mints a child span under the same trace id
- a hash-chained audit log recording the authenticated principal (never
  self-reported clientInfo), method, tool, redacted arguments, result
  channel, request id, and trace id
- tamper detection for both a naive edit and an edit that rewrites its
  own hash, caught one entry later through prev_hash
- denied attempts (unknown tool, bad bearer token) logged as protocol
  errors, not dropped
- deprecated logging replaced by stderr and OpenTelemetry, with the
  per-request logLevel still producing notifications/message

* feat(mcpa): long-running work and the tasks extension lesson

- the tasks extension (SEP-2663) negotiated per request: the client
  declares io.modelcontextprotocol/tasks in clientCapabilities and the
  server advertises it in server/discover
- a tool that runs synchronously without the extension and returns
  resultType task (CreateTaskResult) with it
- tasks/get polling through working, input_required, and completed,
  with tasks/get itself always a complete result
- mid-flight input supplied with tasks/update and inputResponses;
  cooperative cancellation with tasks/cancel
- -32602 for an unknown or expired taskId and -32021 naming the
  extension for a client that never declared it
- what changed from the 2025-11-25 experimental tasks: no tasks/result,
  no tasks/list, and no task parameter on tools/call
- choosing among a plain call, an MRTR round trip, and a task

* feat(mcpa): OAuth authorization lesson

- the three roles: the MCP server as resource server, the MCP client,
  and the authorization server
- Protected Resource Metadata discovery (RFC 9728): the
  WWW-Authenticate resource_metadata URL first, then the path-specific
  and root well-known fallbacks
- authorization server metadata discovery order for path and root
  issuers, with the issuer required to match
- PKCE S256 generated with the standard library, and the client
  refusing to proceed when code_challenge_methods_supported is absent
- the RFC 8707 resource parameter carrying the canonical server URI on
  both the authorization and token requests
- the RFC 9207 iss decision table, compared without normalization and
  applied to error responses too
- bearer tokens in the Authorization header on every request, audience
  validation, no token passthrough, and 401 versus 403 versus 400

* feat(mcpa): operational use-case selection lesson

- the control model as the first design question: model-controlled
  tools, application-driven resources, user-controlled prompts, and the
  skills extension for cataloged multi-step procedures
- local systems on stdio with credentials from the environment; remote
  systems on Streamable HTTP with the interactive OAuth flow, the client
  credentials extension for unattended callers, or enterprise-managed
  authorization for a central identity provider
- the tasks extension for long jobs and MCP Apps for interactive views,
  with a text fallback for hosts that do not declare the ui extension
- cacheScope chosen from data sensitivity, and a no-external-system case
  where MCP is not the right fit
- a decision engine that turns a use-case profile into a primitive,
  transport, auth path, and extension recommendation

* feat(mcpa): deployment roles and adoption lesson

- six deployment roles on top of the host, client, and server
  architecture: server author, host and client developer, platform or
  gateway operator, security and governance owner, registry publisher,
  and end user
- a responsibility matrix built from a catalog of exact MUST and SHOULD
  requirements with page citations, where the same requirement (Origin
  validation) moves from the server author to the gateway operator as
  the deployment shape changes
- gap reporting for a requirement no role owns
- stdio credentials supplied by the operator's environment and never
  carried on the wire
- governance: Agentic AI Foundation stewardship, the maintainer
  hierarchy and contributor ladder minimums, Working Groups versus
  Interest Groups, the SEP status workflow, and the feature lifecycle
- SDK tiers (SEP-1730) as an adoption risk decision: conformance,
  feature, triage, and critical-bug commitments per tier

* feat(mcpa): tool invocation lifecycle lesson

- eight checkpoints from discover and list through select, confirm,
  call, validate, execute, and result, with select and confirm as
  host-only steps that never touch the wire
- validate ends only in a protocol error (-32602 for an unknown tool);
  schema violations, upstream failures, and business-rule failures in
  execute become tool execution errors with isError true
- a genuine internal fault during execute surfaced as -32603
- notifications/progress bound to the request's progressToken with a
  strictly increasing progress value, and a hard timeout that progress
  does not extend
- transport-specific cancellation: closing the stream on Streamable HTTP,
  notifications/cancelled on stdio, and no cancelled or timed-out
  resultType
- reissuing after a broken stream with a new request id, guided by
  idempotentHint as an untrusted hint and a server-minted handle

* feat(mcpa): protocol eras and compatibility lesson

- the revision timeline from 2024-11-05 to 2026-07-28 and the legacy,
  modern, and dual-era terms
- the stdio probe for a dual-era client: server/discover first, a
  DiscoverResult or a recognized modern error such as -32022 means
  modern (retry with a supported version, never fall back), any other
  error or a timeout means legacy
- the fallback never keyed on one specific error code
- the Streamable HTTP probe: reading a 400 body for a recognized modern
  JSON-RPC error before assuming legacy
- era cached per server process on stdio or per origin on HTTP, with a
  re-probe when the cached assumption fails
- a modern-only server naming its supported versions when it rejects a
  legacy initialize, and the legacy opening sequence shown only inside
  wrapped legacy examples

* feat(mcpa): consent and least privilege lesson

- consent gathered through a multi round-trip request: a tools/call
  returns input_required with an elicitation/create request, and the
  retry carries a new id, inputResponses, and the signed requestState
- decline, cancel, and a rejected retry returned as tool execution
  errors with isError true, never an invented consent error code
- consent scoped per tool and bound to the approved arguments, with
  requestState signed, single-use, and rejected when a retry changes
  the arguments
- annotation defaults (destructiveHint and openWorldHint true,
  readOnlyHint false) deciding when to prompt, treated as untrusted
  hints rather than enforcement
- step-up authorization: a 403 insufficient_scope challenge, a new token
  requested for the union of granted and challenged scopes, and a retry
  cap when the scope is never granted
- tools/list filtered by granted scopes and cached with cacheScope
  private because it varies by authorization

* feat(mcpa): extensions framework lesson

- extension identifiers as {vendor-prefix}/{name}: the
  io.modelcontextprotocol/ prefix for official extensions and a reversed
  owned domain for third parties
- negotiation on every request: the client declares extensions in
  clientCapabilities and the server advertises its own in server/discover,
  with the settings object as the value
- an optional extension falling back to core behavior and a mandatory
  one rejected with -32021 naming it in data.requiredCapabilities
- extensions disabled by default and opt-in; a breaking change needs a
  new identifier
- the SEP-2133 process, experimental-ext- repositories owned by a
  Working Group or Interest Group, and the official roster: tasks, MCP
  Apps, skills, and the two authorization extensions
- what changed from initialize-time declaration to per-request
  negotiation

* feat(mcpa): registry, gateways, and SDK tiers lesson

- the MCP Registry as a preview metadata index for public servers:
  server.json with packages (npm, pypi, nuget, cargo, oci, mcpb),
  remotes (streamable-http, deprecated sse), or both
- reverse-DNS namespaces admitted only after GitHub, DNS, or HTTP
  verification; spoofed namespaces and private servers rejected
- exact version strings (ranges prohibited), and the server.json schema
  version kept separate from the protocol version
- aggregators and subregistries built on top of the registry
- gateways that route on the Mcp-Method and Mcp-Name headers, reject a
  header and body mismatch with -32020 before any backend is touched,
  partition private cache entries by caller, and never pass the
  caller's token through to a backend
- SDK tiers (SEP-1730): conformance, feature, triage, and critical-bug
  commitments, and relegation after four weeks of continuous failures

* docs(mcpa): pin MCP Apps metadata shapes in the protocol brief

The extension summary named _meta.ui.csp and _meta.ui.permissions
without their shape or location. Checked against the MCP Apps
specification (ext-apps 2026-01-26, draft agrees):

- a tool's _meta.ui carries resourceUri and visibility, which defaults
  to model and app and gates what the agent lists and what an app may
  call
- the UI resource's _meta.ui carries csp, an object of connectDomains,
  resourceDomains, frameDomains, and baseUriDomains origin lists, and
  permissions, an object of camera, microphone, geolocation, and
  clipboardWrite flags
- the host builds CSP from declared domains only, may restrict further,
  applies a restrictive default when csp is omitted, and may honor
  permissions that apps must not assume
- the app's ui/initialize is unrelated to the removed core initialize

* feat(mcpa): capstone that reads one exchange end to end

One incident-console server driven through a single transcript that
exercises every domain:

- server/discover with cache hints after an unsupported version is
  corrected from the -32022 data.supported list
- protocol version and client capabilities in _meta on every request
- a schema-invalid call returned with isError true and corrected
- a consent round trip with an HMAC-signed, principal-bound
  requestState, a retry on a new id, and a tampered state rejected
- -32021 when the client never declared elicitation or the tasks
  extension
- a long diagnostic run as a task, polled to completion, and a second
  task cancelled
- request-scoped notifications/progress and the Mcp-Method and Mcp-Name
  header contract
- a wrong-audience token rejected with 401 before any JSON-RPC body is
  read
- one traceparent trace id carried through every hop, and a hash-chained
  audit log that pinpoints a tampered entry
- a readiness checklist mapping every objective to what a candidate must
  be able to do

* docs(mcpa): teach the tutor and guides the stateless 34-lesson route

- tutor skill (and its Claude Code mirror): the 2026-07-28 revision is
  taught as current, with no initialize handshake, no sessions, and
  per-request _meta plus server/discover; older revisions appear only as
  what changed, and deprecated features as still working until removal
- the protocol brief and the wire-shape checker join the tutor's source
  list, and the tutor runs the checker when a learner edits a transcript
- three full mocks with distinct emphasis, and a fresh mock for every
  retake so a second score measures readiness rather than recall
- the capstone description now matches lesson 33's integrated exchange
- learner guide: 34-lesson route, 30-question diagnostic, three mocks,
  the multi round-trip lesson as the worked example, and the wire
  checker in the local verification suite
- root README: the MCPA summary describes the new route, and the skills
  table and install list include mcpa-certification

* ci(certifications): gate MCPA transcripts on the 2026-07-28 wire shape

Runs the wire-shape checker's own tests and then the checker over every
MCPA lesson transcript, so a lesson that reintroduces the legacy
handshake, drops the per-request _meta fields, omits resultType, or
emits an undefined error code fails CI. Both scripts are added to the
workflow's path filters.

* feat(mcpa): MCP Apps interactive interfaces lesson

- the io.modelcontextprotocol/ui extension negotiated per request, with
  a text-only fallback for hosts that never declare it
- a tool's _meta.ui carrying resourceUri and visibility: the agent's tool
  list excludes tools without "model", and an app's tools/call is refused
  for tools without "app", before any consent prompt
- the ui:// resource fetched with an ordinary resources/read and
  recognized only by the text/html;profile=mcp-app mime type
- the resource's _meta.ui csp object (connectDomains, resourceDomains,
  frameDomains, baseUriDomains) turned into a Content Security Policy
  that starts from default-src 'none' and admits only declared domains,
  with the specification's restrictive default when csp is omitted
- permissions flags (camera, microphone, geolocation, clipboardWrite)
  honored at the host's discretion and never assumed by the app
- the sandboxed iframe, a sandbox proxy on its own origin for web hosts,
  and the app's ui/ JSON-RPC dialect over postMessage, including a
  ui/initialize unrelated to the removed core initialize

* feat(mcpa): register all 34 lesson figures and index the route

- site/figures-mcpa-certifications.js rebuilt from each lesson's figure
  snippet: 34 mechanism figures (blueprint weights through the capstone
  flow), each with its own CSS prefix and marker ids, all rendering an
  SVG under a stub DOM; the ten figures of the legacy route are gone
- certifications/mcpa/README.md: the lesson index now lists all 34
  lessons by title, the route diagram follows the new domains, and the
  overview names the protocol brief, the wire-shape checker, the
  30-question diagnostic, and the three mocks

* docs(mcpa): give each deprecated feature its own source and removal floor

The brief grouped all six deprecated features under SEP-2577 with one
2027-07-28 removal floor. The specification's deprecated registry
records them separately:

- Roots, Sampling, and Logging: SEP-2577, deprecated in 2026-07-28,
  earliest removal on or after 2027-07-28
- Dynamic Client Registration: PR #2858, deprecated in 2026-07-28,
  earliest removal on or after 2027-07-28
- includeContext "thisServer" and "allServers": SEP-2596, deprecated in
  2025-11-25, removal follows Sampling
- HTTP+SSE: SEP-2596, deprecated in 2025-03-26, earliest removal three
  months after SEP-2596 reaches Final

Lesson 15 already teaches the registry's values; this brings the brief
in line so later questions do not inherit the grouped floor.

* docs(mcpa): state the x-mcp-header limits with their RFC 2119 strength

The brief said x-mcp-header was "never for secrets" and that clients
drop invalid tools. The tools page is more precise:

- only integer, string, and boolean parameters that are statically
  reachable from the schema root can be mirrored; number is excluded
- server developers SHOULD NOT mark sensitive parameters (passwords,
  API keys, tokens, PII), because header values are visible to
  intermediaries; this is a SHOULD NOT, not a prohibition
- HTTP clients MUST reject a tool whose x-mcp-header values violate the
  constraints by excluding it from their tools/list result

* docs(mcpa): teach what a revision date means and how hosts load skills

Two facts the assessments test were not taught by any lesson:

- lesson 01 and the brief: a revision identifier is a YYYY-MM-DD date
  marking the last backwards incompatible change, so a Current revision
  can take compatible fixes without being renamed (Draft, Current,
  Final)
- lesson 30 and the brief: reading a skill's SKILL.md through
  resources/read only retrieves text; skill content is untrusted input
  tagged with its origin, explicit user policy decides whether a skill
  is loaded, hosts should let users inspect a skill first, may not let
  it trigger host-side execution without per-skill approval, and ignore
  permission-widening frontmatter such as allowed-tools unless the user
  approved it (SEP-2640)

* docs(mcpa): teach the x-mcp-header mirroring limits in the transports lesson

The transports lesson introduced x-mcp-header mirroring and the base64
sentinel but not its limits, which the assessments test:

- only integer, string, and boolean parameters statically reachable
  from the schema root can be mirrored, never number
- a Streamable HTTP client must exclude a tool whose x-mcp-header
  values break those constraints from its tools/list result
- server developers should not mark sensitive parameters such as API
  keys or tokens, because header values are visible to intermediaries

* fix(site): refresh the figure manifest cache key for the rebuilt MCPA figures

lesson.html pins figure-manifest.js with a content hash. Rebuilding
site/figures-mcpa-certifications.js with the 34 new figures changed the
generated manifest, so the pinned key moves from 0980822a99ac to
339ef88a8714; test_build_artifacts.js checks that the committed page
matches the manifest the build produces.

* docs(i18n): regenerate the translated READMEs with the MCPA sections

The English README gained the MCPA route row, the MCPA section, the
mcpa-certification skill row, and the updated install list, but the
twelve i18n/<lang>/README.md files were never regenerated, so
build_readme_i18n.py --check failed. Regenerated with the script; the
new blocks have no hand-authored translations yet and fall back to
English as the generator documents, and no existing translation was
dropped.

* feat(mcpa): diagnostic and three full-length mocks on the 2026-07-28 protocol

Four original assessments declared on the mcpa-f track, 210 questions,
every key checked against the protocol brief and the specification
pages it cites:

- diagnostic (30 questions, 30 minutes): one concept per item across
  all 18 objectives, for placing a learner by domain
- mock 1 (60 questions, 90 minutes): operational scenarios, such as
  scaling a stateful tool behind a load balancer, one-click local
  server consent, registry preview planning, and DCR application_type
- mock 2 (60 questions, 90 minutes): wire-level messages, such as
  missing _meta fields, -32021 and -32022 error data, progress
  monotonicity, HTTP cancellation by closing the stream, all-zero
  traceparent ids, and MCP Apps mime types
- mock 3 (60 questions, 90 minutes): design and security trade-offs,
  such as where state lives, requestState protection, token audience,
  skill loading consent, and SDK tier risk

Each mock follows the blueprint split (10/8/16/14/12 across the five
domains), uses every objective, and references all 34 lessons. Items
that repeated another file's scenario were rewritten to test a
different fact, and answer positions are balanced by
debias_certification_questions.py.

* chore(mcpa): balance answer positions across the 34 lesson quizzes

Applies debias_certification_questions.py to the MCPA lesson quizzes so
correct answers follow the per-file balanced position cycle the CI
check enforces. Only option order changes; every question, option, and
explanation is unchanged, and no Claude certification file is touched.

* fix(mcpa): answer unknown methods with -32601 in four lesson servers

The dispatchers in lessons 04, 11, 12, and 14 fell through to -32602
(invalid params) for a JSON-RPC method they do not implement. An
unknown method is -32601 (method not found); -32602 is reserved for an
unknown tool, missing _meta, and other invalid parameters. Every other
lesson dispatcher already used -32601, so a learner reading these four
servers top to bottom would have learned the wrong code.

Also adds a test for lesson 03's validate_error_shape, the one shape
validator its tests never exercised.

* fix(mcpa): align three labs with what their lessons say they show

- lesson 30: the text pointed learners at the sixth exchange for the
  malformed no-slash-here identifier, but it is the seventh
- lesson 32: the demo printed "prefer remote" while calling
  resolve_install_target with prefer="package"
- lesson 33: acknowledge_incident is listed by tools/list but a plain
  tools/call answered "Unknown tool" (-32602); it now returns a tool
  execution error saying the tool is served only over the authorized
  HTTP endpoint, with a test

* fix(scripts): flag legacy and violation wrappers that do not wrap a message

check_mcpa_wire.py skipped every entry marked legacy or violation, and
when the wrapped message was not an object (a typo such as None or a
bare method string) the entry produced no finding at all, hiding an
authoring mistake. Such wrappers are now reported, matching how every
other malformed entry is handled, with a regression test for both
wrapper kinds.

* docs(mcpa): give every lab the source header the lesson contract requires

AGENTS.md asks each lesson's code/main.py to open with a 4 to 6 line
header citing its docs/en.md path and its spec or RFC sources. The MCPA
labs had a one-line docstring. Each header now names the lesson's full
docs/en.md path, what the lab does, and the specification pages, SEPs,
RFCs, or W3C documents it implements.

* fix(mcpa): require an explicit approval before consent-gated work runs

Two labs treated any elicitation answer with action "accept" as consent,
even when the form content said no:

- lesson 25: accepting the delete_file confirmation with approved false
  still deleted the file; consent now needs action accept and a content
  object whose approved field is exactly true, and a non-object answer
  is handled safely
- lesson 20: accepting the vault read with proceed false still returned
  the private note; the vault now needs proceed exactly true

Each fix has a test showing the refused path leaves no side effect.

* fix(mcpa): turn malformed wire input into refusals instead of exceptions

Six labs raised on attacker-controlled or malformed input instead of
answering it:

- lesson 05: a -32022 error with an empty supported list raised
  IndexError; the probe now records a modern era with no confirmed
  version and only caches a retry version after a successful discover
- lesson 09: an x-mcp-header value that is not a string crashed the
  linter, and object or array properties were not flagged; only string,
  integer, and boolean parameters may be mirrored
- lesson 14: a non-ASCII requestState raised during signature checks; it
  is now rejected as malformed
- lesson 19: a base64 sentinel header with invalid base64 or UTF-8
  raised; it now decodes to a mismatch and gets the HeaderMismatch reply
- lessons 27 and 33: a malformed traceparent raised ValueError; lesson
  27 restarts the trace with a new root and lesson 33 continues without
  a trace id

Each case has a regression test.

* fix(mcpa): close gaps in the labs' own security controls

- lesson 22: only the first CALL server.tool instruction in server
  content was checked, so a harmless first instruction could hide a
  cross-server one; every match is now quarantined and checked on relay
- lesson 26: the pinned tool descriptor left out annotations, so a
  silent destructiveHint flip was not caught as a rug pull, although the
  lesson's own threat matrix pins annotations; they are now part of the
  hashed descriptor
- lesson 32: the gateway log stored the caller's raw bearer token as the
  principal and printed it in the transcript; it now records a
  non-reversible principal reference, and the backend answers an
  unimplemented method with -32601 instead of -32602

Each fix has a test.

* fix(mcpa): expire idle baskets from their last activity

Lesson 04 tells learners a basket expires after five idle ticks, but
is_expired measured from creation, so an actively used basket still
expired. Baskets now record their last activity, adding an item
refreshes it, and a test shows an item added inside the window keeps
the basket alive.

* docs(mcpa): match run locations and rejection rules to the labs and spec

- lessons 09, 10, and 12 told learners to run python3 code/main.py from
  the repository root, where that relative path does not exist; they
  now say to run it from the lesson directory like the other labs
- the lesson 14 checklist said a protocol error is the wrong channel for
  a tampered or expired requestState; the spec requires rejection but
  does not prescribe the channel, so the checklist now names the lab's
  choice, a tool execution error, and a fresh input_required result as
  valid options
- the lesson 21 patterns sheet said a tools/call without the tasks
  extension always gets -32021; per SEP-2663 the server returns an
  ordinary result when it can finish within the request and -32021 only
  when it cannot, and tasks/get, tasks/update, and tasks/cancel without
  the declaration get -32021

* docs(mcpa): say the MCPA track is on GitHub only for now

The website build renders only the first certification program, so no
MCPA page is published there yet. GETTING_STARTED, both copies of the
tutor skill, and the README certification row pointed learners at
website routes that do not exist. They now send learners to the GitHub
lesson and assessment paths and say the website does not carry MCPA
yet. The translated READMEs are regenerated from the English one.

* fix(audit): stop treating the MCPA practice mock size as an exam fact

The official MCPA page does not publish an item count, and the fact
ledger records it as not published. The audit's MCPA-F exam facts listed
items 60, so check_track verified the practice mock size as if it were
official. The item count is removed from the verified facts; the track
keeps 60 as the curriculum's chosen mock length.

* fix(mcpa): refuse differently cased cross-server instructions in lesson 22

The embedded-instruction pattern matches CALL server.tool without regard
to letter case, but the relay check compared the captured names exactly,
so CALL Tickets.delete_all_tickets in untrusted server content still let
a relay to tickets.delete_all_tickets through as allowed. Quarantine and
relay checks now compare server and tool names case-insensitively, so a
case variant is refused, and a test covers the case-shifted instruction
and the same-server case.

* feat(site): render every certification program, not only the first

parseCertifications read only the first folder under certifications/, so
the website showed Claude alone: the catalog, track, assessment, and
lesson pages, the homepage spotlight, search, sitemap, and llms.txt never
saw MCPA. The build now loads every program, tags each track, lesson, and
assessment with its programId, derives each program's learner guide and
tutor skill paths, and fails on duplicate program, track, or assessment
ids across programs.

- the catalog renders one section per program with its own access
  notice, GitHub tutor links, and track grid, and the no-JavaScript
  discovery block follows the same structure
- track, assessment, and lesson pages show the disclaimer and scoring
  notice of the track's own program instead of a hardcoded Anthropic one
- lesson data loading, the language picker, and the api lesson and
  certification routes accept any certifications/<program>/lessons path
- the homepage spotlight shows both badges and names both providers, and
  build.js keeps its track, lesson, and question counts in sync
- program.json gains shortName and accessNoticeTitle, and the MCPA badge
  is marked square so it is not clipped to the Claude badge outline
- exam facts read "Not published" when a track marks its item count or
  passing score unpublished, so the MCPA card no longer presents the
  practice mock size as the official question count
- a track with four assessments lays them out two by two
- tests derive track counts from the data and cover a second program in
  the lesson and certification routes

* docs(mcpa): send learners to the MCPA track now that the site renders it

The GitHub-only wording existed because the website could not show a
second program. With the multi-program build, GETTING_STARTED, both
copies of the tutor skill, and the README point to the MCPA track page
again, the certifications index lists both programs with their onboarding
guides, and the translated READMEs are regenerated.

* docs(i18n): translate the MCPA goal row in the Portuguese and Russian READMEs

The Portuguese and Russian READMEs translate every row of the "choose
what you want to build" table, but the new MCPA row had no entry in
readme_translations.py, so it rendered in English between translated
rows. Both languages now carry the row, with the onboarding guide and
the MCPA track links unchanged, and the translated READMEs are
regenerated.

* feat(site): add a Sponsor us page rendered from SPONSORS.md

SPONSORS.md promises a sponsor page on the curriculum site, but the site
had no sponsor page and no route to one. build.js now renders SPONSORS.md
into site/sponsors.html on every build, so the page cannot drift from the
file the maintainer edits.

- the renderer covers what SPONSORS.md uses: headings with GitHub-style
  anchor ids, paragraphs, lists with wrapped items, aligned tables, bold,
  inline code, and links, where repository files resolve to GitHub and
  javascript: or parent-directory links render as plain labels
- raw HTML is limited to a, picture, source, and img with https-only
  URLs; any other tag is escaped, and a closing tag is kept only when its
  opening tag was kept
- the SerpApi logo follows the site's theme toggle instead of the
  operating system color scheme
- "Sponsor us" joins the hamburger menu through header.js and the footer
  of every page, and is added to the interface strings for translation
- /sponsors is served from sponsors.html with the same markdown
  negotiation as the other public pages, and the page is listed in the
  sitemap and llms.txt
- tests pin the rendered page to SPONSORS.md, the HTML allowlist, and
  the menu and footer links

* fix(site): give the privacy, contact, and developer pages the shared header

The three trust pages shipped a stripped header with no header.js, no
search or theme controls, no site fonts, and an old stylesheet key. At
1400px and below the stylesheet renders the header nav as a dropdown
panel that header.js normally hides behind the menu button, so on these
pages the panel stayed open over the content.

They now use the same header markup as the other pages, load the site
fonts, the current stylesheet, and the shared theme, progress, header,
and search scripts, gain a skip link to main, and style their eyebrow
line. The shared asset test now covers all three pages so their cache
keys cannot fall behind again.
2026-09-25 10:20:11 +05:30
2026-04-22 14:49:03 +01:00
2026-03-18 20:53:00 +05:30
2026-03-18 20:53:00 +05:30
2026-03-18 20:49:10 +05:30
2026-03-19 15:22:47 +05:30

AI Engineering from Scratch — reference manual banner

Read in your language: Español · Français · Português · Deutsch · Italiano · 简体中文 · 日本語 · 한국어 · हिन्दी · العربية · Русский · Türkçe
Translated landing pages, committed to the repo. English is canonical; lesson pages are machine-translated on the translations branch. See docs/i18n.md.

MIT License 523 lessons 20 phases GitHub stars Website

From the creator of Agent Memory - #1 Persistent memory ⭐ GitHub stars which naturally works with any agents or chat assistants.

░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒

84% of students already use AI tools. Only 18% feel prepared to use them professionally. This curriculum closes that gap.

523 lessons. 20 phases. ~342 hours. Python, TypeScript, Rust, Julia. Every lesson ships a reusable artifact: a prompt, a skill, an agent, an MCP server. Free, open source, MIT.

You don't just learn AI. You build it. End-to-end. By hand.

114,584 readers  ·  181,995 page views in the last 30 days  ·  as of 2026-08-29

Start here: choose what you want to build

You do not need to scan 523 lessons before beginning. Pick one goal. Each link opens the same curriculum on GitHub or the website, and both versions use the same lesson code.

Your goal Learn on GitHub Learn on the website
I am new and want the complete foundation Phase 0: Setup and Tooling Dev Environment
I know Python and want math plus ML foundations Phase 1: Math Foundations Linear Algebra Intuition
I want to build production LLM applications Phase 11: LLM Engineering Prompt Engineering
I want to build agents Phase 14: Agent Engineering The Agent Loop
I want to use coding agents on real repositories Agent-Assisted Engineering path Agent-Assisted Engineering
I want to shape the right build before implementation Product Judgment and Delivery path Product Judgment and Delivery
I want to build with Model Context Protocol (MCP) Model Context Protocol (MCP) route Model Context Protocol (MCP) path
I want to write and ship Agent Skills Focused Agent Skills route Agent Skills path
I want to prepare for a Claude certification Certification onboarding Certification Academy
I want to prepare for the MCP Associate (MCPA) MCPA onboarding MCPA track

Not sure where you fit? Use the start-learning placement tutor or the website prerequisites guide.

Compare four core domains and six career routes in the AI Engineering Learning Paths.

Sponsors

SerpApi. Web Search API for your AI apps. Available in Markdown and JSON for any integration.


Thank you to our sponsors.

Your support keeps every lesson free and open source.

See all supporters
Become a sponsor

Use every lesson the same way

  1. Read docs/en.md and explain the core idea in your own words.
  2. Type and build the important code instead of treating the code block as decoration.
  3. Run the lesson command from the repository root, the directory containing README.md and phases/.
  4. Keep evidence: the command, working directory, exit code, meaningful output, and the artifact you changed or produced.
  5. Continue only when you can explain the output and make one small change without guessing.

Commands in lesson pages are paths from the repository root unless the lesson explicitly says to change directories. If a lesson offers several languages, run the implementation for the language you are learning.

Clone it and produce your first evidence

git clone https://github.com/rohitg00/ai-engineering-from-scratch.git
cd ai-engineering-from-scratch
python3 phases/00-setup-and-tooling/01-dev-environment/code/verify.py --route beginner
python3 phases/01-math-foundations/01-linear-algebra-intuition/code/vectors.py

The preflight separates requirements needed now from tools needed later. Every required failure includes the detected reason and a corrective command. The second command is a dependency-free lesson and ends by showing that a matrix times a vector is the operation inside a neural network layer. Save that terminal output as your first evidence.

Add the AI tutor in 30 seconds

If Node.js, npx, and a skill-capable coding agent are already installed, your coding agent can become your tutor in two commands. A repository clone is not needed to install or read the tutor. Runnable focused-path labs need python3. Agent Skills host labs also need a selected host and a writable user or project skill scope.

Check the local requirements first:

node --version
npx --version
python3 --version

Then install the curriculum skills and choose the host and scope you intend to use when the installer asks:

npx skills add rohitg00/ai-engineering-from-scratch

Invocation syntax belongs to the host, not to the portable SKILL.md format:

Host Start the course Start Model Context Protocol (MCP) Start Agent Skills Run a phase quiz
Codex start-learning, or choose it from /skills learn-mcp, or choose it from /skills learn-agent-skills, or choose it from /skills check-understanding 13, or choose it from /skills
Claude Code /start-learning /learn-mcp /learn-agent-skills /check-understanding 13
Other compatible hosts Use start-learning to begin the course. Use learn-mcp to start the Model Context Protocol (MCP) path. Use learn-agent-skills to start the Agent Skills Engineering path. Use check-understanding to quiz me on Phase 13.

A ten-question placement quiz maps what you already know to a starting phase and saves a personalized study plan to LEARNING.md. From there, the learn skill teaches one lesson per session: concept, math, code, quiz. It streams lessons straight from this repo, and the course-guide skill jumps you to the exact lesson that covers anything you are stuck on. In Codex, invoke these skills with learn and course-guide; in Claude Code, use /learn and /course-guide; in other compatible hosts, ask to use the skill by name.

Only want Model Context Protocol (MCP)? Use the MCP invocation for your host. It creates MCP-LEARNING.md and follows one 17-lesson route through stateless requests, transports, bidirectional work, security, reliability, registry governance, and conformance evidence. The exact order and checkpoints live in the Model Context Protocol (MCP) manifest.

Only want Agent Skills? Use the Agent Skills invocation for your host. It creates AGENT-SKILLS-LEARNING.md and follows one coherent five-lesson route: contract, discovery, invocation, sandbox boundaries, then release evals and real-host portability. Start on the web with the Agent Skills path.

The installer lists the hosts it can configure and asks where to install. If you do not have Node.js, npx, python3, a supported host, or a writable scope yet, use the website or read docs/en.md manually. That path teaches the concepts, but real-host discovery, invocation, script, and uninstall evidence remains pending until the preflight is available. Read the lessons at aiengineeringfromscratch.com.

How this works

Most AI material teaches in scattered pieces. A paper here, a fine-tuning post there, a flashy agent demo somewhere else. The pieces rarely line up. You ship a chatbot but can't explain its loss curve. You hook a function to an agent but can't say what attention does inside the model that's calling it.

This curriculum is the spine. 20 phases, 523 lessons, four languages: Python, TypeScript, Rust, Julia. Linear algebra at one end, autonomous swarms at the other. Every algorithm gets built from raw math first. Backprop. Tokenizer. Attention. Agent loop. By the time PyTorch shows up, you already know what it's doing under the hood.

Each lesson runs the same loop: read the problem, derive the math, write the code, run the test, keep the artifact. No five-minute videos, no copy-paste deploys, no hand-holding. Free, open source, and built to run on your own laptop.

░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒

The shape of the curriculum

Twenty phases stack on top of each other. Math is the floor. Agents and production are the roof. Skip ahead if you already know the lower layers, but don't skip and then wonder why something at the top is breaking.

%%{init: {'theme':'base','themeVariables':{'primaryColor':'#fafaf5','primaryTextColor':'#1a1a1a','primaryBorderColor':'#3553ff','lineColor':'#3553ff','fontFamily':'JetBrains Mono','fontSize':'12px'}}}%%
flowchart TB
  P0["Phase 0 — Setup &amp; Tooling"] --> P1["Phase 1 — Math Foundations"]
  P1 --> P2["Phase 2 — ML Fundamentals"]
  P2 --> P3["Phase 3 — Deep Learning Core"]
  P3 --> P4["Phase 4 — Vision"]
  P3 --> P5["Phase 5 — NLP"]
  P3 --> P6["Phase 6 — Speech &amp; Audio"]
  P3 --> P9["Phase 9 — RL"]
  P5 --> P7["Phase 7 — Transformers"]
  P7 --> P8["Phase 8 — GenAI"]
  P7 --> P10["Phase 10 — LLMs from Scratch"]
  P10 --> P11["Phase 11 — LLM Engineering"]
  P10 --> P12["Phase 12 — Multimodal"]
  P11 --> P13["Phase 13 — Tools &amp; Protocols"]
  P13 --> P14["Phase 14 — Agent Engineering"]
  P14 --> P15["Phase 15 — Autonomous Systems"]
  P15 --> P16["Phase 16 — Multi-Agent &amp; Swarms"]
  P14 --> P17["Phase 17 — Infrastructure &amp; Production"]
  P15 --> P18["Phase 18 — Ethics &amp; Alignment"]
  P16 --> P19["Phase 19 — Capstone Projects"]
  P17 --> P19
  P18 --> P19
░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒

The shape of a lesson

Each lesson lives in its own folder, with the same structure across the entire curriculum:

phases/<NN>-<phase-name>/<NN>-<lesson-name>/
├── code/      runnable implementations (Python, TypeScript, Rust, Julia)
├── docs/
│   └── en.md  lesson narrative
└── outputs/   prompts, skills, agents, or MCP servers this lesson produces

Every lesson follows six beats. The Build It / Use It split is the spine — you implement the algorithm from scratch first, then run the same thing through the production library. You understand what the framework is doing because you wrote the smaller version yourself.

%%{init: {'theme':'base','themeVariables':{'primaryColor':'#fafaf5','primaryTextColor':'#1a1a1a','primaryBorderColor':'#3553ff','lineColor':'#3553ff','fontFamily':'JetBrains Mono','fontSize':'13px'}}}%%
flowchart LR
  M["MOTTO<br/><sub>one-line core idea</sub>"] --> Pr["PROBLEM<br/><sub>concrete pain</sub>"]
  Pr --> C["CONCEPT<br/><sub>diagrams &amp; intuition</sub>"]
  C --> B["BUILD IT<br/><sub>raw math, no frameworks</sub>"]
  B --> U["USE IT<br/><sub>same thing in PyTorch / sklearn</sub>"]
  U --> S["SHIP IT<br/><sub>prompt · skill · agent · MCP</sub>"]

Getting started

Three ways in. Pick one.

Option A — learn in your terminal (recommended). After the Node.js, npx, host, and scope preflight above, install the learning skills into a compatible agent and let the course drive itself:

npx skills add rohitg00/ai-engineering-from-scratch

Use the host-specific invocation table above. The installed skills provide start-learning, learn, course-guide, and the focused learn-mcp and learn-agent-skills routes. Lesson prose can stream from this repository without a clone. A local clone is required for copied repository code commands and executable MCP or Agent Skills labs. Progress lives in LEARNING.md, MCP-LEARNING.md, or AGENT-SKILLS-LEARNING.md in your project, so every session can resume.

Option B — read. Open any completed lesson on aiengineeringfromscratch.com or expand a phase under Contents. No setup, no cloning.

Option C — clone and run.

git clone https://github.com/rohitg00/ai-engineering-from-scratch.git
cd ai-engineering-from-scratch
python3 phases/01-math-foundations/01-linear-algebra-intuition/code/vectors.py

Cloning also auto-loads the learning skills in Claude Code, and gives every lesson's code to the learn tutor for real execution instead of read-along.

Prerequisites

  • You can write code (any language; Python helps).
  • You want to understand how AI actually works, not just call APIs.

Prepare for Claude certifications

The Claude Certification Academy is a free, open-source preparation program for all four official Claude certification tracks: Associate Foundations, Developer Foundations, Architect Foundations, and Architect Professional. Each route combines blueprint-mapped lessons, runnable labs, a diagnostic, capstone work, and a full-length original practice exam.

Use the AI-native GitHub onboarding guide with Claude Code, Codex, ChatGPT, Cursor, or another agent. Run claude-certification in Codex, /claude-certification in Claude Code, or ask another host to use claude-certification. It chooses a track, creates a persistent route in CLAUDE-CERTIFICATION.md, teaches one step at a time, runs the real labs, and gives artifact-based feedback. The same curriculum remains available on the certification website.

The academy is independent study material based on public exam objectives. It is not affiliated with Anthropic, does not reproduce live exam questions, and cannot guarantee a passing score.

Prepare for the MCP Associate (MCPA) certification

The MCPA Certification Curriculum is a free, open-source preparation program for the Model Context Protocol Associate exam from the Agentic AI Foundation, delivered through Linux Foundation Training. Its 34 lessons teach the stateless 2026-07-28 protocol across the five exam domains: per-request _meta and server/discover in place of the old handshake, multi round-trip requests, subscriptions, caching, the tasks and MCP Apps extensions, OAuth authorization, and the registry and SDK tiers. Every lesson ships a runnable standard-library lab whose transcript is checked for the current wire shape, and the track adds a diagnostic, a capstone, and three full-length original practice exams whose question mix follows the published blueprint weights.

Use the AI-native GitHub onboarding guide with Claude Code, Codex, ChatGPT, Cursor, or another agent. Run mcpa-certification in Codex, /mcpa-certification in Claude Code, or ask another host to use mcpa-certification. It creates a persistent route in MCPA-CERTIFICATION.md, teaches one step at a time, runs the real labs, and gives artifact-based feedback. The same curriculum is available on the MCPA track page.

This curriculum is independent study material based on public exam objectives. It is not affiliated with the Agentic AI Foundation or the Linux Foundation, does not reproduce live exam questions, and cannot guarantee a passing score.

The learning skills

Skill What it does
start-learning One-time onboarding: why you're learning, placement quiz, personalized plan saved to LEARNING.md.
learn The tutor loop. Warm-up recall, then the next lesson taught interactively, then its quiz; records progress and a review queue.
course-guide Topic router. "Where do I learn attention?" or "my loss is NaN" → the exact lessons, with links.
learn-mcp Focused Model Context Protocol (MCP) tutor. Creates MCP-LEARNING.md, follows the 17-lesson manifest, and records wire, security, reliability, and conformance evidence.
learn-agent-skills Focused Agent Skills tutor. Creates AGENT-SKILLS-LEARNING.md, teaches lessons 22, 24, 25, 26, and 27, and records real-host evidence.
claude-certification Certification tutor. Chooses CCAO-F, CCDV-F, CCAR-F, or CCAR-P; teaches each lesson; runs labs; reviews artifacts; administers diagnostics and mocks; saves progress.
mcpa-certification MCPA tutor. Follows the 34-lesson mcpa-f route on the 2026-07-28 protocol; teaches each lesson; runs labs and the wire checker; administers the diagnostic and three mocks; saves progress.
find-your-level Ten-question placement quiz. Maps your knowledge to a starting phase and produces a personalized path with hour estimates.
check-understanding <phase> Per-phase quiz, eight questions, with feedback and specific lessons to review. Use the Codex, Claude Code, or natural-language form in the invocation table above.
░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒

Read the core curriculum as a book

The 20-phase core curriculum under phases/ compiles into a six-volume book series. EPUB and PDF are built by CI from the same core lesson sources and attached to every GitHub release; the links below always resolve to the newest release. Volume numbers index the series, not versions: each copy carries a dated edition stamp, and older editions stay downloadable from their release.

Certification curricula are intentionally not converted into the books. Their AI tutor state, runnable labs, interactive figures, diagnostics, and timed mocks remain first-class on GitHub and the website.

Vol Title Phases Download
1 Foundations · Math, Tooling, and Classical Machine Learning 00-02 EPUB · PDF
2 Deep Learning · Networks, Vision, and Speech 03, 04, 06 EPUB · PDF
3 Language · NLP Foundations and the Transformer 05, 07 EPUB · PDF
4 Large Language Models · Generation, Reinforcement, Pretraining, and Engineering 08-11 EPUB · PDF
5 Agents · Multimodality, Protocols, Autonomy, and Swarms 12-16 EPUB · PDF
6 Production · Infrastructure, Safety, and Capstones 17-19 EPUB · PDF

The book is the snapshot; this repository is the living edition. Every chapter ends with links back to the lesson's animated figures, quiz, and runnable code. Build locally with python3 scripts/build_book.py (pandoc required); pipeline details in book/README.md.

░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒

Every lesson ships something

Other curricula end with "congratulations, you learned X." Each lesson here ends with a reusable tool you can install or paste into your daily workflow.

FIG_001.A prompts
FIG_001 · A
PROMPTS
FIG_001.B skills
FIG_001 · B
SKILLS
FIG_001.C agents
FIG_001 · C
AGENTS
FIG_001.D MCP servers
FIG_001 · D
MCP SERVERS
Paste into any AI assistant for expert-level help on a narrow task. Drop into Claude, Cursor, Codex, OpenClaw, Hermes, or any agent that reads SKILL.md. Deploy as autonomous workers — you wrote the loop yourself in Phase 14. Plug into any MCP-compatible client. Built end-to-end in Phase 13.

Install the lot with python3 scripts/install_skills.py <target>. Real tools, not homework. By the end of the curriculum, you have a portfolio of 523 artifacts you actually understand because you built them.

FIG_002 · A worked sample

Phase 14, lesson 1: the agent loop. ~120 lines of pure Python, no dependencies.

code/agent_loop.py   build it

def run(query, tools):
    history = [user(query)]
    for step in range(MAX_STEPS):
        msg = llm(history)
        if msg.tool_calls:
            for call in msg.tool_calls:
                result = tools[call.name](**call.args)
                history.append(tool_result(call.id, result))
            continue
        return msg.content
    raise StepLimitExceeded

outputs/skill-agent-loop.md   ship it

---
name: agent-loop
description: ReAct-style loop for any tool list
phase: 14
lesson: 01
---

Implement a minimal agent loop that...

outputs/prompt-debug-agent.md

You are an agent debugger. Given the trace
of an agent run, identify the step where
the agent went wrong and explain why...
░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒

Contents

Twenty phases. Click any phase to expand its lesson list.

Phase 0: Setup & Tooling 12 lessons

Get your environment ready for everything that follows.

# Lesson Type Lang
01 Dev Environment Build Python
02 Git & Collaboration Learn —
03 GPU Setup & Cloud Build Python
04 APIs & Keys Build Python
05 Jupyter Notebooks Build Python
06 Python Environments Build Shell
07 Docker for AI Build Docker
08 Editor Setup Build —
09 Data Management Build Python
10 Terminal & Shell Learn —
11 Linux for AI Learn —
12 Debugging & Profiling Build Python
Phase 1 — Math Foundations  22 lessons  The intuition behind every AI algorithm, through code.
# Lesson Type Lang
01 Linear Algebra Intuition Learn Python, Julia
02 Vectors, Matrices & Operations Build Python, Julia
03 Matrix Transformations & Eigenvalues Build Python, Julia
04 Calculus for ML: Derivatives & Gradients Learn Python
05 Chain Rule & Automatic Differentiation Build Python
06 Probability & Distributions Learn Python
07 Bayes' Theorem & Statistical Thinking Build Python
08 Optimization: Gradient Descent Family Build Python
09 Information Theory: Entropy, KL Divergence Learn Python
10 Dimensionality Reduction: PCA, t-SNE, UMAP Build Python
11 Singular Value Decomposition Build Python, Julia
12 Tensor Operations Build Python
13 Numerical Stability Build Python
14 Norms & Distances Build Python
15 Statistics for ML Build Python
16 Sampling Methods Build Python
17 Linear Systems Build Python
18 Convex Optimization Build Python
19 Complex Numbers for AI Learn Python
20 The Fourier Transform Build Python
21 Graph Theory for ML Build Python
22 Stochastic Processes Learn Python
Phase 2 — ML Fundamentals  18 lessons  Classical ML — still the backbone of most production AI.
# Lesson Type Lang
01 What Is Machine Learning Learn Python
02 Linear Regression from Scratch Build Python
03 Logistic Regression & Classification Build Python
04 Decision Trees & Random Forests Build Python
05 Support Vector Machines Build Python
06 KNN & Distance Metrics Build Python
07 Unsupervised Learning: K-Means, DBSCAN Build Python
08 Feature Engineering & Selection Build Python
09 Model Evaluation: Metrics, Cross-Validation Build Python
10 Bias, Variance & the Learning Curve Learn Python
11 Ensemble Methods: Boosting, Bagging, Stacking Build Python
12 Hyperparameter Tuning Build Python
13 ML Pipelines & Experiment Tracking Build Python
14 Naive Bayes Build Python
15 Time Series Fundamentals Build Python
16 Anomaly Detection Build Python
17 Handling Imbalanced Data Build Python
18 Feature Selection Build Python
Phase 3 — Deep Learning Core  13 lessons  Neural networks from first principles. No frameworks until you build one.
# Lesson Type Lang
01 The Perceptron: Where It All Started Build Python
02 Multi-Layer Networks & Forward Pass Build Python
03 Backpropagation from Scratch Build Python
04 Activation Functions: ReLU, Sigmoid, GELU & Why Build Python
05 Loss Functions: MSE, Cross-Entropy, Contrastive Build Python
06 Optimizers: SGD, Momentum, Adam, AdamW Build Python
07 Regularization: Dropout, Weight Decay, BatchNorm Build Python
08 Weight Initialization & Training Stability Build Python
09 Learning Rate Schedules & Warmup Build Python
10 Build Your Own Mini Framework Build Python
11 Introduction to PyTorch Build Python
12 Introduction to JAX Build Python
13 Debugging Neural Networks Build Python
Phase 4 — Computer Vision  28 lessons  From pixels to understanding — image, video, 3D, VLMs, and world models.
# Lesson Type Lang
01 Image Fundamentals: Pixels, Channels, Color Spaces Learn Python
02 Convolutions from Scratch Build Python
03 CNNs: LeNet to ResNet Build Python
04 Image Classification Build Python
05 Transfer Learning & Fine-Tuning Build Python
06 Object Detection — YOLO from Scratch Build Python
07 Semantic Segmentation — U-Net Build Python
08 Instance Segmentation — Mask R-CNN Build Python
09 Image Generation — GANs Build Python
10 Image Generation — Diffusion Models Build Python
11 Stable Diffusion — Architecture & Fine-Tuning Build Python
12 Video Understanding — Temporal Modeling Build Python
13 3D Vision: Point Clouds, NeRFs Build Python
14 Vision Transformers (ViT) Build Python
15 Real-Time Vision: Edge Deployment Build Python
16 Build a Complete Vision Pipeline Build Python
17 Self-Supervised Vision — SimCLR, DINO, MAE Build Python
18 Open-Vocabulary Vision — CLIP Build Python
19 OCR & Document Understanding Build Python
20 Image Retrieval & Metric Learning Build Python
21 Keypoint Detection & Pose Estimation Build Python
22 3D Gaussian Splatting from Scratch Build Python
23 Diffusion Transformers & Rectified Flow Build Python
24 SAM 3 & Open-Vocabulary Segmentation Build Python
25 Vision-Language Models (ViT-MLP-LLM) Build Python
26 Monocular Depth & Geometry Estimation Build Python
27 Multi-Object Tracking & Video Memory Build Python
28 World Models & Video Diffusion Build Python
Phase 5 — NLP: Foundations to Advanced  29 lessons  Language is the interface to intelligence.
# Lesson Type Lang
01 Text Processing: Tokenization, Stemming, Lemmatization Build Python
02 Bag of Words, TF-IDF & Text Representation Build Python
03 Word Embeddings: Word2Vec from Scratch Build Python
04 GloVe, FastText & Subword Embeddings Build Python
05 Sentiment Analysis Build Python
06 Named Entity Recognition (NER) Build Python
07 POS Tagging & Syntactic Parsing Build Python
08 Text Classification — CNNs & RNNs for Text Build Python
09 Sequence-to-Sequence Models Build Python
10 Attention Mechanism — The Breakthrough Build Python
11 Machine Translation Build Python
12 Text Summarization Build Python
13 Question Answering Systems Build Python
14 Information Retrieval & Search Build Python
15 Topic Modeling: LDA, BERTopic Build Python
16 Text Generation Build Python
17 Chatbots: Rule-Based to Neural Build Python
18 Multilingual NLP Build Python
19 Subword Tokenization: BPE, WordPiece, Unigram, SentencePiece Learn Python
20 Structured Outputs & Constrained Decoding Build Python
21 NLI & Textual Entailment Learn Python
22 Embedding Models Deep Dive Learn Python
23 Chunking Strategies for RAG Build Python
24 Coreference Resolution Learn Python
25 Entity Linking & Disambiguation Build Python
26 Relation Extraction & Knowledge Graph Construction Build Python
27 LLM Evaluation: RAGAS, DeepEval, G-Eval Build Python
28 Long-Context Evaluation: NIAH, RULER, LongBench, MRCR Learn Python
29 Dialogue State Tracking Build Python
Phase 6 — Speech & Audio  17 lessons  Hear, understand, speak.
# Lesson Type Lang
01 Audio Fundamentals: Waveforms, Sampling, FFT Learn Python
02 Spectrograms, Mel Scale & Audio Features Build Python
03 Audio Classification Build Python
04 Speech Recognition (ASR) Build Python
05 Whisper: Architecture & Fine-Tuning Build Python
06 Speaker Recognition & Verification Build Python
07 Text-to-Speech (TTS) Build Python
08 Voice Cloning & Voice Conversion Build Python
09 Music Generation Build Python
10 Audio-Language Models Build Python
11 Real-Time Audio Processing Build Python
12 Build a Voice Assistant Pipeline Build Python
13 Neural Audio Codecs — EnCodec, SNAC, Mimi, DAC Learn Python
14 Voice Activity Detection & Turn-Taking Build Python
15 Streaming Speech-to-Speech — Moshi, Hibiki Learn Python
16 Voice Anti-Spoofing & Audio Watermarking Build Python
17 Audio Evaluation — WER, MOS, MMAU, Leaderboards Learn Python
Phase 7 — Transformers Deep Dive  16 lessons  The architecture that changed everything.
# Lesson Type Lang
01 Why Transformers: The Problems with RNNs Learn Python
02 Self-Attention from Scratch Build Python
03 Multi-Head Attention Build Python
04 Positional Encoding: Sinusoidal, RoPE, ALiBi Build Python
05 The Full Transformer: Encoder + Decoder Build Python
06 BERT — Masked Language Modeling Build Python
07 GPT — Causal Language Modeling Build Python
08 T5, BART — Encoder-Decoder Models Learn Python
09 Vision Transformers (ViT) Build Python
10 Audio Transformers — Whisper Architecture Learn Python
11 Mixture of Experts (MoE) Build Python
12 KV Cache, Flash Attention & Inference Optimization Build Python
13 Scaling Laws Learn Python
14 Build a Transformer from Scratch Build Python
15 Attention Variants — Sliding Window, Sparse, Differential Build Python
16 Speculative Decoding — Draft, Verify, Repeat Build Python
Phase 8 — Generative AI  15 lessons  Create images, video, audio, 3D, and more.
# Lesson Type Lang
01 Generative Models: Taxonomy & History Learn Python
02 Autoencoders & VAE Build Python
03 GANs: Generator vs Discriminator Build Python
04 Conditional GANs & Pix2Pix Build Python
05 StyleGAN Build Python
06 Diffusion Models — DDPM from Scratch Build Python
07 Latent Diffusion & Stable Diffusion Build Python
08 ControlNet, LoRA & Conditioning Build Python
09 Inpainting, Outpainting & Editing Build Python
10 Video Generation Build Python
11 Audio Generation Build Python
12 3D Generation Build Python
13 Flow Matching & Rectified Flows Build Python
14 Evaluation: FID, CLIP Score Build Python
19 Visual Autoregressive Modeling (VAR): Next-Scale Prediction Build Python
Phase 9 — Reinforcement Learning  12 lessons  The foundation of RLHF and game-playing AI.
# Lesson Type Lang
01 MDPs, States, Actions & Rewards Learn Python
02 Dynamic Programming Build Python
03 Monte Carlo Methods Build Python
04 Q-Learning, SARSA Build Python
05 Deep Q-Networks (DQN) Build Python
06 Policy Gradients — REINFORCE Build Python
07 Actor-Critic — A2C, A3C Build Python
08 PPO Build Python
09 Reward Modeling & RLHF Build Python
10 Multi-Agent RL Build Python
11 Sim-to-Real Transfer Build Python
12 RL for Games Build Python
Phase 10 — LLMs from Scratch  24 lessons  Build, train, and understand large language models.
# Lesson Type Lang
01 Tokenizers: BPE, WordPiece, SentencePiece Build Python, Rust
02 Building a Tokenizer from Scratch Build Python
03 Data Pipelines for Pre-Training Build Python
04 Pre-Training a Mini GPT (124M) Build Python
05 Distributed Training, FSDP, DeepSpeed Build Python
06 Instruction Tuning — SFT Build Python
07 RLHF — Reward Model + PPO Build Python
08 DPO — Direct Preference Optimization Build Python
09 Constitutional AI & Self-Improvement Build Python
10 Evaluation — Benchmarks, Evals Build Python
11 Quantization: INT8, GPTQ, AWQ, GGUF Build Python
12 Inference Optimization Build Python
13 Building a Complete LLM Pipeline Build Python
14 Open Models: Architecture Walkthroughs Learn Python
15 Speculative Decoding and EAGLE-3 Build Python
16 Differential Attention (V2) Build Python
17 Native Sparse Attention (DeepSeek NSA) Build Python
18 Multi-Token Prediction (MTP) Build Python
19 DualPipe Parallelism Learn Python
20 DeepSeek-V3 Architecture Walkthrough Learn Python
21 Jamba — Hybrid SSM-Transformer Learn Python
22 Async and Hogwild! Inference Build Python
25 Speculative Decoding and EAGLE Build Python
34 Gradient Checkpointing and Activation Recomputation Build Python
Phase 11 — LLM Engineering  17 lessons  Put LLMs to work in production.
# Lesson Type Lang
01 Prompt Engineering: Techniques & Patterns Build Python
02 Few-Shot, CoT, Tree-of-Thought Build Python
03 Structured Outputs Build Python
04 Embeddings & Vector Representations Build Python
05 Context Engineering Build Python
06 RAG: Retrieval-Augmented Generation Build Python
07 Advanced RAG: Chunking, Reranking Build Python
08 Fine-Tuning with LoRA & QLoRA Build Python
09 Function Calling & Tool Use Build Python
10 Evaluation & Testing Build Python
11 Caching, Rate Limiting & Cost Build Python
12 Guardrails & Safety Build Python
13 Building a Production LLM App Build Python
14 Model Context Protocol (MCP) Build Python
15 Prompt Caching & Context Caching Build Python
16 Agent State Machines — Graphs, Nodes, Checkpoints Build Python
17 Agent Framework Tradeoffs Learn Python
Phase 12 — Multimodal AI  25 lessons  See, hear, read, and reason across modalities — from ViT patches to computer-use agents.
# Lesson Type Lang
01 Vision Transformers and the Patch-Token Primitive Learn Python
02 CLIP and Contrastive Vision-Language Pretraining Build Python
03 BLIP-2 Q-Former as Modality Bridge Build Python
04 Flamingo and Gated Cross-Attention Learn Python
05 LLaVA and Visual Instruction Tuning Build Python
06 Any-Resolution Vision — Patch-n'-Pack and NaFlex Build Python
07 Open-Weight VLM Recipes: What Actually Matters Learn Python
08 LLaVA-OneVision: Single, Multi, Video Build Python
09 Qwen-VL Family and Dynamic-FPS Video Learn Python
10 InternVL3 Native Multimodal Pretraining Learn Python
11 Chameleon Early-Fusion Token-Only Build Python
12 Emu3 Next-Token Prediction for Generation Learn Python
13 Transfusion Autoregressive + Diffusion Build Python
14 Show-o Discrete-Diffusion Unified Learn Python
15 Janus-Pro Decoupled Encoders Build Python
16 MIO Any-to-Any Streaming Learn Python
17 Video-Language Temporal Grounding Build Python
18 Long-Video at Million-Token Context Build Python
19 Audio-Language Models: Whisper to AF3 Build Python
20 Omni Models: Thinker-Talker Streaming Build Python
21 Embodied VLAs: RT-2, OpenVLA, π0, GR00T Learn Python
22 Document and Diagram Understanding Build Python
23 ColPali Vision-Native Document RAG Build Python
24 Multimodal RAG and Cross-Modal Retrieval Build Python
25 Multimodal Agents and Computer-Use (Capstone) Build Python
Phase 13 — Tools & Protocols  31 lessons  The interfaces between AI and the real world.
# Lesson Type Lang
01 The Tool Interface Learn Python
02 Function Calling Deep Dive Build Python
03 Parallel and Streaming Tool Calls Build Python
04 Structured Output Build Python
05 Tool Schema Design Learn Python
06 MCP Fundamentals: Stateless Requests and JSON-RPC Learn Python
07 Building an MCP Server: Stateless Python and TypeScript Build Python, TypeScript
08 Building an MCP Client: Discovery, Routing, and Dual-Era Fallback Build Python
09 MCP Transports: stdio and Stateless Streamable HTTP Learn Python
10 MCP Resources and Prompts: Addressable Context for Stateless Servers Build Python
11 MCP Model Input: Sampling Migration and Stateless MRTR Build Python
12 Explicit Scope and Stateless Elicitation Build Python
13 MCP Tasks Extension: Durable Work on a Stateless Core Build Python
14 MCP Apps on the Stateless Protocol Build Python
15 MCP Security: Poisoned Metadata, Routing, and MRTR State Learn Python
16 MCP Authorization: CIMD, Issuer Binding, PKCE, and Step-Up Build Python
17 Stateless MCP Gateways and Registry Admission Learn Python
18 MCP Auth in Production: Issuer-Bound Enrollment and Tokens Build Python
19 A2A Protocol Build Python
20 OpenTelemetry GenAI Build Python
21 LLM Routing Layer Learn Python
22 Agent Skills: Portable Contract and Runtime Boundary Build Python
23 Capstone: Stateless Tool Ecosystem Build Python
24 Skill Discovery and Progressive Disclosure Build Python
25 Skill Invocation and Routing Build Python
26 Skill Permissions, Sandboxes, and Trust Build Python
27 Skill Evals, Packaging, and Portability Build Python
28 MCP Tool Contracts and Content Build Python
29 MCP Reliability, Cancellation, and Flow Control Build Python
30 MCP Registry Supply Chain: Admission, Drift, and Rollback Build Python
31 MCP Conformance Engineering: Versioning, Evidence, and Operations Build Python

Lessons 06-18 and 28-31 form the focused Model Context Protocol (MCP) path. Its manifest order is 06, 07, 08, 09, 10, 11, 12, 13, 14, 15, 16, 18, 17, 28, 29, 30, 31. Start it with the host-specific learn-mcp invocation above. Lesson 23 is its only optional capstone and also requires Lessons 19 and 20.

Lessons 22 and 24-27 form the focused Agent Skills learning path, from package contract through real-host release gates. Start it with the host-specific learn-agent-skills invocation shown above; do not follow numeric next navigation from 22 to 23.

Phase 14 — Agent Engineering  54 lessons  Build agents from first principles, use coding agents reliably, and shape the work before implementation.
# Lesson Type Lang
01 The Agent Loop Build Python
02 ReWOO and Plan-and-Execute Build Python
03 Reflexion and Verbal Reinforcement Learning Build Python
04 Tree of Thoughts and LATS Build Python
05 Self-Refine and CRITIC Build Python
06 Tool Use and Function Calling Build Python
07 Agent Memory — Virtual Context and Memory Paging Build Python
08 Memory Blocks and Sleep-Time Compute Build Python
09 Hybrid Memory — Vector + Graph + KV Build Python
10 Skill Libraries and Lifelong Learning (Voyager) Build Python
11 Planning with HTN and Evolutionary Search Build Python
12 Anthropic's Workflow Patterns Build Python
13 Stateful Graph Orchestration — Durable Execution and Checkpoints Build Python
14 The Actor Model for Agents Build Python
15 Role-Based Agent Teams — Roles, Tasks, Processes Build Python
16 OpenAI Agents SDK — Handoffs, Guardrails, Tracing Build Python
17 The Harness as a Library — Subagents and Session Store Build Python
18 Production Agent Runtimes Learn Python
19 Benchmarks — SWE-bench, GAIA, AgentBench Learn Python
20 Benchmarks — WebArena and OSWorld Learn Python
21 Computer Use — Claude, OpenAI CUA, Gemini Build Python
22 Voice Agents — Pipecat and LiveKit Build Python
23 OpenTelemetry GenAI Semantic Conventions Build Python
24 Agent Observability — Langfuse, Phoenix, Opik Learn Python
25 Multi-Agent Debate and Collaboration Build Python
26 Failure Modes — Why Agents Break Build Python
27 Prompt Injection and the PVE Defense Build Python
28 Orchestration Patterns — Supervisor, Swarm, Hierarchical Build Python
29 Production Runtimes — Queue, Event, Cron Learn Python
30 Eval-Driven Agent Development Build Python
31 Agent Workbench: Why Capable Models Still Fail Learn Python
32 The Minimal Agent Workbench Build Python
33 Agent Instructions as Executable Constraints Build Python
34 Repo Memory and Durable State Build Python
35 Initialization Scripts for Agents Build Python
36 Scope Contracts and Task Boundaries Build Python
37 Runtime Feedback Loops Build Python
38 Verification Gates Build Python
39 Reviewer Agent: Separate Builder from Marker Build Python
40 Multi-Session Handoff Build Python
41 The Workbench on a Real Repo Build Python
42 Capstone: Ship a Reusable Agent Workbench Pack Build Python
43 Frame the Task Before the Agent Writes Code Build Python
44 Build an Evidence-Backed Execution Plan Build Python
45 Delegate Agent Work with Isolation and Merge Contracts Build Python
46 Turn Every Agent Correction into a System Improvement Build Python
47 Define the Outcome Before You Choose the Output Build Python
48 Discover the Workflow People Actually Perform Build Python
49 Map Assumptions and Resolve the Riskiest One First Build Python
50 Choose the Smallest Slice That Can Change the Decision Build Python
51 Write Specifications That Preserve Judgment Build Python
52 Design Success Metrics Before the Result Exists Build Python
53 Choose Prototype, Pilot, or Production Deliberately Build Python
54 Build a Feedback Ratchet with Ownership and Retirement Build Python

Each Phase 14 workbench lesson (31-42) ships a mission.md briefing the agent before it opens the full lesson docs.

Lessons 31-46 form the Agent-Assisted Engineering path. Its manifest order combines the workbench foundation with task framing, planning, delegation, and durable feedback. Lessons 47-54 form the Product Judgment and Delivery path, from outcome framing through evidence, risk, scope, measurement, staged release, and feedback ownership.

Phase 15 — Autonomous Systems  22 lessons  Long-horizon agents, self-improvement, and the 2026 safety stack.
# Lesson Type Lang
01 From Chatbots to Long-Horizon Agents (METR) Learn Python
02 STaR, V-STaR, Quiet-STaR: Self-Taught Reasoning Learn Python
03 AlphaEvolve: Evolutionary Coding Agents Learn Python
04 Darwin Gödel Machine: Self-Modifying Agents Learn Python
05 AI Scientist v2: Workshop-Level Research Learn Python
06 Automated Alignment Research (Anthropic AAR) Learn Python
07 Recursive Self-Improvement: Capability vs Alignment Learn Python
08 Bounded Self-Improvement Designs Learn Python
09 Autonomous Coding Agent Landscape (SWE-bench, CodeAct) Learn Python
10 Permission Modes for Autonomous Agents Learn Python
11 Browser Agents and Indirect Prompt Injection Learn Python
12 Durable Execution for Long-Running Agents Learn Python
13 Action Budgets, Iteration Caps, Cost Governors Learn Python
14 Kill Switches, Circuit Breakers, Canary Tokens Learn Python
15 HITL: Propose-Then-Commit Learn Python
16 Checkpoints and Rollback Learn Python
17 Constitutional AI and Rule Overrides Learn Python
18 Llama Guard and Input/Output Classification Learn Python
19 Anthropic Responsible Scaling Policy v3.0 Learn Python
20 OpenAI Preparedness Framework and DeepMind FSF Learn Python
21 METR Time Horizons and External Evaluation Learn Python
22 CAIS, CAISI, and Societal-Scale Risk Learn Python
Phase 16 — Multi-Agent & Swarms  25 lessons  Coordination, emergence, and collective intelligence.
# Lesson Type Lang
01 Why Multi-Agent Learn TypeScript
02 FIPA-ACL Heritage and Speech Acts Learn Python
03 Communication Protocols Build TypeScript
04 The Multi-Agent Primitive Model Learn Python
05 Supervisor / Orchestrator-Worker Pattern Build Python
06 Hierarchical Architecture and Decomposition Drift Learn Python
07 Society of Mind and Multi-Agent Debate Build Python
08 Role Specialization — Planner / Critic / Executor / Verifier Build Python
09 Parallel Swarm and Networked Architectures Build Python
10 Group Chat and Speaker Selection Build Python
11 Handoffs and Routines (Stateless Orchestration) Build Python
12 A2A — The Agent-to-Agent Protocol Build Python
13 Shared Memory and Blackboard Patterns Build Python
14 Consensus and Byzantine Fault Tolerance Build Python
15 Voting, Self-Consistency, and Debate Topology Build Python
16 Negotiation and Bargaining Build Python
17 Generative Agents and Emergent Simulation Build Python
18 Theory of Mind and Emergent Coordination Build Python
19 Swarm Optimization (PSO, ACO) Build Python
20 MARL — MADDPG, QMIX, MAPPO Learn Python
21 Agent Economies, Token Incentives, Reputation Learn Python
22 Production Scaling — Queues, Checkpoints, Durability Build Python
23 Failure Modes — MAST, Groupthink, Monoculture Learn Python
24 Evaluation and Coordination Benchmarks Learn Python
25 Case Studies and 2026 State of the Art Learn Python
Phase 17 — Infrastructure & Production  28 lessons  Ship AI to the real world.
# Lesson Type Lang
01 Managed LLM Platforms — Bedrock, Azure OpenAI, Vertex AI Learn Python
02 Inference Platform Economics — Fireworks, Together, Baseten, Modal Learn Python
03 GPU Autoscaling on Kubernetes — Karpenter, KAI Scheduler Learn Python
04 Serving Engine Internals — PagedAttention, Continuous Batching, Chunked Prefill Learn Python
05 EAGLE-3 Speculative Decoding in Production Learn Python
06 Prefix-Cache Serving — RadixAttention and KV Reuse Learn Python
07 Hardware-Specialized Inference Compilation — FP8 and NVFP4 on Blackwell Learn Python
08 Inference Metrics — TTFT, TPOT, ITL, Goodput, P99 Learn Python
09 Production Quantization — AWQ, GPTQ, GGUF, FP8, NVFP4 Learn Python
10 Cold Start Mitigation for Serverless LLMs Learn Python
11 Multi-Region LLM Serving and KV Cache Locality Learn Python
12 Edge Inference — ANE, Hexagon, WebGPU, Jetson Learn Python
13 LLM Observability Stack Selection Learn Python
14 Prompt Caching and Semantic Caching Economics Learn Python
15 Batch APIs — the 50% Discount as Industry Standard Learn Python
16 Model Routing as a Cost-Reduction Primitive Learn Python
17 Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d Learn Python
18 Production Serving Stack — KV Offloading and Cache-Aware Routing Learn Python
19 AI Gateways — LiteLLM, Portkey, Kong, Bifrost Learn Python
20 Shadow, Canary, and Progressive Deployment Learn Python
21 A/B Testing LLM Features — GrowthBook and Statsig Learn Python
22 Load Testing LLM APIs — k6, LLMPerf, GenAI-Perf Build Python
23 SRE for AI — Multi-Agent Incident Response Learn Python
24 Chaos Engineering for LLM Production Learn Python
25 Security — Secrets, PII Scrubbing, Audit Logs Learn Python
26 Compliance — SOC 2, HIPAA, GDPR, EU AI Act, ISO 42001 Learn Python
27 FinOps for LLMs — Unit Economics and Multi-Tenant Attribution Learn Python
28 Self-Hosted Serving Selection — Matching Engine to Hardware and Scale Learn Python
Phase 18 — Ethics, Safety & Alignment  30 lessons  Build AI that helps humanity. Not optional.
# Lesson Type Lang
01 Instruction-Following as Alignment Signal Learn Python
02 Reward Hacking & Goodhart's Law Learn Python
03 Direct Preference Optimization Family Learn Python
04 Sycophancy as RLHF Amplification Learn Python
05 Constitutional AI & RLAIF Learn Python
06 Mesa-Optimization & Deceptive Alignment Learn Python
07 Sleeper Agents — Persistent Deception Learn Python
08 In-Context Scheming in Frontier Models Learn Python
09 Alignment Faking Learn Python
10 AI Control — Safety Despite Subversion Learn Python
11 Scalable Oversight & Weak-to-Strong Learn Python
12 Red-Teaming: PAIR & Automated Attacks Build Python
13 Many-Shot Jailbreaking Learn Python
14 ASCII Art & Visual Jailbreaks Build Python
15 Indirect Prompt Injection Build Python
16 Red-Team Tooling: Garak, Llama Guard, PyRIT Build Python
17 WMDP & Dual-Use Capability Evaluation Learn Python
18 Frontier Safety Frameworks — RSP, PF, FSF Learn Python
19 Model Welfare Research Learn Python
20 Bias & Representational Harm Build Python
21 Fairness Criteria: Group, Individual, Counterfactual Learn Python
22 Differential Privacy for LLMs Build Python
23 Watermarking: SynthID, Stable Signature, C2PA Build Python
24 Regulatory Frameworks: EU, US, UK, Korea Learn Python
25 EchoLeak & CVEs for AI Learn Python
26 Model, System & Dataset Cards Build Python
27 Data Provenance & Training-Data Governance Learn Python
28 Alignment Research Ecosystem: MATS, Redwood, Apollo, METR Learn Python
29 Moderation Systems: OpenAI, Perspective, Llama Guard Build Python
30 Dual-Use Risk: Cyber, Bio, Chem, Nuclear Learn Python
Phase 19 — Capstone Projects  85 lessons  17 end-to-end products + 9 deep-build tracks. 20-40 hours per project; 4-12 lessons per track.
# Project Combines Lang
01 Terminal-Native Coding Agent P0 P5 P7 P10 P11 P13 P14 P15 P17 P18 Python
02 RAG over Codebase (Cross-Repo Semantic Search) P5 P7 P11 P13 P17 Python
03 Real-Time Voice Assistant (ASR → LLM → TTS) P6 P7 P11 P13 P14 P17 Python
04 Multimodal Document QA (Vision-First) P4 P5 P7 P11 P12 P17 Python
05 Autonomous Research Agent (AI-Scientist Class) P0 P2 P3 P7 P10 P14 P15 P16 P18 Python
06 DevOps Troubleshooting Agent for Kubernetes P11 P13 P14 P15 P17 P18 Python
07 End-to-End Fine-Tuning Pipeline P2 P3 P7 P10 P11 P17 P18 Python
08 Production RAG Chatbot (Regulated Vertical) P5 P7 P11 P12 P17 P18 Python
09 Code Migration Agent (Repo-Level Upgrade) P5 P7 P11 P13 P14 P15 P17 Python
10 Multi-Agent Software Engineering Team P11 P13 P14 P15 P16 P17 Python
11 LLM Observability & Eval Dashboard P11 P13 P17 P18 Python
12 Video Understanding Pipeline (Scene → QA) P4 P6 P7 P11 P12 P17 Python
13 Stateless MCP Server with Registry and Governance P11 P13 P14 P17 P18 Python
14 Speculative-Decoding Inference Server P3 P7 P10 P17 Python
15 Constitutional Safety Harness + Red-Team Range P10 P11 P13 P14 P18 Python
16 GitHub Issue-to-PR Autonomous Agent P11 P13 P14 P15 P17 Python
17 Personal AI Tutor (Adaptive, Multimodal) P5 P6 P11 P12 P14 P17 P18 Python

Deep-build tracks — multi-lesson series that build a complete subsystem from scratch.

# Project Combines Lang
20 Agent Harness Loop Contract A. Agent harness Python
21 Tool Registry with Schema Validation A. Agent harness Python
22 JSON-RPC 2.0 Over Newline-Delimited Stdio A. Agent harness Python
23 Function Call Dispatcher A. Agent harness Python
24 Plan-Execute Control Flow A. Agent harness Python
25 Verification Gates and Observation Budget A. Agent harness Python
26 Sandbox Runner with Denylist and Path Jail A. Agent harness Python
27 Eval Harness with Fixture Tasks A. Agent harness Python
28 Observability with OTel GenAI Spans and Prometheus Metrics A. Agent harness Python
29 End-to-End Coding Agent on the Harness A. Agent harness Python
30 BPE Tokenizer From Scratch B. NLP LLM Python
31 Tokenized Dataset with Sliding Window B. NLP LLM Python
32 Token and Positional Embeddings B. NLP LLM Python
33 Multi-Head Self-Attention B. NLP LLM Python
34 Transformer Block from Scratch B. NLP LLM Python
35 GPT Model Assembly B. NLP LLM Python
36 Training Loop and Evaluation B. NLP LLM Python
37 Loading Pretrained Weights B. NLP LLM Python
38 Classifier Fine-Tuning by Head Swap B. NLP LLM Python
39 Instruction Tuning by Supervised Fine-Tuning B. NLP LLM Python
40 Direct Preference Optimization from Scratch B. NLP LLM Python
41 Full Evaluation Pipeline B. NLP LLM Python
42 Large Corpus Downloader C. Train end-to-end Python
43 HDF5 Tokenized Corpus C. Train end-to-end Python
44 Cosine LR with Linear Warmup C. Train end-to-end Python
45 Gradient Clipping and Mixed Precision C. Train end-to-end Python
46 Gradient Accumulation C. Train end-to-end Python
47 Checkpoint Save and Resume C. Train end-to-end Python
48 Distributed Data Parallel and FSDP from Scratch C. Train end-to-end Python
49 Language Model Evaluation Harness C. Train end-to-end Python
50 Hypothesis Generator D. Auto research Python
51 Literature Retrieval D. Auto research Python
52 Experiment Runner D. Auto research Python
53 Result Evaluator D. Auto research Python
54 Paper Writer D. Auto research Python
55 Critic Loop D. Auto research Python
56 Iteration Scheduler D. Auto research Python
57 End-to-End Research Demo D. Auto research Python
58 Vision Encoder Patches E. Multimodal VLM Python
59 Vision Transformer Encoder E. Multimodal VLM Python
60 Projection Layer for Modality Alignment E. Multimodal VLM Python
61 Cross-Attention Fusion E. Multimodal VLM Python
62 Vision-Language Pretraining E. Multimodal VLM Python
63 Multimodal Evaluation E. Multimodal VLM Python
64 Chunking Strategies, Compared F. Advanced RAG Python
65 Hybrid Retrieval with BM25 and Dense Embeddings F. Advanced RAG Python
66 Cross-Encoder Reranker F. Advanced RAG Python
67 Query Rewriting: HyDE, Multi-Query, and Decomposition F. Advanced RAG Python
68 RAG Evaluation: Precision, Recall, MRR, nDCG, Faithfulness, Answer Relevance F. Advanced RAG Python
69 End-to-End RAG System F. Advanced RAG Python
70 Task Spec Format G. Eval framework Python
71 Classical Metrics G. Eval framework Python
72 Code Exec Metric G. Eval framework Python
73 Perplexity and Calibration G. Eval framework Python
74 Leaderboard Aggregation G. Eval framework Python
75 End-to-End Eval Runner G. Eval framework Python
76 Collective Ops From Scratch H. Distributed train Python
77 Data Parallel DDP From Scratch H. Distributed train Python
78 ZeRO Optimizer State Sharding H. Distributed train Python
79 Pipeline Parallel and Bubble Analysis H. Distributed train Python
80 Sharded Checkpoint and Atomic Resume H. Distributed train Python
81 End-to-End Distributed Training H. Distributed train Python
82 Jailbreak Taxonomy I. Safety harness Python
83 Prompt Injection Detector I. Safety harness Python
84 Refusal Evaluation I. Safety harness Python
85 Content Classifier Integration I. Safety harness Python
86 Constitutional Rules Engine I. Safety harness Python, YAML
87 End-to-End Safety Gate I. Safety harness Python
░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒

The toolkit

Every lesson produces a reusable artifact. By the end you have:

outputs/
├── prompts/      prompt templates for every AI task
└── skills/       SKILL.md files for AI coding agents

Plug them into Claude, Cursor, Codex, OpenClaw, Hermes, or any agent that reads a SKILL.md / AGENTS.md directory. Real tools, not homework.

Install course skills into your agent

Two skill sets, two installers:

The learning skills (start-learning, learn, course-guide, learn-mcp, learn-agent-skills, claude-certification, mcpa-certification, find-your-level, and check-understanding) live under skills/ and install into a supported skill-capable host with one command. Installation needs Node.js and npx, but not a repository clone or Python:

npx skills add rohitg00/ai-engineering-from-scratch

skills writes to the host and scope selected during installation, such as .claude/skills/, .cursor/skills/, .codex/skills/, or another supported skills folder. Verify that the selected host discovers that exact destination.

The lesson artifacts. The repo ships 396 skills and 99 prompts under phases/**/outputs/; install them via scripts/install_skills.py. Requires cloning the repo. Supports tag filters, dry-runs, and per-agent layouts:

python3 scripts/install_skills.py <target>                                 # every skill, default --layout skills (nested)
python3 scripts/install_skills.py <target> --layout skills                 # same as above, explicit
python3 scripts/install_skills.py <target> --type all                      # skills + prompts + agents
python3 scripts/install_skills.py <target> --phase 14                      # one phase only
python3 scripts/install_skills.py <target> --tag rag                       # filter by tag
python3 scripts/install_skills.py <target> --layout flat                   # flat files
python3 scripts/install_skills.py <target> --dry-run                       # preview without writing
python3 scripts/install_skills.py <target> --force                         # overwrite existing files

<target> is the skills directory for your agent (examples: ~/.claude/skills/, ~/.cursor/skills/, ~/.config/openclaw/skills/, .skills/, or any path your agent reads).

By default the script refuses to overwrite an existing destination and exits with code 1 after listing every colliding path. Use --dry-run to preview collisions or --force to overwrite. Every non-dry-run run writes a manifest.json in the target with the full inventory grouped by type and phase. Pick the layout your agent reads:

--layout Path written
skills <target>/<name>/SKILL.md (nested convention, supported by Claude / Cursor / Codex / OpenClaw / Hermes)
by-phase <target>/phase-NN/<name>.md
flat <target>/<name>.md

Drop the agent workbench into your own repo

The Phase 14 capstone ships a reusable Agent Workbench pack (AGENTS.md, schemas, init / verify / handoff scripts). Scaffold it into any repo with:

python3 scripts/scaffold_workbench.py path/to/your-repo            # full pack + seeds
python3 scripts/scaffold_workbench.py path/to/your-repo --minimal  # skip docs/
python3 scripts/scaffold_workbench.py path/to/your-repo --dry-run  # preview only
python3 scripts/scaffold_workbench.py path/to/your-repo --force    # overwrite

You get the seven workbench surfaces wired up, a starter task_board.json, and a fresh agent_state.json at schema_version: 1. From there: edit the task, edit AGENTS.md, run scripts/init_agent.py, hand the contract to your agent. The pack source lives at phases/14-agent-engineering/42-agent-workbench-capstone/outputs/agent-workbench-pack/.

Browse the entire course as JSON

scripts/build_catalog.py walks every phase, every lesson, every artifact on disk and writes catalog.json at the repo root. One file, every course truth.

python3 scripts/build_catalog.py               # writes <repo>/catalog.json
python3 scripts/build_catalog.py --stdout      # to stdout, do not touch repo
python3 scripts/build_catalog.py --out path/to/file.json

The catalog is filesystem-derived, not README-derived, so counts always match what is actually on disk. Use it for site builds, downstream tooling, or to verify the README counts have not drifted. Schema is documented at the top of the script.

A GitHub Action (.github/workflows/curriculum.yml) rebuilds catalog.json on every PR and fails the build if the committed file is stale. After editing any lesson, run python3 scripts/build_catalog.py and commit the result, or CI will reject the PR. The same workflow runs audit_lessons.py in warn-only mode (so existing drift does not block contributors).

Smoke-check every lesson's Python code

scripts/lesson_run.py byte-compiles every .py file under each lesson's code/ directory. Default mode is syntax-check only — no execution, no API keys, no heavy ML deps required. Catches the regressions contributors introduce most often (bad indentation, broken f-strings, stray edits).

python3 scripts/lesson_run.py                  # syntax-check the whole curriculum
python3 scripts/lesson_run.py --phase 14       # one phase only
python3 scripts/lesson_run.py --json           # JSON report on stdout
python3 scripts/lesson_run.py --strict         # exit 1 if any lesson fails
python3 scripts/lesson_run.py --execute        # actually run, 10s timeout per lesson

--execute runs each lesson's code/main.py (or the first .py file) with a 10-second timeout. Lessons whose entry file starts with a # requires: pkg1, pkg2 comment listing non-stdlib deps are skipped with reason needs <deps>. The script is opt-in and not wired into CI.

Stdlib only, Python 3.10+. Set LINK_CHECK_SKIP=domain1,domain2 to override the default skip-list (twitter.com, x.com, linkedin.com, instagram.com, medium.com — domains that aggressively block automated HEAD/GET).

Where to start

Background Start at Estimated time
New to programming and AI Phase 0 — Setup ~306 hours
Know Python, new to ML Phase 1 — Math Foundations ~270 hours
Know ML, new to deep learning Phase 3 — Deep Learning Core ~200 hours
Know deep learning, want LLMs and agents Phase 10 — LLMs from Scratch ~100 hours
Senior engineer, only want agent engineering Phase 14 — Agent Engineering ~60 hours
Only want to build production MCP systems Model Context Protocol (MCP) path ~23 hours 15 min
Only want to build production Agent Skills Agent Skills Engineering path ~9.5 hours
░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒

Why this matters now

FIG_003 · A
THE INDUSTRY SIGNAL
FIG_003 · B
FOUNDATIONAL PAPERS COVERED

"The hottest new programming language is English."
— Andrej Karpathy (tweet)

"Software engineering is being remade in front of our eyes."
— Boris Cherny, creator of Claude Code

"Models will keep getting better. The skill that compounds is knowing what to build."
— Industry consensus, 2026

  • Attention Is All You Need — Vaswani et al., 2017 → Phase 7
  • Language Models are Few-Shot Learners (GPT-3) → Phase 10
  • Denoising Diffusion Probabilistic Models → Phase 8
  • InstructGPT / RLHF → Phase 10
  • Direct Preference Optimization → Phase 10
  • Chain-of-Thought Prompting → Phase 11
  • ReAct: Reasoning + Acting in LLMs → Phase 14
  • Model Context Protocol — Anthropic → Phase 13
░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒

Contributing

Goal Read
Contribute a lesson or fix CONTRIBUTING.md
Fork for your team or school FORKING.md
Lesson template LESSON_TEMPLATE.md
Track progress ROADMAP.md
Glossary glossary/terms.md
Code of conduct CODE_OF_CONDUCT.md

Before submitting a lesson, run the invariant check:

python3 scripts/audit_lessons.py           # full curriculum
python3 scripts/audit_lessons.py --phase 14  # single phase
python3 scripts/audit_lessons.py --json    # CI-friendly output

Exit code is non-zero when any rule fails. Rules (L001–L010) validate directory shape, docs/en.md presence + H1, code/ non-emptiness, quiz.json schema (rejects the legacy q/choices/answer keys that caused issue #102), and relative links inside lesson docs.

░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒

Sponsor the work

Free, MIT-licensed, 523 lessons. Thank you to the sponsors and backers who make the work possible. See all sponsors and backers.

Want to support the work? See sponsorship options, including hardware sponsorships, or sponsor on GitHub.

░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒░░░▒▒▒

If this manual helped you, star the repo. It keeps the project alive.

License

MIT. Use it however you want — fork it, teach it, sell it, ship it. Attribution appreciated, not required.

Maintained by Rohit Ghumare and the community.

@ghumare64  ·  aiengineeringfromscratch.com  ·  Report / Suggest
Languages
Python 51.7%
JavaScript 31.1%
TypeScript 6%
HTML 5.4%
Rust 1.8%
Other 3.9%