Files
Jason ZhouandClaude Opus 5 d09f6b8d40 feat(seo): robots, sitemap, a crawlable catalog, and social cards
The catalog had no URLs. ~2,630 endpoints across 80 platform shelves — the whole
substance of the product — existed only as hash routes (/app#platform/<slug>)
behind a login, so a crawler could reach six thin marketing pages and nothing
else. On top of that: no robots.txt, no sitemap.xml, HEAD answering 405 on every
page, no og/twitter tags or image, no structured data, and /docs serving
FastAPI's stock Swagger shell — a kilobyte of JavaScript to anything that does
not run scripts.

Crawler plumbing:
- /robots.txt (bundled, {BASE}-templated) and /sitemap.xml (generated — 80 of
  its 88 URLs come from the catalog, lastmod from file and catalog mtimes).
- HEAD widened onto every GET route. FastAPI's APIRoute pins methods to {"GET"}
  and never adds HEAD, unlike Starlette's plain Route.
- Canonicals on /support (collapsing /contact and /help), /terms, /privacy,
  /tutorial; noindex on the dashboard and on /vendor-listing.md's duplicate URL.
- robots.txt and sitemap.xml 301 from the legacy host, so its name resolves to
  one crawlable site rather than a duplicate of it.

Crawlable surfaces:
- /catalog and /catalog/<slug>: server-rendered, no JavaScript, every capability
  and price as real text. Registered after the JSON routes, with the reserved
  names refused explicitly.
- /docs: a real API reference built from app.openapi(). Swagger UI moves to
  /docs/api and is disallowed; the public catalog routes join the schema.

Metadata:
- og/twitter cards everywhere, backed by a new 1200x630 card rendered from
  assets/brand/og-card.html. Every brand on it is a real provider.
- SoftwareApplication + Offer + Organization on the landing, ItemList +
  BreadcrumbList on the catalog pages, FAQPage over the five questions already
  written on /support.

Two things this had to be careful about. Widening HEAD put 58 duplicate
operations into the public openapi.json, so _openapi_without_head() narrows the
widened routes for the duration of schema generation. And landing/legal/tutorial
pages now read-and-substitute {BASE} instead of being served as bare files: a
hardcoded treg.to canonical tells a self-hosted registry's crawler that the real
page lives on someone else's domain.

Counts reconciled to the live catalog (2,630/47); the landing said 2,617/42 and
llms.txt said ~2,600/~48.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 21:45:10 +10:00
..