docs: make pages MDX-compatible

This commit is contained in:
Hamed Rabah
2026-08-24 17:10:59 -07:00
parent 59babb88e4
commit 7f706af22b
6 changed files with 8 additions and 8 deletions
+1 -1
View File
@@ -151,7 +151,7 @@ pair's current restack, base dev `292ed4c`):
stamp already resolves to `int8-row` since the #528 INVERSION, and a
stamp's role there is instead letting a genuinely-stamped `fmt=8` tensor
override that default — see "The metadata stamp" below for the exact
rule in both cases. FMT_NAMES table (name string <-> fmt int):
rule in both cases. FMT_NAMES table (`name string` to `fmt int`):
`c/colibri.c:1316`.
- **no ordinal** (`int4-rans256-g0`, merged tools-only tier — line numbers
at dev `7fb1159`, post-#671 merge `a3a5a75`, not at this PR pair's
+2 -2
View File
@@ -276,7 +276,7 @@ telemetry stack — hardware, scheduler, tier bar, per-turn time breakdown, tok/
trend and per-GPU expert counts:
<p align="center">
<img src="media/colibri-mobile.png" width="270" alt="the dashboard on a phone-sized viewport">
<img src="media/colibri-mobile.png" width="270" alt="the dashboard on a phone-sized viewport" />
&nbsp;&nbsp;
<img src="media/colibri-metrics.png" width="300" alt="the telemetry sidebar">
<img src="media/colibri-metrics.png" width="300" alt="the telemetry sidebar" />
</p>
+1 -1
View File
@@ -223,7 +223,7 @@ existing disk/wait numbers. `[METAL] residency-set: on` / the two fallback stder
across cap1/cap16 may legitimately differ (different dispatch composition, per the
fix-plan's "Determinism side-finding").
5. **`[METAL] residency-set: on` line present in stderr** at flag-on startup, and absent
(or the OS<15/create-failed fallback line) otherwise — cheap sanity check that a run
(or the `OS < 15`/create-failed fallback line) otherwise — cheap sanity check that a run
actually exercised the intended path before trusting its numbers. Also read the
**`METAL-RESSET: flush` line** (gate-on only): if that number is large, the deferred
set-commit cost is eating the stall win from the dispatch side.
+2 -2
View File
@@ -84,9 +84,9 @@ help in proportion to the RAM you can give them.
| Invocation | What it does |
|---|---|
| `-p "text" [-n N]` | streaming greedy generation (stops at eos or N tokens) |
| `--chat -p "text"` | wraps the prompt in Inkling's chat template (role tokens + `<|content_text|>`, `<|message_model|>` as the generation prompt). Instruct models fed raw text are out of distribution and answer badly. `THINK=<0..1>` raises the reasoning effort (default 0) |
| `--chat -p "text"` | wraps the prompt in Inkling's chat template (role tokens plus the content and model-message markers as the generation prompt). Instruct models fed raw text are out of distribution and answer badly. `THINK=0..1` raises the reasoning effort (default 0) |
| `-f prompts.txt [-n N]` | one prompt per line (`#` comments skipped), single model load, state reset between prompts — the cache-warming workflow below |
| `--audio file.dmel [-p "text"]` | spoken input: raw u8 DMel frames `[n_frames, 80]`, one `<|audio|>` position per frame (implies `--chat`) |
| `--audio file.dmel [-p "text"]` | spoken input: raw u8 DMel frames `[n_frames, 80]`, one audio-token position per frame (implies `--chat`) |
| `[cap] [bits] [ref.json]` | token-exact oracle harness against a `tools/make_tiny_inkling.py` fixture (CI-style validation; `tools/make_tiny_inkling_audio.py` for the audio path) |
`coli chat` / `coli serve` / `coli web` render the same template through the
+1 -1
View File
@@ -53,7 +53,7 @@ for equal-split scheduling — so including all 18 helps rather than drags.
Swept 8→96. The **cache-driven metrics are monotonic**: hit rate 4→53%, bytes
streamed and `eload` fall as the budget grows. But `eload` hits its floor
(~10 s) and hit-rate gains flatten around **44 GB** — beyond that you buy <1% hit
(~10 s) and hit-rate gains flatten around **44 GB** — beyond that you buy less than 1% hit
for more memory. On the CPU path the equivalent knee is ~48 GB; Metal sits a
touch lower because GPU-wired buffers trim the headroom.
+1 -1
View File
@@ -149,7 +149,7 @@ Ported all operations already implemented in the shared backend (backend_metal).
- CPU fallbacks preserved in both paths when Metal unavailable
**Standalone op battery test (Phase 5 expanded coverage):**
All ops verified against CPU reference — 33 standalone op tests, all pass (maxAbs <= 4.77e-7, MAE <= 5.87e-8 across all RMSNorm cases; exact match on all add cases; maxAbs <= 3.58e-7, MAE <= 1.35e-8 on silu_mul).
All ops verified against CPU reference — 33 standalone op tests, all pass (maxAbs ≤ 4.77e-7, MAE ≤ 5.87e-8 across all RMSNorm cases; exact match on all add cases; maxAbs ≤ 3.58e-7, MAE ≤ 1.35e-8 on silu_mul).
| Op | Shapes tested | MaxAbs | MAE |
|---|---|---|---|