mirror of
https://github.com/JustVugg/colibri.git
synced 2026-10-02 02:54:37 +08:00
docs: make pages MDX-compatible
This commit is contained in:
+1
-1
@@ -151,7 +151,7 @@ pair's current restack, base dev `292ed4c`):
|
||||
stamp already resolves to `int8-row` since the #528 INVERSION, and a
|
||||
stamp's role there is instead letting a genuinely-stamped `fmt=8` tensor
|
||||
override that default — see "The metadata stamp" below for the exact
|
||||
rule in both cases. FMT_NAMES table (name string <-> fmt int):
|
||||
rule in both cases. FMT_NAMES table (`name string` to `fmt int`):
|
||||
`c/colibri.c:1316`.
|
||||
- **no ordinal** (`int4-rans256-g0`, merged tools-only tier — line numbers
|
||||
at dev `7fb1159`, post-#671 merge `a3a5a75`, not at this PR pair's
|
||||
|
||||
+2
-2
@@ -276,7 +276,7 @@ telemetry stack — hardware, scheduler, tier bar, per-turn time breakdown, tok/
|
||||
trend and per-GPU expert counts:
|
||||
|
||||
<p align="center">
|
||||
<img src="media/colibri-mobile.png" width="270" alt="the dashboard on a phone-sized viewport">
|
||||
<img src="media/colibri-mobile.png" width="270" alt="the dashboard on a phone-sized viewport" />
|
||||
|
||||
<img src="media/colibri-metrics.png" width="300" alt="the telemetry sidebar">
|
||||
<img src="media/colibri-metrics.png" width="300" alt="the telemetry sidebar" />
|
||||
</p>
|
||||
|
||||
@@ -223,7 +223,7 @@ existing disk/wait numbers. `[METAL] residency-set: on` / the two fallback stder
|
||||
across cap1/cap16 may legitimately differ (different dispatch composition, per the
|
||||
fix-plan's "Determinism side-finding").
|
||||
5. **`[METAL] residency-set: on` line present in stderr** at flag-on startup, and absent
|
||||
(or the OS<15/create-failed fallback line) otherwise — cheap sanity check that a run
|
||||
(or the `OS < 15`/create-failed fallback line) otherwise — cheap sanity check that a run
|
||||
actually exercised the intended path before trusting its numbers. Also read the
|
||||
**`METAL-RESSET: flush` line** (gate-on only): if that number is large, the deferred
|
||||
set-commit cost is eating the stall win from the dispatch side.
|
||||
|
||||
+2
-2
@@ -84,9 +84,9 @@ help in proportion to the RAM you can give them.
|
||||
| Invocation | What it does |
|
||||
|---|---|
|
||||
| `-p "text" [-n N]` | streaming greedy generation (stops at eos or N tokens) |
|
||||
| `--chat -p "text"` | wraps the prompt in Inkling's chat template (role tokens + `<|content_text|>`, `<|message_model|>` as the generation prompt). Instruct models fed raw text are out of distribution and answer badly. `THINK=<0..1>` raises the reasoning effort (default 0) |
|
||||
| `--chat -p "text"` | wraps the prompt in Inkling's chat template (role tokens plus the content and model-message markers as the generation prompt). Instruct models fed raw text are out of distribution and answer badly. `THINK=0..1` raises the reasoning effort (default 0) |
|
||||
| `-f prompts.txt [-n N]` | one prompt per line (`#` comments skipped), single model load, state reset between prompts — the cache-warming workflow below |
|
||||
| `--audio file.dmel [-p "text"]` | spoken input: raw u8 DMel frames `[n_frames, 80]`, one `<|audio|>` position per frame (implies `--chat`) |
|
||||
| `--audio file.dmel [-p "text"]` | spoken input: raw u8 DMel frames `[n_frames, 80]`, one audio-token position per frame (implies `--chat`) |
|
||||
| `[cap] [bits] [ref.json]` | token-exact oracle harness against a `tools/make_tiny_inkling.py` fixture (CI-style validation; `tools/make_tiny_inkling_audio.py` for the audio path) |
|
||||
|
||||
`coli chat` / `coli serve` / `coli web` render the same template through the
|
||||
|
||||
@@ -53,7 +53,7 @@ for equal-split scheduling — so including all 18 helps rather than drags.
|
||||
|
||||
Swept 8→96. The **cache-driven metrics are monotonic**: hit rate 4→53%, bytes
|
||||
streamed and `eload` fall as the budget grows. But `eload` hits its floor
|
||||
(~10 s) and hit-rate gains flatten around **44 GB** — beyond that you buy <1% hit
|
||||
(~10 s) and hit-rate gains flatten around **44 GB** — beyond that you buy less than 1% hit
|
||||
for more memory. On the CPU path the equivalent knee is ~48 GB; Metal sits a
|
||||
touch lower because GPU-wired buffers trim the headroom.
|
||||
|
||||
|
||||
@@ -149,7 +149,7 @@ Ported all operations already implemented in the shared backend (backend_metal).
|
||||
- CPU fallbacks preserved in both paths when Metal unavailable
|
||||
|
||||
**Standalone op battery test (Phase 5 expanded coverage):**
|
||||
All ops verified against CPU reference — 33 standalone op tests, all pass (maxAbs <= 4.77e-7, MAE <= 5.87e-8 across all RMSNorm cases; exact match on all add cases; maxAbs <= 3.58e-7, MAE <= 1.35e-8 on silu_mul).
|
||||
All ops verified against CPU reference — 33 standalone op tests, all pass (maxAbs ≤ 4.77e-7, MAE ≤ 5.87e-8 across all RMSNorm cases; exact match on all add cases; maxAbs ≤ 3.58e-7, MAE ≤ 1.35e-8 on silu_mul).
|
||||
|
||||
| Op | Shapes tested | MaxAbs | MAE |
|
||||
|---|---|---|---|
|
||||
|
||||
Reference in New Issue
Block a user