* feat(cli): add --memory-percent and --ram-percent overrides
Limit VRAM or system RAM to a percentage of the detected capacity, so a
caller can reserve headroom without detecting the hardware itself.
--memory-percent scales every detected GPU by the same factor, so the
pooled total scales with it, and errors when no VRAM was detected.
--ram-percent scales total RAM and, on unified memory, caps VRAM at the
reduced pool. Both keep live free-memory readings, capped at the new
capacity, since they describe this machine rather than simulated
hardware. Each conflicts with its absolute counterpart and with
--profile, and both are forwarded to the auto-launched dashboard.
Closes#1086
* fix(hardware): recompute pooled free VRAM after --memory-percent
* fix(hardware): cap grouped free VRAM at the whole group after --memory-percent
* fix(hardware): cap each grouped card's free VRAM after --memory-percent
* feat: recognize native ternary (1.58-bit) models
Add an i2_s quantization tier and a bitnet.cpp runtime so BitNet-style 1.58-bit models are modeled as native ternary CPU models instead of generic Q4_K_M GGUFs. Detection keys off the HF architecture field and repo-name conventions; bf16/unpacked/MLX variants are excluded.
* review: wire BitNet runtime through CLI/API vocab, docs, and CPU fit
Addresses the maintainer + Greptile review on ternary support:
- CLI/API: accept and filter `bitnetcpp` in `--force-runtime`, `--runtime`,
the `/api/v1/models` runtime/force_runtime params, and the MCP runtime
parser. Ternary models are now selectable/filterable and are no longer
silently dropped by `runtime=llamacpp`.
- fit: route native-ternary to the CPU path (bitnet.cpp is CPU-only) so fit
is scored against system RAM, not VRAM; select BitNet before MLX on Apple
unified memory; estimate_tps uses a CPU throughput constant for BitNet on
any GPU backend instead of a GPU constant.
- download: dispatch `bitnetcpp` pulls through the GGUF provider so the
emitted runtime is no longer rejected as unknown.
- docs: add `bitnetcpp` to the runtime vocabulary in API.md and docs/cli.md.
- test: native-ternary on a GPU box resolves to CpuOnly.
* review: pull ternary GGUF artifacts for bitnet.cpp downloads
The bitnetcpp download path routed through the generic GGUF resolver, whose
quant preference omits i2_s/TQ, so a repository shipping both ordinary
k-quants and ternary weights could report a successful pull of a k-quant
file the bitnet.cpp runtime cannot load (and a repo without GGUFs failed
immediately either way).
Add `LlamaCppProvider::select_best_ternary_gguf` + `start_pull_ternary`,
which select only i2_s/TQ artifacts and error clearly when a repository has
none, and dispatch `bitnetcpp` downloads through that path.
* review: validate an explicitly-named bitnet.cpp GGUF is ternary
start_pull_ternary honoured an explicit `org/repo/file.gguf` without
checking its quantization, so `bitnetcpp` with a named k-quant could still
pull a file bitnet.cpp cannot load. Factor the ternary quant tags into a
shared const + `is_ternary_gguf_filename`, and reject a non-ternary
explicit filename with a clear error instead of downloading it.
* review: enforce the bitnet.cpp <-> native-ternary runtime invariant
bitnet.cpp loads only native-ternary (i2_s) weights, and a native-ternary
model loads only under bitnet.cpp. Route runtime selection through that single
invariant so a force_runtime override can no longer advertise a config that
cannot load the model:
- forcing bitnet.cpp on a non-ternary model is ignored (falls back to
auto-detection, with a note) instead of producing an i2_s CPU analysis;
- forcing a general runtime on a native-ternary model keeps bitnet.cpp (with a
note) instead of advertising mlx/llama.cpp/vLLM for i2_s weights.
Every downstream use (run-mode, quant hierarchy, tps, download) already keys
off `runtime == BitNet`, so enforcing it here fixes the whole class. Tests
cover both override directions.
* review: complete the bitnetcpp runtime vocabulary across docs + help
Add bitnetcpp to every remaining runtime-vocabulary enumeration so the
advertised contract matches the code: the MCP recommend_models schema doc, the
CLI --runtime help, API.md's force_runtime line, and the JA/ZH README
query-parameter tables (which were also missing vllm).
* style: cargo fmt to satisfy the Rustfmt CI check
Rustfmt wanted the two long assert chains in the ternary tests wrapped
(fit.rs, providers.rs). Formatting only, no logic change.
* review: exclude standard-quant repacks of 1.58-bit models from ternary detection
@AlexsJones flagged tiiuae/Falcon3-10B-Base-1.58bit-prequantized: a Q4_K_M GGUF
(architecture llama) that the name-only heuristic classified as native ternary
(BitNet/I2_S), which bitnet.cpp cannot load.
Root cause + data note: the scraped catalogue defaults `quantization` to "Q4_K_M"
for EVERY entry, including genuine i2_s models (microsoft/bitnet-b1.58-2B-4T,
1bitLLM/bitnet_b1_58-3B, ...), so a quant-field guard would wrongly exclude every
real bitnet model and break the feature. The only per-entry artifact signal is
the repo name, so this extends the existing repack-exclusion list (bf16 /
unpacked / -mlx) to also drop "-prequantized" (k-quant repack) and
"mlx-community" N-bit repacks — the same mechanism, respecting the declared
artifact marker before the name heuristic.
Adds the flagged catalogue entry as a regression fixture (loaded from the
embedded catalogue) alongside a genuine i2_s model, plus unit negatives.
* fix(bench): identify Ferrum and vLLM by endpoint owner
* fix(bench): normalize equivalent loopback endpoints
* fix(bench): reject malformed endpoint ports
* fix(bench): consume complete fixture HTTP requests
Read headers and the Content-Length body before sending the fixture response, avoiding Windows connection aborts when a POST arrives in multiple TCP reads. Cover short reads, bodies larger than the old buffer, and truncated requests while retaining the Ferrum provider attribution assertions.
* feat: hardware profiles, MoE Tier-2 fixes, and estimate confidence (#969)
Add spec-based --profile / hardware subcommand, per-arch MoE fallback
constants, memory-ratio fit verdicts, confidence labels, and catalog
sanitization so unreleased hardware and sparse MoE estimates are honest.
Co-Authored-By: Cursor <cursoragent@cursor.com>
* fix: thread profile CalcConfig through all estimate surfaces
Route rankable_models through every fit sweep, derive estimate_confidence
at serialize time, correct gpt-oss published active params, and make
plan/TUI/serve/MCP honor --profile bandwidth so the #969 repro hits ~50 tok/s.
Co-Authored-By: Cursor <cursoragent@cursor.com>
* docs: walkthrough for simulating hardware with --profile
Make the unreleased-SKU path obvious in cli.md (write → validate → score)
and link it from the hardware data README.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Saman Badakhshan <52349+saman-mb@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
The global --memory flag must precede the recommend subcommand; placing it after recommend fails with "unexpected argument '--memory'". Update the identical example in docs/platform-support.md, README.zh.md, and README.ja.md to the flag-first form already used by the sibling fit line.
Signed-off-by: Junhuan Zheng <3373484735@qq.com>
Adds docs/benchmarking.md walking the full TUI journey — download a
model, serve it, run a live benchmark, and share results as an
auto-opened PR — with four annotated screenshots in assets/.
README banner and Documentation table now point to the guide. The TUI
guide gains the previously undocumented "Benchmark this model" popup
keys, and the leaderboard docs are reframed community-first: llmfit
community submissions are the primary data source, localmaxxing.com is
supplemental, with the trust order spelled out to match cli.md.
Sets DEFAULT_CLIENT_ID to the registered llmfit OAuth App (device flow
enabled; a public client id is not a secret — the device flow needs no
client secret, which is why it was chosen). oauth_client_id() no longer
treats the default as an unregistered placeholder: the shipped id is
used unless LLMFIT_GH_CLIENT_ID overrides it, and an explicitly empty
override opts out of interactive login (restricted CI).
Verified live: with no GITHUB_TOKEN/GH_TOKEN, no cached token, and no
env override, `llmfit bench --share` presents a real GitHub device
code. Completes the last item on #712's pre-merge checklist.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Closes the read side of the contribution loop (RFC #710): community
submissions merged under llmfit-core/data/community/ are aggregated by a
new llmfit-core build.rs into an embedded JSON array — a merged
benchmark ships in the next release with no CI step or network fetch.
On hardware identical to a submission's (CPU + GPU fingerprint):
- benchmark page shows the runs as 'llmfit community' rows (distinct
color), pinned below the user's own 'you (local)/(shared)' rows and
deduped against them (your own shared runs also live in the embed)
- fit rows get measured tok/s from a new CommunityBenchIndex, slotted
between the user's own runs and localmaxxing preset medians; a new
MeasuredSource::CommunityLlmfit variant keeps provenance visible
('Measured on Identical Hardware' in the estimate detail)
- community anchors feed estimate calibration, so a fresh install gets
corrected estimates before its user ever benchmarks anything
hardware_payload_matches moves to benchmarks.rs and is shared with the
local store's matches_hardware. Verified end-to-end: with a temp
submission for llama3.1:8b @ 4.0 tok/s and an empty local store
(simulated fresh machine), the fit table shows 4.0 measured
(community_llmfit) and every estimate calibrated x0.37.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Step 1 — catalog hygiene: remove 100 CI/test-stub entries from
hf_models.json (tiny-random/*, *-tiny-random, tiny-dummy, ci-random-*,
test-*/testing-* repos). These are randomly initialized micro-models
whose names shadow real families — e.g. tiny-random/gemma-3 (9M params)
captured a local 2.5B gemma-3 GGUF and produced an 880 tok/s estimate
for a model that measures 3.8. The scraper now skips them at discovery
and at scrape time (is_test_stub), so they cannot return.
Step 2 — local calibration: one trustworthy benchmark now corrects the
whole estimate column for this machine. Anchors are locally-benched
fits on real catalog models (>= 1B params, dense); the median
measured/estimated ratio (clamped to [0.05, 3.0]) scales every row's
estimate, recorded in estimate_basis.local_calibration and printed in
the estimate-basis detail. Application is idempotent, so post-bench
refreshes never compound.
Stored runs are now also filtered by hardware fingerprint (CPU + GPU
name): measurements from a previous machine configuration neither
override nor calibrate the current one.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Before opening a PR, submit_stored now looks for the user's already-open
bench/* PR upstream and appends new results to it instead of opening one
PR per bench run. Upstream file names mirror the local store entry
(record timestamp + content hash), so a retry after a partial failure —
including a failed pending→shared move — skips files that already landed
(GitHub's 422 on create-over-existing) rather than duplicating them.
When everything is already upstream, no empty PR is attempted.
submit_stored/share_all_pending now return a SubmitOutcome (pr_url,
reused_existing_pr, uploaded, skipped) and the CLI/TUI report each case
distinctly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Benchmark runs are now always recorded locally as ready-to-upload,
schema-valid submission payloads (data dir, overridable with
LLMFIT_BENCH_STORE), so declining to share never discards data:
- Sharing uploads the entire pending store in one PR; uploaded files
move to shared/ so they persist as history but are never sent twice.
A failed or cancelled upload leaves everything pending.
- Bare `llmfit bench --share` contributes the stored backlog without
benchmarking again. The confirm prompt now offers to share all local
benchmarks, and the TUI toggle shows "this run + N stored".
- Credentials are resolved and verified against the GitHub API before
any benchmark starts: the CLI fails fast on a missing/expired token,
the TUI runs the device flow up front and downgrades to save-locally
with a note; a stale cached token is discarded and re-acquired.
- Stored results render pinned at the top of the TUI leaderboard as
"you (local)" / "you (shared)", including when the API is
unreachable (the full-page error now only shows when there are no
rows at all).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Previously `llmfit bench` only probed Ollama, vLLM, and MLX. A running
llama-server was either missed entirely (custom port) or silently
benchmarked under the MLX label when on the shared default port 8080 —
which would have polluted community benchmark data submitted via --share.
- New BenchTarget::LlamaCpp, positively identified via llama.cpp's
/props endpoint (MLX and vLLM 404 it), probed before the MLX check.
- Respects LLAMA_SERVER_HOST (full URL) and LLAMA_SERVER_PORT
(default 8080, matching the existing providers.rs convention).
- `--provider llamacpp` (aliases: llama.cpp, llama-server) selects it
explicitly for both performance and quality benchmarks.
- discover_all_targets skips the MLX probe when llama-server already
claimed the same URL, so models aren't double-counted.
- Results are labeled "llamacpp"; community schema enum updated and the
share payload test now covers it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds `llmfit bench --share`, which runs a benchmark sweep and opens a pull
request contributing the results to llmfit-core/data/community/ — with no
`gh` CLI and no server infrastructure.
Authentication uses the GitHub OAuth device flow (the same mechanism
`gh auth login` uses) with a public, env-overridable client id
(LLMFIT_GH_CLIENT_ID); a GITHUB_TOKEN/GH_TOKEN env var or a cached token
short-circuits the interactive step, making --share usable in CI. The
fork/branch/commit/open-PR steps go through the GitHub REST API via ureq
(already a dependency).
- New llmfit-core/src/share.rs: device flow, token resolution/caching,
payload building, and REST helpers.
- --share/--dry-run/--yes flags wired into the bench command.
- community/ data dir scaffold with README + JSON schema; a unit test
validates generated payloads against that schema.
Interactive login is gated on a registered OAuth App client id being
supplied via LLMFIT_GH_CLIENT_ID; until then, env tokens work.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The README had grown to ~1050 lines covering everything from vim
keybindings to crates.io publishing. Following the sympozium pattern:
the README keeps the hook (banner, install, quickstart, short how-it-
works) at ~190 lines, and deep content moves verbatim to focused
guides under docs/ (tui, cli, how-it-works, providers, platform
support, custom models, development, openclaw), linked from the
Documentation matrix at the top.
Also removes the stale /docs entry from .gitignore (added in 7f35522
with no generator behind it — nothing in CI or the web build writes
to /docs).
All relative links and anchors validated programmatically; root-
relative links in moved content rewritten with ../. README.zh.md and
README.ja.md are intentionally untouched until the English layout
settles.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>