17 Commits
Author SHA1 Message Date
Parman Mohammadalizadeh 17d45bd667 feat(cli): add --memory-percent and --ram-percent overrides (#1089)
* feat(cli): add --memory-percent and --ram-percent overrides

Limit VRAM or system RAM to a percentage of the detected capacity, so a
caller can reserve headroom without detecting the hardware itself.

--memory-percent scales every detected GPU by the same factor, so the
pooled total scales with it, and errors when no VRAM was detected.
--ram-percent scales total RAM and, on unified memory, caps VRAM at the
reduced pool. Both keep live free-memory readings, capped at the new
capacity, since they describe this machine rather than simulated
hardware. Each conflicts with its absolute counterpart and with
--profile, and both are forwarded to the auto-launched dashboard.

Closes #1086

* fix(hardware): recompute pooled free VRAM after --memory-percent

* fix(hardware): cap grouped free VRAM at the whole group after --memory-percent

* fix(hardware): cap each grouped card's free VRAM after --memory-percent
2026-10-01 06:38:00 +01:00
Ali Nabipour 42c2641b75 feat(storage): add disk planning for model libraries (#1023)
* feat(plan): include estimated disk size

* feat(storage): plan multi-model library capacity
2026-09-16 10:50:09 +01:00
ghc 569a9ac6cf feat: recognize native ternary (1.58-bit) models (#886)
* feat: recognize native ternary (1.58-bit) models

Add an i2_s quantization tier and a bitnet.cpp runtime so BitNet-style 1.58-bit models are modeled as native ternary CPU models instead of generic Q4_K_M GGUFs. Detection keys off the HF architecture field and repo-name conventions; bf16/unpacked/MLX variants are excluded.

* review: wire BitNet runtime through CLI/API vocab, docs, and CPU fit

Addresses the maintainer + Greptile review on ternary support:

- CLI/API: accept and filter `bitnetcpp` in `--force-runtime`, `--runtime`,
  the `/api/v1/models` runtime/force_runtime params, and the MCP runtime
  parser. Ternary models are now selectable/filterable and are no longer
  silently dropped by `runtime=llamacpp`.
- fit: route native-ternary to the CPU path (bitnet.cpp is CPU-only) so fit
  is scored against system RAM, not VRAM; select BitNet before MLX on Apple
  unified memory; estimate_tps uses a CPU throughput constant for BitNet on
  any GPU backend instead of a GPU constant.
- download: dispatch `bitnetcpp` pulls through the GGUF provider so the
  emitted runtime is no longer rejected as unknown.
- docs: add `bitnetcpp` to the runtime vocabulary in API.md and docs/cli.md.
- test: native-ternary on a GPU box resolves to CpuOnly.

* review: pull ternary GGUF artifacts for bitnet.cpp downloads

The bitnetcpp download path routed through the generic GGUF resolver, whose
quant preference omits i2_s/TQ, so a repository shipping both ordinary
k-quants and ternary weights could report a successful pull of a k-quant
file the bitnet.cpp runtime cannot load (and a repo without GGUFs failed
immediately either way).

Add `LlamaCppProvider::select_best_ternary_gguf` + `start_pull_ternary`,
which select only i2_s/TQ artifacts and error clearly when a repository has
none, and dispatch `bitnetcpp` downloads through that path.

* review: validate an explicitly-named bitnet.cpp GGUF is ternary

start_pull_ternary honoured an explicit `org/repo/file.gguf` without
checking its quantization, so `bitnetcpp` with a named k-quant could still
pull a file bitnet.cpp cannot load. Factor the ternary quant tags into a
shared const + `is_ternary_gguf_filename`, and reject a non-ternary
explicit filename with a clear error instead of downloading it.

* review: enforce the bitnet.cpp <-> native-ternary runtime invariant

bitnet.cpp loads only native-ternary (i2_s) weights, and a native-ternary
model loads only under bitnet.cpp. Route runtime selection through that single
invariant so a force_runtime override can no longer advertise a config that
cannot load the model:

- forcing bitnet.cpp on a non-ternary model is ignored (falls back to
  auto-detection, with a note) instead of producing an i2_s CPU analysis;
- forcing a general runtime on a native-ternary model keeps bitnet.cpp (with a
  note) instead of advertising mlx/llama.cpp/vLLM for i2_s weights.

Every downstream use (run-mode, quant hierarchy, tps, download) already keys
off `runtime == BitNet`, so enforcing it here fixes the whole class. Tests
cover both override directions.

* review: complete the bitnetcpp runtime vocabulary across docs + help

Add bitnetcpp to every remaining runtime-vocabulary enumeration so the
advertised contract matches the code: the MCP recommend_models schema doc, the
CLI --runtime help, API.md's force_runtime line, and the JA/ZH README
query-parameter tables (which were also missing vllm).

* style: cargo fmt to satisfy the Rustfmt CI check

Rustfmt wanted the two long assert chains in the ternary tests wrapped
(fit.rs, providers.rs). Formatting only, no logic change.

* review: exclude standard-quant repacks of 1.58-bit models from ternary detection

@AlexsJones flagged tiiuae/Falcon3-10B-Base-1.58bit-prequantized: a Q4_K_M GGUF
(architecture llama) that the name-only heuristic classified as native ternary
(BitNet/I2_S), which bitnet.cpp cannot load.

Root cause + data note: the scraped catalogue defaults `quantization` to "Q4_K_M"
for EVERY entry, including genuine i2_s models (microsoft/bitnet-b1.58-2B-4T,
1bitLLM/bitnet_b1_58-3B, ...), so a quant-field guard would wrongly exclude every
real bitnet model and break the feature. The only per-entry artifact signal is
the repo name, so this extends the existing repack-exclusion list (bf16 /
unpacked / -mlx) to also drop "-prequantized" (k-quant repack) and
"mlx-community" N-bit repacks — the same mechanism, respecting the declared
artifact marker before the name heuristic.

Adds the flagged catalogue entry as a regression fixture (loaded from the
embedded catalogue) alongside a genuine i2_s model, plus unit negatives.
2026-09-14 07:43:20 +01:00
Carl b3e09fd2d8 fix(bench): identify Ferrum and vLLM by endpoint owner (#994)
* fix(bench): identify Ferrum and vLLM by endpoint owner

* fix(bench): normalize equivalent loopback endpoints

* fix(bench): reject malformed endpoint ports

* fix(bench): consume complete fixture HTTP requests

Read headers and the Content-Length body before sending the fixture response, avoiding Windows connection aborts when a POST arrives in multiple TCP reads. Cover short reads, bodies larger than the old buffer, and truncated requests while retaining the Ferrum provider attribution assertions.
2026-09-09 21:30:24 +01:00
a8a1a93f7f feat: hardware profiles, MoE Tier-2 fixes, and estimate confidence (#969) (#971)
* feat: hardware profiles, MoE Tier-2 fixes, and estimate confidence (#969)

Add spec-based --profile / hardware subcommand, per-arch MoE fallback
constants, memory-ratio fit verdicts, confidence labels, and catalog
sanitization so unreleased hardware and sparse MoE estimates are honest.

Co-Authored-By: Cursor <cursoragent@cursor.com>

* fix: thread profile CalcConfig through all estimate surfaces

Route rankable_models through every fit sweep, derive estimate_confidence
at serialize time, correct gpt-oss published active params, and make
plan/TUI/serve/MCP honor --profile bandwidth so the #969 repro hits ~50 tok/s.

Co-Authored-By: Cursor <cursoragent@cursor.com>

* docs: walkthrough for simulating hardware with --profile

Make the unreleased-SKU path obvious in cli.md (write → validate → score)
and link it from the hardware data README.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Saman Badakhshan <52349+saman-mb@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-30 13:08:10 +01:00
hwan 5b023e2815 docs: place --memory before the recommend subcommand in Termux/Android examples (#932) (#937)
The global --memory flag must precede the recommend subcommand; placing it after recommend fails with "unexpected argument '--memory'". Update the identical example in docs/platform-support.md, README.zh.md, and README.ja.md to the flag-first form already used by the sibling fit line.

Signed-off-by: Junhuan Zheng <3373484735@qq.com>
2026-08-25 13:29:58 +01:00
Alex Jones f744292f86 Merge pull request #728 from mvanhorn/fix/288-lmstudio-download-status-endpoint
fix: poll LM Studio's real download status endpoint (/api/v1/models/download/status/:job_id) and correct stale docs
2026-07-23 05:49:23 +01:00
Alex Jones 3528c1f3a7 docs: add step-by-step benchmarking guide with screenshots
Adds docs/benchmarking.md walking the full TUI journey — download a
model, serve it, run a live benchmark, and share results as an
auto-opened PR — with four annotated screenshots in assets/.

README banner and Documentation table now point to the guide. The TUI
guide gains the previously undocumented "Benchmark this model" popup
keys, and the leaderboard docs are reframed community-first: llmfit
community submissions are the primary data source, localmaxxing.com is
supplemental, with the trust order spelled out to match cli.md.
2026-07-20 13:24:58 +01:00
Matt Van Horn e6a07aa045 fix: address self-review findings 2026-07-10 03:45:09 -07:00
Alex JonesandClaude Fable 5 325f8a3ff3 feat(share): ship the registered OAuth App client id — interactive login enabled
Sets DEFAULT_CLIENT_ID to the registered llmfit OAuth App (device flow
enabled; a public client id is not a secret — the device flow needs no
client secret, which is why it was chosen). oauth_client_id() no longer
treats the default as an unregistered placeholder: the shipped id is
used unless LLMFIT_GH_CLIENT_ID overrides it, and an explicitly empty
override opts out of interactive login (restricted CI).

Verified live: with no GITHUB_TOKEN/GH_TOKEN, no cached token, and no
env override, `llmfit bench --share` presents a real GitHub device
code. Completes the last item on #712's pre-merge checklist.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 12:02:40 +01:00
Alex JonesandClaude Fable 5 5c46f92a24 feat(bench): ship merged community benchmarks to every user in the next release
Closes the read side of the contribution loop (RFC #710): community
submissions merged under llmfit-core/data/community/ are aggregated by a
new llmfit-core build.rs into an embedded JSON array — a merged
benchmark ships in the next release with no CI step or network fetch.

On hardware identical to a submission's (CPU + GPU fingerprint):
- benchmark page shows the runs as 'llmfit community' rows (distinct
  color), pinned below the user's own 'you (local)/(shared)' rows and
  deduped against them (your own shared runs also live in the embed)
- fit rows get measured tok/s from a new CommunityBenchIndex, slotted
  between the user's own runs and localmaxxing preset medians; a new
  MeasuredSource::CommunityLlmfit variant keeps provenance visible
  ('Measured on Identical Hardware' in the estimate detail)
- community anchors feed estimate calibration, so a fresh install gets
  corrected estimates before its user ever benchmarks anything

hardware_payload_matches moves to benchmarks.rs and is shared with the
local store's matches_hardware. Verified end-to-end: with a temp
submission for llama3.1:8b @ 4.0 tok/s and an empty local store
(simulated fresh machine), the fit table shows 4.0 measured
(community_llmfit) and every estimate calibrated x0.37.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 11:49:18 +01:00
Alex JonesandClaude Fable 5 fb049c210c feat(fit): purge test-stub catalog entries; calibrate estimates from local benchmarks
Step 1 — catalog hygiene: remove 100 CI/test-stub entries from
hf_models.json (tiny-random/*, *-tiny-random, tiny-dummy, ci-random-*,
test-*/testing-* repos). These are randomly initialized micro-models
whose names shadow real families — e.g. tiny-random/gemma-3 (9M params)
captured a local 2.5B gemma-3 GGUF and produced an 880 tok/s estimate
for a model that measures 3.8. The scraper now skips them at discovery
and at scrape time (is_test_stub), so they cannot return.

Step 2 — local calibration: one trustworthy benchmark now corrects the
whole estimate column for this machine. Anchors are locally-benched
fits on real catalog models (>= 1B params, dense); the median
measured/estimated ratio (clamped to [0.05, 3.0]) scales every row's
estimate, recorded in estimate_basis.local_calibration and printed in
the estimate-basis detail. Application is idempotent, so post-bench
refreshes never compound.

Stored runs are now also filtered by hardware fingerprint (CPU + GPU
name): measurements from a previous machine configuration neither
override nor calibrate the current one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 11:07:42 +01:00
Alex JonesandClaude Fable 5 52f3afb10b feat(share): reuse open benchmark PR and make submissions idempotent
Before opening a PR, submit_stored now looks for the user's already-open
bench/* PR upstream and appends new results to it instead of opening one
PR per bench run. Upstream file names mirror the local store entry
(record timestamp + content hash), so a retry after a partial failure —
including a failed pending→shared move — skips files that already landed
(GitHub's 422 on create-over-existing) rather than duplicating them.
When everything is already upstream, no empty PR is attempted.

submit_stored/share_all_pending now return a SubmitOutcome (pr_url,
reused_existing_pr, uploaded, skipped) and the CLI/TUI report each case
distinctly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 10:43:08 +01:00
Alex JonesandClaude Fable 5 0e22a75abc feat(bench): local benchmark store — save runs, share the backlog later
Benchmark runs are now always recorded locally as ready-to-upload,
schema-valid submission payloads (data dir, overridable with
LLMFIT_BENCH_STORE), so declining to share never discards data:

- Sharing uploads the entire pending store in one PR; uploaded files
  move to shared/ so they persist as history but are never sent twice.
  A failed or cancelled upload leaves everything pending.
- Bare `llmfit bench --share` contributes the stored backlog without
  benchmarking again. The confirm prompt now offers to share all local
  benchmarks, and the TUI toggle shows "this run + N stored".
- Credentials are resolved and verified against the GitHub API before
  any benchmark starts: the CLI fails fast on a missing/expired token,
  the TUI runs the device flow up front and downgrades to save-locally
  with a note; a stale cached token is discarded and re-acquired.
- Stored results render pinned at the top of the TUI leaderboard as
  "you (local)" / "you (shared)", including when the API is
  unreachable (the full-page error now only shows when there are no
  rows at all).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 10:33:14 +01:00
Alex JonesandClaude Fable 5 6dc367be97 feat(bench): detect llama-server as a first-class provider
Previously `llmfit bench` only probed Ollama, vLLM, and MLX. A running
llama-server was either missed entirely (custom port) or silently
benchmarked under the MLX label when on the shared default port 8080 —
which would have polluted community benchmark data submitted via --share.

- New BenchTarget::LlamaCpp, positively identified via llama.cpp's
  /props endpoint (MLX and vLLM 404 it), probed before the MLX check.
- Respects LLAMA_SERVER_HOST (full URL) and LLAMA_SERVER_PORT
  (default 8080, matching the existing providers.rs convention).
- `--provider llamacpp` (aliases: llama.cpp, llama-server) selects it
  explicitly for both performance and quality benchmarks.
- discover_all_targets skips the MLX probe when llama-server already
  claimed the same URL, so models aren't double-counted.
- Results are labeled "llamacpp"; community schema enum updated and the
  share payload test now covers it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 07:54:39 +01:00
Alex JonesandClaude Opus 4.8 721eb666ab feat(bench): add --share to contribute benchmarks via PR without gh CLI
Adds `llmfit bench --share`, which runs a benchmark sweep and opens a pull
request contributing the results to llmfit-core/data/community/ — with no
`gh` CLI and no server infrastructure.

Authentication uses the GitHub OAuth device flow (the same mechanism
`gh auth login` uses) with a public, env-overridable client id
(LLMFIT_GH_CLIENT_ID); a GITHUB_TOKEN/GH_TOKEN env var or a cached token
short-circuits the interactive step, making --share usable in CI. The
fork/branch/commit/open-PR steps go through the GitHub REST API via ureq
(already a dependency).

- New llmfit-core/src/share.rs: device flow, token resolution/caching,
  payload building, and REST helpers.
- --share/--dry-run/--yes flags wired into the bench command.
- community/ data dir scaffold with README + JSON schema; a unit test
  validates generated payloads against that schema.

Interactive login is gated on a registered OAuth App client id being
supplied via LLMFIT_GH_CLIENT_ID; until then, env tokens work.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 07:43:19 +01:00
Alex JonesandClaude Fable 5 1bac5e0dba docs: break README sections out into docs/ guides, sympozium-style
The README had grown to ~1050 lines covering everything from vim
keybindings to crates.io publishing. Following the sympozium pattern:
the README keeps the hook (banner, install, quickstart, short how-it-
works) at ~190 lines, and deep content moves verbatim to focused
guides under docs/ (tui, cli, how-it-works, providers, platform
support, custom models, development, openclaw), linked from the
Documentation matrix at the top.

Also removes the stale /docs entry from .gitignore (added in 7f35522
with no generator behind it — nothing in CI or the web build writes
to /docs).

All relative links and anchors validated programmatically; root-
relative links in moved content rewritten with ../. README.zh.md and
README.ja.md are intentionally untouched until the English layout
settles.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 10:59:09 +01:00