CPU preset leads with Whisper small; CLIs and loaders share omnivoice.utils.dtype
(honours OMNIVOICE_CPU_DTYPE, no import cycle); the Windows-on-ARM notice and the
disk-space message are localized (setup_space takes {{gib}}: 5 for CPU installs, 9
otherwise); ARM64 payloads without a manifest fail the release check; Windows on
ARM is labelled experimental; test imports resolve at run time.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
NOTE blocks stay non-speech in frontend and backend VTT parsing; audio_info
keeps declared sizes for float WAV and restores stream position; ffmpeg shim
uses absolute targets, byte-verified copies, PATH priority, no path logging;
transcription guidance is localized through failure codes in all 21 locales.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
- model_manager: queued callers wait in short slices and re-check abandonment
- backend.ts/backend-gate: scrub child output (tokens, home paths) at source and at agent hand-off
- failure/error_journal: recognise OSError.winerror == 8 and [WinError 8] without English text
- subprocess_backend: generating heartbeats re-arm recv only, not model-load grace
- tests: event-based sync for voxcpm2 heartbeat test
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Install staged audio with a restorable backup before committing its row,
restore on any failure, serialize lock/unlock/consent file swaps, name the
FileNotFoundError skip (CodeQL), and drop the real-timer race in the offline
delete test.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Require tokenizer, feature-extractor settings and all weight shards before a
cached Whisper snapshot is used, including the first main-ref hit; speak the
last line-start SAY: of unclosed reasoning; keep quoted closing tags without
disabling reasoning removal when the source also contains the tag.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Adds a windows-11-arm leg to the packaging rehearsal and the release matrix
(experimental: a failed leg does not block the four established targets; the
release-asset check verifies the arm64 feed in full whenever any trace of it is
published). install.ps1 installs the native ARM64 build and falls back to x64
under emulation. README and the Windows/script docs gain a hardware table, CPU-only
and Windows-on-ARM guidance, and the external-drive recipe for #2436.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
OmniVoice was always loaded as float16, which has no fast CPU GEMM (and is
unimplemented for some kernels), so a host without a GPU loaded the model and
then crawled. The dtype now follows the device (float16 on cuda/xpu/mps/DirectML,
float32 on cpu; OMNIVOICE_CPU_DTYPE=bfloat16 opt-in) in the in-process loader, the
sidecar and the CLIs. Setup also explains CPU-only and Windows-on-ARM hosts,
and the CPU preset curates the small Whisper model.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The lock pins the +cu128 build for every Linux/Windows x64 host, so a laptop
with integrated graphics downloaded several GB of CUDA wheels (plus ~3 GB of
nvidia-* packages on Linux) it can never use, and needed 9 GiB free. Setup now
picks the torch flavour from the host (OMNIVOICE_TORCH_VARIANT=auto|cuda|cpu|rocm):
no NVIDIA driver -> frozen sync without torch/nvidia-*, then the +cpu pins. Existing
CUDA environments stay valid on inferred-CPU hosts; a CPU environment is reinstalled
when a driver later appears. Windows on ARM requests the emulated x64 interpreter
(no win_arm64 torchaudio/torchvision wheels exist) with CPU torch.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The pre-launch import probe killed at 30 s and treated a timeout as a broken
runtime (false setup_required / 'environment incomplete' on slow disks and
scanner-contended hosts). It now gets 180 s and a pure timeout is inconclusive.
The default startup budget doubles on small hosts and a backend that is still
printing extends the wait (capped at 3 budgets).
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The liveness probe was a sync route that imported torch and queried the CUDA
driver on every probe from the shared worker pool, so a busy but healthy
backend missed its 1.5 s deadline and was reported as unresponsive.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
snapshot_download(repo, local_files_only=True) only resolves the main ref, which
catalogue installs never write, so cloning reported an installed speech-to-text
model as missing. The model fallback now also tries every cached snapshot of the
configured PyTorch Whisper, large-v3-turbo and large-v3, skipping truncated
ones. Adds end-to-end resolution tests (short, long, supplied, no ASR).
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Classify private per-chunk ASR exceptions into fixed public replies, raise a
typed MediaToolUnavailableError from the validated decoder, and expose the
resolved ffmpeg under the bare name so dependencies that exec 'ffmpeg' find
imageio-ffmpeg's renamed binary on every platform.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
torchaudio 2.9 removed info() and routes load/save through TorchCodec. Add
audio_io.audio_info (header-declared WAV values, soundfile fallback), move the
dub cache checks and gpu_sandbox onto the audio_io helpers, and gate the whole
class with an AST test.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Review findings on #2435: the 180 s recv deadline killed valid CPU/slow-GPU
renders (no frames during generate) and the 30-step clamp silently reduced
the 32/64-step quality presets.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Same class as #2483: clear/delete/prune history and unlock unlinked audio
before the transaction committed; consent re-record deleted the previous
recording first; lock copied over the live locked take in place. Stage files
and swap/unlink after commit.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Review findings on #2467: a late-failing abandoned load left _load_abandoned set;
a queued caller's timeout could brand another caller's live load as stuck; the
warm-path test had no assertions.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Consolidate community engine, workflow, dictation, and setup fixes with the Electron composer, sidebar, and voice UI.
Fix review findings in engine residency, remote exports, backup cleanup, bounded compressed-audio decoding, and reference preprocessing. Preserve contributor credits in CHANGELOG.md and leave the app version unchanged.
Supersedes #2325, #2338, #2368, #2377, #2379, #2380, #2383, #2384, #2387, #2390, #2391, #2392, #2393, #2395, #2400, #2401, #2402, #2409, #2410, and #2412.