mirror of
https://github.com/debpalash/VoiceStudio.git
synced 2026-10-02 01:26:35 +08:00
OmniVoice was always loaded as float16, which has no fast CPU GEMM (and is unimplemented for some kernels), so a host without a GPU loaded the model and then crawled. The dtype now follows the device (float16 on cuda/xpu/mps/DirectML, float32 on cpu; OMNIVOICE_CPU_DTYPE=bfloat16 opt-in) in the in-process loader, the sidecar and the CLIs. Setup also explains CPU-only and Windows-on-ARM hosts, and the CPU preset curates the small Whisper model. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>