Default Branch

ad8cd63847 · Share one CUDA encoder per IQ family (#2615) · Updated 2026-10-02 03:14:42 +08:00

Branches

99338d57fe · Deploying to gh-pages from @ NVIDIA/Model-Optimizer@fadbf74d31 🚀 · Updated 2026-10-02 03:06:06 +08:00

1310
1

5a8c939898 · Fix Megatron vLLM fakequant export state · Updated 2026-10-02 03:05:51 +08:00

9
18

8009e713c0 · docs(qwen3.6): explain why the penalty speeds up SciCode and slows GPQA · Updated 2026-10-02 02:43:34 +08:00

8
8

8009e713c0 · docs(qwen3.6): explain why the penalty speeds up SciCode and slows GPQA · Updated 2026-10-02 02:43:34 +08:00

8
8

d765177226 · megatron_bridge: store the local copy as local_checkpoint_path, the ID as hub_model_id · Updated 2026-10-02 02:25:47 +08:00

3
7

f2b9ecbd8a · llm_sparsity: store the local copy as local_checkpoint_path, the ID as hub_model_id · Updated 2026-10-02 02:25:47 +08:00

3
9

137a128a8c · HF exporters write off-index safetensors and non-model files · Updated 2026-10-02 02:25:46 +08:00

3
3

9ce3570eb9 · hf_ptq: store the local copy as local_checkpoint_path, the ID as hub_model_id · Updated 2026-10-02 02:25:46 +08:00

3
5

2d71a30ff7 · Name ensure_local_checkpoint's results hub_model_id and local_checkpoint_path · Updated 2026-10-02 02:25:45 +08:00

3
2

439ad4bcf5 · Add vLLM NVFP4 MLA KV-cache fake quant; fix FP8 and CUDA-graph serving · Updated 2026-10-02 02:15:17 +08:00

9
1

a7efcc9f55 · Merge branch 'main' into chenjiel/iq-kernel-dedup · Updated 2026-10-02 02:01:47 +08:00

3
3

a7efcc9f55 · Merge branch 'main' into chenjiel/iq-kernel-dedup · Updated 2026-10-02 02:01:47 +08:00

3
3

4fa387bd0a · test(export): consolidate reverse conversion regression coverage · Updated 2026-10-02 01:49:40 +08:00

36
8

4fa387bd0a · test(export): consolidate reverse conversion regression coverage · Updated 2026-10-02 01:49:40 +08:00

36
8

2ad0e70ae8 · Address review: free teacher fp32 copy early, input-only skip_lm_loss, dtype-tied ghost floor · Updated 2026-10-02 01:12:44 +08:00

9
11

2ad0e70ae8 · Address review: free teacher fp32 copy early, input-only skip_lm_loss, dtype-tied ghost floor · Updated 2026-10-02 01:12:44 +08:00

9
11

2beb70ffab · Merge main and resolve the GGML format registry conflict · Updated 2026-10-02 00:06:54 +08:00

4
8

2beb70ffab · Merge main and resolve the GGML format registry conflict · Updated 2026-10-02 00:06:54 +08:00

4
8

9e1ede83f8 · refactor(speculative): project the base distribution only where a loss reads it, family-wide · Updated 2026-10-01 23:30:43 +08:00

15
28

8f85d402da · Merge branch 'main' into lilicorr-streaming-cal · Updated 2026-10-01 21:24:02 +08:00

4
5