mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Improve technique-specific documentation links in the main README (OM… (#2584)
### What does this PR do? Improve technique-specific documentation links in the main README (OMNIML-5944): The Techniques table in ./README.md contains links that do not lead directly to the relevant documentation: - The Docs link for Quantization Aware Training / Distillation points to the general quantization documentation. It should point to ./guides/quantization_aware_training.html. - The Megatron Bridge links for Post Training Quantization, Quantization Aware Training / Distillation, Pruning, and Distillation all point to the same folder. Users must then find the relevant section themselves. Update the Megatron Bridge links to target the corresponding sections in ./examples/megatron_bridge/README.md: - In ./README.md’s Techniques table, rename Docs to Getting started and reverse the order of the two link columns: currently Examples → Docs, proposed Getting started → Examples. ### Usage just see the table of techniques in the main readme.me ### Testing tested manually ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - Did you write any new necessary tests?: ❌ , only manual testing, only docs changes <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the Techniques table with “Getting started” and “Examples” columns, replacing the former “Examples” and “Docs” columns. * Added technique-specific guide links, including a link to the README for pruning. * Updated some example destinations and anchors to point to relevant workflow sections. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
This commit is contained in:
co-authored by
Keval Morabia
parent
8990897c56
commit
94d7272e8b
@@ -106,14 +106,14 @@ more fine-grained control on installed dependencies or for alternative docker im
|
||||
|
||||
<div align="center">
|
||||
|
||||
| **Technique** | **Description** | **Examples** | **Docs** |
|
||||
| **Technique** | **Description** | **Getting started** | **Examples** |
|
||||
| :------------: | :------------: | :------------: | :------------: |
|
||||
| Post Training Quantization | Compress model size by 2x-4x, speeding up inference while preserving model quality! | \[[HF LLMs / VLMs](./examples/hf_ptq/)\] \[[Megatron-Bridge LLMs / VLMs](./examples/megatron_bridge/)\] \[[Diffusers](./examples/diffusers/)\] \[[ONNX](./examples/onnx_ptq/)\] \[[Windows](./examples/windows/)\] | \[[docs](https://nvidia.github.io/Model-Optimizer/guides/1_quantization.html)\] |
|
||||
| Quantization Aware Training / Distillation | Refine accuracy of quantized models even further with a few training steps! | \[[Hugging Face](./examples/llm_qat/)\] \[[Megatron-Bridge](./examples/megatron_bridge)\] | \[[docs](https://nvidia.github.io/Model-Optimizer/guides/1_quantization.html)\] |
|
||||
| Pruning | Reduce your model parameters or memory footprint and accelerate inference by removing unnecessary weights! | \[[General](./examples/pruning/)\] \[[Megatron-Bridge](./examples/megatron_bridge/)\] | |
|
||||
| Distillation | Reduce deployment model size by teaching small models to behave like larger models! | \[[Hugging Face](./examples/llm_distill/)\] \[[Megatron-Bridge](./examples/megatron_bridge/)\] \[[Megatron-LM](./examples/llm_distill/README.md#knowledge-distillation-kd-in-nvidia-megatron-lm-framework)\] | \[[docs](https://nvidia.github.io/Model-Optimizer/guides/4_distillation.html)\] |
|
||||
| Speculative Decoding | Train draft modules to predict extra tokens during inference! | \[[Hugging Face](./examples/speculative_decoding/)\] \[[Megatron-LM](./examples/speculative_decoding#mlm-example)\] | \[[docs](https://nvidia.github.io/Model-Optimizer/guides/5_speculative_decoding.html)\] |
|
||||
| Sparsity | Efficiently compress your model by storing only its non-zero parameter values and their locations | \[[Hugging Face](./examples/llm_sparsity/)\] | \[[docs](https://nvidia.github.io/Model-Optimizer/guides/6_sparsity.html)\] |
|
||||
| Post Training Quantization | Compress model size by 2x-4x, speeding up inference while preserving model quality! | \[[Start here](https://nvidia.github.io/Model-Optimizer/guides/1_quantization.html)\] | \[[HF LLMs / VLMs](./examples/hf_ptq/)\] \[[Megatron-Bridge LLMs / VLMs](./examples/megatron_bridge/README.md#post-training-quantization)\] \[[Diffusers](./examples/diffusers/)\] \[[ONNX](./examples/onnx_ptq/)\] \[[Windows](./examples/windows/)\] |
|
||||
| Quantization Aware Training / Distillation | Refine accuracy of quantized models even further with a few training steps! | \[[Start here](https://nvidia.github.io/Model-Optimizer/guides/quantization_aware_training.html)\] | \[[Hugging Face](./examples/llm_qat/)\] \[[Megatron-Bridge](./examples/megatron_bridge/README.md#quantization-aware-distillation-qad)\] |
|
||||
| Pruning | Reduce your model parameters or memory footprint and accelerate inference by removing unnecessary weights! | \[[Start here](./examples/pruning/README.md)\] | \[[General](./examples/pruning/)\] \[[Megatron-Bridge](./examples/megatron_bridge/README.md#pruning)\] |
|
||||
| Distillation | Reduce deployment model size by teaching small models to behave like larger models! | \[[Start here](https://nvidia.github.io/Model-Optimizer/guides/4_distillation.html)\] | \[[Hugging Face](./examples/llm_distill/)\] \[[Megatron-Bridge](./examples/megatron_bridge/README.md#distillation)\] \[[Megatron-LM](./examples/llm_distill/README.md#knowledge-distillation-kd-in-nvidia-megatron-lm-framework)\] |
|
||||
| Speculative Decoding | Train draft modules to predict extra tokens during inference! | \[[Start here](https://nvidia.github.io/Model-Optimizer/guides/5_speculative_decoding.html)\] | \[[Hugging Face](./examples/speculative_decoding/)\] \[[Megatron-LM](./examples/speculative_decoding#mlm-example)\] |
|
||||
| Sparsity | Efficiently compress your model by storing only its non-zero parameter values and their locations | \[[Start here](https://nvidia.github.io/Model-Optimizer/guides/6_sparsity.html)\] | \[[Hugging Face](./examples/llm_sparsity/)\] |
|
||||
|
||||
</div>
|
||||
|
||||
|
||||
Reference in New Issue
Block a user