[doc] fix typos

This commit is contained in:
Zilin Zhu
2025-09-06 10:43:02 -07:00
parent 7feded2599
commit d4a774126f
6 changed files with 54 additions and 46 deletions
Binary file not shown.

After

Width:  |  Height:  |  Size: 182 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 596 KiB

+28 -24
View File
@@ -11,16 +11,16 @@ In a nutshell, this version can be summarized as follows:
Specifically, this version brings the following improvements:
- **Performance**:
- Provides **efficient inference for MoE models**, especially with `fp8` rollout + `deepep` + `mtp`.
- Provides **efficient inference for MoE models**, especially with **fp8 rollout + deepep + mtp**.
- Designed a generic **training framework memory offload solution** to save more KV Cache space, thus increasing inference concurrency.
- **Faster parameter updates**.
- Achieves more training with fewer GPUs through **CPU Adam**.
- Supports **all of Megatron's parallel strategies** as well as `deepep`.
- Supports **all of Megatron's parallel strategies** as well as **deepep**.
- **Features**:
- Added support for **GSPO** for MoE model training.
- Added support for **TIS** for `fp8` rollout.
- Added support for **TIS** for fp8 rollout.
- **Correctness**:
- Implemented **Dense and MoE model CI** to strictly check metrics like `kl`.
- Implemented **Dense and MoE model CI** to strictly check metrics like kl.
We hope to use slime v0.1.0 to demonstrate our understanding of high-performance RL training frameworks and have it become a baseline for future performance comparisons.
@@ -30,7 +30,7 @@ Next, I'll elaborate on the design philosophy behind these features.
## Performance Optimization: Pushing the Limits of RL Training Speed
In traditional deep learning training, there's a universal solution for speedup: **add more cards**. By reducing the amount of data processed per card, you can significantly lower end-to-end training latency.
In traditional deep learning training, there's a universal solution for speedup: **add more GPUs**. By reducing the amount of data processed per GPU, you can significantly lower end-to-end training latency.
However, this method doesn't work for RL training because **inference latency cannot be reduced by adding more GPUs**. Even with more GPUs, we still have to wait for the longest sample to finish decoding. While increasing throughput can improve the amount of training data per rollout, the off-policy issues caused by an excessively large inference batch size still have some limitations.
@@ -40,9 +40,11 @@ I believe this is the biggest challenge for infrastructure under the current RL
The decoding speed of a single data point determines the upper limit of RL training speed. For larger MoE models, there are currently three common optimization methods to push this limit, and we've tried all of them:
1. **Reduce memory access through quantization**: Considering that long calibration is not feasible in RL training, slime opts for `fp8` quantization.
2. **Use `deepep` low-latency mode to reduce `all2all` latency across machines**: To work with `deepep`, slime recommends using blockwise quantization with `fp8` to enable related SGLang configurations.
3. **Enable Speculative Sampling**: slime allows loading any `draft model` for the inference part (currently, it doesn't support updating the `draft model` during training).
1. **Reduce memory access through quantization**: Considering that long calibration is not feasible in RL training, slime opts for fp8 quantization.
2. **Use deepep low-latency mode to reduce all2all latency across machines**: To work with deepep, slime recommends using blockwise quantization with fp8 to enable related SGLang configurations.
3. **Enable Speculative Sampling**: slime allows loading any draft model for the inference part (currently, it doesn't support updating the draft model during training).
![](../../_static/image/blogs/release_v0.1.0/overrall.png)
By using the three optimizations mentioned above, we can increase a model like **GLM4.5 355B-A32B** from less than 10 tokens/s for a single data point to **60-70 tokens/s**, which significantly raises the upper limit of RL training speed.
@@ -50,36 +52,38 @@ In addition to monitoring inference throughput, slime also monitors `perf/longes
-----
## Doing More Experiments with Fewer Cards: Fully Offloading Megatron
## Doing More Experiments with Fewer GPUs: Fully Offloading Megatron
After optimizing the upper limit, we noticed another characteristic of RL training: as long as the **KV Cache doesn't overflow**, increasing the inference batch size doesn't significantly affect training latency.
**KV Cache overflow** occurs during inference when the response lengths of the data are all very long, leading to insufficient KV Cache space. This requires kicking out some half-generated data from the queue and then re-running `prefill` and subsequent inference steps after other data has been processed and freed up space. If a data point with a response length of 64k has to wait for 32k tokens to be decoded by other data during its inference, its total time is equivalent to decoding 96k tokens. This greatly impacts the RL training speed.
**KV Cache overflow** occurs during inference when the response lengths of the data are all very long, leading to insufficient KV Cache space. This requires kicking out some half-generated data from the queue and then re-running prefill and subsequent inference steps after other data has been processed and freed up space. If a data point with a response length of 64k has to wait for 32k tokens to be decoded by other data during its inference, its total time is equivalent to decoding 96k tokens. This greatly impacts the RL training speed.
Therefore, a more suitable training configuration is to calculate the minimum number of GPUs needed to prevent KV Cache overflow based on the inference batch size, the average response length, and the available KV Cache space on a single server. A group of these GPUs is then used for training. For example, if we have 512 cards and the calculation shows that 256 cards provide enough KV Cache, we should run two experiments in parallel instead of launching one experiment with all 512 cards.
Therefore, a more suitable training configuration is to calculate the minimum number of GPUs needed to prevent KV Cache overflow based on the inference batch size, the average response length, and the available KV Cache space on a single server. A group of these GPUs is then used for training. For example, if we have 512 GPUs and the calculation shows that 256 GPUs provide enough KV Cache, we should run two experiments in parallel instead of launching one experiment with all 512 GPUs.
Based on this consideration, we noticed two points for optimization:
1. **The optimal number of cards may not be sufficient to load the training part**. Inference only needs to load the `fp8` parameters, while training generally requires more than 18 times the parameter size of GPU memory (bf16 param, fp32 grad, fp32 master param, fp32 m and v). To solve this, slime uses **Megatron's built-in CPU Adam** to save GPU memory for the training part. This strategy allowed us to provide solutions for training GLM 4.5 355B-A32B with 8 nodes and DeepSeek R1 with 16 nodes.
2. **Increase the KV Cache space available per SGLang Server**, which means increasing `--mem-fraction`. For the more common integrated training and inference tasks, the main limitation for `mem_fraction` is the residual GPU memory after offloading the training part to the CPU. Therefore, we need to find a generic way to offload the GPU memory used by the Megatron part.
1. **The optimal number of GPUs may not be sufficient to load the training part**. Inference only needs to load the fp8 parameters, while training generally requires more than 18 times the parameter size of GPU memory (bf16 param, fp32 grad, fp32 master param, fp32 m and v). To solve this, slime uses **Megatron's built-in CPU Adam** to save GPU memory for the training part. This strategy allowed us to provide solutions for training GLM 4.5 355B-A32B with 8 nodes and DeepSeek R1 with 16 nodes.
2. **Increase the KV Cache space available per SGLang Server**, which means increasing `mem_fraction`. For the more common integrated training and inference tasks, the main limitation for a larger `mem_fraction` is the residual GPU memory after offloading the training part to the CPU. Therefore, we need to find a generic way to offload the GPU memory used by the Megatron part.
### How to Offload GPU Tensors Generically
One crude approach is to find all the GPU Tensors allocated by Megatron and call `.to("cpu")` on all of them. This method has three difficulties:
- It's hard to capture all GPU Tensors allocated by Megatron.
- Because Megatron's `distributed optimizer` reorganizes all parameters into some contiguous GPU buffers and then divides them with various `slices`, it's difficult to properly handle all references to correctly free the GPU Tensors.
- Because Megatron's distributed optimizer reorganizes all parameters into some contiguous GPU buffers and then divides them with various slices, it's difficult to properly handle all references to correctly free the GPU Tensors.
- It requires checking the source code again with every new Megatron version, which is hard to maintain.
Is there a more generic solution?
We noticed that SGLang's `torch_memory_saver` and VLLM's `cumem_allocator` provide a more general offload solution. Their principle is that CUDA 10.2 provides a **series of Virtual Memory Management APIs**, similar to an operating system's virtual and physical addresses (VA and PA). When allocating GPU memory, they return a handle to a memory mapping instead of the actual physical address. Therefore, when offloading, we only need to "secretly" release the memory corresponding to this mapping and reallocate it when this memory is needed. The upper-level application doesn't need to be aware of this.
![](../../_static/image/blogs/release_v0.1.0/cuda_vmm.png)
A natural idea is to use this method to take over the entire training process in RL. However, this prevents the reuse of PyTorch's `CUDACachingAllocator`, and without the cache, memory fragmentation becomes more pronounced, easily leading to **OOM** during training.
To continue reusing the native, cached allocator, we cannot use `CUDAPluggableAllocator`. Noticing again that slime's architecture has training and inference in different processes, we only need to **directly replace `cudaMalloc` and `cudaFree` used by `CUDACachingAllocator` in the training process with VMM APIs via `LD_PRELOAD`**. This allows us to completely and generically offload all GPU Tensors allocated by PyTorch.
At the same time, we must also note one detail: VMM APIs and `cudaIPC APIs` (such as `cudaIpcGetMemHandle`) are incompatible. Therefore, for integrated training and inference tasks and DeepEP, we need to disable the `LD_PRELOAD` replacement and switch back to `cudaMalloc`.
At the same time, we must also note one detail: VMM APIs and cudaIPC APIs (such as `cudaIpcGetMemHandle`) are incompatible. Therefore, for integrated training and inference tasks and DeepEP, we need to disable the `LD_PRELOAD` replacement and switch back to `cudaMalloc`.
With help from the SGLang community, we updated `torch_memory_saver` for slime's needs, implementing this offload solution.
@@ -87,7 +91,7 @@ With help from the SGLang community, we updated `torch_memory_saver` for slime's
After thoroughly offloading the GPU Tensors in Megatron, we found that a large amount of GPU memory still remained, which was caused by **NCCL**. In PyTorch, each NCCL group involved in communication allocates a substantial buffer. This issue is particularly noticeable for larger MoE models due to the various parallel strategies, potentially taking up **more than 10GB**.
The `LD_PRELOAD` solution mentioned above doesn't handle the NCCL issue well, and we don't want to modify the NCCL source code to avoid having to maintain a separate NCCL fork in addition to slime. So, slime's approach is to use `destroy_process_group` to destroy the NCCL group when offloading Megatron and then recreate it before loading Megatron. To do this, we mimicked the VMM API and `monkey patched` `dist.new_group` to add a layer of `ReloadableProcessGroup`.
The `LD_PRELOAD` solution mentioned above doesn't handle the NCCL issue well, and we don't want to modify the NCCL source code to avoid having to maintain a separate NCCL fork in addition to slime. So, slime's approach is to use `destroy_process_group` to destroy the NCCL group when offloading Megatron and then recreate it before loading Megatron. To do this, we mimicked the VMM API and monkey patched `dist.new_group` to add a layer of `ReloadableProcessGroup`.
In this way, we achieved a generic **NCCL offload**. However, because we need to rebuild the NCCL group, this operation has a slight impact on the speed of the first communication in each training iteration. But we believe this approach offers a significant advantage in terms of maintainability and the GPU memory it saves.
@@ -101,7 +105,7 @@ Parameter update is another special step in RL training. For this, slime v0.1.0
- [Efficient Reinforcement Learning Training - Optimizing Weight Synchronization in slime](https://hebiao064.github.io/rl-weight-sync)
Currently, slime can complete weight synchronization for a GLM4.5 355B-A32B model with bf16 weights in **48s** and complete `fp8` blockwise quantization + parameter update in **100s** (the `fp8` branch is still being optimized).
Currently, slime can complete weight synchronization for a GLM4.5 355B-A32B model with bf16 weights in **48s** and complete fp8 blockwise quantization + parameter update in **100s** (the fp8 branch is still being optimized).
-----
@@ -109,7 +113,7 @@ Currently, slime can complete weight synchronization for a GLM4.5 355B-A32B mode
For the pure training part of slime, we believe Megatron already provides ample optimizations, so our main focus was to **ensure compatibility with all of Megatron's parallel strategies**.
During this adaptation, we found an interesting bugfix: we discovered that when SGLang enabled `mtp`, the Megatron part couldn't start DeepEP. It turned out that when `mtp` is enabled, SGLang `disables the overlap schedule`, which causes a certain metadata communication to use `nccl` instead of `gloo` after being offloaded to the CPU, and this conflicts with DeepEP.
During this adaptation, we found an interesting bugfix: we discovered that when SGLang enabled mtp, the Megatron part couldn't start DeepEP. It turned out that when mtp is enabled, SGLang disables the overlap schedule, which causes a certain metadata communication to use nccl instead of gloo after being offloaded to the CPU, and this conflicts with DeepEP.
-----
@@ -122,9 +126,9 @@ My understanding of benchmarks is that they should not be used as a weapon for f
I also believe that before running benchmarks, you can analyze a framework's focus on performance from a qualitative perspective. Here's a basic feature check list for optimizations:
- Does it support MoE training? (Currently, large-scale experiments are focused on MoE)
- Can the internal `sglang mem fraction` or `vllm gpu utilization` be adjusted to over 0.7? (Ensures KV Cache space)
- Does it support `fp8` or lower precision inference? (Reduces inference memory access, boosts speed)
- Does it support enabling `deepep` for both training and inference? (Optimizes MoE `all2all` communication)
- Can the internal sglang `mem_fraction` or vllm `gpu_utilization` be adjusted to over 0.7? (Ensures KV Cache space)
- Does it support fp8 or lower precision inference? (Reduces inference memory access, boosts speed)
- Does it support enabling deepep for both training and inference? (Optimizes MoE all2all communication)
- Does it support speculative sampling? (Improves inference latency and throughput)
- Does it have an efficient training backend, such as Megatron or torchtitan, and support all necessary parallelization strategies? (Reuses mature training optimizations)
@@ -134,7 +138,7 @@ slime v0.1.0 has made preliminary attempts at all the above optimizations, and t
## New Algorithm Support
To better train MoE models and perform `fp8` rollouts, we implemented **GSPO** and **TIS**. Additionally, community experts have helped implement algorithms like `reinforce++` and `reinforce++ baseline`.
To better train MoE models and perform fp8 rollouts, we implemented **GSPO** and **TIS**. Additionally, community experts have helped implement algorithms like reinforce++ and reinforce++ baseline.
-----
@@ -142,8 +146,8 @@ To better train MoE models and perform `fp8` rollouts, we implemented **GSPO** a
slime v0.1.0 adds **end-to-end CI**: we run single-machine GLM4 9B and Qwen3 30B-A3B training for each PR, ensuring correctness through strict checks. For example, we explicitly require:
- The recomputed `log prob` of the first rollout must be exactly equal to the `log prob` of the `reference model`.
- The `ppo_kl` of the first training step within each rollout must be exactly 0.
- The recomputed log prob of the first rollout must be exactly equal to the log prob of the reference model.
- The ppo_kl of the first training step within each rollout must be exactly 0.
Such precise verification is rarely achieved in training frameworks, and it's something we are very proud of.
+1 -1
View File
@@ -61,5 +61,5 @@ slime is an LLM post-training framework for RL scaling, providing two core capab
:maxdepth: 1
:caption: Blogs
blogs/introducing_slime.md
blogs/release_v0.1.0.md
blogs/introducing_slime.md
+22 -18
View File
@@ -11,16 +11,16 @@
具体来说,这个版本带来了以下改进:
- **性能**:
- 提供了 **MoE 模型的高效推理**,特别是 `fp8` rollout + `deepep` + `mtp`。
- 提供了 **MoE 模型的高效推理**,特别是 fp8 rollout + deepep + mtp。
- 设计了通用的**训练框架显存 offload 方案**,节省出更多 KV Cache 空间,从而提升推理并发度。
- **更快的参数更新**。
- 通过 **CPU Adam** 实现用更少的 GPU 也能进行更多训练。
- 支持了 **Megatron 的全部并行策略**以及 `deepep`。
- 支持了 **Megatron 的全部并行策略**以及 deepep。
- **功能**:
- 针对 MoE 模型训练支持了 **GSPO**。
- 针对 `fp8` rollout 支持了 **TIS**。
- 针对 fp8 rollout 支持了 **TIS**。
- **正确性**:
- 加入了 **Dense 与 MoE 模型 CI**,严格检查 `kl` 等指标。
- 加入了 **Dense 与 MoE 模型 CI**,严格检查 kl 等指标。
我们希望通过 slime v0.1.0 展示我们对高性能 RL 训练框架的理解,并有机会成为未来性能对比的基准(baseline)。
@@ -40,9 +40,11 @@
单条数据的解码速度决定了 RL 训练速度的上限。对于较大的 MoE 模型,目前主要有以下 3 种常规优化方案来提升这一上限,我们也在每个方向上都做了尝试:
1. **通过量化来降低访存**:考虑到 RL 训练中不能进行长时间的 calibration,slime 选择进行 `fp8` 量化。
2. **使用 `deepep` low latency 模式来降低跨机 `all2all` 的时延**:为了配合 `deepep`,slime 推荐使用 `fp8` 的 blockwise 量化来开启 SGLang 的相关配置。
3. **开启投机采样(Speculative Sampling)**:slime 允许推理部分加载任意的 `draft model`(目前还不支持训练中更新 `draft model`)。
1. **通过量化来降低访存**:考虑到 RL 训练中不能进行长时间的 calibration,slime 选择进行 fp8 量化。
2. **使用 deepep low latency 模式来降低跨机 all2all 的时延**:为了配合 deepep,slime 推荐使用 fp8 的 blockwise 量化来开启 SGLang 的相关配置。
3. **开启投机采样(Speculative Sampling)**:slime 允许推理部分加载任意的 draft model(目前还不支持训练中更新 draft model)。
![](../../_static/image/blogs/release_v0.1.0/overrall.png)
使用上述 3 种优化,我们可以将 **GLM4.5 355B-A32B** 这样的模型,从单条数据小于 10 token/s,提升至 **60~70 token/s**,从而极大提升 RL 训练速度的上限。
@@ -52,13 +54,13 @@ slime 也会在监控推理的吞吐之外,监控 `perf/longest_sample_tokens_
在对上限进行优化后,我们注意到 RL 训练的另外一个特性:只要 **KV Cache 不溢出**,推理 batch size 的提升并不会明显影响训练的延迟。
**KV Cache 溢出**指的是推理过程中,当数据的回复长度都很长时,KV Cache 空间不足,需要将某些生成到一半的数据先踢出队列,等其他数据推理完并腾出空间后,再重新进行 `prefill` 和后续的推理。如果一条回复长度为 64k 的数据在推理过程中等待了其他数据解码 32k token,相当于它的总时长对应了解码 96k token,这对 RL 训练速度影响很大。
**KV Cache 溢出**指的是推理过程中,当数据的回复长度都很长时,KV Cache 空间不足,需要将某些生成到一半的数据先踢出队列,等其他数据推理完并腾出空间后,再重新进行 prefill 和后续的推理。如果一条回复长度为 64k 的数据在推理过程中等待了其他数据解码 32k token,相当于它的总时长对应了解码 96k token,这对 RL 训练速度影响很大。
因此,一个比较合适的训练配置是根据推理部分的 batch size、数据的平均回复长度和单个 server 能预留的 KV Cache 空间,计算一个在 KV Cache 不溢出情况下的最少 GPU 数量,并以这些 GPU 为一组进行训练。例如,我们有 512 张卡,计算出来 256 卡的 KV Cache 就已经充足,那么就应该并行运行 2 个实验,而不是用 512 卡一起启动实验。
基于这样的考量,我们注意到了 2 个优化点:
1. **最佳卡数可能不足以支持加载训练部分**。推理只需要加载 `fp8` 参数就够了,而训练一般需要 18 倍参数量以上的显存(bf16 param、fp32 grad、fp32 master param、fp32 m 和 v)。为了解决这个问题,slime 选择开启 **Megatron 自带的 CPU Adam** 来节省训练部分的显存。我们也是基于这样的策略提供了 8 节点训练 GLM 4.5 355B-A32B 以及 16 节点训练 DeepSeek R1 的方案。
1. **最佳卡数可能不足以支持加载训练部分**。推理只需要加载 fp8 参数就够了,而训练一般需要 18 倍参数量以上的显存(bf16 param、fp32 grad、fp32 master param、fp32 m 和 v)。为了解决这个问题,slime 选择开启 **Megatron 自带的 CPU Adam** 来节省训练部分的显存。我们也是基于这样的策略提供了 8 节点训练 GLM 4.5 355B-A32B 以及 16 节点训练 DeepSeek R1 的方案。
2. **提升每个 SGLang Server 能预留的 KV Cache**,也就是开大 `--mem-fraction`。对于现在更为常见的训推一体的训练任务,限制 `mem_fraction` 的主要是将训练部分 offload 至 CPU 后的残留显存。因此,我们需要找到一个通用的方法,将 Megatron 部分占用的显存 offload 到 GPU。
### 如何通用地 Offload GPU Tensor
@@ -66,13 +68,15 @@ slime 也会在监控推理的吞吐之外,监控 `perf/longest_sample_tokens_
一个比较粗暴的方式是找到 Megatron 分配的所有 GPU Tensor,然后把它们全部 `.to("cpu")`。这种做法有 3 个难点:
- 很难捕获 Megatron 分配的全部 GPU Tensor。
- 由于 Megatron 的 `distributed optimizer` 会把所有参数重新整理到一些连续的 GPU Buffer 里,然后再通过各种 `slice` 划分出去,很难处理好所有的引用从而正确释放 GPU Tensor。
- 由于 Megatron 的 distributed optimizer 会把所有参数重新整理到一些连续的 GPU Buffer 里,然后再通过各种 slice 划分出去,很难处理好所有的引用从而正确释放 GPU Tensor。
- Megatron 每次版本更新都要重新查一遍,不太好维护。
有没有一个更通用的方案呢?
我们注意到 SGLang 中的 `torch_memory_saver` 和 VLLM 中的 `cumem_allocator` 提供了一种较为通用的 offload 方案。它们的原理大致是,CUDA 10.2 提供了**一系列 Virtual Memory Management API**,类似于操作系统的虚拟地址与物理地址(VA 和 PA)。在分配显存时会返回一个显存映射的句柄(handle),而不是实际的物理地址。因此,我们在 offload 时只需要偷偷释放这个映射对应的显存,然后在需要这段显存时重新分配就可以了,上层的应用无需感知。
![](../../_static/image/blogs/release_v0.1.0/cuda_vmm.png)
一个自然的想法就是用这个方式接管 RL 中训练部分的全流程。但是,这会导致无法复用 PyTorch 的 `CUDACachingAllocator`,没有缓存会导致显存碎片更明显,训练过程很容易 **OOM**。
为了能继续复用原生的带缓存的 allocator,我们不能使用 `CUDAPluggableAllocator` 了。再次注意到 slime 的架构中训练和推理是在不同进程的,所以我们只需要通过 `LD_PRELOAD` **直接替换训练进程中 `CUDACachingAllocator` 使用的 `cudaMalloc` 和 `cudaFree` 为 VMM API**。这样我们就可以完整且通用地 offload PyTorch 分配的所有 GPU Tensor。
@@ -97,13 +101,13 @@ slime 也会在监控推理的吞吐之外,监控 `perf/longest_sample_tokens_
- [高效强化学习训练 - 优化 slime 中的权重同步](https://hebiao064.github.io/rl-weight-sync)
目前 slime 可以做到 **48s** 完成训推一体下 GLM4.5 355B-A32B 模型 bf16 权重的参数同步,以及 **100s** 完成 `fp8` blockwise 量化 + 参数更新(`fp8` 分支还在优化中)。
目前 slime 可以做到 **48s** 完成训推一体下 GLM4.5 355B-A32B 模型 bf16 权重的参数同步,以及 **100s** 完成 fp8 blockwise 量化 + 参数更新(fp8 分支还在优化中)。
## 训练优化
对于 slime 的纯训练部分,我们认为 Megatron 已经提供了充足的优化,所以我们主要是**保证了对 Megatron 全部并行策略的适配**。
适配过程中有一个有趣的 bugfix:我们发现在 SGLang 开启 `mtp` 后,Megatron 部分无法启动 DeepEP。后来发现是因为在开启 `mtp` 时,SGLang 会 `disable overlap schedule`,导致在被 offload 到 CPU 后仍用 `nccl` 而非 `gloo` 进行某个 metadata 的通信,并与 DeepEP 产生了冲突。
适配过程中有一个有趣的 bugfix:我们发现在 SGLang 开启 mtp 后,Megatron 部分无法启动 DeepEP。后来发现是因为在开启 mtp 时,SGLang 会 disable overlap schedule,导致在被 offload 到 CPU 后仍用 nccl 而非 gloo 进行某个 metadata 的通信,并与 DeepEP 产生了冲突。
## 性能优化 Check List
@@ -114,9 +118,9 @@ slime 发布以来,我经常被问到它与其他框架的性能对比。
同时,我也认为在跑分之前,可以从很多定性的角度来分析一个框架对性能的重视程度。这里提供一个基础优化的 feature check list:
- 是否支持 MoE 的训练?(目前大规模实验集中于 MoE)
- 是否支持内部的 `sglang mem fraction` 或 `vllm gpu utilization` 能调至 0.7 以上?(保证 KV Cache 空间)
- 是否支持 `fp8` 或更低精度量化的推理?(降低推理访存,提升速度)
- 是否支持训练和推理均开启 `deepep`?(优化 MoE `all2all` 通信)
- 是否支持内部的 sglang `mem_fraction` 或 vllm `gpu_utilization` 能调至 0.7 以上?(保证 KV Cache 空间)
- 是否支持 fp8 或更低精度量化的推理?(降低推理访存,提升速度)
- 是否支持训练和推理均开启 deepep?(优化 MoE all2all 通信)
- 是否支持投机采样?(提升推理时延与吞吐)
- 是否有高效的训练 backend,如 Megatron、torchtitan,并支持各项必要的并行?(复用成熟的训练优化)
@@ -124,14 +128,14 @@ slime v0.1.0 对上述的所有优化都进行了初步的尝试,当然提升
## 新算法支持
为了更好地训练 MoE 模型以及进行 `fp8` rollout,我们实现了 **GSPO** 与 **TIS**。同时,社区的大佬也帮忙实现了例如 `reinforce++`、`reinforce++ baseline` 这样的算法。
为了更好地训练 MoE 模型以及进行 fp8 rollout,我们实现了 **GSPO** 与 **TIS**。同时,社区的大佬也帮忙实现了例如 reinforce++、reinforce++ baseline 这样的算法。
## 正确性验证
slime v0.1.0 增加了**端到端 CI**:我们会对每个 PR 运行单机的 GLM4 9B 和 Qwen3 30B-A3B 训练,通过一些严格的检查来保证更新的正确性。例如,我们会明确要求:
- 第一个 rollout 的重算 `log prob` 和 `reference model` 的 `log prob` 完全相等。
- 每个 rollout 内的第一个训练步的 `ppo_kl` 严格为 0。
- 第一个 rollout 的重算 log prob 和 reference model 的 log prob 完全相等。
- 每个 rollout 内的第一个训练步的 ppo_kl 严格为 0。
这样的精确验证是训练框架中很少能做到的,也是我们非常自豪的一点。
+3 -3
View File
@@ -1,7 +1,7 @@
slime文档
slime 文档
====================
slime 是一个面向 RL 扩展的 LLM 后训练框架,提供两大核心能力:
slime 是一个面向 RL Scaling 的 LLM 后训练框架,提供两大核心能力:
- 高性能训练:通过连接 Megatron 与 SGLang,支持多种模式下的高效训练;
- 灵活的数据生成:通过自定义数据生成接口与基于服务器的引擎,实现任意训练数据生成流程。
@@ -55,5 +55,5 @@ slime 是一个面向 RL 扩展的 LLM 后训练框架,提供两大核心能
:maxdepth: 1
:caption: 博客
blogs/introducing_slime.md
blogs/release_v0.1.0.md
blogs/introducing_slime.md