### What does this PR do?
Type of change: Bug fix
Fixes synthetic conversation generation that assumed alternating user
and assistant messages. That assumption skipped user turns in prompt
skeletons and mishandled
leading system messages.
• Preserve conversation history. Regenerate every user turn while
retaining system messages and generated reasoning for subsequent
requests.
• Expose generation controls. Support model-specific request parameters,
configurable timeouts, and server-managed response budgets.
• Handle failures explicitly. Reject empty final answers and unsupported
tool calls. Failed conversations remain retryable without duplicating
saved output.
• Identify incomplete outputs. Mark length- and repetition-stopped
conversations as truncated, preserve stop metadata, and stop generating
follow-up turns.
### Usage
Run from the repository root against a compatible Qwen server with
reasoning parsing enabled:
python examples/speculative_decoding/scripts/server_generate.py \
--data_path input_conversations/train.jsonl \
--output_path synthetic/train.jsonl \
--url http://localhost:8000/v1 \
--model model \
--max_tokens 0 \
--request_timeout 3600 \
--extra_body
'{"chat_template_kwargs":{"enable_thinking":true,"preserve_thinking":true}}'
The model name must match the server’s configured name. Filter truncated
conversations before training.
### Testing
Focused regression tests: 15 passed.
The tests execute the command-line entry point using the real OpenAI
client library with mocked HTTP transport.
Coverage includes multi-turn generation, system prompts, reasoning
preservation, request parameters, failure recovery, resume
deduplication, truncation, and invalid
responses.
python -m pytest \
--confcutdir=tests/examples/speculative_decoding \
tests/examples/speculative_decoding/test_server_generate.py -q
The isolated test configuration avoids an unrelated parent configuration
import failure. All applicable pre-commit checks passed for the
generator, tests, and
documentation.
### Before your PR is "Ready for review"
• Is this change backward compatible?: ✅ Existing valid inputs,
defaults, conversation output structure, and resume behavior remain
supported. Invalid inputs and
failed requests now raise errors instead of being silently accepted.
• Copied code or new PIP dependencies?: N/A. No new third-party code or
dependencies were added.
• Did you write any new necessary tests?: ✅ Added focused command-line
regression tests.
• Did you update Changelog?: N/A. These are example-script correctness
fixes, not critical released library fixes.
• Did you get Claude approval on this PR?: ❌ Not yet obtained.
### Additional Information
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Data generation supports `conversations` and `messages` inputs,
preserves reasoning content, and accepts additional chat settings and
configurable request timeouts.
* Failed conversations are recorded separately, with options to retry
failures or exit when errors occur. Resume behavior distinguishes
retryable failures from rejected inputs.
* Outputs identify conversations truncated by length or repetition
limits.
* **Documentation**
* Updated data preparation guides with generation setup, input formats,
failure handling, resuming, and training guidance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
3.6 KiB
SLURM Prepare Data
For basic parallelization of synthetic data generation we provide some SLURM support.
Assuming a $SLURM_JOB_ID is present and nodes, n1, n2, n3, n4 are selected the following is achievable.
Example of allocating 4 nodes for 120 minutes
salloc -N4 -A <account> -p <partition> -J <account>-synthetic:data-gen -t 120
Create shards of some given size
python3 distributed_generate/sharding_utils.py --input_path /data/train.jsonl --output_dir /data/train/ --max_lines_per_shard 10000
Run workers on SLURM
bash distributed_generate/launch.sh $SLURM_JOB_ID vllm TinyLlama/TinyLlama-1.1B-Chat-v1.0 /data/train/ /data/output /scripts/ 0 10 n1,n2,n3,n4 "\"You are a helpful assistant.\""
/scripts/ is the absolute path to modelopt/examples/speculative_decoding which contains server_generate.py and distributed_generate.
This will launch a vllm server (sglang is also available) on each node. Each node will work through 10 shards of data (10*max_lines_per_shard number of samples).
In this case, the first 40 shards of data will be processed.
To process the next 40 shards
bash distributed_generate/launch.sh $SLURM_JOB_ID vllm TinyLlama/TinyLlama-1.1B-Chat-v1.0 /data/train/ /data/output /scripts/ 40 10 n1,n2,n3,n4
Failures and resuming
Workers continue to later shards after per-conversation failures by default. Successful
conversations stay in the output shard, while failures are logged to stderr and recorded in
<output_shard>.failures. Failed conversations and partial answers are never written as
training rows. The journal records the conversation ID, error, and whether ordinary resume
will retry it.
Rerun the launch command for the same shard range and output directory to resume. Keep shard names and input ordering unchanged, since conversation IDs are positions within each shard. Completed IDs are skipped. Temporary connection, timeout, rate-limit, and server errors remain retryable. Empty final answers are retryable at positive temperatures, but recorded as rejected at temperature zero. Unsupported tool roles or calls, malformed conversations, and HTTP 400/422 responses are also recorded as rejected and skipped on ordinary resume. Inspect these rejections before training; they can indicate bad input or incompatible request settings.
Set RETRY_FAILED=1 when launching to retry rejected conversations after fixing their input
or request settings. Set FAIL_ON_ERROR=1 to make any unresolved failures return a nonzero
status after the current shard finishes, stopping that worker before later shards. These
variables enable the generator's --retry_failed and --fail_on_error flags:
RETRY_FAILED=1 FAIL_ON_ERROR=1 bash distributed_generate/launch.sh $SLURM_JOB_ID vllm TinyLlama/TinyLlama-1.1B-Chat-v1.0 /data/train/ /data/output /scripts/ 0 10 n1,n2,n3,n4
Authentication failures, missing endpoints or models, and unexpected internal or output-write
errors stop workers even without strict mode. The worker enables --log_empty_conversations; its
finished marker means all inputs are saved or recorded as rejected, not that every input
succeeded. No new marker is written while retryable failures remain. A retry can append rows
after an older marker, so the generator reads the entire output when resuming.
Combining shards
To combine the shards back
python3 distributed_generate/sharding_utils.py --input_dir /data/output/ --output_path /data/output.jsonl --combine
The combiner ignores the .failures journals and completion markers. Review the journals and
filter conversations marked truncated: true before using the combined data for training.