mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
[Examples]: MiniMax-M3 DSpark (#1965)
### What does this PR do?
Type of change: new example + bug fixes
Adds a **MiniMax-M3 DSpark streaming training recipe** under
`tools/launcher/examples/MiniMaxAI/MiniMax-M3/`, plus the four fixes it
needs to actually run. Each fix addresses a failure mode that is silent
or misleading without it:
1. **Gemma-style final norm for the fake base**
(`modeling_final_norm.py`, `modeling_fakebase.py`): M3 uses a
gemma-style final RMSNorm (`(1 + weight)` scale, fp32
multiply-then-cast). Selecting the norm by `model_type` alone picks
plain `rmsnorm` — MiniMax's VL remote code coerces its `model_type`-less
`text_config` to **mixtral** — which silently drops the `+1` and
corrupts the distillation target. New `_FinalGemmaRMSNorm` + selection
by the explicit `use_gemma_norm` config flag (only MiniMax sets it).
2. **Loud failure for un-maskable `answer_only_loss`**
(`hf_streaming_dataset.py`): with a fast tokenizer whose chat template
has no `{% generation %}` tags,
`apply_chat_template(return_assistant_tokens_mask=True)` only warns and
returns an **all-zero mask** — training runs at zero loss on every
sample with no other symptom. Now raises at tokenization with an
actionable message.
3. **`SERVE_BLOCK_SIZE` knob** (`train_eagle_streaming.sh`): nemo_run
exports env values unquoted, so a multi-token
`SERVE_EXTRA_ARGS="--trust-remote-code --block-size 128"` loses
everything after the first token. M3's MSA sparse attention requires KV
block 128 (`ValueError: No common block size for 16` at engine init
otherwise), so `--block-size` gets a dedicated single-token knob.
4. **Relax the speculative_decoding `transformers` pin to `<5.13`**
(match `pyproject.toml`): the old `<5.4` pin downgrades recent vLLM
containers (e.g. transformers 5.12.1, which also provides in-tree
`minimax_m3_vl`) and breaks `vllm serve` (`ALLOWED_LAYER_TYPES` needs
>=5.5.3).
The example itself encodes the validated M3 specifics: generation-tagged
chat template copy (required for `answer_only_loss` — see fix 2), draft
dims + base `rope_theta=5e6` set explicitly (not inherited), mask token
200063 (reserved slot; added tokens end at 200060), `EAGLE_CAPTURE_IDS`
= draft default `target_layer_ids+1` + final layer, `trust_remote_code`
at serve/export, and AWS-EFA NIXL notes (UCX segfaults at agent init on
EFA nodes; LIBFABRIC required there).
### Usage
```bash
cd tools/launcher
export SLURM_HOST=localhost SLURM_ACCOUNT=<account> SLURM_PARTITION=<partition> \
SLURM_HF_LOCAL=<hf_models_dir> SLURM_JOB_DIR=<experiments_dir> NEMORUN_HOME=$PWD
uv run launch.py --yaml examples/MiniMaxAI/MiniMax-M3/hf_streaming_dspark_multi_node.yaml \
identity=$HOME/.ssh/id_ecdsa detach=True --yes
```
### Testing
- `tests/unit/torch/speculative/plugins/`: **129 passed** inside a
current vLLM x86_64 nightly container (transformers 5.12.1), including 3
new tests for the assistant-mask guard.
- Generation-tagged template verified against real corpus samples:
`input_ids` identical to the original template, mask covers exactly the
assistant turns (think prefix + content + eos), contiguous, no
user-prompt leak.
- The recipe is exercised end-to-end by a live M3 DSpark training run
(this yaml modulo cluster paths): streaming serve + NIXL transport +
resume all healthy; drafter MT-Bench AL exceeds our Kimi-K2.6 DSpark
reference by ~8k steps.
- `ruff check` / `ruff format` (0.15.20) clean.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (norm selection only changes
models with `use_gemma_norm=True`, previously mis-normed; the guard
turns a silent zero-loss run into an error; `SERVE_BLOCK_SIZE` is
opt-in; the pin relax widens the allowed range)
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ (can add if desired)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added Gemma-style RMS normalization support for speculative decoding.
- Added MiniMax-M3 multi-node training configuration, including a full
Jinja chat template for tool calls, multimodal content, and thinking
modes.
- Added optional `SERVE_BLOCK_SIZE` support for vLLM serve launches.
- **Bug Fixes**
- Improved `answer_only_loss` masking validation: now fails fast with
clear errors when required `{% generation %}` markers are missing or
when using a non-fast tokenizer.
- **Compatibility**
- Expanded the supported Transformers version range for speculative
decoding examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
This commit is contained in:
@@ -47,6 +47,9 @@
|
||||
# SERVE_CPU_OFFLOAD_GB GB/GPU offloaded to host RAM (fits big models on too-few GPUs; slower)
|
||||
# SERVE_MAX_MODEL_LEN cap context length (trims KV/activation)
|
||||
# SERVE_MAX_NUM_SEQS cap concurrent sequences (trims KV/activation)
|
||||
# SERVE_BLOCK_SIZE KV-cache block size (e.g. 128 for MiniMax-M3 MSA sparse attention).
|
||||
# Needs its own knob: multi-token SERVE_EXTRA_ARGS values are mangled
|
||||
# by nemo_run's unquoted env export, so "--block-size 128" cannot ride it.
|
||||
# SERVE_HOST single-node: bind/connect host. default 127.0.0.1
|
||||
# SERVE_GPU single-node: CUDA_VISIBLE_DEVICES for vllm. default "0"
|
||||
# SERVE_TP tensor-parallel size. default 1 single-node / all serve-node GPUs
|
||||
@@ -153,6 +156,7 @@ launch_vllm() {
|
||||
[ -n "${SERVE_CPU_OFFLOAD_GB:-}" ] && opt_args+=(--cpu-offload-gb "$SERVE_CPU_OFFLOAD_GB")
|
||||
[ -n "${SERVE_MAX_MODEL_LEN:-}" ] && opt_args+=(--max-model-len "$SERVE_MAX_MODEL_LEN")
|
||||
[ -n "${SERVE_MAX_NUM_SEQS:-}" ] && opt_args+=(--max-num-seqs "$SERVE_MAX_NUM_SEQS")
|
||||
[ -n "${SERVE_BLOCK_SIZE:-}" ] && opt_args+=(--block-size "$SERVE_BLOCK_SIZE")
|
||||
# --no-enable-chunked-prefill / --no-enable-prefix-caching: connector captures hidden states during prefill; both skip recomputing cached/partial prefixes, yielding short/empty hidden_states. Required.
|
||||
# --no-enable-flashinfer-autotune: on NVFP4 MoE the autotuner re-tunes on the first serving step and stalls a worker past vLLM's execute-model timeout, killing EngineCore.
|
||||
# Hidden states move serve -> trainer over NIXL RDMA (no disk round-trip): one
|
||||
|
||||
@@ -0,0 +1,134 @@
|
||||
# DSpark streaming speculative-decoding training for MiniMax-M3 (multi-node).
|
||||
# DSpark = the DFlash backbone + a lightweight Markov head + a confidence head,
|
||||
# generating a causal block semi-autoregressively; see dspark.yaml for the head
|
||||
# and loss config. Runs the shared streaming pipeline
|
||||
# (common/eagle3/train_eagle_streaming.sh) with the M3-specific base, draft dims,
|
||||
# mask token and chat template; trained from scratch. A starting point for
|
||||
# reproduction — tune node counts, batch, steps and serve limits for your cluster.
|
||||
#
|
||||
# MiniMax-M3 specifics this yaml encodes (each was a silent failure mode):
|
||||
# * SERVE_BLOCK_SIZE=128: M3's MSA sparse attention (sparse_block_size=128)
|
||||
# requires KV block 128 or vLLM dies at engine init ("No common block size").
|
||||
# * data.chat_template: M3 ships a FAST tokenizer whose template has no
|
||||
# {% generation %} tags -> assistant_masks come back ALL-ZERO and
|
||||
# answer_only_loss training silently runs at zero loss. The tagged template
|
||||
# copy next to this yaml wraps the assistant turn (think prefix + content +
|
||||
# tool calls + eos) in {% generation %} tags.
|
||||
# * The draft does NOT inherit base GQA/FFN dims (set explicitly below), and
|
||||
# M3's base rope_theta (5e6) is pinned onto the draft.
|
||||
# * EAGLE_CAPTURE_IDS = draft default target_layer_ids+1 (6 aux) + final (60).
|
||||
# * M3 is a VLM wrapper (text_config nested): the base's gemma-style final
|
||||
# norm is selected via the config's use_gemma_norm flag (see
|
||||
# modeling_final_norm.py) — its text_config coerces to a mixtral model_type,
|
||||
# so the model_type table alone would pick the wrong norm.
|
||||
#
|
||||
# Run ON the cluster login node (paramiko can't reach it through the login proxy):
|
||||
# export SLURM_HOST=localhost SLURM_ACCOUNT=<your_account> \
|
||||
# SLURM_PARTITION=<multi_node_partition> \
|
||||
# SLURM_HF_LOCAL=<hf_models_dir> \
|
||||
# SLURM_JOB_DIR=<experiments_dir> \
|
||||
# NEMORUN_HOME=$PWD
|
||||
# uv run launch.py --yaml examples/MiniMaxAI/MiniMax-M3/hf_streaming_dspark_multi_node.yaml \
|
||||
# identity=$HOME/.ssh/id_ecdsa detach=True --yes
|
||||
#
|
||||
# The export lands in /scratchspace/export.
|
||||
|
||||
job_name: MiniMax-M3_DSpark_streaming_multi_node
|
||||
pipeline:
|
||||
allow_to_fail: false
|
||||
skip: false
|
||||
note:
|
||||
|
||||
global_vars:
|
||||
hf_model: /hf-local/MiniMaxAI/MiniMax-M3
|
||||
|
||||
# Build /scratchspace/data/train.jsonl. Point data.data_path at the full
|
||||
# Spec-Decoding-Dataset-v2 corpus to reproduce; eagle_utils also accepts a
|
||||
# directory of *.jsonl shards directly.
|
||||
task_0:
|
||||
script: common/eagle3/make_dataset.sh
|
||||
args:
|
||||
- -f modules/Model-Optimizer/examples/dataset/example_data_config.yaml
|
||||
- --full-conversations
|
||||
slurm_config:
|
||||
_factory_: "slurm_factory"
|
||||
nodes: 1
|
||||
ntasks_per_node: 1
|
||||
gpus_per_node: 8
|
||||
container: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc10
|
||||
|
||||
task_1:
|
||||
script: common/eagle3/train_eagle_streaming.sh
|
||||
args:
|
||||
- --config modules/Model-Optimizer/modelopt_recipes/general/speculative_decoding/dspark.yaml
|
||||
- model.model_name_or_path=<<global_vars.hf_model>>
|
||||
- model.use_fake_base_for_offline=true
|
||||
- model.trust_remote_code=true
|
||||
- data.mode=streaming
|
||||
- data.data_path=/scratchspace/data/train.jsonl
|
||||
# M3's own template has no {% generation %} tags; without this tagged copy
|
||||
# answer_only_loss trains on an all-zero mask (see header).
|
||||
- data.chat_template=examples/MiniMaxAI/MiniMax-M3/m3_chat_template_generation.jinja
|
||||
- training.output_dir=/scratchspace/dspark
|
||||
- training.training_seq_len=4096
|
||||
- training.disable_tqdm=true
|
||||
- training.ar_validate_steps=500000
|
||||
- training.num_train_epochs=1
|
||||
- training.per_device_train_batch_size=4
|
||||
- training.gradient_accumulation_steps=1
|
||||
- training.save_steps=1000
|
||||
- training.logging_steps=20
|
||||
- training.learning_rate=1.0e-4
|
||||
- training.warmup_steps=2000
|
||||
- training.answer_only_loss=true
|
||||
# The vLLM serve container has no tensorboard -> trainer init crash.
|
||||
- training.report_to=none
|
||||
# The DSpark draft does NOT inherit the base GQA/FFN dims, so set them
|
||||
# explicitly to match the M3 backbone (else a silently wrong-shape draft).
|
||||
# intermediate_size matches M3's dense FFN (dense_intermediate_size).
|
||||
- dflash.dflash_architecture_config.num_hidden_layers=6
|
||||
- dflash.dflash_architecture_config.num_key_value_heads=8
|
||||
- dflash.dflash_architecture_config.intermediate_size=12288
|
||||
# Pin the base's rope_theta onto the draft (M3 uses 5e6, not the Qwen3
|
||||
# default 1e6; a mismatch trains rope into the weights and caps AL).
|
||||
- dflash.dflash_architecture_config.rope_theta=5000000
|
||||
# Semi-AR generation block (dspark.yaml ships 16; Kimi/M3 runs use 8).
|
||||
- dflash.dflash_block_size=8
|
||||
# M3 has no mask token; vocab 200064, added tokens end at 200060 -> 200063 free.
|
||||
- dflash.dflash_mask_token_id=200063
|
||||
environment:
|
||||
- HF_MODEL_CKPT: <<global_vars.hf_model>>
|
||||
# 6 aux capture ids = the draft's default target_layer_ids+1, plus the true
|
||||
# final hidden (60). Requires the aux-capture fix vllm#46788 (in-tree in
|
||||
# recent nightlies); mismatched ids silently skew train vs inference.
|
||||
- EAGLE_CAPTURE_IDS: "[2,13,24,36,47,58,60]"
|
||||
- SERVE_NODES: "4"
|
||||
- SERVE_TP: "8"
|
||||
- STREAMING_NUM_WORKERS: "4"
|
||||
# M3's custom-modeling base needs trust_remote_code at export and serve.
|
||||
- EXPORT_EXTRA_ARGS: "--trust_remote_code"
|
||||
- SERVE_EXTRA_ARGS: "--trust-remote-code"
|
||||
# REQUIRED for M3: MSA sparse attention needs KV block 128 (dedicated knob —
|
||||
# multi-token SERVE_EXTRA_ARGS values are mangled by nemo_run's unquoted
|
||||
# env export, so "--block-size 128" cannot ride it).
|
||||
- SERVE_BLOCK_SIZE: "128"
|
||||
- SERVE_MAX_MODEL_LEN: "4160"
|
||||
- SERVE_MAX_NUM_SEQS: "32"
|
||||
- SERVE_GPU_MEM_UTIL: "0.9"
|
||||
- SERVE_READY_TIMEOUT: "3600"
|
||||
- VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS: "1200"
|
||||
- VLLM_ENGINE_ITERATION_TIMEOUT_S: "1200"
|
||||
# RDMA transport is UCX (InfiniBand) by default. On AWS EFA, uncomment —
|
||||
# and note UCX SEGFAULTS at agent init on EFA nodes (it detects the EFA
|
||||
# devices), so LIBFABRIC is required there even for single-node runs:
|
||||
# - NIXL_BACKENDS: "LIBFABRIC"
|
||||
# - FI_PROVIDER: "efa"
|
||||
# - NCCL_IB_DISABLE: "1"
|
||||
slurm_config:
|
||||
_factory_: "slurm_factory"
|
||||
nodes: 6
|
||||
ntasks_per_node: 1
|
||||
gpus_per_node: 8
|
||||
# vLLM x86_64 build with native MiniMax-M3 support (vllm/models/minimax_m3)
|
||||
# and the aux-capture fix (vllm#46788).
|
||||
container: <vllm-image-with-native-minimax-m3>
|
||||
@@ -0,0 +1,255 @@
|
||||
{# MiniMax-M3 chat template with {% generation %} tags for answer_only_loss training.
|
||||
Adapted from https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/chat_template.jinja
|
||||
with {% generation %} / {% endgeneration %} wrapping assistant content (think prefix,
|
||||
content, and eos), so apply_chat_template can return the assistant token mask.
|
||||
-#}
|
||||
{# ---------- special token variables ---------- #}
|
||||
{%- set ns_token = ']<]minimax[>[' -%}
|
||||
{%- set bod_token = ']~!b[' -%}
|
||||
{%- set bos_token = ']~b]' -%}
|
||||
{%- set eos_token = '[e~[' -%}
|
||||
{%- set toolcall_begin_token = ns_token ~ '<tool_call>' -%}
|
||||
{%- set toolcall_end_token = ns_token ~ '</tool_call>' -%}
|
||||
{%- set think_begin_token = '<mm:think>' -%}
|
||||
{%- set think_end_token = '</mm:think>' -%}
|
||||
{%- set image_token = ']<]image[>[' -%}
|
||||
{%- set video_token = ']<]video[>[' -%}
|
||||
{#- Thinking mode: "enabled" / "disabled" / "adaptive" / not defined -#}
|
||||
{#- Recursive XML renderer for tool_call arguments ======================== -#}
|
||||
{#- None values are intentionally skipped in mapping iteration so that
|
||||
`<key>null</key>` (which would round-trip to the literal string "null")
|
||||
never appears in the rendered tool_call. The convention is: omit the
|
||||
field entirely. The top-level `_args` loop applies the same rule.
|
||||
The `val is none` branch below is a safety net only — upstream cleaning
|
||||
(drop_none_in_tool_arguments) should ensure no None ever reaches here. -#}
|
||||
{%- macro to_xml(val, ns) -%}
|
||||
{%- if val is mapping -%}
|
||||
{%- for k, v in val.items() if v is not none -%}
|
||||
{{ ns }}<{{ k }}>{{ to_xml(v, ns) }}{{ ns }}</{{ k }}>
|
||||
{%- endfor -%}
|
||||
{%- elif val is iterable and val is not string -%}
|
||||
{%- for item in val -%}
|
||||
{{ ns }}<item>{{ to_xml(item, ns) }}{{ ns }}</item>
|
||||
{%- endfor -%}
|
||||
{%- elif val is none -%}
|
||||
{#- Should be unreachable when upstream cleaning is applied. -#}
|
||||
{%- elif val is boolean -%}
|
||||
{{ val | tojson }}
|
||||
{%- else -%}
|
||||
{{ val }}
|
||||
{%- endif -%}
|
||||
{%- endmacro -%}
|
||||
{#- Tool Rendering Functions ============================================== -#}
|
||||
{%- macro render_tool_namespace(namespace_name, tool_list) -%}
|
||||
{%- for tool in tool_list -%}
|
||||
<tool>{{ tool.function | tojson(ensure_ascii=False) }}</tool>
|
||||
{% endfor -%}
|
||||
{%- endmacro -%}
|
||||
{%- macro visible_text(content) -%}
|
||||
{%- if content is string -%}
|
||||
{{ content }}
|
||||
{%- elif content is iterable and content is not mapping -%}
|
||||
{%- for item in content -%}
|
||||
{%- if item is mapping and item.type == 'text' -%}
|
||||
{{- item.text }}
|
||||
{%- elif item is mapping and item.type == 'image' -%}
|
||||
{{- image_token }}
|
||||
{%- elif item is mapping and item.type == 'video' -%}
|
||||
{{- video_token}}
|
||||
{%- elif item is string -%}
|
||||
{{- item }}
|
||||
{%- endif -%}
|
||||
{%- endfor -%}
|
||||
{%- elif content is none -%}
|
||||
{{- '' }}
|
||||
{%- else -%}
|
||||
{{- content }}
|
||||
{%- endif -%}
|
||||
{%- endmacro -%}
|
||||
{#- System Message Construction ============================================ -#}
|
||||
{%- macro build_system_message(system_message) -%}
|
||||
{%- if system_message and system_message.content -%}
|
||||
{{- visible_text(system_message.content) }}
|
||||
{%- else -%}
|
||||
{{- 'Your model version is MiniMax-M3, developed by MiniMax. Knowledge cutoff: January 2026. Founded in early 2022, MiniMax is a global AI foundation model company committed to advancing the frontiers of AI towards AGI.' }}
|
||||
{%- endif -%}
|
||||
|
||||
{#- Thinking mode instructions -#}
|
||||
{{- '\n\n<thinking_instructions>\n' }}
|
||||
{{- 'You have a thinking capability that allows you to reason step by step before responding. When thinking is enabled, wrap your reasoning in ' ~ think_begin_token ~ think_end_token ~ ' tags before your response. When thinking is disabled, begin your response directly after the ' ~ think_end_token ~ ' prefix. When thinking is adaptive, decide on your own whether to think for the current turn.\n' }}
|
||||
{%- if thinking_mode is defined -%}
|
||||
{%- if thinking_mode == "enabled" -%}
|
||||
{{- 'Current thinking mode: enabled. You MUST think step by step before every response, including after receiving function/tool results.\n' }}
|
||||
{%- elif thinking_mode == "disabled" -%}
|
||||
{{- 'Current thinking mode: disabled. Do not output any thinking process.\n' }}
|
||||
{%- elif thinking_mode == "adaptive" -%}
|
||||
{{- 'Current thinking mode: adaptive. You are encouraged to think for complex decision-making, multi-step reasoning, or when analyzing function/tool results.\n' }}
|
||||
{%- endif -%}
|
||||
{%- else -%}
|
||||
{{- 'Current thinking mode: adaptive. You are encouraged to think for complex decision-making, multi-step reasoning, or when analyzing function/tool results.\n' }}
|
||||
{%- endif -%}
|
||||
{{- '</thinking_instructions>' }}
|
||||
{%- endmacro -%}
|
||||
{%- macro build_developer_message(developer_message) -%}
|
||||
{%- if developer_message and developer_message.content -%}
|
||||
{{- visible_text(developer_message.content) }}
|
||||
{%- else -%}
|
||||
{%- if model_identity is not defined -%}
|
||||
{%- set model_identity = "You are a helpful assistant." -%}
|
||||
{%- endif -%}
|
||||
{{- model_identity }}
|
||||
{%- endif -%}
|
||||
{%- endmacro -%}
|
||||
{#- Main Template Logic ================================================= -#}
|
||||
{#- Role mapping: root -> system sp (high priority), system/developer -> developer sp (low priority) -#}
|
||||
{%- set system_message = none -%}
|
||||
{%- set developer_message = none -%}
|
||||
{%- set conversation_messages = messages -%}
|
||||
{%- if messages and messages[0].role == "root" -%}
|
||||
{%- set system_message = messages[0] -%}
|
||||
{%- set conversation_messages = messages[1:] -%}
|
||||
{%- if conversation_messages and conversation_messages[0].role in ["system", "developer"] -%}
|
||||
{%- set developer_message = conversation_messages[0] -%}
|
||||
{%- set conversation_messages = conversation_messages[1:] -%}
|
||||
{%- endif -%}
|
||||
{%- elif messages and messages[0].role in ["system", "developer"] -%}
|
||||
{%- set developer_message = messages[0] -%}
|
||||
{%- set conversation_messages = messages[1:] -%}
|
||||
{%- endif -%}
|
||||
{#- Render system sp (higher priority, root role only) -#}
|
||||
{{- bod_token ~ bos_token ~ 'system' ~ '\n' }}
|
||||
{{- build_system_message(system_message) }}
|
||||
{{- eos_token ~ '\n' }}
|
||||
|
||||
{#- Render developer sp (lower priority: system/developer role + tools) -#}
|
||||
{{- bos_token ~ 'developer' ~ '\n' }}
|
||||
{{- build_developer_message(developer_message) }}
|
||||
{%- if tools -%}
|
||||
{{- '\n\n' ~ '# Tools' ~ '\n' ~ 'You may call one or more tools to assist with the user query.\nHere are the tools available in JSONSchema format:' ~ '\n' }}
|
||||
{{- '\n' ~ '<tools>' ~ '\n' }}
|
||||
{{- render_tool_namespace("functions", tools) }}
|
||||
{{- '</tools>' ~ '\n\n' }}
|
||||
{{- 'To call tools, wrap all invocations in a single ' ~ toolcall_begin_token ~ toolcall_end_token ~ ' block. Parameter values containing nested objects or arrays are recursively expanded into XML elements. Example:\n' }}
|
||||
{{- '\n' ~ toolcall_begin_token ~ '\n' }}
|
||||
{{- ns_token + '<invoke name="tool-name-1">' }}
|
||||
{{- ns_token + '<param-1>value-1' + ns_token + '</param-1>' }}
|
||||
{{- ns_token + '<param-2>' }}
|
||||
{{- ns_token + '<item>' }}
|
||||
{{- ns_token + '<key-a>val-a' + ns_token + '</key-a>' }}
|
||||
{{- ns_token + '<key-b>val-b' + ns_token + '</key-b>' }}
|
||||
{{- ns_token + '</item>' }}
|
||||
{{- ns_token + '</param-2>' }}
|
||||
{{- ns_token + '</invoke>\n' }}
|
||||
{{- ns_token + '<invoke name="tool-name-2">' }}
|
||||
{{- ns_token + '<param-1>value-1' + ns_token + '</param-1>' }}
|
||||
{{- ns_token + '</invoke>\n' }}
|
||||
{{- toolcall_end_token }}
|
||||
{%- endif -%}
|
||||
{{- eos_token ~ '\n' }}
|
||||
|
||||
{#- Render messages -#}
|
||||
{%- set last_tool_call = namespace(name=none) -%}
|
||||
{%- for message in conversation_messages -%}
|
||||
{%- if message.role == 'assistant' -%}
|
||||
{{- bos_token ~ 'ai' ~ '\n' }}
|
||||
{%- generation -%}
|
||||
|
||||
{%- set reasoning_content = '' %}
|
||||
{%- set content = visible_text(message.content) %}
|
||||
{%- if message.reasoning_content is string %}
|
||||
{%- set reasoning_content = message.reasoning_content %}
|
||||
{%- else %}
|
||||
{%- if think_end_token in content %}
|
||||
{%- set reasoning_content = content.split(think_end_token)[0].strip('\n').split(think_begin_token)[-1].strip('\n') %}
|
||||
{%- set content = content.split(think_end_token)[-1].strip('\n') %}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
|
||||
{%- if reasoning_content -%}
|
||||
{#- Render thinking for every assistant turn (all-turn visible) -#}
|
||||
{{- think_begin_token ~ reasoning_content ~ think_end_token }}
|
||||
{%- else -%}
|
||||
{#- No thinking rendered → prefix with think_end_token -#}
|
||||
{{- think_end_token }}
|
||||
{%- endif -%}
|
||||
|
||||
{%- if content -%}
|
||||
{{- content }}
|
||||
{%- endif -%}
|
||||
{%- if message.tool_calls -%}
|
||||
{{- toolcall_begin_token ~ '\n' }}
|
||||
|
||||
{%- for tool_call in message.tool_calls -%}
|
||||
{%- if tool_call.function -%}
|
||||
{%- set tool_call = tool_call.function -%}
|
||||
{%- endif -%}
|
||||
{{- ns_token + '<invoke name="' + tool_call.name + '">' }}
|
||||
{%- set _args = tool_call.arguments -%}
|
||||
{%- for k, v in _args.items() if v is not none %}
|
||||
{{- ns_token + '<' + k + '>' -}}
|
||||
{{- to_xml(v, ns_token) -}}
|
||||
{{- ns_token + '</' + k + '>' }}
|
||||
{%- endfor -%}
|
||||
{{- ns_token + '</invoke>' ~ '\n' }}
|
||||
{%- endfor -%}
|
||||
|
||||
{{- toolcall_end_token }}
|
||||
{%- if message.tool_calls[-1].function -%}
|
||||
{%- set last_tool_call.name = message.tool_calls[-1].function.name -%}
|
||||
{%- else -%}
|
||||
{%- set last_tool_call.name = message.tool_calls[-1].name -%}
|
||||
{%- endif -%}
|
||||
{%- else -%}
|
||||
{%- set last_tool_call.name = none -%}
|
||||
{%- endif -%}
|
||||
{{- eos_token }}
|
||||
{%- endgeneration -%}
|
||||
{{- '\n' }}
|
||||
|
||||
{%- elif message.role == 'tool' -%}
|
||||
{%- if last_tool_call.name is none -%}
|
||||
{{- raise_exception("Message has tool role, but there was no previous assistant message with a tool call!") }}
|
||||
{%- endif -%}
|
||||
{%- if loop.first or (conversation_messages[loop.index0 - 1].role != 'tool') -%}
|
||||
{{- bos_token ~ 'tool' }}
|
||||
{%- endif -%}
|
||||
{{- '\n<response>' }}
|
||||
{%- if message.content is string -%}
|
||||
{{- message.content }}
|
||||
{%- else -%}
|
||||
{%- for tr in message.content -%}
|
||||
{%- if tr is mapping and tr.type is defined and tr.type == 'image' -%}
|
||||
{{- image_token }}
|
||||
{%- elif tr is mapping and tr.type is defined and tr.type == 'video' -%}
|
||||
{{- video_token }}
|
||||
{%- else -%}
|
||||
{{- tr.output if tr.output is defined else (tr.text if tr.type == 'text' and tr.text is defined else tr) }}
|
||||
{%- endif -%}
|
||||
{%- endfor -%}
|
||||
{%- endif -%}
|
||||
{{- '</response>' }}
|
||||
{%- if loop.last or (conversation_messages[loop.index0 + 1].role != 'tool') -%}
|
||||
{{- eos_token ~ '\n' -}}
|
||||
{%- endif -%}
|
||||
|
||||
{%- elif message.role == 'user' -%}
|
||||
{{- bos_token ~ 'user' ~ '\n' }}
|
||||
{{- visible_text(message.content) }}
|
||||
{{- eos_token ~ '\n' }}
|
||||
{%- endif -%}
|
||||
{%- endfor -%}
|
||||
|
||||
{#- Generation prompt -#}
|
||||
{%- if add_generation_prompt -%}
|
||||
{{- bos_token ~ 'ai' ~ '\n' }}
|
||||
{%- if thinking_mode is defined and thinking_mode == "disabled" -%}
|
||||
{{- think_end_token }}
|
||||
{%- elif thinking_mode is defined and thinking_mode == "adaptive" -%}
|
||||
{#- adaptive: no prefix, let model decide -#}
|
||||
{%- elif thinking_mode is defined and thinking_mode == "enabled" -%}
|
||||
{#- enabled or not defined: default to think -#}
|
||||
{{- think_begin_token }}
|
||||
{%- else -%}
|
||||
{#- adaptive: no prefix, let model decide -#}
|
||||
{%- endif -%}
|
||||
{%- endif -%}
|
||||
Reference in New Issue
Block a user