mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
## What does this PR do? **Type of change:** New example **Overview:** Modify launch_train.sh script to enable multi-node training. Provide a slurm template script. ## Usage Add required fields in slurm.sh. Then use the command below to submit multi-node job: ```bash bash slurm.sh ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added multi-node training support with new configuration options for distributed setups. * Introduced Slurm batch script for streamlined job submission to cluster environments. * **Improvements** * Enhanced GPU resource management for multi-GPU and distributed training configurations. * Updated speculative decoding model support and validation handling. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com>
58 lines
1.8 KiB
Bash
58 lines
1.8 KiB
Bash
#!/bin/bash
|
|
|
|
# SPDX-FileCopyrightText: Copyright (c) 2024 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
|
# SPDX-License-Identifier: Apache-2.0
|
|
#
|
|
# Licensed under the Apache License, Version 2.0 (the "License");
|
|
# you may not use this file except in compliance with the License.
|
|
# You may obtain a copy of the License at
|
|
#
|
|
# http://www.apache.org/licenses/LICENSE-2.0
|
|
#
|
|
# Unless required by applicable law or agreed to in writing, software
|
|
# distributed under the License is distributed on an "AS IS" BASIS,
|
|
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
# See the License for the specific language governing permissions and
|
|
# limitations under the License.
|
|
|
|
#SBATCH -A {account}
|
|
#SBATCH --job-name={job_name}
|
|
#SBATCH --nodes={num_nodes} --ntasks-per-node=1 --gpus-per-node={num_gpus_per_node}
|
|
#SBATCH -p {partition}
|
|
#SBATCH -t {time_limit}
|
|
|
|
CONTAINER_IMAGE={container_image}
|
|
WORK_DIR={path_to_modelopt}
|
|
|
|
CONTAINER_MOUNT="${WORK_DIR}:/modelopt"
|
|
|
|
OUTPUT_DIR={path_to_output_dir}
|
|
MODEL={path_to_model_dir}
|
|
DATA={path_to_data_dir}
|
|
OFFLINE_DATA={path_to_offline_data_dir}
|
|
|
|
CMD="./launch_train.sh --model $MODEL \
|
|
--output_dir $OUTPUT_DIR \
|
|
--data $DATA \
|
|
--num_epochs 1 \
|
|
--train_bs 1 \
|
|
--lr 1e-4 \
|
|
--eagle_config eagle_config.json \
|
|
--training_seq_len 4096 \
|
|
--save_steps 1000 \
|
|
--estimate_ar True \
|
|
--disable_tqdm True \
|
|
--offline-data $OFFLINE_DATA \
|
|
--num_nodes $SLURM_NNODES \
|
|
--head_node_ip $head_node_ip \
|
|
"
|
|
|
|
srun -l \
|
|
--mpi=pmix \
|
|
--output=%x_%j_$DATETIME.log \
|
|
--container-workdir "/modelopt/examples/speculative_decoding" \
|
|
--container-image ${CONTAINER_IMAGE} --container-mounts ${CONTAINER_MOUNT} \
|
|
bash -lc "$CMD"
|
|
|
|
set +x
|