中文版

什么是分布式梯度累积(Distributed Gradient Accumulation)?

在深度学习训练中,梯度累积(Gradient Accumulation)是指将多个小批次(mini-batch)上的梯度累加起来,然后再进行一次梯度更新的技术。这样做可以有效缓解由于显存限制导致的小批次大小受限的问题,并提高训练的稳定性。

分布式梯度累积(Distributed Gradient Accumulation) 是指在多台机器(或多张GPU)上进行的梯度累积。在分布式训练中,每个GPU都计算自己本地小批次的梯度,然后进行累积。通常,所有GPU会共享梯度累积的状态,直到达到预设的累积步数后,再进行一次全局的梯度更新。

gradient_accumulation_steps 参数

在训练过程中,gradient_accumulation_steps 是一个关键的超参数,它指定了在执行一次梯度更新之前,进行多少个小批次(mini-batch)梯度的累积。

假设在一次训练中,你使用的小批次(batch size)为 B,而你设置的 gradient_accumulation_steps 为 N,那么实际的有效批次大小(effective batch size)就是:

effective batch size=B×N \text{effective batch size} = B \times N effective batch size=B×N

例如,如果你的 train_batch_size 设置为 8,并且 gradient_accumulation_steps 设置为 4,那么每次模型更新时,实际的有效批次大小将是 32(8 × 4)。这个技术能够让你以较小的批次大小进行计算,从而避免显存溢出的问题,同时又能模拟较大的批次大小,从而提高模型训练的稳定性和泛化能力。

为什么需要使用分布式梯度累积?

在微调大模型(如 LLaMA 2 7B)时,单个GPU的显存可能不足以容纳大的批次或模型参数。为了在有限的显存中运行训练,我们可以利用 梯度累积 技术,使得每个GPU使用较小的批次来计算梯度,并多次累积后再执行梯度更新。

gradient_accumulation_steps 设置建议

1. 显存的限制

使用 3090 显卡进行 LLaMA 2 7B 模型的微调时,显存限制是一个重要的考虑因素。3090显卡有24GB的显存,对于大型模型,显存往往是瓶颈。如果你想使用较大的有效批次而不改变单个小批次大小,你可以通过增加 gradient_accumulation_steps 来实现。

2. 设置示例

假设你的 LLaMA 2 7B 模型需要的批次大小较大,但单个 3090 显卡无法直接容纳,你可以通过以下方式调整:

  • 设定单个小批次的大小为 2(train_micro_batch_size_per_gpu = 2),这将确保每个GPU只处理两个样本。
  • 然后,你可以将 gradient_accumulation_steps 设置为 16,这样就会在 16 个小批次后进行一次梯度更新。

这样,尽管每个小批次的大小只有 2,但实际的有效批次大小将会是 32(2 × 16)。通过累积梯度,你可以模拟一个较大的批次大小,而不会超出显存的限制。

{
    "train_batch_size": 32,  # 总的批次大小,通常在分布式训练中由 NUM_PROCESSES * PER_DEVICE_TRAIN_BATCH_SIZE 来决定
    "gradient_accumulation_steps": 16,  # 累积 16 个小批次的梯度后再进行一次更新
    "train_micro_batch_size_per_gpu": 2  # 每个GPU的小批次大小
}
3. 根据显存调整

在 3090 显卡上训练 LLaMA 2 7B 模型时,显存可能会因使用不同的微调数据集和训练配置而有所不同。为了找到最合适的 gradient_accumulation_steps,你可以通过实验调整不同的 train_micro_batch_size_per_gpu 和 gradient_accumulation_steps 参数:

  • 如果你的显存还充裕,可以试着增加 train_micro_batch_size_per_gpu,并减少 gradient_accumulation_steps。
  • 如果显存不够用,可以减少 train_micro_batch_size_per_gpu,同时增加 gradient_accumulation_steps 来保持较大的有效批次大小。

如何选择最优的 gradient_accumulation_steps?

  1. 观察显存使用情况:在训练过程中监控显存使用情况,确保不会超过显存的限制。如果显存即将溢出,可以考虑减小 train_micro_batch_size_per_gpu,增加 gradient_accumulation_steps。
  2. 适当调整批次大小:虽然较大的有效批次有助于提高模型性能,但也需要根据你的显存来选择合适的批次大小。
  3. 通过实验找到平衡点:没有固定的规则,最佳的 gradient_accumulation_steps 通常通过多次实验来确定。你可以尝试不同的设置,观察训练速度和性能。

代码示例

假设我们要微调 LLaMA 2 7B 模型,并且已经配置了分布式训练,下面是如何设置 gradient_accumulation_steps 和 train_micro_batch_size_per_gpu 的一个例子:

# 设置训练参数
MODEL_NAME=meta-llama/Llama-2-7b-hf
PER_DEVICE_TRAIN_BATCH_SIZE=2
GRADIENT_ACCUMULATION_STEPS=16
TRAIN_BATCH_SIZE=32  # 总批次大小 = NUM_PROCESSES * PER_DEVICE_TRAIN_BATCH_SIZE

# 使用 accelerate 启动训练
accelerate launch \
  --mixed_precision bf16 \
  --num_machines 1 \
  --num_processes 4 \
  --machine_rank 0 \
  --main_process_ip 127.0.0.1 \
  --main_process_port 29400 \
  --use_deepspeed \
  --deepspeed_config_file configs/deepspeed_config.json \
  --model_name_or_path $MODEL_NAME \
  --per_device_train_batch_size $PER_DEVICE_TRAIN_BATCH_SIZE \
  --gradient_accumulation_steps $GRADIENT_ACCUMULATION_STEPS \
  --train_batch_size $TRAIN_BATCH_SIZE \
  --learning_rate 5e-6 \
  --output_dir output/llama2-7b \
  --num_train_epochs 3 \
  --dataset_name allenai/tulu-3-sft-mixture

总结

  • 分布式梯度累积 可以帮助我们在显存有限的情况下训练大模型,通过累积多个小批次的梯度来模拟一个大的批次。
  • gradient_accumulation_steps 控制了梯度累积的步数,较大的累积步数可以模拟较大的批次,从而提高模型训练的稳定性和效果。
  • 在训练 LLaMA 2 7B 模型时,合适的设置 gradient_accumulation_steps 和 train_micro_batch_size_per_gpu 可以有效地利用显存,避免显存溢出,同时保持较好的训练性能。

英文版

What is Distributed Gradient Accumulation?

In deep learning training, gradient accumulation refers to the technique of accumulating gradients over multiple mini-batches before updating the model weights. This helps mitigate the issue of memory limitations when training on large models or datasets by effectively simulating larger batch sizes while using smaller mini-batches.

Distributed gradient accumulation refers to the same technique applied in distributed training across multiple GPUs or nodes. Each GPU computes gradients for its local mini-batch, and these gradients are accumulated over multiple steps. After the specified accumulation steps are reached, the gradients are averaged and applied to update the model weights across all GPUs.

What is gradient_accumulation_steps?

The parameter gradient_accumulation_steps specifies how many mini-batches’ worth of gradients should be accumulated before performing an update. If you are using a batch size B and set gradient_accumulation_steps to N, the effective batch size becomes:

Effective batch size=B×N \text{Effective batch size} = B \times N Effective batch size=B×N

For example, if your train_batch_size is 8 and gradient_accumulation_steps is set to 4, the effective batch size becomes 32 (8 × 4). This technique allows for simulating a larger batch size without running into memory limitations.

Why is Distributed Gradient Accumulation Needed?

When training large models (like LLaMA 2 7B), single GPUs often do not have enough memory to fit large batches or even the model parameters themselves. Gradient accumulation allows you to accumulate gradients over multiple smaller mini-batches, and then perform a single update. This way, you can simulate a large effective batch size without exceeding memory limits.

How to Set gradient_accumulation_steps for LLaMA 2 7B Model on a 3090 GPU

Let’s consider the case where you are fine-tuning a 7B parameter model (like LLaMA 2 7B) on an NVIDIA 3090 GPU, which has 24 GB of memory. Due to the size of the model and the available memory, you may not be able to fit a large batch into memory. Instead, you can adjust gradient_accumulation_steps to accumulate gradients over multiple small batches.

1. Memory Constraints

The 3090 GPU has a limited amount of memory. When training large models like LLaMA 2 7B, you will often hit memory limits with a large batch size. To handle this, you can use gradient accumulation to simulate a larger batch size while keeping the individual mini-batch size small.

2. Example Settings

Let’s say you decide to use a small mini-batch size of 2 (i.e., train_micro_batch_size_per_gpu = 2) for each GPU, but you want an effective batch size of 32. To achieve this, you can set gradient_accumulation_steps to 16. This means you will accumulate gradients over 16 mini-batches before performing a single update.

{
    "train_batch_size": 32,  # Total batch size, which is NUM_PROCESSES * PER_DEVICE_TRAIN_BATCH_SIZE
    "gradient_accumulation_steps": 16,  # Accumulate gradients over 16 mini-batches before updating
    "train_micro_batch_size_per_gpu": 2  # Mini-batch size per GPU
}

In this setup, despite each GPU processing only 2 samples at a time, the effective batch size for the model is 32 (2 × 16), which is a much larger batch size than what would fit into memory.

3. Adjust Based on Memory Usage

You can experiment with different train_micro_batch_size_per_gpu and gradient_accumulation_steps values based on the available memory. If the memory usage is too high, reduce the train_micro_batch_size_per_gpu and increase gradient_accumulation_steps. Conversely, if there’s ample memory, you can try increasing the mini-batch size and reducing the accumulation steps.

How to Choose the Right gradient_accumulation_steps?

  1. Monitor Memory Usage: Keep an eye on memory usage during training. If memory is close to being exhausted, you can either reduce the mini-batch size (train_micro_batch_size_per_gpu) or increase the number of gradient accumulation steps (gradient_accumulation_steps).

  2. Balance Batch Size and Memory: While larger batch sizes can improve performance and stability, it is essential to balance them with the available GPU memory.

  3. Experiment: There is no one-size-fits-all formula for the best gradient_accumulation_steps. Typically, you will need to experiment with different values and observe the impact on training speed, memory usage, and model performance.

Example Code

Let’s assume that you are fine-tuning the LLaMA 2 7B model and you want to configure gradient accumulation to fit within the memory constraints of your 3090 GPU. Here is an example configuration for launching training with Accelerate:

# Set training parameters
MODEL_NAME=meta-llama/Llama-2-7b-hf
PER_DEVICE_TRAIN_BATCH_SIZE=2
GRADIENT_ACCUMULATION_STEPS=16
TRAIN_BATCH_SIZE=32  # Total batch size = NUM_PROCESSES * PER_DEVICE_TRAIN_BATCH_SIZE

# Launch training with accelerate
accelerate launch \
  --mixed_precision bf16 \
  --num_machines 1 \
  --num_processes 4 \
  --machine_rank 0 \
  --main_process_ip 127.0.0.1 \
  --main_process_port 29400 \
  --use_deepspeed \
  --deepspeed_config_file configs/deepspeed_config.json \
  --model_name_or_path $MODEL_NAME \
  --per_device_train_batch_size $PER_DEVICE_TRAIN_BATCH_SIZE \
  --gradient_accumulation_steps $GRADIENT_ACCUMULATION_STEPS \
  --train_batch_size $TRAIN_BATCH_SIZE \
  --learning_rate 5e-6 \
  --output_dir output/llama2-7b \
  --num_train_epochs 3 \
  --dataset_name allenai/tulu-3-sft-mixture

Conclusion

  • Distributed gradient accumulation helps overcome memory limitations by allowing us to simulate larger batch sizes while using smaller mini-batches.
  • The gradient_accumulation_steps parameter controls how many mini-batches’ worth of gradients are accumulated before performing a weight update. By adjusting this parameter, you can effectively manage GPU memory usage and train large models even with limited memory.
  • When training large models like LLaMA 2 7B on a 3090 GPU, experimenting with different gradient_accumulation_steps and train_micro_batch_size_per_gpu values will help you balance memory constraints and training efficiency.

With these adjustments, you can effectively train large models without exceeding GPU memory limitations, improving both training stability and performance.

后记

2024年11月29日13点17分于上海,在GPT4o大模型辅助下完成。

Logo

北京人形旗下天工造物具身智能开源社区,聚焦具身天工与慧思开物两大平台

更多推荐