Distributed Data Parallel (DDP)

DDP is a PyTorch module that allows you to parallelize your model across multiple machines, making it perfect for large-scale deep learning applications.

Resources

# create model and move it to GPU with id rank
model = ToyModel().to(rank)
ddp_model = DDP(model, device_ids=[rank])
  • a Broadcast is done to make sure the initial model parameters are the same

During training, gradients are synchronized (averaged) across processes after each backward pass, so all models remain in sync.

During training (after intialization), do we ever check again that the values are still the same?

DDP does not periodically compare the parameters on every rank to make sure they are still equal. It establishes equality initially, then relies on the distributed algorithm to preserve that invariant