Skip to content

Topic

Distributed training

Splitting model training across multiple GPUs and nodes using data, tensor or sharded parallelism, together with the checkpointing and memory techniques that make it practical.

Current clusters