Skip to content

Topic

Neural network optimizers

Algorithms that turn gradients into weight updates during training, including SGD, Adam and AdamW, along with the research on their convergence behaviour, memory cost and per-step compute.

Current clusters