Skip to content

Topic

AI inference efficiency

Work aimed at reducing the compute cost and latency of running trained models in production, as distinct from the cost of training them.

Current clusters