Skip to content

project

tpu-inference

vLLM's TPU backend, which runs models on Google Cloud TPUs through JAX.

Current clusters