Skip to content

Topic

Serverless LLM inference

Running language-model generation inside short-lived, per-invocation cloud functions billed by memory and duration, instead of on dedicated GPU servers or through a hosted model API.

Current clusters