Skip to content

Topic

LLM cost and latency optimization

Techniques for reducing the tokens, model calls and wall-clock time an application spends per request, such as prompt trimming, caching, budget controls and routing work to cheaper model tiers.

Current clusters