Skip to content

Topic

Model inference cost

What model API calls cost for a unit of work, measured per task or per run rather than per token, including the gap between larger and smaller models.

Current clusters