Skip to content

Topic

Inference Parameters

The runtime settings a caller passes to a language model API or serving stack to shape generation, including temperature, top_k, top_p, min_p and repetition penalty.

Current clusters

build1 publisher

Halving temperature turns a 4:1 token preference into 16:1

Shrijith Venkatramana's walkthrough of sampling puts the odds-ratio arithmetic behind temperature on the page. It shows how much of the difference between two runs of one prompt is settled after the model has finished computing.

Publishers:dev.to

Reality

Evidence58
Adoption
Insufficient
Hype gap+8
Incentives30
Confidence55