Skip to content

Topic

Inference determinism

Why identical prompts to the same model can yield different outputs, covering decoding settings, seeds, floating-point non-associativity on GPUs and request batching on hosted endpoints.

Current clusters