build1 publisher
Routing on the prompt's first tokens cut AWS's median time to first token by up to 77 percent
SageMaker Inference now sends requests that begin with the same tokens to the same instance. On AWS's seven-node Llama 3.1 70B benchmark that lifted the KV cache hit rate from roughly 25 percent to 82 percent.
Publishers:aws.amazon.com
Reality
- Evidence48
- Adoption22
- Hype gap+22
- Incentives85
- Confidence52