Amazon SageMaker Inference now offers prefix-aware routing, a routing strategy that sends requests sharing the same prompt prefix to the same instance so the KV cache stays warm. In benchmarks on Llama 3.1 70B, it reduc…
Compute supply, energy and data-center capacity decide how cheaply AI can run. Infrastructure shifts show up in inference costs weeks later.
Companies and models mentioned in this story — open their pages and live prices
Summaries are aggregated for information only — follow the source link for the full story. Demo entries are illustrative.