You’re read­ing The Brief­ing, Michael Wald­­­­­man’s weekly news­­­­­­­­­let­ter. Click here to receive it in your inbox. A year ago we warned that Donald Trump had a concerted strategy to undermine ...
Model weights keep growing, and the KV cache scales with context length multiplied by batch size. So 64 users at 1M context can mean roughly 935 GB of KV cache. Weights and cache together create a ...