Faster, cheaper AI
How much of a large model’s memory can be dropped within a set error budget?
Running large models is expensive, and most ways of cutting the cost make no promise about what you lose. We decide, prompt by prompt, how much of a model’s memory can be dropped within a set error budget, and compress weights in ways that lose less.
What we’ve found
Compressing a model’s memory with a guarantee
Deciding, prompt by prompt, how much of a model’s memory (its key-value cache) can be dropped while staying within a set error budget.
Better 4-bit models
Rearranging a model’s weights in mathematically exact ways so it loses less quality when compressed to 4 bits.
Triton on NVIDIA’s GB10
An open-source add-on that makes Triton, a widely used tool for writing fast GPU programs, work on NVIDIA’s GB10 (Blackwell) desktop hardware until official support arrives.
Where it stops working
- Recent work from other groups uses the same family of weight rearrangements, so our next step is a head-to-head comparison before we claim an advantage.
Projects
Compressing a model’s memory with a guarantee
CompletedDeciding, prompt by prompt, how much of a model’s memory can be dropped while staying within a set error budget.
Better 4-bit models with exact rearrangements
CompletedRearranging a model’s weights in mathematically exact ways so it loses less quality when compressed.
Running Triton on NVIDIA’s GB10
CompletedAn open-source add-on that makes Triton work on NVIDIA’s GB10 (Blackwell) desktop hardware.