Faster, cheaper AI

How much of a large model’s memory can be dropped within a set error budget?

Running large models is expensive, and most ways of cutting the cost make no promise about what you lose. We decide, prompt by prompt, how much of a model’s memory can be dropped within a set error budget, and compress weights in ways that lose less.

Compressing a model’s memory within a budget A diagram. A row of memory entries, one per token; some are kept and some dropped. Below, the error caused by dropping entries stays under a set budget. Memory: one entry per token read kept dropped Error from dropped entries budget
A diagram, not data. A model’s memory gains an entry for every token it reads. We decide, prompt by prompt, how many entries can be dropped while the error stays within a set budget.

What we’ve found

  • Compressing a model’s memory with a guarantee

    Deciding, prompt by prompt, how much of a model’s memory (its key-value cache) can be dropped while staying within a set error budget.

  • Better 4-bit models

    Rearranging a model’s weights in mathematically exact ways so it loses less quality when compressed to 4 bits.

  • Triton on NVIDIA’s GB10

    An open-source add-on that makes Triton, a widely used tool for writing fast GPU programs, work on NVIDIA’s GB10 (Blackwell) desktop hardware until official support arrives.

Where it stops working

  • Recent work from other groups uses the same family of weight rearrangements, so our next step is a head-to-head comparison before we claim an advantage.

Projects

  • Compressing a model’s memory with a guarantee

    Completed

    Deciding, prompt by prompt, how much of a model’s memory can be dropped while staying within a set error budget.

  • Better 4-bit models with exact rearrangements

    Completed

    Rearranging a model’s weights in mathematically exact ways so it loses less quality when compressed.

  • Running Triton on NVIDIA’s GB10

    Completed

    An open-source add-on that makes Triton work on NVIDIA’s GB10 (Blackwell) desktop hardware.