Sensitivity
Some changes matter more than others.
A model does not react equally to every small error in its weights. Sensitivity measures how much a change in one part can affect the model's predictions. It helps identify where compression needs care.
Fragility is only part of the cost.
A sensitive matrix may still be easy to represent in a compact format. Another may be less sensitive but much larger. The allocator considers the error from the format and the bytes it saves, as well as sensitivity.
The temporary values flowing into a layer matter too. Compressing those activations can change the output even when two choices reconstruct identical weights.
Read the views together.
Technical detail opens an interactive panel. Compare the measured sensitivity, the estimated cost of a format, and the assignment in a published checkpoint. These are different quantities; a bright cell does not by itself mean a bad choice.
The Qwen3.8-27B panel pairs a cost stage with the allocation it produced. The Qwen3-4B panel is an earlier research cost stage with no paired published allocation.
loading…
The cost-to-sensitivity ratio is a diagnostic, not an isolated reconstruction error for the AURA objective. The Qwen3.8 panel includes activation-side cost. The Qwen3-4B residuals use round-to-nearest rather than the production render.
A changed decision still needs a served comparison.
In the historical PrismaAQUA experiment, including activation cost changed the allocation. The weight-only control remained a solve: it was not exported and served. That demonstrates a different decision, not a served quality gain caused by the extra term.