Method
Measure the tradeoff before making the cut.
PrismaQuant treats each weight matrix as a choice. It estimates the damage from compressing that matrix, accounts for the space saved, then checks the whole model.
Start with representative text.
The model processes a calibration set: examples used to measure how its parts behave. A separate held-out set checks choices on text the estimate did not see. If the two sets overlap, the check cannot tell us whether a choice generalises.
Weights are only part of the calculation. The temporary values passed between layers, called activations, can also lose precision. The cost estimate must account for the calculation the serving software will actually perform.
Spend memory where it helps.
The allocator compares each option's estimated damage and stored size. A small matrix can be cheap to leave untouched. A large one may justify more compression. Some matrices must share a format because the runtime combines them into one operation.
The target is a byte budget. Parts the recipe keeps unchanged still occupy space in the download, even when they are outside the allocator's choices.
What the estimate means
AURA uses end-to-end KL-Fisher probes and production-rendered residuals. KL measures changes in the model's predicted token distribution; it is not a task score. Activation-aware costs must use the executed activation contract, rather than assuming the stored weight format determines it.
Bits per parameter (bpp) is reported over quantizable parameters only. Fixed source-precision regions are excluded from that denominator but included in the file size. Historical artifact labels can use older accounting conventions.
Check the file people will run.
A lower predicted error is a reason to test a candidate. It is not proof of a better model. Export checks must cover the stored bytes, the runtime that reads them, and the route used to calculate an answer.
Quality checks include the model's predictions on held-out text and downstream tasks where available. An average can hide a small set of badly damaged prompts, so tail behaviour matters too. Past failures explain these rules.
Serving gates are lane-specific
compressed-tensors targets vanilla vLLM. GGUF adds packing checks against gguf-py. Tessera uses an exact producer/runtime commit and packaged-contract digest. Evidence is scoped to the selected format, shape, hardware, residency and execution mode; unsupported combinations fail closed.
Per-layer Tessera rate allocation also needs a served comparison against a byte-matched uniform control. A tensor reconstruction screen cannot clear that gate.
Keep execution out of the operator's head.
The campaign declares its inputs and dependencies. PrismaBuild dispatches the work and records the results. An operator can fix a failed implementation, but should not have to decide by hand which layer runs next.
Sources: PrismaQuant architecture and design guidelines.