Limits

Keep the result no broader than the test.

A smaller file can still produce worse answers. A good average can hide a failure on the prompt that matters. These are the limits to keep in mind when reading this site.

Different checks answer different questions.

A tensor screen

Tests the error in selected parts of a model. Useful for comparing encoders, but it does not measure the whole model's answers or serving speed.

A load and generation check

Shows that a runtime can open a file and produce text under the tested settings. Coherent text alone does not establish retained quality.

A served quality comparison

Compares the running artifact against a reference on specified text or tasks. Its result belongs to that setup, not to every model using the same format.

A size projection

Estimates what a future artifact might occupy. Until the bytes are exported and checked, it is not a shipped saving.

Results the project withdrew.

Local quality screens have sometimes improved while the served model got worse. Grouped cost estimates, staged rendering and damping sweeps all produced claims that failed a stronger comparison. Those claims were withdrawn rather than promoted.

The original Tessera-versus-EXL3 headline also used a reference measured on a different probe. The corrected experiment favoured EXL3. Passing structural checks on the new Tessera export does not establish a quality or speed advantage.

Read the corrected comparison and its conditions.

What remains unproven.

  • Full GLM-5.3-Flash serving comparison: in progress. The recorded EXL3 repeatability issue prevents treating its aggregate score as a settled baseline.
  • Missing teachers limit quality claims. Some large historical artifacts could be loaded but lacked a full-precision reference that fit the target hardware. Their load checks are not quality scores.
  • Activation-aware allocation needs an isolated comparison. The historical PrismaAQUA solve changed when activation cost was included. Its weight-only control was not served, so that exercise did not establish a served advantage caused by the extra cost term.
  • Historical support is not current qualification. A changed encoder, serving route or runtime needs fresh evidence. A plugin supporting one combination does not qualify all combinations.

Metric and accounting limits

KL based on a truncated teacher distribution is a lower bound, not full-vocabulary KL. Perplexity measures prediction of the evaluation text, not every downstream task. Report the corpus and the tail as well as the mean.

Historical bits-per-parameter labels use different denominators. Current accounting excludes fixed regions from quantizable parameters; total file size still includes them. Compare actual bytes and matched evaluation conditions.

Cross-layer solvers and post-frontier polish described in older pages are archived research, not the current production stage graph.

Records before headlines.

Measurements need a source revision, an artifact identity and the conditions under which they ran. PrismaBuild receipts make execution inspectable; they do not turn a narrow experiment into a broad claim.

For downloads, use the individual model card's requirements. For the format's current serving scope, use Tessera's packaged runtime contract.

Sources: project working agreement and Tessera architecture.