Tessera

Smaller weight files, built for the GPU.

Tessera is Robert Tand's weight format, with its own vLLM plugin. It stores a compact description of a sequence of weights, then reconstructs the values needed for calculation. The expensive search happens when the model is encoded.

Storage precision and calculation precision are different choices.

A weight used in a higher-precision calculation does not have to occupy that many bits in the file. Tessera separates those choices. This gives the allocator more room to adjust size without forcing every part of the model to use the same calculation format.

The technique is called trellis coding. Instead of rounding each weight independently, the encoder searches for a good sequence of reconstructed values. Decoding follows that sequence. It does not recover the original weights losslessly.

Formats and runtime boundary

The families reconstruct NVFP4, FP8 or BF16 values, with W4A4, W8A8 or W16A16 arithmetic respectively. W and A name weight and activation precision during computation, not stored bits per weight. The wire accountant includes tables, scales, padding and metadata.

The checkpoint's quant_method: "tessera" selects tessera.serving. TESSERA_SERVE_MODE=resident|streamed is explicit. PrismaQuant pins an exact Tessera commit and the digest of its packaged runtime_contract.json; a package name alone is not qualification.

A separate serving lane, with its own checks.

Tessera sits alongside compressed-tensors and GGUF in PrismaQuant. It installs its plugin into stock vLLM rather than requiring core patches. That does not make it a universal reader: each supported combination needs evidence for its hardware and runtime.

Historical small-model serving tests have demonstrated useful compression and low measured disagreement with a source model. Those results belong to the tested artifacts. They do not qualify every current encoder or serving route.

The packaged contract is the authority on supported combinations. A route test and a whole-model quality result answer different questions.

Tessera and EXL3 · evidence as of September 26, 2026

Similar file size. More to establish.

EXL3 is another trellis-based weight format. Tessera's new GLM-5.3-Flash export is slightly smaller than the published EXL3 reference when both are counted the same way: every file, including model weights and supporting metadata.

Quality and speed superiority over EXL3 are not established. Export checks concern the saved artifact. They do not show that the model gives better answers or runs faster.

Like-for-like size comparison

All regular files, in bytes; not just tensor payload or weight shards
ArtifactAll-file size
Tessera GLM-5.3-Flash, new R1024-MTP export175,643,087,583 B
GLM-5.3-Flash EXL3 reference175,790,111,275 B

The difference is 147,023,692 B. This compares stored files, not per-rank serving memory or inference workspace.

The current Tessera body mostly uses compressed BF16-grid weights. That does not imply an FP4 calculation speed advantage. Its routed MTP projections, which help draft future tokens, also use compressed BF16-grid weights; they are not uncompressed BF16 storage. The EXL3 reference's routed MTP projections are EXL3-compressed.

What the original experiment found

The historical shared-row experiment favoured EXL3. It measured how closely selected expert matrices reproduced their original outputs. An expert is a part of the model used for selected tokens. This was a test of individual matrices, not a whole-model serving comparison, and it does not measure the new export above.

The corrected September 1 screen used six routed-expert gate/up projections, layers 5, 20 and 42, expert 0, with the last 1,024 capture rows held out. Tessera-4's output-error ratio to EXL3 K4 was 1.176× on the weight leg and 1.070× with both evaluated under the experimental A4 projection. Lower is better; EXL3 was ahead.

The EXL3 arm freshly encoded the source weights. Its A4 column was an experimental activation perturbation, not EXL3's serving contract. The earlier headline based on a different reference probe is retired and is not used here.

Corrected original measurement.

Full serving comparison: in progress.

The head-to-head serving comparison on the full GLM-5.3-Flash artifact is still in progress. The recorded EXL3 harness work found request-dependent outputs inside one running engine, so its aggregate score is not a valid comparison result.

A fair result needs a stable reference, the same evaluation text, and checked serving settings on the same hardware. We do not yet have an admissible end-to-end speed comparison. The older projected size saving is not used here; the file counts above belong to the new export.

For now, the supported conclusion is narrow: a Tessera export that passed its structural checks, at a similar all-file size, with quality and speed still to compare.