Artifacts

Start with a model card.

An artifact is a saved model you can download and run. Its card names the runtime and hardware it needs. This catalogue links the published checkpoints represented in the explorer; it is not a quality ranking or a complete live account inventory.

Checkpoints with allocation maps.

These checkpoints use compressed-tensors with vLLM. Sizes below count the stored weights and their supporting numbers, including parts left unchanged. They come from the checkpoint's file headers, exclude accompanying files, and are not total serving-memory requirements.

Historical bpp labels identify the published recipes, not comparable denominators. Current bpp accounting uses quantizable parameters only. Use actual bytes and the same calibration contract for comparisons. The Explorer's tooltips use binary byte units; this table uses decimal GB.

Other serving paths.

The Hy3 GGUF model card documents a separate historical build. A load or coherent-generation check without a full-precision teacher is not a quality comparison.

Tessera is the third lane, with its own plugin and pinned runtime contract. The Tessera page covers the GLM-5.3-Flash export and its structural checks. Its similar file size does not establish better quality or speed than EXL3; the full serving comparison remains in progress. Read the Tessera evidence.

Allow room for the running model.

Weights are only part of memory use. The runtime also needs temporary work space and a cache for the conversation. A file that fits on disk may not fit the intended serving setup.

Quality measurements belong to their evaluation text and runtime. Different models, teachers or scoring methods should not be ranked by a shared column of numbers.

When measuring perplexity, use a serving setup without speculative decoding unless the scorer explicitly establishes which model's log probabilities it reads. A draft-model score is not a target-model score.