BENCHMARKS · AUGUST 2026

Every number, measured.

Compression is lossy by construction. So a performance figure is published here only if the same run proves retrieval intact.

One missed passkey and the result is discarded.

Reproduce on your hardware → python fraqtl_repro_receipts.py  # rented A100 is enough
9
users at 128K each, one A100
fp16 holds 2 · FP8 receipt-valid rung: 4 · 14.9 tok/s/user (fp16: ~33 for its 2)
2.42×
more KV tokens per GPU
997,200 vs 412,544 · Mistral-7B
1.34×
fp16 decode speed at 128K
95–99% of fp16 at 8K–32K
405/405
needle cells retrieved, exact match
3 models · every arm · zero misses
THE GATE WHY IT HOLDS R1 · VLLM SERVING R2 · MODEL MATRIX R3 · LLAMA.CPP R4 · QUALITY & FOOTPRINT ARTIFACTS
Receipt index · pick your depth · nothing forced on scroll
2026-08-14 SERVING CAPACITY
Nine 128K users on one A100 — 134.1 tok/s aggregate, 2× fp16, 9/9 needles
ap-pBrqbRbG2UJqb5HvQDbMg4 · ap-U44I0iXVh2vrYX9ALCqIeB
2026-07-03 SERVING CAPACITY
The 128K inversion: 1.34× fp16 decode, single user (decode-1024 arbiter)
ap-QuvidvohDU0DVYEKUomCKt
2026-07-05 SERVING CAPACITY
189/189 needle grids — 7 depths × 3 keys, all arms, 8K/32K/128K
ap-zjyaiI4b7Rl4VujTuozW7O · ap-utQpsypNBJQrwrXHzRqhTN
2026-05-26 HI-FI · WEIGHTS
Qwen3.6-35B at 23.8 GB — MMLU within 0.16pp of FP16, passkey 30/30 @125K
huggingface.co/fraQtl/Qwen3.6-35B-A3B-compressed
2026-07-22 HI-FI · WEIGHTS
Gemma-4 E2B + 26B-A4B Hi-Fi GGUFs — calibrated, KLD receipts vs bf16
huggingface.co/fraQtl · Hi-Fi Quants collection
2026 H1 KV SUBSTRATE RESEARCH
KIVI / KVQuant / TOVA / SnapKV bake-offs, cross-arch NIAH + PPL tables
C44B ledger · full tables in the archive
2026-05 LEGACY · LLAMA.CPP
D1: Mistral-7B 128K below Q8 VRAM with 5/5 retrieval where Q8 drops to 1/5
llama.cpp lane · superseded for speed by the vLLM runtime
Boundaries — what is NOT claimed here
  • FP8 beyond 4 users: higher rungs did not pass the receipt gate (r2 audit) — no claim either way.
  • 256K: capacity verified, quality gate still open — no claims yet.
  • H100 / consumer GPUs: runtime is SM80 (A100) only today.
  • GPTQ-on-Phi-3, AWQ 3/5-bit, MoE matched-bits vs AWQ/GPTQ: not attempted — no implicit claim. Details in the archive.
  • Blanks in any table are runs not done, not failures.
THE GATE

Low-bit KV buys memory by losing the answer. Ours doesn't.

Mistral-7B at 128K in llama.cpp. Same model, same prompts — only the KV representation changes. Five needles buried at five depths; each dot is one retrieved.

fp16 KV
22,657 MiB
5 / 5
Q8 KV
15,437 MiB · −31.9%
1 / 5
Q4 KV
11,287 MiB · −50.2%
0 / 5
fraQtl D1
13,261 MiB · −41.5%
5 / 5

Mistral-7B-Instruct · 128K context · llama.cpp · NIAH = 5-needle retrieval · live VRAM is measured resident KV footprint.
fraQtl D1 sits 14.1% below Q8 on memory and keeps 5/5 where Q8 keeps 1/5. Lower memory and intact retrieval at once.

WHY IT HOLDS

Quantize the wrong basis and you quantize the answer. So we change basis first.

Per-token integer quantization spends its bits uniformly across a basis where the signal is not uniform — which is why retrieval is the first thing to go. fraQtl rotates each head's K and V into their own eigenbasis, protects the leading directions, and quantizes the rest. The packed result is read fragment-native into the tensor cores: no dequantization pass, no reconstruction buffer.

ONE HEAD, ONE PAGE · WHAT THE MEMBRANE DOES TO IT
01 · RAW KV, fp16

Signal spread across every direction. Nothing to cut cheaply.

→
02 · EIGENBASIS ROTATION
v₁ v₂

Per-head rotation, calibrated in 0.3s. Variance collapses onto a few axes.

→
03 · PACKED MEMBRANE
.

Protected rows kept at rank. The tail goes to INT3 with sign correction.

k = 16
protected directions per head, V+K INT3
0.3 s
calibration per model — weights untouched
0
reconstruction passes — read fragment-native
σ = 0
bit-identical across seeds 7 / 123 / 2024

The V/K theorem is validated at architecture level on Mistral 7B GQA-4, Qwen 2.5 3B GQA-2, Llama 3.2 3B and Phi-3 — the same recipe transfers because the rotation is derived per head, not per model. Eviction methods (TOVA, SnapKV, StreamingLLM, H2O) drop tokens; per-token integer methods (KIVI, KVQuant) keep every token in the original basis. fraQtl keeps every token and changes the basis.

R1 · VLLM RUNTIME FLAGSHIP

Alone it matches fp16. Under load it pulls away.

The runtime reads packed KV pages straight into the tensor cores — no reconstruction step. At one user, weights dominate and everyone lands within a few percent. Add users and fp16 runs out of KV memory first, fp8 second. Two models, one kernel, zero per-model changes.

CONCURRENT USERS · 128K CONTEXT EACH · ONE A100-80GB
Qwen3-4B-2507 · bit-plane V2 · Aug 14
fraQtl
134.1 tok/s
9/9 needles
fp8 KV
117.6 tok/s
peak at 4 users
fp16
66.6 tok/s
ceiling at 2 users

Each block is one user holding a full 128K context (130,048-token prompt + 1,024-token decode) with an independent passkey at a distinct depth. Prefill at 128K: ~3,950 tok/s, at parity with fp16.

KV POOL CAPACITY · MISTRAL-7B
fraQtl 997,200
fp8 KV 825,104
fp16 412,544

Measured pool tokens, 8K config, one A100-80GB. 2.42× fp16.

DECODE SPEED VS CONTEXT
% of fp16, same hardware — the speed win grows with context
fp16 = 100% 134% 123% 8K 32K 128K

Solid: Qwen3-4B-2507. Mid: Llama-3.1-8B. Faint: Mistral-7B (32K native).

FULL SERVING TABLE
CUDA graphs on · temperature 0 · prefix caching off · fp8 always shown
Cell fraQtl fp16 fp8 KV Needles
Mistral-7B · 8K · batch 1 · decode tok/s 85.7 90.39 90.05 ✓
Mistral-7B · 32K · batch 1 · decode tok/s 76.84 77.98 82.81 ✓
Qwen3-4B-2507 · 128K · batch 1 · decode tok/s 69.5 52.59 72.7 ✓
Mistral-7B · KV pool tokens · one A100 997,200 412,544 825,104 —
Qwen3-4B-2507 · users @128K each · bit-plane V2 9 2 5 9/9
Qwen3-4B-2507 · aggregate tok/s at 9 users 134.1 66.6 117.6 peak, at 4 ✓
Qwen3-4B-2507 · users @128K each · V1 layout 6 2 5 every user
Qwen3-4B-2507 · 32K · 24 concurrent users · agg tok/s 429.7 246.0 432.3 24/24
RETRIEVAL VERIFICATION · EVERY ARM, EVERY RECEIPT

Exact-match passkey grids (7 depths × 3 keys per context) run on every arm of every receipt; a single missed passkey discards the result. Qwen3-4B-2507: 189/189 cells across 8K/32K/128K. Mistral-7B-v0.3: 126/126 across 8K/32K. Llama-3.1-8B: 90/90 — 405/405 total, plus the Aug 14 nine-user run's 9/9 per-user passkeys at depths spanning 0–100%. Same kernel for all three models; a new model's sidecar builds in under an hour.

METHOD & CAVEATS — full conditions, known bugs, what is not claimed expand ▾

BF16/FP16 is the quality reference; public Q4 is the practical size baseline teams already deploy. Matched Q4_K_M bakeoff in progress — results published when locked. Method: vLLM real paged serving, CUDA graphs on, prefix caching off, two-pass subtractive decode timing, temperature 0, exact-match passkey retrieval per arm. fp8 = fp16 weights + fp8-E4M3 KV (FlashInfer), the strong baseline; weights identical across all arms. fraQtl = rank-protected eigenbasis KV read fragment-native into the tensor cores, zero reconstruction. 128K receipts run with chunked prefill disabled (known perf bug under chunking, fix in progress; stock arms unaffected). A batch-routing bug that dropped needles at batch ≥2 was found 07-03, root-caused, fixed, and fully re-receipted — pre-fix numbers were never published. No 256K claims yet: capacity verified, quality gate open. Qwen3-4B KV pool: 845,136 tokens vs fp16 361,776 (fp8 709,488) — 2.3× capacity. Aug 14 nine-user cells use bit-plane V2 KV tiering (+45% KV pool vs V1); the July V1-layout receipt remains the per-layout bandwidth roof — V2 trades ~5% aggregate for 3 more users. Per-user decode at 9 users is 14.9 tok/s (fp16: ~33, for its 2 users). All cells reproducible from the public Hugging Face sidecar repos with fraqtl_repro_receipts.py; run IDs available on request.

R2 · MODEL MATRIX

One kernel, every model, zero per-model changes.

Each model gets a sidecar from the factory in under an hour, weights untouched, and runs on the identical membrane kernel. Same command, same recipe, three architectures, one pattern.

Qwen3-4B-Instruct-2507
GQA · native 256K
189 / 189 NEEDLE CELLS
Llama-3.1-8B-Instruct
GQA-8 · native 128K
90 / 90 NEEDLE CELLS
Mistral-7B-v0.3
GQA-4 · native 32K
126 / 126 NEEDLE CELLS
Model KV vs fp16 vs fp8 8K 32K 128K
Qwen3-4B-Instruct-2507
2025 · native 256K
2.39× 1.21× 92% 97% 1.34×
Llama-3.1-8B-Instruct
native 128K
2.35× 1.20× 96% 97% 1.23×
Mistral-7B-v0.3
native 32K
2.42× 1.21× 95% 99% n/a — collapses in all arms incl. fp16

Decode columns are % of fp16 at that context. Needle cells = exact-match passkey grids, 7 depths × 3 keys per context (Llama 5 × 3), all three arms. Capacity = measured KV pool tokens, 8K config. Cells are receipts; blanks are runs not done, not failures.

R4 · QUALITY & FOOTPRINT HI-FI WEIGHTS

fp16 OOMs at 64K. We run 128K with 28 GB spare.

The public Qwen 3.6 35B-A3B artifact on a single A100-80GB. Same model, same hardware, eight times the context — the difference between needing one GPU and needing two.

VRAM FOOTPRINT · QWEN 3.6 35B-A3B · ONE A100-80GB
—— 80 GB CEILING
16K CONTEXT both fit on one GPU
fp16 ~71 GB
25.6 GB
64K CONTEXT fp16 needs 2 GPUs · fraQtl 1
82.9 GB → OOM
36.8 GB
128K CONTEXT 28 GB still free
85+ GB → 2 GPUs
51.7 GB

fp16 at 64K = 82.86 GB measured (OOMs on the 80 GB ceiling). 16K and 128K fp16 figures are model + KV extrapolations, conservative lower bounds. fraQtl figures measured with the public artifact at huggingface.co/fraQtl/Qwen3.6-35B-A3B-compressed.

QUALITY VS FP16 BASELINE
Qwen 3.6 35B-A3B compressed · single seed
Benchmark fp16 fraQtl Delta
MMLU 0-shot 82.40% 82.24% −0.16 pp
∞Bench Passkey @ 125,315 tokens 30 / 30 30 / 30 parity
HumanEval pass@1 reference 100% retention within noise
Wikitext-2 PPL Δ — +0.033 tight

3-seed bit-identical verification on PPL: σ = 0 across seeds 7 / 123 / 2024. Architecture-level 3-seed validation (V/K theorem) on Mistral 7B GQA-4, Qwen 2.5 3B GQA-2, Llama 3.2 3B, Phi-3.

RESEARCH BACKING — partial-stack bake-off vs KIVI / TOVA / SnapKV / H2O expand ▾

A separate research comparison, not the customer claim. Partial-stack, Mistral / Llama, sub-4-bit, C44 bake-off.

Metric fraQtl Nearest published Difference
NIAH retention 1080 trials 98.5% 97.8% · TOVA +0.7 pp
PPL Δ at 3.5× compression +0.012 +0.214 · SnapKV ~18× tighter
NIAH at 128K Llama 3.1 8B 100% 0% · KIVI-2 +100 pp
GPUs needed at 128K 35B MoE 1× A100-80GB 2× A100-80GB half the hardware

Honest note: at matched 4-bit, fraQtl ties KIVI-4 / KVQuant-4 within sampling noise. The wins above are at sub-4-bit, at 128K, against eviction methods, and on hardware footprint. Sources: Mistral-7B-Instruct C44 bake-off (1080 NIAH trials, 8K–31K, 3 needle types); Llama 3.1 8B C44d/e (128K, 3 needles × 5 trials per depth).

QUALITY EVIDENCE · QUOTED VERBATIM “On the 2026-06-11 hardened paired packet, D2 located 63/63 NIAH needles (fp16-parity); observed failures are rare exact-ID transcription noise (~5%) and synthetic 3-hop state-tracking flips (~17%, symmetric), both root-caused; real multi-hop QA (LongBench) shows no F1 gap.” D2-family recipe lineage · direct evidence for the current recipe: the 9/9 grid above
PUBLIC ARTIFACTS

Everything above is downloadable. Nothing is a demo.

Every receipt traces to a public repo on huggingface.co/fraQtl — KV sidecar kits with the runtime and repro script, and calibrated Hi-Fi GGUFs with KLD receipts.

qwen3-4b-instruct-2507-kv-sidecars KV KIT
The sidecars behind the 9-user flagship · 189/189 grids · ships with fraqtl_repro_receipts.py
mistral-7b-instruct-v0.3-kv-sidecars KV KIT
2.4× fp16 KV capacity · 126/126 grids · independently reproduced Jul 3
fraqtl-sm80-runtime RUNTIME
The vLLM runtime wheel (A100/SM80) plus repro script — free for evaluation
Qwen3.6-35B-A3B Hi-Fi · compressed HI-FI GGUF
35B MoE at 23.8 GB · 128K on one A100 · MMLU within 0.16 pp of fp16
Gemma-4-E2B-it · Gemma-4-26B-A4B-it Hi-Fi HI-FI GGUF
Calibrated Hi-Fi GGUFs with KLD receipts vs bf16 — the phone-demo lineage
Collections: Hi-Fi Quants · KV-Cache Receipts ORG
Two curated collections — calibrated GGUFs with KLD receipts, and the KV receipts lineage
ARCHIVE — research bake-offs, cross-architecture weight matrix, 2026-H1 receipts expand ▾

Superseded numbers preserved for the record: the matched-4-bit weight-compression matrix across MHA + GQA-2 + GQA-3 + GQA-4, the 9/9 matched-bits peer wins, per-row scripts and commit hashes, and the H1 cross-architecture tables. Restyled to this system in the production file — kept short here so the mock stays readable.

Run it against your own workload.

The runtime wheel and repro script are public. A rented A100 is enough to check every number on this page.

Get a pilot Browse the artifacts