Research direction II · proposal v7 · updated 20 September 2026
Ternary weights cost quality. The question is where the information goes instead.
BitFuse is a staged research programme on the Qwen3 architecture — trained at the real Qwen3-0.6B scale (~596M), with Qwen3-4B retired — under MLP-only 1.58-bit ternary FFN weights. Its core thesis: when the static parameter channel is aggressively compressed, sequence mixing, residual pathways and conditional computation stop being independent plug-ins and become alternative information carriers competing for the same freed budget. The code, the fairness audit, the instrumentation and the Azure runbook exist; what is missing is GPU hours.
What changed in v7. Ternarization is now MLP-only —
the FFN/SwiGLU weight matrices go to 1.58-bit ternary and nothing else does, so the
base variant is ternary_mlp, not a BitNet reproduction — and the campaign
trains at the real Qwen3-0.6B scale (~596M) on every machine, with
Qwen3-4B retired. The variant set grew to thirty-six. The DeepSeek
work is two branches that never share a table: the sliding-window +
CSA + HCA schedule with a Lightning Indexer is the V4 legacy control
(G, H, deepseek_impl="v4_hybrid"), and the DeepSeek-V4.1-Flash
family (DS41-1…DS41-7) isolates CSA2, the Causal Encoder-Decoder, the
hierarchical sparse indexer, single-pass mHC and a research-scale Engram one mechanism
at a time. Three more families join as their own comparison sets: the Kimi
K3 line (NoPE Gated MLA, SiTU-GLU, Attention Residuals, Stable LatentMoE), the
Qwen3.8-Flash-Next line (an n-gram conditional memory on the GDN+QSA
mixer) and a KDA×QSA cross-family that pairs Kimi's recurrent memory
with Qwen's sparse retrieval. The residual sweep is six first-class families
(channel-gated, Qwen GR, DeepSeek mHC in iterative R3 and single-pass R5 forms, Kimi
AttnRes). Two ternary-recipe refinements — a learned per-channel
scale and magnitude-aware int8 activations — and a sparse top-2 MoE FFN
over eight ternary experts join as Phase-6 ablations, each in its own table. And an
Apple-Silicon (MPS) dev backend plus a reduced-width benchmark harness
let the whole 36-variant sweep run locally before any GPU spend.
The hypothesis, stated so it can lose
Ternarization flattens the network's internal representations. Where a full-precision layer has a few large activations carrying most of the signal, a ternary layer's activations are more uniform. That loss of contrast is the mechanism behind the quality drop. If so, mechanisms that restore contrast — a state-based attention with per-channel gating (KDA), a state-space mixer (Mamba-2), Qwen's fixed-size memory paired with block-indexed sparse retrieval, compressed long-range attention, a widened or manifold-constrained residual stream, or an attention over depth itself — should recover quality, and the recovery should be visible in the statistics, not just in the loss curve.
BitNet b1.58 constrains every weight to \( \{-1, 0, +1\} \) — about 1.58 bits each instead of sixteen, so a weight costs roughly a tenth of a byte to store. The published recipe leaves one slot conspicuously untouched: BitNet still runs full softmax attention with RoPE, and its own technical report names “investigating efficient attention mechanisms suitable for low-bit models” as future work. BitFuse is that investigation — but it borrows only the ternary technique: it quantizes the FFN/SwiGLU weight matrices alone (MLP-only) and keeps attention, embeddings, the tied head and every sequence mixer full-precision Qwen, so it is not a BitNet reproduction. And the newest literature keeps strengthening the original thesis rather than replacing it: Qwen, DeepSeek and Kimi have each shipped a different answer to “which dynamic carrier should do the work”, at full precision, without anyone testing whether the ranking survives ternarization.
Six carriers, one of them crippled
It helps to treat the network as a set of distinct information carriers, and to notice that ternarization is not symmetric across them. Ternarization is primarily a reduction in per-weight resolution, not a total loss of magnitude information — the model keeps a full-precision latent parameter and derives its ternary forward weight through a scale factor. So the useful question is not “how much accuracy is lost?” but “which channels become more valuable when static weight resolution is constrained?”
FFN weights — heavily degraded
Persistent learned knowledge — and only the FFN/SwiGLU weights are ternary; attention, embeddings and the sequence mixer stay full-precision. A static report reads the {−1,0,+1} distribution of those FFN weights straight from any checkpoint — zero density, sign balance, ternary and sign entropy, scale factors, per-layer quantization error, and the BITCOS storage rate — per parameter, per module and per depth, no forward pass required.
Activations — untouched
Still continuous, still expressive. Magnitude, variance, contrast, effective rank. Not a control knob — normalization eats deliberate rescaling, so these are read as evidence, never tuned.
Attention — untouched and scalable
Which tokens interact and what gets retrieved. Entropy, max probability, top-k mass, output norm — and now QSA and CSA2 indexer recall and coverage, selected-block mass, and compressed-pool visibility statistics.
Recurrent state — a fixed-size memory
What a KDA, Mamba-2 or Gated DeltaNet layer has absorbed. State norm, update ratio, cosine retention, and GDN/KDA retention-and-forgetting curves.
Residual — promoted to an axis
No longer one arm but six families. Input/update norm ratio, plus GR branch usage, mHC mixing entropy and doubly-stochastic deviation (iterative R3 and single-pass R5), and AttnRes depth weights.
Prediction & modality — the newest axes
Two carriers the core six do not touch. Auxiliary-token losses, acceptance rate and speculative speedup for MTP; cross-modal probe accuracy, image-token saliency and modality-conditioned norms for vision.
That asymmetry is the opening. A ternary transformation cannot express fine distinctions in its weights, but the dynamic carriers can still create sharp differences in what each token reads, remembers and preserves. If a low-precision network is short of contrast, those are where it can still be manufactured — and the organizing principle of the whole programme follows: compress static weight expressivity, measure where the information moves, then reinvest the freed budget in whichever mechanism delivers the most quality per unit of memory or compute.
The hypotheses, H1–H8
Eight claims, written before the runs.
Three cover the core ablation; five more follow from the residual programme, the predictive objective and the multimodal extension, each making a claim the original three do not. Each one is scoped so a null result is informative rather than merely disappointing.
Dynamic reallocation
Ternarization changes where useful information is carried. If the sequence mixer is also too restrictive, quality may fall further than necessary — richer dynamic computation may compensate for constrained weights.
Dynamic-capacity compensation
A ternary model may want a different allocation of global, local, recurrent and compressed-attention capacity than a full-precision one. This has to be measured, not assumed.
Residual preservation
Residual topology can preserve or re-route information that would otherwise be attenuated, overwritten or diluted by a deep sequence-mixing stack.
Heterogeneous context roles
Local retrieval, compressed retrieval, fixed-size recurrent memory and precise global attention may divide the representational work more efficiently than any single mixer.
Depth-wise information selection
Kimi-style Attention Residuals may help ternary models by letting each layer select which earlier representations deserve to contribute — turning depth into an additional dynamic channel.
Topology-stable multi-stream residuals
DeepSeek mHC and Qwen Gated Residuals may preserve multiple pathways while avoiding the instability of unconstrained hyper-connections.
Predictive compression and acceleration
MTP can make more of the same backbone useful per forward pass, potentially improving speculative-decoding acceptance and backbone quality without altering the carrier thesis.
Modality as a carrier
Multimodal inputs test whether a ternarized backbone can preserve cross-modal information while leaning harder on dynamic representations, connectors and residual pathways.
The falsification rule, made concrete
Every mechanism has to predict where the information moves.
This is what separates an explanation from a coincidence, and it is the single most important methodological commitment on the page. Each mechanism enters with a prediction about which carriers should gain and which should give way. If quality improves and the carriers do not move as predicted, the explanation is treated as refuted — the gain is real but unexplained, which is a different and weaker claim.
The trade, in bytes
Why dynamic capacity has to be very good to be worth buying.
Before any experiment runs, the arithmetic already says something uncomfortable. Ternary weights are so cheap that every attention, residual or prediction-head parameter costs many backbone parameters — and an order of magnitude more again if it stays at higher precision, as an indexer or a residual gate has to. That is the bar every hypothesis on this page has to clear, and it is why BitFuse holds the shared core budget fixed to within ~0.1% on the primary chain: otherwise the better architecture would just be the bigger one.
This control is a design space, not a result. The horizontal axis of the real experiment is this same split; the vertical axis — quality — is what the measurements have to supply.
The experiment
Thirty-six models. One class. One switch at a time.
All of them share one Qwen3 skeleton — per-head QK-norm, a decoupled head_dim of 128, unbiased attention projections, SwiGLU, RMSNorm, RoPE, tied embeddings — and all of them are one implementation driven by config switches, not thirty-six model files. Divergent copies are how “identical except for X” quietly stops being true. The campaign trains at the real Qwen3-0.6B scale (~596M) on every machine — Apple-silicon and GPU alike, with Qwen3-4B retired — so the geometry and every per-variant count below are the live 0.6B numbers: 28 layers, hidden 1024, 16 query heads, 8 KV heads, intermediate 3072. Each mixer's head dims are re-solved for that width so core-matching holds. New mechanisms enter as separate expansion sets, so the primary table never fills up with partially controlled comparisons. Every backbone is matched within ~0.1% on core parameters on the primary chain (the Kimi K3 family lands a touch under, −1.0%, and the KDA×QSA cross-family a touch over, +2.6%, reported as-is) — the MoE variants (Q, R, K3-5, Q38-3) and the Engram / n-gram–bearing variants (DS41-6, DS41-7, Q38-1…3) on their active count, since their expert banks and hash tables are stored but only sparsely accessed.
| Set | Name | Sequence mixer | Residual topology | Primary question | |
|---|---|---|---|---|---|
| A | core | fp_baseline |
softmax, all layers · FP | standard | the full-precision reference |
| B | core | ternary_mlp |
softmax, all layers | standard | what MLP-only ternarization costs |
| C | core | ternary_mlp_kda |
3 KDA : 1 global | standard | does KDA recover the ternary loss? |
| D | core | ternary_mlp_kda_attnres |
3 KDA : 1 global | channel-gated | do low-cost residual gates add anything? |
| E | core | ternary_mlp_hybrid |
24 Mamba-2 : 4 global (Nemotron-H) | standard | does Mamba-2 survive ternarization? |
| F | core | ternary_mlp_hybrid_attnres |
24 Mamba-2 : 4 global | channel-gated | do residuals help Mamba-2 too? |
| G | 4-B | ternary_mlp_deepseek |
swa / csa / hca + Lightning Indexer (V4 legacy) | standard | does V4 legacy compressed attention help? |
| H | 4-B | ternary_mlp_mamba_deepseek |
Mamba-2 + csa / hca (V4 legacy) | standard | does V4 legacy attention complement recurrence? |
| I | 4-Q | ternary_mlp_qwen_hybrid |
3 GDN : 1 QSA (Qwen3.8) | standard | does recurrent memory + selective retrieval fit ternarization? |
| J | 4-Q | ternary_mlp_qwen_hybrid_res |
3 GDN : 1 QSA | channel-gated (R1) | the GDN + QSA residual-sweep baseline |
| K | 4-C | ternary_mlp_res_gr |
3 GDN : 1 QSA | Qwen GR (R2) | does a widened gated residual stream help? |
| L | 4-C | ternary_mlp_res_mhc |
3 GDN : 1 QSA | DeepSeek mHC (R3) | does constrained multi-stream routing beat gating? |
| M | 4-C | ternary_mlp_res_attnres |
3 GDN : 1 QSA | Kimi AttnRes (R4) | does depth-wise selection add a separate carrier? |
| N | 5-P | ternary_mlp_mtp |
best Phase-4 mixer | best topology | does predictive supervision improve quality or only serving? |
| O | 6-T | ternary_mlp_learned_scale |
softmax, all layers | standard | does a learned per-channel scale recover the ternary loss? |
| P | 6-T | ternary_mlp_magact |
softmax, all layers | standard | do magnitude-aware int8 activations help? |
| Q | 6-E | ternary_mlp_moe |
softmax + sparse MoE FFN | standard | is freed ternary budget better spent on experts? |
| R | 6-E | ternary_mlp_gdn_moe |
3 GDN : 1 QSA + MoE FFN | standard | does MoE stack with the best dynamic mixer? |
| DS41-2 | 4-D | ternary_mlp_csa2 |
CSA2 · full/reindex/reuse, flat indexer | standard | does CSA2 alone beat plain softmax? (RQ20) |
| DS41-1 | 4-D | ternary_mlp_ced |
Causal Encoder-Decoder · shared global KV | standard | does the encoder-decoder split help? (RQ19) |
| DS41-3 | 4-D | ternary_mlp_ced_csa2 |
CED + CSA2 · flat indexer | standard | the CED × CSA2 interaction (RQ21) |
| DS41-4 | 4-D | ternary_mlp_ced_csa2_hsi |
CED + CSA2 · hierarchical indexer | standard | hierarchical vs flat sparse indexer (RQ22) |
| DS41-5 | 4-D | ternary_mlp_ced_csa2_mhcsp |
CED + CSA2 · hier. | Single-Pass mHC (R5) | does the single-pass residual add on top? (RQ23) |
| DS41-6 | 4-D | ternary_mlp_ds41_engram |
softmax + Engram memory | standard | does n-gram conditional memory help? (RQ24) |
| DS41-7 | 4-D | ternary_mlp_ds41 |
CED + CSA2 (hier.) + Engram | Single-Pass mHC (R5) | the integrated V4.1-Flash model (RQ25) |
| R5 | 4-C | ternary_mlp_res_mhc_sp |
3 GDN : 1 QSA | Single-Pass mHC (R5) | single-pass vs iterative mHC on a frozen mixer (RQ23) |
| K3-1 | K3 | ternary_mlp_k3_gmla |
3 KDA : 1 NoPE Gated MLA | standard | swap the KDA global layer for a latent-KV Gated MLA — vs ternary_mlp_kda |
| K3-2 | K3 | ternary_mlp_k3_attnres |
3 KDA : 1 Gated MLA | Kimi AttnRes (R4) | + depth-wise Attention Residuals on the gmla base |
| K3-3 | K3 | ternary_mlp_k3_situ |
3 KDA : 1 Gated MLA · SiTU-GLU | standard | + the SiTU-GLU bounded activation on the gmla base |
| K3-4 | K3 | ternary_mlp_k3_core |
3 KDA : 1 Gated MLA · SiTU-GLU | Kimi AttnRes (R4) | the integrated dense Kimi K3 model |
| K3-5 | K3 | ternary_mlp_k3_latentmoe |
3 KDA : 1 Gated MLA · SiTU-GLU + LatentMoE | Kimi AttnRes (R4) | dense FFN → Stable LatentMoE (own table) |
| Q38-1 | Q38 | ternary_mlp_qwen_ngram |
3 GDN : 1 QSA + n-gram memory | standard | n-gram conditional memory alone on the GDN + QSA base |
| Q38-2 | Q38 | ternary_mlp_qwen_core |
3 GDN : 1 QSA + n-gram memory | Qwen GR (R2) | the integrated dense Qwen3.8-Flash-Next model |
| Q38-3 | Q38 | ternary_mlp_qwen_moe |
3 GDN : 1 QSA + n-gram memory | Qwen GR (R2) | + a sparse MoE FFN (own table) |
| XF-1 | XF | ternary_mlp_kda_qsa |
3 KDA : 1 QSA | standard | KDA remembering + QSA retrieving — beat either parent mixer? |
| XF-2 | XF | ternary_mlp_kda_qsa_attnres |
3 KDA : 1 QSA | Kimi AttnRes (R4) | + depth-wise Attention Residuals on the cross-family mixer |
the chain
Each step changes exactly one component: full precision → ternary → KDA → attention residual. Anything that moves is attributable to the step that moved it.
the branch
E and F branch off B rather than continuing from C: they swap in a different linear-attention family under identical ternary weights. So C and E are rivals, not successive steps — C vs E asks which family survives ternarization better.
Phase 4 — four expansion sets, reported separately
4-B is the DeepSeek-V4 legacy control (G, H — swa/csa/hca + Lightning Indexer), kept as a historical baseline; 4-Q adds Qwen's GDN + QSA (I, J); 4-C freezes that mixer and sweeps the residual topology (K/L/M, plus R5's single-pass mHC); 4-D is the new DeepSeek-V4.1-Flash family (DS41-*), isolating CSA2, the Causal Encoder-Decoder, the hierarchical indexer, single-pass mHC and Engram one at a time. The V4 and V4.1 branches never share a table. Each set has its own comparison set and report, so the six-way core ablation cannot quietly become a many-way one asking a different question.
Phases 5–6 — extensions and ablations, not rivals to the core
N adds multi-token prediction on the winning dense architecture, changing the training objective. O–R are Phase-6 ablations: two ternary-recipe refinements (O, P) and a sparse MoE FFN (Q, R). None belongs in the Phase 2–4 perplexity table — each gets its own outcome.
Three more families, each on its own comparison set
The Kimi K3 line (K3-1…K3-5) rebuilds the KDA stack toward Kimi's
design — a NoPE Gated MLA global layer, SiTU-GLU, Attention Residuals and a Stable
LatentMoE — scored against ternary_mlp_kda. The
Qwen3.8-Flash-Next line (Q38-1…Q38-3) adds an n-gram conditional
memory, then the Gated Residual, on the GDN + QSA base. The KDA×QSA
cross-family (XF-1, XF-2) pairs Kimi's recurrent “remember” with Qwen's
sparse “retrieve”. Each is its own set, never folded into the six-way core.
Why the Mamba ratio is not 3:1. The KDA variants use 3 KDA : 1
global, following Kimi Linear. The Mamba-2 variants deliberately do not
mirror that, because copying NVIDIA's architecture means copying their placement.
Nemotron-H's published hybrid_override_pattern is 52 single-purpose
layers — 24 Mamba-2, 4 attention, 24 MLP, so attention is only 7.7% of layers. A
BitFuse block holds a mixer and an MLP, so counting mixers only gives 28 —
which lines up one-to-one with BitFuse's 28-layer depth, so attention lands at
mixer indices 4, 10, 16, 22 (4 of 28). Matching NVIDIA matters more here
than matching Kimi.
Where the mechanisms come from
Six borrowed mechanisms, each with the experiment that tests it.
These are external literature inputs, used to motivate experiments. None of them is a claim that BitFuse has already reproduced the effect. What each card says is: this is what the source contributes, why it bears on the carrier thesis, the carrier signature that has to move for its explanation to hold — a quality gain without that movement is evidence against the mechanism — and the specific comparison that would settle it here.
Gated DeltaNet + Qwen Sparse Attention
GDN compresses history into a fixed-size recurrent state; QSA runs a lightweight indexer that aggregates tokens into micro-blocks and selects the important regions for sparse global attention. Qwen reports a 3:1 GDN:attention pattern.
This is the BitFuse reallocation question stated in someone else's architecture: how should representational work split between recurrent memory and selective retrieval?
carrier → GDN state norm and retention, the QSA selected-token fraction, and the n-gram store'sqwen_ngram/hit_rate and qwen_ngram/gate_mean — the recall or retrieval gain is unexplained unless these move.
test → ternary GDN + QSA vs ternary KDA, plain attention and DeepSeek CSA2, matched parameters and tokens
CSA2 — Compressed Sparse Attention
One attention that reads a local sliding window at full resolution and a compressed global pool chosen by a hierarchical indexer — a coarse block filter, then a fine top-64 — all in a single softmax. Adjacent layers share the compressed KV, the candidate pool and the top-k across a Full / Reindex / Reuse schedule.
The new DeepSeek branch — the Phase 4-D V4.1-Flash family (DS41-*), kept separate from the DeepSeek-V4 legacy control (G, H). It isolates CSA2, the Causal Encoder-Decoder, the hierarchical indexer, single-pass mHC and Engram one at a time: does compressed sparse attention stay valuable after ternarization?
carrier → thecsa2/* indexer recall and selected-block mass, the Full/Reindex/Reuse reuse rates, the shared encoder→decoder KV, and engram/hit_rate — with the genuine-sparsity invariant guarding against selection quietly going dense.
test → run the V4.1-Flash family against plain ternary softmax, Qwen GDN+QSA and the KDA/Mamba hybrids at matched parameters and tokens
Gated Residual (GR)
The residual stream is widened to four branches with dynamic, content-dependent read gating and per-branch write control.
Adds multiple residual information channels without requiring a separate attention mechanism to do it.
carrier → GR branch usage and per-branch norm share, read-gate entropy — branch specialization has to appear, not just a lower loss. test → replace the channel-gated residual with a parameter-matched GR branch; probe whether branch specialization tracks the carrier metricsmHC — manifold-constrained hyper-connections
Hyper-connections widen the residual stream and route information dynamically; mHC constrains the residual mixing matrix to the doubly-stochastic manifold, restoring identity-map behaviour and stabilizing training.
A principled multi-stream topology whose stability is enforced rather than hoped for — which is exactly what makes it a fair rival to GR.
carrier → residual mixing entropy and deviation from double stochasticity — iterative R3 vs single-pass R5 row deviation — stable multi-stream transport must show up in the residual carrier. test → mHC vs GR vs baseline at matched capacity; log residual mixing entropy and singular-value/norm behaviourAttention Residuals
AttnRes replaces fixed residual accumulation with input-dependent attention over preceding layer outputs. Kimi also pairs KDA with global MLA.
This makes depth itself an information-selection mechanism, which is the carrier hypothesis applied along an axis the other four families do not touch.
carrier → KDA state magnitude and decay, the MLA gate opening (gated_mla/gate_mean), and AttnRes depth entropy — depth has to become an active selection channel, not a passive skip.
test → full or block AttnRes as a separate residual family; measure depth-wise contribution concentration, output norms, gradient distribution
Multi-token prediction, and native multimodality
MTP trains future-token prediction heads; Qwen reports multi-step training aimed at consistency with speculative decoding, and Kimi K3 retains an MTP layer. Both models are also multimodal — Kimi trains language and vision jointly in one shared context.
Together they ask whether the carrier framework survives a change of training objective and a change of input space.
carrier → auxiliary-token loss and speculative-decoding acceptance for MTP; cross-modal probe accuracy and image-token saliency for vision — the new carrier must carry signal, not merely change the objective. test → Phase 5-A: 1/2/4 future-token heads on the best dense model · Phase 5-B: frozen vision encoder and projector first, earlier fusion only if the backbone stays stablePhase 4-Q · remember + retrieve
Qwen's split: one layer kind to remember, another to retrieve.
Qwen3.8-Flash-Next puts Gated DeltaNet in most layers and Qwen Sparse Attention in periodic global ones, and frames it plainly as “remember” plus “retrieve”: GDN continuously compresses history into a fixed-size state, while QSA's indexer first aggregates tokens into micro-blocks and then selects the important regions to attend over. That is what makes it a genuinely different comparator from the existing KDA setup — the retrieval side is not dense global attention but sparse, block-indexed retrieval.
Implementation stance. Begin at the published 3:1 placement, keep the 28-layer budget and the existing parameter-matching framework, and run a ternary GDN path for the projection matrices wherever the quantization contract permits it. Architectural attribution comes first; aggressive kernel optimization can wait, because a fast implementation of an uninterpretable comparison is worth nothing.
The indexer is declared, not assumed. QSA's indexer parameters
are separated from the main ternary projection budget and the precision is
stated explicitly — the same reasoning that keeps DeepSeek's
index_q/index_k out of the name-matched ternary
conversion pass. Quantizing a ranking corrupts selection for a
negligible memory saving, and a corrupted ranking would be indistinguishable
from the mechanism simply not working.
Three families, one controlled axis. Experiment 4.1 puts KDA + global against GDN + QSA at identical token budget, parameter budget, optimizer and seed. Experiment 4.2 puts GDN + QSA against DeepSeek CSA2 at the same 0.6B scale. Then 4.3 sweeps attention share (DS-25/50/75), 4.4 sweeps placement with the attention count fixed, and 4.5 sweeps head and indexer budget with the structure fixed — because “how much retrieval capacity” and “where the retrieval layers sit” are two different questions and answering them together answers neither. And if GDN + QSA wins, the claim to make is about fixed-size memory plus selective retrieval under ternary weights — not about Qwen being better.
Phase 4-B legacy · Phase 4-D DeepSeek-V4.1-Flash family
Two DeepSeek branches, kept apart: a V4 legacy control, and the V4.1-Flash family.
The DeepSeek work is now two separate branches that never share a table.
The V4 legacy pair (G, H, deepseek_impl="v4_hybrid") keeps
the original sliding-window + Compressed + Heavily-Compressed attention with a Lightning
Indexer — a historical control, Phase 4-B. The Phase 4-D V4.1-Flash family
(DS41-*) is the current branch, adding one mechanism at a time on the ternary backbone:
CSA2 (illustrated below), the Causal Encoder-Decoder
(14 encoder layers, then decoder layers reading one shared global KV projected from the
encoder's final state), the hierarchical sparse indexer,
single-pass mHC (R5) and a research-scale Engram
n-gram memory. Each CSA2 layer reads two things in one softmax: a local
sliding window at full resolution, and a compressed global pool — four
tokens summarised into one latent — chosen by a two-level indexer (a coarse block filter,
then a fine top-64). What changes between CSA2 layers is not the compression level but how
much structure they build versus reuse: a full layer builds and publishes the
compressed KV, candidate pool and top-k; a reindex layer reuses the KV and
pool but recomputes its own top-k; a reuse layer reuses everything, which is
where the compute and memory saving comes from.
These layers are not “attention over fewer tokens.” A compressed path alone cannot resolve fine-grained local structure, so removing the local branch would change what the layer is for. What each one actually does is read a short local window at full resolution and a long context at reduced resolution, in one softmax:
┌── local window (full resolution, w=128) ───┐
hidden → ─┤ ├→ one softmax → out
└── compressed pool · hierarchical indexer ──┘
Causality is structural, not assumed. A compressed entry covering source positions \([\text{start}, \text{end})\) summarises token \(\text{end}-1\), so it is made readable only from that position onwards — and the indexer ranks after that mask, which is what guarantees its selection is always a subset of what dense causal attention would have seen.
Only complete blocks are emitted: a partial trailing block would summarise fewer tokens than its siblings and make the entries incomparable. Tests assert that perturbing the tail of a sequence leaves every earlier logit bit-identical, that the local window alone cannot see distant tokens while the compressed pool can, and that cached decoding matches a full forward pass exactly.
index_q/index_k so the name-matching ternary conversion pass does
not claim them — quantizing a ranking would corrupt selection for a
negligible memory saving. The single-pass mHC hyper-connection (R5) is a residual
question and gets its own place in the Phase 4-C sweep as well — see below.
Phase 4-C · a first-class axis
Six answers to “what should a layer be allowed to read?”
Since 2017 the residual pathway has been a single fixed addition from the layer directly beneath — a write, not a read. Under ternary weights every one of those transformations is lossy by construction, which is why the slot becomes worth reopening at 1.58 bits and not before. BitFuse tests six of them, and treats them as different algorithmic families rather than settings of one idea — because the distinction is what makes attribution possible at all. R3 and R5 are two takes on the same multi-stream idea: DeepSeek's mHC, projected either by iterative Sinkhorn (R3) or in a single normalization pass (R5, from V4.1-Flash).
| Family | Information path | Core operation | Expected signal | Main risk | |
|---|---|---|---|---|---|
| R0 | standard residual | 1 stream | x + F(x) |
identity preservation | depth-wise dilution |
| R1 | channel-gated (current) | 1 stream | F(x) + α(x)·x |
content-dependent skip strength | limited path diversity |
| R2 | Qwen Gated Residual | 4 branches | dynamic read gate + per-branch write control | parallel pathways, selective preservation | routing and parameter overhead |
| R3 | DeepSeek mHC | 4+ branches | pre/post routing + doubly-stochastic mixing | stable multi-stream transport, identity-like | projection cost, implementation complexity |
| R4 | Kimi Attention Residuals | depth-indexed | softmax over prior layer outputs | selective depth-wise aggregation | memory/communication cost; may need block approximation |
| R5 | Single-Pass mHC (V4.1-Flash) | 4 branches | one route_down→route_up pass, single row/column normalization | R3's transport at a fraction of the routing cost | looser normalization — rows only approximately unit |
R1 and R4 — reach. The current channel-gated residual modulates a single skip path at 2048 full-precision parameters per layer, initialized to zero so the module starts as a plain fixed residual and cannot destabilise early training. Kimi's AttnRes goes much further and makes the pathway a retrieval:
The full gated matrix variant of R1 would add ~29M parameters (+4.9%)
and make that model the biggest one, confounding the comparison it exists to run —
which is precisely the trap the parameter-control rule below exists to catch.
R2 and R3 — width. Both widen the residual state, but they differ in how information is routed and constrained, so a win for one implies nothing about the other. mHC keeps the mixing matrix on the doubly-stochastic manifold:
The constraint is the contribution, not an implementation detail: it restores
identity-like signal behaviour and keeps the recombination from drifting into an
unstable scaling regime as depth grows. If mHC beats GR, the result to report is
that constrained multi-stream routing trades stability against
expressivity better than unconstrained gating does. R5 is
V4.1-Flash's single-pass take on the same mHC: read/write and the 4×4 mixing
coefficients come from one route_down→route_up projection,
normalized once instead of by R3's iterative Sinkhorn — so R5 vs R3 (and R5 vs L on a
frozen mixer) isolates exactly what the iterative projection buys.
Why this comparison is scientifically useful. Widening the stream and selecting along depth are orthogonal moves — R4 expands what a layer can reach along depth, R2/R3/R5 expand how many parallel routes exist, and neither substitutes for the other. R1 stays in as the necessary low-complexity control, because the question “is multi-stream complexity actually required under ternarization?” has to have an answer. And a residual win only means anything mechanistically if the carrier metrics show the predicted change in norm, contrast, path utilization, depth contribution or branch specialization. To avoid confounding attention with topology, the sweep freezes the best two Phase-4 mixers and changes only the residual design.
RQ11 · the capacity question made testable
How much dynamic retrieval is worth buying at 1.58 bits?
Not an assumption baked into an architecture, but a sweep. Ternarization removes fine-grained expressivity from the weights, so a ternary model may benefit from more dynamic sequence-mixing capacity than a full-precision one would. Three axes, moved one at a time, because changing two at once makes the result uninterpretable — and placement and indexer budget are now swept separately too, since “how much retrieval” and “where it sits” are different questions.
DS-25 / DS-50 / DS-75
What fraction of layers do sequence mixing by attention rather than by recurrence. Attention is placed evenly, never clustered, so “how much attention” does not get confounded with “where the attention is”.
Experiment A — layer count
The number of attention layers moves; head configuration is frozen. Isolates “more attention layers” from “wider attention”.
Experiment B — head count
Head count moves with the layer count fixed, and head_dim is pinned to
hidden_size / heads — so total attention width and its parameter count
stay constant. That is what separates attention diversity from attention
capacity.
Experiment C — placement, then retrieval capacity
Where the dynamic retrieval layers occur with their count fixed, then how much indexer and head capacity they get with the structure fixed. That answers whether retrieval capacity or placement is the binding constraint — the question QSA's micro-block indexer makes newly worth asking.
Residual topology, and MoE
Both change something other than sequence mixing — the residual pathway is its own Phase 4-C sweep, and the sparse MoE FFN changes active-vs-stored parameters plus routing, so it is a Phase-6 ablation (Q, R). Folding either into this sweep would confound it. Kimi, Qwen and DeepSeek are all MoE at full scale; only their sequence mixers and residual topologies are borrowed here, never their expert routing.
Experiment B varies head geometry on purpose, which the fairness rules otherwise
forbid. That needs an explicit opt-in — --phase4 --allow-attention-budget —
and the differences are then recorded as declared, not waived. The comparison
is fair within the sweep, which is exactly why Phase 4 reports on its own
variant set.
What each comparison answers
Twenty-seven questions, each tied to a specific pair.
Written down before the runs, so a result cannot be retrofitted to a question that happens to have been answered.
RQ4 is the one that distinguishes explanation from coincidence, RQ7 is what the two rival families were added to settle, RQ19–RQ25 are the DeepSeek-V4.1-Flash family isolated one mechanism at a time, RQ5 is the Pareto question the others are instruments for, and RQ26–RQ27 look ahead to FP4 KV caching and speculative decoding.
RQ5, drawn
Five competing claims on the same freed memory, and a phase that owns each decision. Drag the split; the numbers are structural, not measured.
Phase 5-A · multi-token prediction
Multi-token prediction, treated as two experiments rather than one.
Qwen reports MTP as a native component of its Next/3.8 lineage, using multi-step training to keep training consistent with speculative decoding while also improving the backbone; Kimi K3 retains an MTP layer too. BitFuse therefore treats it as both a training-objective experiment and an inference-efficiency experiment — never as an ordinary architecture-only ablation, because those two outcomes can move in opposite directions and merging them would hide it.
| Design | Primary outcome | Secondary outcome | |
|---|---|---|---|
| MTP-1 | no MTP — the reference | PPL / benchmark | — |
| MTP-2 | predict 2 future tokens | backbone loss + acceptance rate | decode throughput |
| MTP-3 | predict 4 future tokens | backbone loss + acceptance rate | training overhead |
| MTP-4 | multi-step consistency schedule | speculative-decoding acceptance | stability / calibration |
| MTP-5 | MTP on the Qwen-vs-DeepSeek best mixer | architecture × prediction interaction | Pareto efficiency |
It has to be measured on a real serving engine. Acceptance rate and end-to-end throughput are the practical motivation for MTP, so they need an engine capable of speculative decoding rather than a synthetic estimate. And the report keeps language-model quality strictly separate from serving acceleration — otherwise a faster decoder gets mistaken for better training, which is the most likely way this particular experiment could produce a misleading headline.
Phase 5-B · multimodal
Does the carrier thesis survive leaving language?
Qwen3.8-Flash-Next is multimodal and Kimi K3 trains language and vision jointly in one shared backbone and context. That motivates a broader test — but the first target here is deliberately the smallest experiment that can answer whether a ternarized language backbone preserves multimodal signal at all. Four staged gates, each of which has to pass before the next is worth running.
Frozen vision encoder + trainable projector + ternary LM
Image–text instruction and captioning. Gate: does the backbone train stably and retain its text quality? If it does not, nothing downstream is interpretable.
gate · backbone stability + text-quality retentionPartially trainable vision adapter + ternary LM
Cross-modal reasoning. Gate: does information transfer improve without destabilizing the language model?
gate · transfer gain without LM regressionEarly-fusion multimodal tokens
Joint context modelling. Gate: can the hybrid attention and residual topology actually route visual and textual information together, rather than one crowding out the other?
gate · joint routing evidenceBest Phase-4 architecture, multimodal
The full controlled comparison. Gate: does the architecture that won on long-context language remain the winner cross-modally? A “no” here would be one of the more interesting results in the programme.
gate · does the ranking transfer?Phase 6 · ternary recipe & MoE
Two ways to sharpen the ternary recipe, and one way to spend the savings.
Four ablations sit behind the dense comparison, each reported in its own table so a change to the training objective or the parameter accounting can never leak into the primary perplexity numbers. Two refine the ternary recipe itself without leaving pure \( \{-1, 0, +1\} \) weights; two ask whether the freed budget is better spent on conditional capacity.
ternary_mlp_learned_scale
A trainable per-output-channel scale γ applied after the ternary
matmul, initialised to 1.0 so step 0 is byte-identical to ternary_mlp. The
weights stay pure ternary; only the read-out is rescaled. About +0.2M parameters
(~0.05% of non-embedding), scored against B.
ternary_mlp_magact
Magnitude-aware int8 activations: a per-token, per-channel-group absmax scale (group 128) instead of one scale per token. Zero extra parameters, and kept causal — the scale only shrinks within a token — so KV-cached decode stays bit-identical. Scored against B.
ternary_mlp_moe
The dense FFN becomes a sparse MoE — a full-precision top-2 router over eight ternary SwiGLU experts, each sized so the two active experts equal one dense FFN. Active parameters match B; stored capacity is roughly 2.3× larger and reported separately.
ternary_mlp_gdn_moe
The same MoE FFN on top of the GDN + QSA mixer, asking whether conditional capacity stacks with the leading dynamic sequence mixer or merely duplicates what it already carries.
gate so the
name-matched conversion pass skips it — quantizing a routing decision would corrupt it
for a negligible saving, the same rule that protects the CSA2 and QSA indexers. And
because the MoE variants change the training objective (a Switch-style load-balancing
auxiliary loss) and the parameter accounting (active versus stored), Q and R are
compared only on their active top-2 count and always live in their own table — never in
the Phase 2–4 perplexity comparison.
What “controlled” means here, enforced in code
Every model is randomly initialized. The pretrained Qwen checkpoint is used only to confirm the architecture and tokenizer load correctly — never as a starting point. Held identical across all runs: tokenizer, dataset and revision, data ordering, token budget, sequence length, optimizer, warmup, gradient clipping, validation set, the seed set (42, 43, aggregated as the mean with the spread reported), checkpoint policy, evaluation protocol, hardware class.
That is not enforced by discipline. A fairness audit loads each variant's run manifest
and compares every controlled field against an allow-list of the switches that
are the experiment; any mismatch marks the whole comparison
NON-CONTROLLED_EXPERIMENT and the reporting CLI exits non-zero, so an
invalid comparison cannot pass silently in automation. Comparing the full model config
against an allow-list — rather than a hand-maintained list of things to check — is what
catches a hidden change like rms_norm_eps drifting on one variant.
The training recipe is unified. fp_baseline and every
ternary variant use Qwen's plain cosine LR with constant weight decay, so
fp_baseline → ternary_mlp carries no declared schedule difference —
the MLP-only FFN weight precision is the only change. (A two-stage cosine schedule
remains available as a possible future ablation of the BitNet b1.58 recipe, but
no variant in the study uses it.) When a per-variant difference is intentional,
the audit records it explicitly in a declared_differences list instead of
waiving it silently — and only schedule fields may ever be declared. Declaring an
unequal token budget raises, because that is not a recipe difference, it is a different
experiment.
The shared core budget — embeddings, the per-layer mixer sized to the
attention slot, the uniform SwiGLU MLP and the norms — is held fixed to within
~0.1% on the primary chain at the live 0.6B scale, never shrunk to
compensate; whatever a mechanism adds beyond that core is reported transparently as its
additional budget rather than hidden. Two families sit a little wider and are
reported as-is rather than compensated away — the Kimi K3 family at −1.0% (its low-rank
Gated MLA lands under the shared budget) and the KDA×QSA cross-family at +2.6% (its gated
sparse-attention slot is heavier). That matching took deliberate tuning: KDA
has no grouped-query KV saving and naive settings inflated it by 25%, a Mamba-2 mixer has
an entirely different footprint, and the residual's full gated mode would
have added ~29M parameters (+4.9%) — which is why the default is channel_gated
at 2048 parameters per layer. Without that, “KDA is better” could just be “KDA is bigger”.
Secondary matched-compute comparisons may relax the parameter rule, but they have
to be labelled as such.
Four rules keep the expansion sets honest, and each one exists because a post-core mechanism would otherwise contaminate the primary table. MTP changes the training objective, so it is a separate Phase 5 extension set and is never merged into the Phase 2–4 perplexity comparison. The sparse MoE FFN changes both the training objective — a load-balancing auxiliary loss — and the parameter accounting, so it is compared on its active top-2 count with stored expert capacity reported separately. Multimodal runs report vision encoder and connector parameters separately, with language-backbone comparisons kept matched wherever that is possible at all. And every residual family arrives with a declared capacity budget, paid for outside the residual dimensions before the certified run.
Everything else holds as before: the same FineWeb-Edu revision, split, tokenization, shuffle seed and packing policy; the same optimizer family and schedule unless the experiment is an optimization study; the same precision, CUDA stack, kernel policy and evaluation environment; the same fixed cached validation set and cadence. Fake-quantized ternary measurements stay separate from genuine low-bit kernel measurements. Provenance recorded per run: git SHA, config hash, data revision, hardware, token count, checkpoint lineage, kernel versions, and every eviction and resume.
What this is not
- Not a speed claim. Ternary weights here are simulated (fake-quantized), which is slower than fp16, not faster. Efficiency reports are labelled
algorithmicvskernel_optimizedand the two are never mixed. - Not a state-of-the-art attempt. Optimizing for the best number would defeat the purpose; each architectural change has to stay attributable.
- MoE stays a Phase-6 ablation. The sparse MoE FFN (Q, R) is wired but deliberately sits behind the dense comparison, precisely to preserve attribution: learn which dense dynamic mechanisms are worth having before asking whether ternary savings are better spent on experts.
- Not a reproduction of anyone's frontier model. Mechanisms are borrowed from Qwen3.8, DeepSeek-V4.1-Flash and Kimi K3 one family at a time — the single deliberate exception being the KDA×QSA cross-family (XF-1, XF-2), which pairs Kimi's recurrent memory with Qwen's sparse retrieval as its own controlled comparison. Their fine-grained expert routing, their FP8/FP4 storage and their full multimodal stacks stay out — the Phase-6 MoE here is a plain top-2 Switch-style router, not their expert system.
- Not a claim that these effects are already reproduced. Everything attributed to an external paper is a literature input motivating an experiment, and is labelled as a proposal until BitFuse has run it.
Two bugs that argue for the whole methodology
Both were in the KDA kernel, both produced plausible-looking output, and neither was
visible from a loss curve. The state decay was applied to the wrong axis — the state is
\([d_v, d_k]\) and the gate indexes \(d_k\). And the intra-chunk term needs
\(\exp(c_t - c_s)\) over cumulative log-decay; computing it as k / exp(c)
overflows for long chunks or steep decay, so it is now split around the
chunk midpoint to keep both exponential factors bounded by 1.
A chunk-parallel training kernel is asserted equal to an explicit recurrent reference to
2e-4 across chunk sizes, sequence lengths, initial states and streaming
continuation. Without that reference to diff against, both bugs would have silently
degraded models C and D and been misread as “KDA doesn't help.”
A third one is worth naming because it is invisible by construction: overriding
_init_weights without calling super() leaves the RoPE
inv_freq buffers zero-filled after from_pretrained, because they
are non-persistent and absent from checkpoints. A resumed model then produces different
hidden states from the one that was saved — identical weights, no error. It surfaced only
as a save/load roundtrip test failing at ~1e-3. Every resumed training run would have
been silently corrupt.
Sequencing
Each phase is a gate, not a milestone.
GPU time is the dominant cost, so the campaign is split into phases launched independently. Only continue if the previous phase looks right — and the later scripts refuse to start if an earlier phase's artifacts are missing or were trained at a different token budget, so an unusable comparison cannot be discovered after paying for it.
Do all thirty-six switches execute at all?
Reduced models, one optimizer step — it produces no comparable numbers. It exists to
fail cheaply: forward hooks count real KDA / Mamba-2 / CSA2 / CED / GDN / Engram / MoE
/ residual invocations, so a wired-but-uncalled path fails here rather than on a
provisioned GPU VM. Runs on any CUDA GPU or Apple-Silicon MPS, auto-detected
(device=auto resolves CUDA → MPS → CPU).
Engineering validation, then a local reduced-width sweep
Reference load, forward/backward/step, strict ternarity, KDA equivalence against an explicit recurrent reference, checkpoint round-trip — then all thirty-six variants trained to ~5M tokens at reduced width through the real runner, plus a benchmark harness for throughput and memory. It ranks the variants under one budget, carried with a stated parameter-confound caveat, and is never mistaken for a Phase-1 result.
gate · code trusted enough to spend GPU money onOne variant per family, on real GPUs, on disposable infrastructure
FP, ternary KDA and ternary Mamba-2 at 10M disposable tokens, then the resource group is destroyed so nothing keeps billing — teardown downloads the evidence first and refuses to delete anything if the download is empty. This is also the gate for the carrier instrumentation, because a diagnostic problem and an architecture problem must never be debugged at the same time.
gate · real kernels and instrumentation workFP + ternary + KDA at full budget
A is retrained here rather than reusing Phase 1's throwaway run, because the fairness audit needs a baseline recorded under the identical contract.
gate · certified evidence for ternarization and KDAThe residual arms, then the certified six-way core report
D, E and F trained and evaluated, and the whole core six reported together.
gate · core fairness audit passesQwen GDN + QSA
Pilot at 100M to establish it is numerically sound, then a finalist run at the certified budget.
gate · numerically and causally interpretableDeepSeek-V4 legacy hybrid attention
The V4 legacy swa/csa/hca track with a Lightning Indexer (G, H,
deepseek_impl="v4_hybrid"), kept as the historical control — deliberately
not the current DeepSeek branch, which is the Phase 4-D V4.1-Flash family.
The residual family sweep — R1 / R2 / R3 / R4 / R5
Best two Phase-4 mixers frozen, topology the only thing moving — including R5, the V4.1-Flash single-pass mHC sibling of R3. Screened at 100M, finalists promoted to the certified budget.
gate · best topology identified at matched capacityDeepSeek-V4.1-Flash family (DS41-*)
The current DeepSeek branch, isolated one mechanism at a time on the ternary backbone:
CSA2, the Causal Encoder-Decoder, the hierarchical sparse indexer, single-pass mHC and
research-scale Engram (DS41-1…DS41-7), all reported against ternary_mlp.
Kimi K3, Qwen3.8-Flash-Next, and the KDA×QSA cross-family
Three more comparison sets, each isolating a family's own mechanisms on the ternary
backbone: Kimi K3 (K3-1…K3-5, vs ternary_mlp_kda), Qwen3.8-Flash-Next
(Q38-1…Q38-3, vs ternary_mlp_qwen_hybrid) and the KDA×QSA cross-family
(XF-1, XF-2). Screened at 100M, finalists promoted to the certified budget; none folds
into the six-way core.
Attention share, placement, head and indexer budget
DS-25/50/75, then placement with the count fixed, then capacity with the structure fixed.
gate · optimal dynamic-capacity allocation identifiedMulti-token prediction on the best dense architecture
1/2/4 heads plus the multi-step consistency schedule, measured on a speculative-decoding engine.
gate · quality and speculative-decoding evidence, reported apartMultimodal integration
M1 through M4, expanding only after each gate — vision parameters always reported separately.
gate · cross-modal evidence without backbone confoundingTernary-recipe ablations and a sparse MoE FFN
Only after the dense architecture is frozen. The learned-scale and magnitude-aware
activation refinements (O, P) scored against ternary_mlp, then a top-2 MoE FFN (Q, R) at
roughly equal active compute — routed and stored expert parameters kept separate from
the backbone in the capacity ledger, each in its own table.
The causal order is the point of the sequence, not the schedule: A→B isolates ternarization; B→C and B→E test two rival recurrent compensations; C/E→residual isolates topology; B/E→Phase 4 tests heterogeneous context allocation; the best mixer feeds the residual sweep; the best dense architecture feeds MTP and multimodal; and only the best overall dense architecture feeds Phase 6.
What gets measured
The headline number is not accuracy. It is marginal return per unit of resource, computed separately for each allocation axis:
where \(q\) is task quality and \(R\) is spent resource — weight memory, KV-cache, latency, training FLOPs, or energy where it can be measured honestly.
Alongside that, every run is instrumented internally, because a result without a
mechanism is not worth much: ternary distribution and zero fraction per layer,
activation magnitude, variance, contrast and effective rank, attention entropy,
concentration, span and head diversity, recurrent-state occupancy, and how much
earlier-layer information the residual pathway is actually carrying. The instrumentation
is emitted at three levels — light, standard, full —
and Phase 1's job includes confirming all three hold finite values before anything
commits to a 1B-token run.
Each new mechanism brings its own diagnostics, each chosen so its mechanism could fail visibly rather than quietly. Weights gain layer-wise quantization error and scale drift. Attention gains QSA and CSA2 indexer recall and coverage, selected-block mass, and compressed-pool visibility statistics. Recurrent state gains GDN and KDA retention-and-forgetting curves. Residuals gain GR branch usage, mHC mixing entropy and deviation from double stochasticity, and AttnRes depth weights. MTP gains auxiliary-token losses, token acceptance rate and speculative-decoding speedup. The MoE FFN gains router load-balancing loss and expert-utilization histograms. Multimodal runs gain cross-modal retrieval and probe accuracy, image-token saliency, and modality-conditioned norm and entropy profiles.
A research-grade telemetry layer sits underneath all of it. Every run also records the per-step optimization trajectory and loss/gradient stability, a fixed probe set, and the internals of each moving part — MoE routing, the MTP heads, the hybrid mixers, attention logits, and the residual and gradient paths — plus efficiency benchmarks and post-hoc interaction, component-attribution, information-flow and stress-test analyses. It exists so a quality difference can be traced to a mechanism rather than asserted.
Reading the ternary weight carrier, straight from the checkpoint
The weight carrier gets its own static report — computed from the saved
{−1, 0, +1} weights alone, with no data and no forward pass, so it can be recomputed
from any checkpoint (including the Phase-0 ones already on disk, with no retraining) and
kept as a lightweight result long after the full weights are reclaimed. One duck-typed interface —
a module that exposes quantized_weight() — discovers the ternary layers in every
family, so fp_baseline simply reports is_ternary: false and each new
mechanism is picked up automatically. Every number is tagged static or
dynamic, so its provenance — the weights alone versus a forward pass — is never
ambiguous.
For every ternary tensor it records the full distribution — zero / positive / negative density, the positive and negative fraction among the non-zeros, and a sign balance in \([-1, +1]\) (\(+1\) and \(-1\) are never assumed equiprobable) — plus the ternary entropy in bits (max \(\log_2 3 \approx 1.585\)), the sign entropy, the per-tensor scale statistics, and that same distribution pooled per parameter, per module category and per decoder depth, so “which layers went sparse, and where” is answerable rather than averaged away. It is the evidence for the study's core question — where information relocates once ternarization removes the weights' precision — read from the weight channel itself.
2 − zero_density, against a conventional 2-bit baseline
and a five-trit-per-byte packing (\(3^5 = 243 \le 256\), so \(8/5 = 1.6\) bpw). When a tensor is
sparse enough (zero_density > 0.415) that rate dips below the ternary information
bound \(\log_2 3\) — which is compression of a sparse symbol stream, not a change of
alphabet: the model is still ternary over {−1, 0, +1}. Ternary entropy is the
information content of the distribution; BITCOS bpw is a storage rate under one
coding scheme — reported side by side, never conflated, and no BITCOS inference kernel is claimed.
Resumable training, bounded storage
A campaign this size — the thirty-six variants, more once scale profiles and budget sweeps are
counted — cannot keep every finished model's full weights on disk. Training is fully
resumable: each checkpoint stores the complete state (model, optimizer, scheduler,
AMP scaler, token/step counters, the data-stream cursor and all RNG), written atomically with a
checksummed manifest, so a crash, restart or Spot eviction continues exactly rather than
restarting from zero — and every checkpoint is stamped with its experiment_id, so one
run can never resume from another's state. But a run is marked COMPLETED only once
training and metric collection have finished, and only then does it persist its lightweight
ternary results and reclaim the large weight payloads, keeping the provenance sidecars and
a cleanup.json marker. So disk grows with the number of running experiments,
not finished ones — a fixed ceiling of one latest + one best + a bounded
periodic history per active run — while the final results never depend on retaining the trained
weights. Each run advertises its lifecycle — RUNNING, RESUMABLE,
COMPLETED, FAILED — and a completed one is skipped, never silently
retrained.
Four outcomes, all of them informative
- Ternary scaling wins. Dynamic capacity is not an efficient substitute for parameters at low precision — which retires a plausible idea and tightens the case for current practice.
- Allocation wins. A smaller ternary backbone with richer retrieval, residual or predictive machinery beats a larger one at matched resources, and low-bit scaling has been leaving quality on the table.
- They are complementary. Both axes help, and the useful contribution becomes the shape of the joint frontier.
- Allocation is task-dependent. Short context favors weights, long context favors selective retrieval, reasoning favors conditional capacity — the most interesting outcome, because it means there is no single correct low-bit architecture, only a correct allocation for a workload.
What counts as a strong result
A strong result is emphatically not “the newest architecture wins.” It is a controlled trade-off that survives matched-budget comparison and has a plausible mechanistic explanation behind it. So the claims are pre-committed with their interpretations attached:
- If Qwen GDN + QSA wins, the claim is about the usefulness of fixed-size memory plus selective retrieval under ternary weights — not about Qwen being superior.
- If mHC beats Gated Residual, the result is that constrained multi-stream routing offers a better stability/expressivity trade than unconstrained gating.
- If AttnRes wins, the contribution is evidence that depth-wise information selection is a carrier separate from parallel residual streams.
- If MTP improves serving but not validation loss, that is still a meaningful efficiency result, and the paper labels it as exactly that.
- If multimodal performance improves only with a large vision-side budget, the result is not attributed to the ternary backbone without a matched control.
- Any fairness-audit failure marks the corresponding delta
NON-CONTROLLED_EXPERIMENTand removes it from the primary causal claims entirely.
Working thesis
Extremely low-precision models should probably not be scaled by uniformly adding more low-precision parameters. Because ternarization redistributes where representational capacity actually lives, the more efficient strategy may be to spend the budget on the mechanisms that select, route, preserve and predict information — recurrent memory, selective retrieval, residual topology, prediction heads, and eventually conditional computation — with ternary weights providing the cheap substrate underneath. The shift under test is from scaling the weights to allocating the information pathways.