Research direction II · proposal v7 · updated 20 September 2026

Ternary weights cost quality. The question is where the information goes instead.

BitFuse is a staged research programme on the Qwen3 architecture — trained at the real Qwen3-0.6B scale (~596M), with Qwen3-4B retired — under MLP-only 1.58-bit ternary FFN weights. Its core thesis: when the static parameter channel is aggressively compressed, sequence mixing, residual pathways and conditional computation stop being independent plug-ins and become alternative information carriers competing for the same freed budget. The code, the fairness audit, the instrumentation and the Azure runbook exist; what is missing is GPU hours.

Not a leaderboard exercise. Every model here is trained from scratch on a fixed 1B-token certified budget, so all of them are weak in absolute terms. That is expected and fine: the absolute perplexity is not the result — the differences between the models are. Which is why the interesting clause in the hypothesis is the last one: a quality improvement with no matching change in carrier statistics would falsify the explanation even if the number went up.
0
variants — the core six, the Phase-4 sets incl. the DeepSeek-V4.1-Flash, Kimi K3, Qwen3.8-Flash-Next and KDA×QSA cross-family lines, a Phase-5 extension and two Phase-6 ablation sets
0
residual topologies treated as separate algorithms, not settings
0
research questions, each tied to a named pair before any run
0B
token certified budget for every primary comparison

What changed in v7. Ternarization is now MLP-only — the FFN/SwiGLU weight matrices go to 1.58-bit ternary and nothing else does, so the base variant is ternary_mlp, not a BitNet reproduction — and the campaign trains at the real Qwen3-0.6B scale (~596M) on every machine, with Qwen3-4B retired. The variant set grew to thirty-six. The DeepSeek work is two branches that never share a table: the sliding-window + CSA + HCA schedule with a Lightning Indexer is the V4 legacy control (G, H, deepseek_impl="v4_hybrid"), and the DeepSeek-V4.1-Flash family (DS41-1…DS41-7) isolates CSA2, the Causal Encoder-Decoder, the hierarchical sparse indexer, single-pass mHC and a research-scale Engram one mechanism at a time. Three more families join as their own comparison sets: the Kimi K3 line (NoPE Gated MLA, SiTU-GLU, Attention Residuals, Stable LatentMoE), the Qwen3.8-Flash-Next line (an n-gram conditional memory on the GDN+QSA mixer) and a KDA×QSA cross-family that pairs Kimi's recurrent memory with Qwen's sparse retrieval. The residual sweep is six first-class families (channel-gated, Qwen GR, DeepSeek mHC in iterative R3 and single-pass R5 forms, Kimi AttnRes). Two ternary-recipe refinements — a learned per-channel scale and magnitude-aware int8 activations — and a sparse top-2 MoE FFN over eight ternary experts join as Phase-6 ablations, each in its own table. And an Apple-Silicon (MPS) dev backend plus a reduced-width benchmark harness let the whole 36-variant sweep run locally before any GPU spend.

The hypothesis, stated so it can lose

Ternarization flattens the network's internal representations. Where a full-precision layer has a few large activations carrying most of the signal, a ternary layer's activations are more uniform. That loss of contrast is the mechanism behind the quality drop. If so, mechanisms that restore contrast — a state-based attention with per-channel gating (KDA), a state-space mixer (Mamba-2), Qwen's fixed-size memory paired with block-indexed sparse retrieval, compressed long-range attention, a widened or manifold-constrained residual stream, or an attention over depth itself — should recover quality, and the recovery should be visible in the statistics, not just in the loss curve.

BitNet b1.58 constrains every weight to \( \{-1, 0, +1\} \) — about 1.58 bits each instead of sixteen, so a weight costs roughly a tenth of a byte to store. The published recipe leaves one slot conspicuously untouched: BitNet still runs full softmax attention with RoPE, and its own technical report names “investigating efficient attention mechanisms suitable for low-bit models” as future work. BitFuse is that investigation — but it borrows only the ternary technique: it quantizes the FFN/SwiGLU weight matrices alone (MLP-only) and keeps attention, embeddings, the tied head and every sequence mixer full-precision Qwen, so it is not a BitNet reproduction. And the newest literature keeps strengthening the original thesis rather than replacing it: Qwen, DeepSeek and Kimi have each shipped a different answer to “which dynamic carrier should do the work”, at full precision, without anyone testing whether the ranking survives ternarization.

Six carriers, one of them crippled

It helps to treat the network as a set of distinct information carriers, and to notice that ternarization is not symmetric across them. Ternarization is primarily a reduction in per-weight resolution, not a total loss of magnitude information — the model keeps a full-precision latent parameter and derives its ternary forward weight through a scale factor. So the useful question is not “how much accuracy is lost?” but “which channels become more valuable when static weight resolution is constrained?”

FFN weights — heavily degraded

Persistent learned knowledge — and only the FFN/SwiGLU weights are ternary; attention, embeddings and the sequence mixer stay full-precision. A static report reads the {−1,0,+1} distribution of those FFN weights straight from any checkpoint — zero density, sign balance, ternary and sign entropy, scale factors, per-layer quantization error, and the BITCOS storage rate — per parameter, per module and per depth, no forward pass required.

Activations — untouched

Still continuous, still expressive. Magnitude, variance, contrast, effective rank. Not a control knob — normalization eats deliberate rescaling, so these are read as evidence, never tuned.

Attention — untouched and scalable

Which tokens interact and what gets retrieved. Entropy, max probability, top-k mass, output norm — and now QSA and CSA2 indexer recall and coverage, selected-block mass, and compressed-pool visibility statistics.

Recurrent state — a fixed-size memory

What a KDA, Mamba-2 or Gated DeltaNet layer has absorbed. State norm, update ratio, cosine retention, and GDN/KDA retention-and-forgetting curves.

Residual — promoted to an axis

No longer one arm but six families. Input/update norm ratio, plus GR branch usage, mHC mixing entropy and doubly-stochastic deviation (iterative R3 and single-pass R5), and AttnRes depth weights.

Prediction & modality — the newest axes

Two carriers the core six do not touch. Auxiliary-token losses, acceptance rate and speculative speedup for MTP; cross-modal probe accuracy, image-token saliency and modality-conditioned norms for vision.

That asymmetry is the opening. A ternary transformation cannot express fine distinctions in its weights, but the dynamic carriers can still create sharp differences in what each token reads, remembers and preserves. If a low-precision network is short of contrast, those are where it can still be manufactured — and the organizing principle of the whole programme follows: compress static weight expressivity, measure where the information moves, then reinvest the freed budget in whichever mechanism delivers the most quality per unit of memory or compute.

The hypotheses, H1–H8

Eight claims, written before the runs.

Three cover the core ablation; five more follow from the residual programme, the predictive objective and the multimodal extension, each making a claim the original three do not. Each one is scoped so a null result is informative rather than merely disappointing.

H1

Dynamic reallocation

Ternarization changes where useful information is carried. If the sequence mixer is also too restrictive, quality may fall further than necessary — richer dynamic computation may compensate for constrained weights.

H2

Dynamic-capacity compensation

A ternary model may want a different allocation of global, local, recurrent and compressed-attention capacity than a full-precision one. This has to be measured, not assumed.

H3

Residual preservation

Residual topology can preserve or re-route information that would otherwise be attenuated, overwritten or diluted by a deep sequence-mixing stack.

H4

Heterogeneous context roles

Local retrieval, compressed retrieval, fixed-size recurrent memory and precise global attention may divide the representational work more efficiently than any single mixer.

H5

Depth-wise information selection

Kimi-style Attention Residuals may help ternary models by letting each layer select which earlier representations deserve to contribute — turning depth into an additional dynamic channel.

H6

Topology-stable multi-stream residuals

DeepSeek mHC and Qwen Gated Residuals may preserve multiple pathways while avoiding the instability of unconstrained hyper-connections.

H7

Predictive compression and acceleration

MTP can make more of the same backbone useful per forward pass, potentially improving speculative-decoding acceptance and backbone quality without altering the carrier thesis.

H8

Modality as a carrier

Multimodal inputs test whether a ternarized backbone can preserve cross-modal information while leaning harder on dynamic representations, connectors and residual pathways.

The falsification rule, made concrete

Every mechanism has to predict where the information moves.

This is what separates an explanation from a coincidence, and it is the single most important methodological commitment on the page. Each mechanism enters with a prediction about which carriers should gain and which should give way. If quality improves and the carriers do not move as predicted, the explanation is treated as refuted — the gain is real but unexplained, which is a different and weaker claim.

The trade, in bytes

Why dynamic capacity has to be very good to be worth buying.

Before any experiment runs, the arithmetic already says something uncomfortable. Ternary weights are so cheap that every attention, residual or prediction-head parameter costs many backbone parameters — and an order of magnitude more again if it stays at higher precision, as an indexer or a residual gate has to. That is the bar every hypothesis on this page has to clear, and it is why BitFuse holds the shared core budget fixed to within ~0.1% on the primary chain: otherwise the better architecture would just be the bigger one.

This control is a design space, not a result. The horizontal axis of the real experiment is this same split; the vertical axis — quality — is what the measurements have to supply.

The experiment

Thirty-six models. One class. One switch at a time.

All of them share one Qwen3 skeleton — per-head QK-norm, a decoupled head_dim of 128, unbiased attention projections, SwiGLU, RMSNorm, RoPE, tied embeddings — and all of them are one implementation driven by config switches, not thirty-six model files. Divergent copies are how “identical except for X” quietly stops being true. The campaign trains at the real Qwen3-0.6B scale (~596M) on every machine — Apple-silicon and GPU alike, with Qwen3-4B retired — so the geometry and every per-variant count below are the live 0.6B numbers: 28 layers, hidden 1024, 16 query heads, 8 KV heads, intermediate 3072. Each mixer's head dims are re-solved for that width so core-matching holds. New mechanisms enter as separate expansion sets, so the primary table never fills up with partially controlled comparisons. Every backbone is matched within ~0.1% on core parameters on the primary chain (the Kimi K3 family lands a touch under, −1.0%, and the KDA×QSA cross-family a touch over, +2.6%, reported as-is) — the MoE variants (Q, R, K3-5, Q38-3) and the Engram / n-gram–bearing variants (DS41-6, DS41-7, Q38-1…3) on their active count, since their expert banks and hash tables are stored but only sparsely accessed.

Set Name Sequence mixer Residual topology Primary question
A core fp_baseline softmax, all layers · FP standard the full-precision reference
B core ternary_mlp softmax, all layers standard what MLP-only ternarization costs
C core ternary_mlp_kda 3 KDA : 1 global standard does KDA recover the ternary loss?
D core ternary_mlp_kda_attnres 3 KDA : 1 global channel-gated do low-cost residual gates add anything?
E core ternary_mlp_hybrid 24 Mamba-2 : 4 global (Nemotron-H) standard does Mamba-2 survive ternarization?
F core ternary_mlp_hybrid_attnres 24 Mamba-2 : 4 global channel-gated do residuals help Mamba-2 too?
G 4-B ternary_mlp_deepseek swa / csa / hca + Lightning Indexer (V4 legacy) standard does V4 legacy compressed attention help?
H 4-B ternary_mlp_mamba_deepseek Mamba-2 + csa / hca (V4 legacy) standard does V4 legacy attention complement recurrence?
I 4-Q ternary_mlp_qwen_hybrid 3 GDN : 1 QSA (Qwen3.8) standard does recurrent memory + selective retrieval fit ternarization?
J 4-Q ternary_mlp_qwen_hybrid_res 3 GDN : 1 QSA channel-gated (R1) the GDN + QSA residual-sweep baseline
K 4-C ternary_mlp_res_gr 3 GDN : 1 QSA Qwen GR (R2) does a widened gated residual stream help?
L 4-C ternary_mlp_res_mhc 3 GDN : 1 QSA DeepSeek mHC (R3) does constrained multi-stream routing beat gating?
M 4-C ternary_mlp_res_attnres 3 GDN : 1 QSA Kimi AttnRes (R4) does depth-wise selection add a separate carrier?
N 5-P ternary_mlp_mtp best Phase-4 mixer best topology does predictive supervision improve quality or only serving?
O 6-T ternary_mlp_learned_scale softmax, all layers standard does a learned per-channel scale recover the ternary loss?
P 6-T ternary_mlp_magact softmax, all layers standard do magnitude-aware int8 activations help?
Q 6-E ternary_mlp_moe softmax + sparse MoE FFN standard is freed ternary budget better spent on experts?
R 6-E ternary_mlp_gdn_moe 3 GDN : 1 QSA + MoE FFN standard does MoE stack with the best dynamic mixer?
DS41-2 4-D ternary_mlp_csa2 CSA2 · full/reindex/reuse, flat indexer standard does CSA2 alone beat plain softmax? (RQ20)
DS41-1 4-D ternary_mlp_ced Causal Encoder-Decoder · shared global KV standard does the encoder-decoder split help? (RQ19)
DS41-3 4-D ternary_mlp_ced_csa2 CED + CSA2 · flat indexer standard the CED × CSA2 interaction (RQ21)
DS41-4 4-D ternary_mlp_ced_csa2_hsi CED + CSA2 · hierarchical indexer standard hierarchical vs flat sparse indexer (RQ22)
DS41-5 4-D ternary_mlp_ced_csa2_mhcsp CED + CSA2 · hier. Single-Pass mHC (R5) does the single-pass residual add on top? (RQ23)
DS41-6 4-D ternary_mlp_ds41_engram softmax + Engram memory standard does n-gram conditional memory help? (RQ24)
DS41-7 4-D ternary_mlp_ds41 CED + CSA2 (hier.) + Engram Single-Pass mHC (R5) the integrated V4.1-Flash model (RQ25)
R5 4-C ternary_mlp_res_mhc_sp 3 GDN : 1 QSA Single-Pass mHC (R5) single-pass vs iterative mHC on a frozen mixer (RQ23)
K3-1 K3 ternary_mlp_k3_gmla 3 KDA : 1 NoPE Gated MLA standard swap the KDA global layer for a latent-KV Gated MLA — vs ternary_mlp_kda
K3-2 K3 ternary_mlp_k3_attnres 3 KDA : 1 Gated MLA Kimi AttnRes (R4) + depth-wise Attention Residuals on the gmla base
K3-3 K3 ternary_mlp_k3_situ 3 KDA : 1 Gated MLA · SiTU-GLU standard + the SiTU-GLU bounded activation on the gmla base
K3-4 K3 ternary_mlp_k3_core 3 KDA : 1 Gated MLA · SiTU-GLU Kimi AttnRes (R4) the integrated dense Kimi K3 model
K3-5 K3 ternary_mlp_k3_latentmoe 3 KDA : 1 Gated MLA · SiTU-GLU + LatentMoE Kimi AttnRes (R4) dense FFN → Stable LatentMoE (own table)
Q38-1 Q38 ternary_mlp_qwen_ngram 3 GDN : 1 QSA + n-gram memory standard n-gram conditional memory alone on the GDN + QSA base
Q38-2 Q38 ternary_mlp_qwen_core 3 GDN : 1 QSA + n-gram memory Qwen GR (R2) the integrated dense Qwen3.8-Flash-Next model
Q38-3 Q38 ternary_mlp_qwen_moe 3 GDN : 1 QSA + n-gram memory Qwen GR (R2) + a sparse MoE FFN (own table)
XF-1 XF ternary_mlp_kda_qsa 3 KDA : 1 QSA standard KDA remembering + QSA retrieving — beat either parent mixer?
XF-2 XF ternary_mlp_kda_qsa_attnres 3 KDA : 1 QSA Kimi AttnRes (R4) + depth-wise Attention Residuals on the cross-family mixer
A→D

the chain

Each step changes exactly one component: full precision → ternary → KDA → attention residual. Anything that moves is attributable to the step that moved it.

E, F

the branch

E and F branch off B rather than continuing from C: they swap in a different linear-attention family under identical ternary weights. So C and E are rivals, not successive steps — C vs E asks which family survives ternarization better.

G–M, DS41-*

Phase 4 — four expansion sets, reported separately

4-B is the DeepSeek-V4 legacy control (G, H — swa/csa/hca + Lightning Indexer), kept as a historical baseline; 4-Q adds Qwen's GDN + QSA (I, J); 4-C freezes that mixer and sweeps the residual topology (K/L/M, plus R5's single-pass mHC); 4-D is the new DeepSeek-V4.1-Flash family (DS41-*), isolating CSA2, the Causal Encoder-Decoder, the hierarchical indexer, single-pass mHC and Engram one at a time. The V4 and V4.1 branches never share a table. Each set has its own comparison set and report, so the six-way core ablation cannot quietly become a many-way one asking a different question.

N, O–R

Phases 5–6 — extensions and ablations, not rivals to the core

N adds multi-token prediction on the winning dense architecture, changing the training objective. O–R are Phase-6 ablations: two ternary-recipe refinements (O, P) and a sparse MoE FFN (Q, R). None belongs in the Phase 2–4 perplexity table — each gets its own outcome.

K3, Q38, XF

Three more families, each on its own comparison set

The Kimi K3 line (K3-1…K3-5) rebuilds the KDA stack toward Kimi's design — a NoPE Gated MLA global layer, SiTU-GLU, Attention Residuals and a Stable LatentMoE — scored against ternary_mlp_kda. The Qwen3.8-Flash-Next line (Q38-1…Q38-3) adds an n-gram conditional memory, then the Gated Residual, on the GDN + QSA base. The KDA×QSA cross-family (XF-1, XF-2) pairs Kimi's recurrent “remember” with Qwen's sparse “retrieve”. Each is its own set, never folded into the six-way core.

Why the Mamba ratio is not 3:1. The KDA variants use 3 KDA : 1 global, following Kimi Linear. The Mamba-2 variants deliberately do not mirror that, because copying NVIDIA's architecture means copying their placement. Nemotron-H's published hybrid_override_pattern is 52 single-purpose layers — 24 Mamba-2, 4 attention, 24 MLP, so attention is only 7.7% of layers. A BitFuse block holds a mixer and an MLP, so counting mixers only gives 28 — which lines up one-to-one with BitFuse's 28-layer depth, so attention lands at mixer indices 4, 10, 16, 22 (4 of 28). Matching NVIDIA matters more here than matching Kimi.

Where the mechanisms come from

Six borrowed mechanisms, each with the experiment that tests it.

These are external literature inputs, used to motivate experiments. None of them is a claim that BitFuse has already reproduced the effect. What each card says is: this is what the source contributes, why it bears on the carrier thesis, the carrier signature that has to move for its explanation to hold — a quality gain without that movement is evidence against the mechanism — and the specific comparison that would settle it here.

Qwen3.8-Flash-Next

Gated DeltaNet + Qwen Sparse Attention

GDN compresses history into a fixed-size recurrent state; QSA runs a lightweight indexer that aggregates tokens into micro-blocks and selects the important regions for sparse global attention. Qwen reports a 3:1 GDN:attention pattern.

This is the BitFuse reallocation question stated in someone else's architecture: how should representational work split between recurrent memory and selective retrieval?

carrier → GDN state norm and retention, the QSA selected-token fraction, and the n-gram store's qwen_ngram/hit_rate and qwen_ngram/gate_mean — the recall or retrieval gain is unexplained unless these move. test → ternary GDN + QSA vs ternary KDA, plain attention and DeepSeek CSA2, matched parameters and tokens
DeepSeek-V4.1-Flash

CSA2 — Compressed Sparse Attention

One attention that reads a local sliding window at full resolution and a compressed global pool chosen by a hierarchical indexer — a coarse block filter, then a fine top-64 — all in a single softmax. Adjacent layers share the compressed KV, the candidate pool and the top-k across a Full / Reindex / Reuse schedule.

The new DeepSeek branch — the Phase 4-D V4.1-Flash family (DS41-*), kept separate from the DeepSeek-V4 legacy control (G, H). It isolates CSA2, the Causal Encoder-Decoder, the hierarchical indexer, single-pass mHC and Engram one at a time: does compressed sparse attention stay valuable after ternarization?

carrier → the csa2/* indexer recall and selected-block mass, the Full/Reindex/Reuse reuse rates, the shared encoder→decoder KV, and engram/hit_rate — with the genuine-sparsity invariant guarding against selection quietly going dense. test → run the V4.1-Flash family against plain ternary softmax, Qwen GDN+QSA and the KDA/Mamba hybrids at matched parameters and tokens
Qwen

Gated Residual (GR)

The residual stream is widened to four branches with dynamic, content-dependent read gating and per-branch write control.

Adds multiple residual information channels without requiring a separate attention mechanism to do it.

carrier → GR branch usage and per-branch norm share, read-gate entropy — branch specialization has to appear, not just a lower loss. test → replace the channel-gated residual with a parameter-matched GR branch; probe whether branch specialization tracks the carrier metrics
DeepSeek

mHC — manifold-constrained hyper-connections

Hyper-connections widen the residual stream and route information dynamically; mHC constrains the residual mixing matrix to the doubly-stochastic manifold, restoring identity-map behaviour and stabilizing training.

A principled multi-stream topology whose stability is enforced rather than hoped for — which is exactly what makes it a fair rival to GR.

carrier → residual mixing entropy and deviation from double stochasticity — iterative R3 vs single-pass R5 row deviation — stable multi-stream transport must show up in the residual carrier. test → mHC vs GR vs baseline at matched capacity; log residual mixing entropy and singular-value/norm behaviour
Kimi K3

Attention Residuals

AttnRes replaces fixed residual accumulation with input-dependent attention over preceding layer outputs. Kimi also pairs KDA with global MLA.

This makes depth itself an information-selection mechanism, which is the carrier hypothesis applied along an axis the other four families do not touch.

carrier → KDA state magnitude and decay, the MLA gate opening (gated_mla/gate_mean), and AttnRes depth entropy — depth has to become an active selection channel, not a passive skip. test → full or block AttnRes as a separate residual family; measure depth-wise contribution concentration, output norms, gradient distribution
Qwen & Kimi

Multi-token prediction, and native multimodality

MTP trains future-token prediction heads; Qwen reports multi-step training aimed at consistency with speculative decoding, and Kimi K3 retains an MTP layer. Both models are also multimodal — Kimi trains language and vision jointly in one shared context.

Together they ask whether the carrier framework survives a change of training objective and a change of input space.

carrier → auxiliary-token loss and speculative-decoding acceptance for MTP; cross-modal probe accuracy and image-token saliency for vision — the new carrier must carry signal, not merely change the objective. test → Phase 5-A: 1/2/4 future-token heads on the best dense model · Phase 5-B: frozen vision encoder and projector first, earlier fusion only if the backbone stays stable

Phase 4-Q · remember + retrieve

Qwen's split: one layer kind to remember, another to retrieve.

Qwen3.8-Flash-Next puts Gated DeltaNet in most layers and Qwen Sparse Attention in periodic global ones, and frames it plainly as “remember” plus “retrieve”: GDN continuously compresses history into a fixed-size state, while QSA's indexer first aggregates tokens into micro-blocks and then selects the important regions to attend over. That is what makes it a genuinely different comparator from the existing KDA setup — the retrieval side is not dense global attention but sparse, block-indexed retrieval.

0 : 1
published GDN : QSA placement, kept as the starting point
0
layer budget held — 21 GDN, 7 QSA
FP
the indexer's own parameters, declared separately from the ternary budget
0
of Qwen3.8's MoE or multimodal stack copied into Phase 4

Implementation stance. Begin at the published 3:1 placement, keep the 28-layer budget and the existing parameter-matching framework, and run a ternary GDN path for the projection matrices wherever the quantization contract permits it. Architectural attribution comes first; aggressive kernel optimization can wait, because a fast implementation of an uninterpretable comparison is worth nothing.

The indexer is declared, not assumed. QSA's indexer parameters are separated from the main ternary projection budget and the precision is stated explicitly — the same reasoning that keeps DeepSeek's index_q/index_k out of the name-matched ternary conversion pass. Quantizing a ranking corrupts selection for a negligible memory saving, and a corrupted ranking would be indistinguishable from the mechanism simply not working.

Three families, one controlled axis. Experiment 4.1 puts KDA + global against GDN + QSA at identical token budget, parameter budget, optimizer and seed. Experiment 4.2 puts GDN + QSA against DeepSeek CSA2 at the same 0.6B scale. Then 4.3 sweeps attention share (DS-25/50/75), 4.4 sweeps placement with the attention count fixed, and 4.5 sweeps head and indexer budget with the structure fixed — because “how much retrieval capacity” and “where the retrieval layers sit” are two different questions and answering them together answers neither. And if GDN + QSA wins, the claim to make is about fixed-size memory plus selective retrieval under ternary weights — not about Qwen being better.

Phase 4-B legacy · Phase 4-D DeepSeek-V4.1-Flash family

Two DeepSeek branches, kept apart: a V4 legacy control, and the V4.1-Flash family.

The DeepSeek work is now two separate branches that never share a table. The V4 legacy pair (G, H, deepseek_impl="v4_hybrid") keeps the original sliding-window + Compressed + Heavily-Compressed attention with a Lightning Indexer — a historical control, Phase 4-B. The Phase 4-D V4.1-Flash family (DS41-*) is the current branch, adding one mechanism at a time on the ternary backbone: CSA2 (illustrated below), the Causal Encoder-Decoder (14 encoder layers, then decoder layers reading one shared global KV projected from the encoder's final state), the hierarchical sparse indexer, single-pass mHC (R5) and a research-scale Engram n-gram memory. Each CSA2 layer reads two things in one softmax: a local sliding window at full resolution, and a compressed global pool — four tokens summarised into one latent — chosen by a two-level indexer (a coarse block filter, then a fine top-64). What changes between CSA2 layers is not the compression level but how much structure they build versus reuse: a full layer builds and publishes the compressed KV, candidate pool and top-k; a reindex layer reuses the KV and pool but recomputes its own top-k; a reuse layer reuses everything, which is where the compute and memory saving comes from.

0
local sliding window, kept by every CSA2 layer
×0
compression ratio — four tokens into one latent, non-overlapping
top-0
fine indexer selection, after a coarse 8-block filter
0
of V4.1-Flash's production scale — 552B MoE, 196B Engram, 1M context, DSpark, FP4/FP8 serving — reproduced

These layers are not “attention over fewer tokens.” A compressed path alone cannot resolve fine-grained local structure, so removing the local branch would change what the layer is for. What each one actually does is read a short local window at full resolution and a long context at reduced resolution, in one softmax:

          ┌── local window (full resolution, w=128) ───┐
hidden → ─┤                                            ├→ one softmax → out
          └── compressed pool · hierarchical indexer ──┘

Causality is structural, not assumed. A compressed entry covering source positions \([\text{start}, \text{end})\) summarises token \(\text{end}-1\), so it is made readable only from that position onwards — and the indexer ranks after that mask, which is what guarantees its selection is always a subset of what dense causal attention would have seen.

Only complete blocks are emitted: a partial trailing block would summarise fewer tokens than its siblings and make the entries incomparable. Tests assert that perturbing the tail of a sequence leaves every earlier logit bit-identical, that the local window alone cannot see distant tokens while the compressed pool can, and that cached decoding matches a full forward pass exactly.

Explicitly not a V4.1-Flash reproduction. The production model is a 552B sparse mixture, a 196B Engram, a 1M-token context, DSpark speculative decoding and FP4/FP8 storage with an entire serving stack. None of that scale is reproduced — the family adapts each mechanism at 0.6B on a matched ternary backbone (the Engram here is a small research analogue, ~201M of stored hash tables held apart on active-parameter matching), because mixing production scale in would make it impossible to attribute a result to the mechanism under test. Every approximation is declared, mechanism by mechanism, in the repo's V4.1 architecture audit. The indexer's projections are named index_q/index_k so the name-matching ternary conversion pass does not claim them — quantizing a ranking would corrupt selection for a negligible memory saving. The single-pass mHC hyper-connection (R5) is a residual question and gets its own place in the Phase 4-C sweep as well — see below.

Phase 4-C · a first-class axis

Six answers to “what should a layer be allowed to read?”

Since 2017 the residual pathway has been a single fixed addition from the layer directly beneath — a write, not a read. Under ternary weights every one of those transformations is lossy by construction, which is why the slot becomes worth reopening at 1.58 bits and not before. BitFuse tests six of them, and treats them as different algorithmic families rather than settings of one idea — because the distinction is what makes attribution possible at all. R3 and R5 are two takes on the same multi-stream idea: DeepSeek's mHC, projected either by iterative Sinkhorn (R3) or in a single normalization pass (R5, from V4.1-Flash).

Family Information path Core operation Expected signal Main risk
R0 standard residual 1 stream x + F(x) identity preservation depth-wise dilution
R1 channel-gated (current) 1 stream F(x) + α(x)·x content-dependent skip strength limited path diversity
R2 Qwen Gated Residual 4 branches dynamic read gate + per-branch write control parallel pathways, selective preservation routing and parameter overhead
R3 DeepSeek mHC 4+ branches pre/post routing + doubly-stochastic mixing stable multi-stream transport, identity-like projection cost, implementation complexity
R4 Kimi Attention Residuals depth-indexed softmax over prior layer outputs selective depth-wise aggregation memory/communication cost; may need block approximation
R5 Single-Pass mHC (V4.1-Flash) 4 branches one route_down→route_up pass, single row/column normalization R3's transport at a fraction of the routing cost looser normalization — rows only approximately unit

R1 and R4 — reach. The current channel-gated residual modulates a single skip path at 2048 full-precision parameters per layer, initialized to zero so the module starts as a plain fixed residual and cannot destabilise early training. Kimi's AttnRes goes much further and makes the pathway a retrieval:

$$ x_{l+1} = x_{l} + F(x_{l}) + \mathrm{Attn}\!\left(x_{l},\, \{x_{0}, \dots, x_{l-1}\}\right) $$

The full gated matrix variant of R1 would add ~29M parameters (+4.9%) and make that model the biggest one, confounding the comparison it exists to run — which is precisely the trap the parameter-control rule below exists to catch.

R2 and R3 — width. Both widen the residual state, but they differ in how information is routed and constrained, so a win for one implies nothing about the other. mHC keeps the mixing matrix on the doubly-stochastic manifold:

$$ x_{l+1} = \sum_{k} \alpha_{k}^{(l)} \, x_{l-k} + \beta^{(l)} F(x_{l}), \qquad A \in \mathcal{DS} $$

The constraint is the contribution, not an implementation detail: it restores identity-like signal behaviour and keeps the recombination from drifting into an unstable scaling regime as depth grows. If mHC beats GR, the result to report is that constrained multi-stream routing trades stability against expressivity better than unconstrained gating does. R5 is V4.1-Flash's single-pass take on the same mHC: read/write and the 4×4 mixing coefficients come from one route_down→route_up projection, normalized once instead of by R3's iterative Sinkhorn — so R5 vs R3 (and R5 vs L on a frozen mixer) isolates exactly what the iterative projection buys.

Why this comparison is scientifically useful. Widening the stream and selecting along depth are orthogonal moves — R4 expands what a layer can reach along depth, R2/R3/R5 expand how many parallel routes exist, and neither substitutes for the other. R1 stays in as the necessary low-complexity control, because the question “is multi-stream complexity actually required under ternarization?” has to have an answer. And a residual win only means anything mechanistically if the carrier metrics show the predicted change in norm, contrast, path utilization, depth contribution or branch specialization. To avoid confounding attention with topology, the sweep freezes the best two Phase-4 mixers and changes only the residual design.

Parameter control rule. Every residual family carries a declared capacity budget. Where a family adds parameters, the compensating reduction is solved in non-residual dimensions before the certified run. A comparison that pairs a more expressive residual with a larger total parameter count is a secondary result, not the primary causal one. The diagnostic is a residual carrier profile — branch norm share, residual update/input ratio, gate entropy, cross-layer cosine retention — and the prediction is not “more residual capacity is better” but that the best topology preserves high-value information without letting redundant or unstable pathways dominate.

RQ11 · the capacity question made testable

How much dynamic retrieval is worth buying at 1.58 bits?

Not an assumption baked into an architecture, but a sweep. Ternarization removes fine-grained expressivity from the weights, so a ternary model may benefit from more dynamic sequence-mixing capacity than a full-precision one would. Three axes, moved one at a time, because changing two at once makes the result uninterpretable — and placement and indexer budget are now swept separately too, since “how much retrieval” and “where it sits” are different questions.

Axis 1 · share

DS-25 / DS-50 / DS-75

What fraction of layers do sequence mixing by attention rather than by recurrence. Attention is placed evenly, never clustered, so “how much attention” does not get confounded with “where the attention is”.

Axis 2 · depth

Experiment A — layer count

The number of attention layers moves; head configuration is frozen. Isolates “more attention layers” from “wider attention”.

Axis 3 · width

Experiment B — head count

Head count moves with the layer count fixed, and head_dim is pinned to hidden_size / heads — so total attention width and its parameter count stay constant. That is what separates attention diversity from attention capacity.

Axis 4 · indexer budget

Experiment C — placement, then retrieval capacity

Where the dynamic retrieval layers occur with their count fixed, then how much indexer and head capacity they get with the structure fixed. That answers whether retrieval capacity or placement is the binding constraint — the question QSA's micro-block indexer makes newly worth asking.

Not an axis here

Residual topology, and MoE

Both change something other than sequence mixing — the residual pathway is its own Phase 4-C sweep, and the sparse MoE FFN changes active-vs-stored parameters plus routing, so it is a Phase-6 ablation (Q, R). Folding either into this sweep would confound it. Kimi, Qwen and DeepSeek are all MoE at full scale; only their sequence mixers and residual topologies are borrowed here, never their expert routing.

Experiment B varies head geometry on purpose, which the fairness rules otherwise forbid. That needs an explicit opt-in — --phase4 --allow-attention-budget — and the differences are then recorded as declared, not waived. The comparison is fair within the sweep, which is exactly why Phase 4 reports on its own variant set.

What each comparison answers

Twenty-seven questions, each tied to a specific pair.

Written down before the runs, so a result cannot be retrofitted to a question that happens to have been answered.

RQ1How much quality is lost when the Qwen FFN weight matrices go to 1.58-bit ternary (MLP-only)?A vs B
RQ2Can KDA recover the degradation from MLP-only ternarization?B vs C
RQ3Do attention residuals add further recovery?C vs D
RQ4Do quality changes coincide with representational work redistributing out of the degraded FFN weight channel?all carriers
RQ5What is the quality/efficiency Pareto frontier?all
RQ6Does Mamba-2 recover the MLP-only ternary loss through a different sequence mixer?B vs E
RQ7Which linear-attention family survives ternarization better?C vs E
RQ8Do attention residuals also help the state-space hybrid?E vs F
RQ9Does local + compressed attention beat plain ternary attention, or the Mamba-2 hybrid?B vs G, E vs G, E vs H
RQ10As attention structure changes, how does representational work redistribute — and is the trade worth its GPU-hours?B, E, G, H
RQ11Does ternarization make more attention capacity worth buying?DS-25/50/75, then depth & heads
RQ12Does Qwen's recurrent memory + sparse retrieval beat KDA and the DeepSeek hybrids under ternary weights?C vs G vs I
RQ13With the mixer held fixed, which residual topology preserves the most information?I vs K vs L vs M
RQ14Is a widened residual stream actually necessary, or does the cheap channel gate capture the effect?R1 vs R2/R3
RQ15Does predictive supervision improve the backbone, and what does it cost to serve?K vs N
RQ16Does a learned per-channel scale recover the weight contrast ternarization removes?B vs O
RQ17Does magnitude-aware (grouped) activation quantization recover the activation carrier?B vs P
RQ18Does a controlled ternary MoE buy quality per active parameter, and at what stored cost?B vs Q, I vs R
RQ19Does the Causal Encoder-Decoder — one shared global KV — recover ternary-weight quality?B vs DS41-1
RQ20Does CSA2 improve ternary models relative to standard attention?B vs DS41-2
RQ21Do CED and CSA2 combine additively, or interact?DS41-1 vs DS41-3
RQ22Does hierarchical sparse indexing preserve quality while cutting attention compute?DS41-3 vs DS41-4
RQ23Does single-pass mHC improve information retention vs iterative-Sinkhorn mHC?L vs R5 · DS41-5
RQ24Can conditional Engram memory compensate for capacity lost to ternarization?B vs DS41-6
RQ25Which information carrier changes when the V4.1-Flash mechanisms recover performance?all V4.1 variants
RQ26What does FP4 KV caching cost in quality, and save in memory/throughput?{ternary_mlp, BitFuse} × {bf16, fp4 KV}
RQ27(future) Does DSpark-style speculative decoding speed generation at equal quality?ternary_mlp vs BitFuse vs + DSpark

RQ4 is the one that distinguishes explanation from coincidence, RQ7 is what the two rival families were added to settle, RQ19–RQ25 are the DeepSeek-V4.1-Flash family isolated one mechanism at a time, RQ5 is the Pareto question the others are instruments for, and RQ26–RQ27 look ahead to FP4 KV caching and speculative decoding.

RQ5, drawn

Five competing claims on the same freed memory, and a phase that owns each decision. Drag the split; the numbers are structural, not measured.

Phase 5-A · multi-token prediction

Multi-token prediction, treated as two experiments rather than one.

Qwen reports MTP as a native component of its Next/3.8 lineage, using multi-step training to keep training consistent with speculative decoding while also improving the backbone; Kimi K3 retains an MTP layer too. BitFuse therefore treats it as both a training-objective experiment and an inference-efficiency experiment — never as an ordinary architecture-only ablation, because those two outcomes can move in opposite directions and merging them would hide it.

Design Primary outcome Secondary outcome
MTP-1 no MTP — the reference PPL / benchmark —
MTP-2 predict 2 future tokens backbone loss + acceptance rate decode throughput
MTP-3 predict 4 future tokens backbone loss + acceptance rate training overhead
MTP-4 multi-step consistency schedule speculative-decoding acceptance stability / calibration
MTP-5 MTP on the Qwen-vs-DeepSeek best mixer architecture × prediction interaction Pareto efficiency

It has to be measured on a real serving engine. Acceptance rate and end-to-end throughput are the practical motivation for MTP, so they need an engine capable of speculative decoding rather than a synthetic estimate. And the report keeps language-model quality strictly separate from serving acceleration — otherwise a faster decoder gets mistaken for better training, which is the most likely way this particular experiment could produce a misleading headline.

Phase 5-B · multimodal

Does the carrier thesis survive leaving language?

Qwen3.8-Flash-Next is multimodal and Kimi K3 trains language and vision jointly in one shared backbone and context. That motivates a broader test — but the first target here is deliberately the smallest experiment that can answer whether a ternarized language backbone preserves multimodal signal at all. Four staged gates, each of which has to pass before the next is worth running.

M1adapter-first

Frozen vision encoder + trainable projector + ternary LM

Image–text instruction and captioning. Gate: does the backbone train stably and retain its text quality? If it does not, nothing downstream is interpretable.

gate · backbone stability + text-quality retention
M2

Partially trainable vision adapter + ternary LM

Cross-modal reasoning. Gate: does information transfer improve without destabilizing the language model?

gate · transfer gain without LM regression
M3

Early-fusion multimodal tokens

Joint context modelling. Gate: can the hybrid attention and residual topology actually route visual and textual information together, rather than one crowding out the other?

gate · joint routing evidence
M4

Best Phase-4 architecture, multimodal

The full controlled comparison. Gate: does the architecture that won on long-context language remain the winner cross-modally? A “no” here would be one of the more interesting results in the programme.

gate · does the ranking transfer?
The vision encoder is reported separately, always. Its parameter count never enters the language-backbone comparison, and it stays frozen initially — because the BitFuse question is how the ternarized language backbone redistributes information when the input space gets richer, not how much a good vision tower is worth. If multimodal performance only improves alongside a large vision-side budget, that improvement must not be attributed to the ternary backbone without a matched control. Full end-to-end vision–language training is a later extension, and only if the controlled adapter experiment succeeds first.

Phase 6 · ternary recipe & MoE

Two ways to sharpen the ternary recipe, and one way to spend the savings.

Four ablations sit behind the dense comparison, each reported in its own table so a change to the training objective or the parameter accounting can never leak into the primary perplexity numbers. Two refine the ternary recipe itself without leaving pure \( \{-1, 0, +1\} \) weights; two ask whether the freed budget is better spent on conditional capacity.

O · 6-T

ternary_mlp_learned_scale

A trainable per-output-channel scale γ applied after the ternary matmul, initialised to 1.0 so step 0 is byte-identical to ternary_mlp. The weights stay pure ternary; only the read-out is rescaled. About +0.2M parameters (~0.05% of non-embedding), scored against B.

P · 6-T

ternary_mlp_magact

Magnitude-aware int8 activations: a per-token, per-channel-group absmax scale (group 128) instead of one scale per token. Zero extra parameters, and kept causal — the scale only shrinks within a token — so KV-cached decode stays bit-identical. Scored against B.

Q · 6-E

ternary_mlp_moe

The dense FFN becomes a sparse MoE — a full-precision top-2 router over eight ternary SwiGLU experts, each sized so the two active experts equal one dense FFN. Active parameters match B; stored capacity is roughly 2.3× larger and reported separately.

R · 6-E

ternary_mlp_gdn_moe

The same MoE FFN on top of the GDN + QSA mixer, asking whether conditional capacity stacks with the leading dynamic sequence mixer or merely duplicates what it already carries.

The router is never ternarized. It is named gate so the name-matched conversion pass skips it — quantizing a routing decision would corrupt it for a negligible saving, the same rule that protects the CSA2 and QSA indexers. And because the MoE variants change the training objective (a Switch-style load-balancing auxiliary loss) and the parameter accounting (active versus stored), Q and R are compared only on their active top-2 count and always live in their own table — never in the Phase 2–4 perplexity comparison.

What “controlled” means here, enforced in code

Every model is randomly initialized. The pretrained Qwen checkpoint is used only to confirm the architecture and tokenizer load correctly — never as a starting point. Held identical across all runs: tokenizer, dataset and revision, data ordering, token budget, sequence length, optimizer, warmup, gradient clipping, validation set, the seed set (42, 43, aggregated as the mean with the spread reported), checkpoint policy, evaluation protocol, hardware class.

That is not enforced by discipline. A fairness audit loads each variant's run manifest and compares every controlled field against an allow-list of the switches that are the experiment; any mismatch marks the whole comparison NON-CONTROLLED_EXPERIMENT and the reporting CLI exits non-zero, so an invalid comparison cannot pass silently in automation. Comparing the full model config against an allow-list — rather than a hand-maintained list of things to check — is what catches a hidden change like rms_norm_eps drifting on one variant.

The training recipe is unified. fp_baseline and every ternary variant use Qwen's plain cosine LR with constant weight decay, so fp_baseline → ternary_mlp carries no declared schedule difference — the MLP-only FFN weight precision is the only change. (A two-stage cosine schedule remains available as a possible future ablation of the BitNet b1.58 recipe, but no variant in the study uses it.) When a per-variant difference is intentional, the audit records it explicitly in a declared_differences list instead of waiving it silently — and only schedule fields may ever be declared. Declaring an unequal token budget raises, because that is not a recipe difference, it is a different experiment.

The shared core budget — embeddings, the per-layer mixer sized to the attention slot, the uniform SwiGLU MLP and the norms — is held fixed to within ~0.1% on the primary chain at the live 0.6B scale, never shrunk to compensate; whatever a mechanism adds beyond that core is reported transparently as its additional budget rather than hidden. Two families sit a little wider and are reported as-is rather than compensated away — the Kimi K3 family at −1.0% (its low-rank Gated MLA lands under the shared budget) and the KDA×QSA cross-family at +2.6% (its gated sparse-attention slot is heavier). That matching took deliberate tuning: KDA has no grouped-query KV saving and naive settings inflated it by 25%, a Mamba-2 mixer has an entirely different footprint, and the residual's full gated mode would have added ~29M parameters (+4.9%) — which is why the default is channel_gated at 2048 parameters per layer. Without that, “KDA is better” could just be “KDA is bigger”. Secondary matched-compute comparisons may relax the parameter rule, but they have to be labelled as such.

Four rules keep the expansion sets honest, and each one exists because a post-core mechanism would otherwise contaminate the primary table. MTP changes the training objective, so it is a separate Phase 5 extension set and is never merged into the Phase 2–4 perplexity comparison. The sparse MoE FFN changes both the training objective — a load-balancing auxiliary loss — and the parameter accounting, so it is compared on its active top-2 count with stored expert capacity reported separately. Multimodal runs report vision encoder and connector parameters separately, with language-backbone comparisons kept matched wherever that is possible at all. And every residual family arrives with a declared capacity budget, paid for outside the residual dimensions before the certified run.

Everything else holds as before: the same FineWeb-Edu revision, split, tokenization, shuffle seed and packing policy; the same optimizer family and schedule unless the experiment is an optimization study; the same precision, CUDA stack, kernel policy and evaluation environment; the same fixed cached validation set and cadence. Fake-quantized ternary measurements stay separate from genuine low-bit kernel measurements. Provenance recorded per run: git SHA, config hash, data revision, hardware, token count, checkpoint lineage, kernel versions, and every eviction and resume.

What this is not

  • Not a speed claim. Ternary weights here are simulated (fake-quantized), which is slower than fp16, not faster. Efficiency reports are labelled algorithmic vs kernel_optimized and the two are never mixed.
  • Not a state-of-the-art attempt. Optimizing for the best number would defeat the purpose; each architectural change has to stay attributable.
  • MoE stays a Phase-6 ablation. The sparse MoE FFN (Q, R) is wired but deliberately sits behind the dense comparison, precisely to preserve attribution: learn which dense dynamic mechanisms are worth having before asking whether ternary savings are better spent on experts.
  • Not a reproduction of anyone's frontier model. Mechanisms are borrowed from Qwen3.8, DeepSeek-V4.1-Flash and Kimi K3 one family at a time — the single deliberate exception being the KDA×QSA cross-family (XF-1, XF-2), which pairs Kimi's recurrent memory with Qwen's sparse retrieval as its own controlled comparison. Their fine-grained expert routing, their FP8/FP4 storage and their full multimodal stacks stay out — the Phase-6 MoE here is a plain top-2 Switch-style router, not their expert system.
  • Not a claim that these effects are already reproduced. Everything attributed to an external paper is a literature input motivating an experiment, and is labelled as a proposal until BitFuse has run it.

Two bugs that argue for the whole methodology

Both were in the KDA kernel, both produced plausible-looking output, and neither was visible from a loss curve. The state decay was applied to the wrong axis — the state is \([d_v, d_k]\) and the gate indexes \(d_k\). And the intra-chunk term needs \(\exp(c_t - c_s)\) over cumulative log-decay; computing it as k / exp(c) overflows for long chunks or steep decay, so it is now split around the chunk midpoint to keep both exponential factors bounded by 1.

A chunk-parallel training kernel is asserted equal to an explicit recurrent reference to 2e-4 across chunk sizes, sequence lengths, initial states and streaming continuation. Without that reference to diff against, both bugs would have silently degraded models C and D and been misread as “KDA doesn't help.”

A third one is worth naming because it is invisible by construction: overriding _init_weights without calling super() leaves the RoPE inv_freq buffers zero-filled after from_pretrained, because they are non-persistent and absent from checkpoints. A resumed model then produces different hidden states from the one that was saved — identical weights, no error. It surfaced only as a save/load roundtrip test failing at ~1e-3. Every resumed training run would have been silently corrupt.

Sequencing

Each phase is a gate, not a milestone.

GPU time is the dominant cost, so the campaign is split into phases launched independently. Only continue if the previous phase looks right — and the later scripts refuse to start if an earlier phase's artifacts are missing or were trained at a different token budget, so an unusable comparison cannot be discovered after paying for it.

Localsmoke

Do all thirty-six switches execute at all?

Reduced models, one optimizer step — it produces no comparable numbers. It exists to fail cheaply: forward hooks count real KDA / Mamba-2 / CSA2 / CED / GDN / Engram / MoE / residual invocations, so a wired-but-uncalled path fails here rather than on a provisioned GPU VM. Runs on any CUDA GPU or Apple-Silicon MPS, auto-detected (device=auto resolves CUDA → MPS → CPU).

gate · every switch executes and is actually called
Phase 0no GPU spend

Engineering validation, then a local reduced-width sweep

Reference load, forward/backward/step, strict ternarity, KDA equivalence against an explicit recurrent reference, checkpoint round-trip — then all thirty-six variants trained to ~5M tokens at reduced width through the real runner, plus a benchmark harness for throughput and memory. It ranks the variants under one budget, carried with a stated parameter-confound caveat, and is never mistaken for a Phase-1 result.

gate · code trusted enough to spend GPU money on
Phase 110M tokens

One variant per family, on real GPUs, on disposable infrastructure

FP, ternary KDA and ternary Mamba-2 at 10M disposable tokens, then the resource group is destroyed so nothing keeps billing — teardown downloads the evidence first and refuses to delete anything if the download is empty. This is also the gate for the carrier instrumentation, because a diagnostic problem and an architecture problem must never be debugged at the same time.

gate · real kernels and instrumentation work
Phase 21B tokens

FP + ternary + KDA at full budget

A is retrained here rather than reusing Phase 1's throwaway run, because the fairness audit needs a baseline recorded under the identical contract.

gate · certified evidence for ternarization and KDA
Phase 31B tokens

The residual arms, then the certified six-way core report

D, E and F trained and evaluated, and the whole core six reported together.

gate · core fairness audit passes
Phase 4A100M → 1B

Qwen GDN + QSA

Pilot at 100M to establish it is numerically sound, then a finalist run at the certified budget.

gate · numerically and causally interpretable
Phase 4B100M → 1B

DeepSeek-V4 legacy hybrid attention

The V4 legacy swa/csa/hca track with a Lightning Indexer (G, H, deepseek_impl="v4_hybrid"), kept as the historical control — deliberately not the current DeepSeek branch, which is the Phase 4-D V4.1-Flash family.

gate · the DeepSeek comparison stays controlled
Phase 4C100M screen

The residual family sweep — R1 / R2 / R3 / R4 / R5

Best two Phase-4 mixers frozen, topology the only thing moving — including R5, the V4.1-Flash single-pass mHC sibling of R3. Screened at 100M, finalists promoted to the certified budget.

gate · best topology identified at matched capacity
Phase 4D100M → 1B

DeepSeek-V4.1-Flash family (DS41-*)

The current DeepSeek branch, isolated one mechanism at a time on the ternary backbone: CSA2, the Causal Encoder-Decoder, the hierarchical sparse indexer, single-pass mHC and research-scale Engram (DS41-1…DS41-7), all reported against ternary_mlp.

gate · each V4.1-Flash mechanism attributable on its own
Family lines100M → 1B

Kimi K3, Qwen3.8-Flash-Next, and the KDA×QSA cross-family

Three more comparison sets, each isolating a family's own mechanisms on the ternary backbone: Kimi K3 (K3-1…K3-5, vs ternary_mlp_kda), Qwen3.8-Flash-Next (Q38-1…Q38-3, vs ternary_mlp_qwen_hybrid) and the KDA×QSA cross-family (XF-1, XF-2). Screened at 100M, finalists promoted to the certified budget; none folds into the six-way core.

gate · each family attributable on its own set
Phase 4 · RQ11targeted

Attention share, placement, head and indexer budget

DS-25/50/75, then placement with the count fixed, then capacity with the structure fixed.

gate · optimal dynamic-capacity allocation identified
Phase 5A100M → 1B

Multi-token prediction on the best dense architecture

1/2/4 heads plus the multi-step consistency schedule, measured on a speculative-decoding engine.

gate · quality and speculative-decoding evidence, reported apart
Phase 5Badapter-first

Multimodal integration

M1 through M4, expanding only after each gate — vision parameters always reported separately.

gate · cross-modal evidence without backbone confounding
Phase 6deferred

Ternary-recipe ablations and a sparse MoE FFN

Only after the dense architecture is frozen. The learned-scale and magnitude-aware activation refinements (O, P) scored against ternary_mlp, then a top-2 MoE FFN (Q, R) at roughly equal active compute — routed and stored expert parameters kept separate from the backbone in the capacity ledger, each in its own table.

gate · the conditional-capacity reinvestment decision

The causal order is the point of the sequence, not the schedule: A→B isolates ternarization; B→C and B→E test two rival recurrent compensations; C/E→residual isolates topology; B/E→Phase 4 tests heterogeneous context allocation; the best mixer feeds the residual sweep; the best dense architecture feeds MTP and multimodal; and only the best overall dense architecture feeds Phase 6.

What gets measured

The headline number is not accuracy. It is marginal return per unit of resource, computed separately for each allocation axis:

$$ \mathrm{Eff}_{\text{attn}} = \frac{\Delta q}{\Delta R_{\text{attn}}}, \qquad \mathrm{Eff}_{\text{ternary}} = \frac{\Delta q}{\Delta R_{\text{ternary}}} $$

where \(q\) is task quality and \(R\) is spent resource — weight memory, KV-cache, latency, training FLOPs, or energy where it can be measured honestly.

Alongside that, every run is instrumented internally, because a result without a mechanism is not worth much: ternary distribution and zero fraction per layer, activation magnitude, variance, contrast and effective rank, attention entropy, concentration, span and head diversity, recurrent-state occupancy, and how much earlier-layer information the residual pathway is actually carrying. The instrumentation is emitted at three levels — light, standard, full — and Phase 1's job includes confirming all three hold finite values before anything commits to a 1B-token run.

Each new mechanism brings its own diagnostics, each chosen so its mechanism could fail visibly rather than quietly. Weights gain layer-wise quantization error and scale drift. Attention gains QSA and CSA2 indexer recall and coverage, selected-block mass, and compressed-pool visibility statistics. Recurrent state gains GDN and KDA retention-and-forgetting curves. Residuals gain GR branch usage, mHC mixing entropy and deviation from double stochasticity, and AttnRes depth weights. MTP gains auxiliary-token losses, token acceptance rate and speculative-decoding speedup. The MoE FFN gains router load-balancing loss and expert-utilization histograms. Multimodal runs gain cross-modal retrieval and probe accuracy, image-token saliency, and modality-conditioned norm and entropy profiles.

A research-grade telemetry layer sits underneath all of it. Every run also records the per-step optimization trajectory and loss/gradient stability, a fixed probe set, and the internals of each moving part — MoE routing, the MTP heads, the hybrid mixers, attention logits, and the residual and gradient paths — plus efficiency benchmarks and post-hoc interaction, component-attribution, information-flow and stress-test analyses. It exists so a quality difference can be traced to a mechanism rather than asserted.

Reading the ternary weight carrier, straight from the checkpoint

The weight carrier gets its own static report — computed from the saved {−1, 0, +1} weights alone, with no data and no forward pass, so it can be recomputed from any checkpoint (including the Phase-0 ones already on disk, with no retraining) and kept as a lightweight result long after the full weights are reclaimed. One duck-typed interface — a module that exposes quantized_weight() — discovers the ternary layers in every family, so fp_baseline simply reports is_ternary: false and each new mechanism is picked up automatically. Every number is tagged static or dynamic, so its provenance — the weights alone versus a forward pass — is never ambiguous.

For every ternary tensor it records the full distribution — zero / positive / negative density, the positive and negative fraction among the non-zeros, and a sign balance in \([-1, +1]\) (\(+1\) and \(-1\) are never assumed equiprobable) — plus the ternary entropy in bits (max \(\log_2 3 \approx 1.585\)), the sign entropy, the per-tensor scale statistics, and that same distribution pooled per parameter, per module category and per decoder depth, so “which layers went sparse, and where” is answerable rather than averaged away. It is the evidence for the study's core question — where information relocates once ternarization removes the weights' precision — read from the weight channel itself.

BITCOS is a storage rate, not a smaller alphabet. The report also computes the BITCOS effective bits/weight, 2 − zero_density, against a conventional 2-bit baseline and a five-trit-per-byte packing (\(3^5 = 243 \le 256\), so \(8/5 = 1.6\) bpw). When a tensor is sparse enough (zero_density > 0.415) that rate dips below the ternary information bound \(\log_2 3\) — which is compression of a sparse symbol stream, not a change of alphabet: the model is still ternary over {−1, 0, +1}. Ternary entropy is the information content of the distribution; BITCOS bpw is a storage rate under one coding scheme — reported side by side, never conflated, and no BITCOS inference kernel is claimed.

Resumable training, bounded storage

A campaign this size — the thirty-six variants, more once scale profiles and budget sweeps are counted — cannot keep every finished model's full weights on disk. Training is fully resumable: each checkpoint stores the complete state (model, optimizer, scheduler, AMP scaler, token/step counters, the data-stream cursor and all RNG), written atomically with a checksummed manifest, so a crash, restart or Spot eviction continues exactly rather than restarting from zero — and every checkpoint is stamped with its experiment_id, so one run can never resume from another's state. But a run is marked COMPLETED only once training and metric collection have finished, and only then does it persist its lightweight ternary results and reclaim the large weight payloads, keeping the provenance sidecars and a cleanup.json marker. So disk grows with the number of running experiments, not finished ones — a fixed ceiling of one latest + one best + a bounded periodic history per active run — while the final results never depend on retaining the trained weights. Each run advertises its lifecycle — RUNNING, RESUMABLE, COMPLETED, FAILED — and a completed one is skipped, never silently retrained.

Four outcomes, all of them informative

  • Ternary scaling wins. Dynamic capacity is not an efficient substitute for parameters at low precision — which retires a plausible idea and tightens the case for current practice.
  • Allocation wins. A smaller ternary backbone with richer retrieval, residual or predictive machinery beats a larger one at matched resources, and low-bit scaling has been leaving quality on the table.
  • They are complementary. Both axes help, and the useful contribution becomes the shape of the joint frontier.
  • Allocation is task-dependent. Short context favors weights, long context favors selective retrieval, reasoning favors conditional capacity — the most interesting outcome, because it means there is no single correct low-bit architecture, only a correct allocation for a workload.

What counts as a strong result

A strong result is emphatically not “the newest architecture wins.” It is a controlled trade-off that survives matched-budget comparison and has a plausible mechanistic explanation behind it. So the claims are pre-committed with their interpretations attached:

  • If Qwen GDN + QSA wins, the claim is about the usefulness of fixed-size memory plus selective retrieval under ternary weights — not about Qwen being superior.
  • If mHC beats Gated Residual, the result is that constrained multi-stream routing offers a better stability/expressivity trade than unconstrained gating.
  • If AttnRes wins, the contribution is evidence that depth-wise information selection is a carrier separate from parallel residual streams.
  • If MTP improves serving but not validation loss, that is still a meaningful efficiency result, and the paper labels it as exactly that.
  • If multimodal performance improves only with a large vision-side budget, the result is not attributed to the ternary backbone without a matched control.
  • Any fairness-audit failure marks the corresponding delta NON-CONTROLLED_EXPERIMENT and removes it from the primary causal claims entirely.

Working thesis

Extremely low-precision models should probably not be scaled by uniformly adding more low-precision parameters. Because ternarization redistributes where representational capacity actually lives, the more efficient strategy may be to spend the budget on the mechanisms that select, route, preserve and predict information — recurrent memory, selective retrieval, residual topology, prediction heads, and eventually conditional computation — with ternary weights providing the cheap substrate underneath. The shift under test is from scaling the weights to allocating the information pathways.

Status: proposal v7, updated 20 September 2026. All thirty-six variants — the core six, the DeepSeek-V4 legacy pair, the Qwen GDN + QSA and DeepSeek-V4.1-Flash (DS41-*) families, the Kimi K3 and Qwen3.8-Flash-Next lines, the KDA×QSA cross-family, the six residual topologies, MTP, the two ternary-recipe refinements and the sparse MoE FFN — are wired on one config class and pass the local smoke gate and the reduced-width Phase-0 sweep, trained at the live Qwen3-0.6B scale (~596M). The fairness audit, the carrier instrumentation — now including the static ternary weight-distribution / BITCOS report and the research-grade telemetry layer — the bounded, resumable checkpoint storage and the phased Azure runbook exist and pass. Nothing has been trained at the certified 1B-token budget yet, and the multimodal Phase 5-B is still specified rather than implemented — so nothing on this page is a result. If you work on low-bit training, linear-attention kernels, residual topology or capacity-matched evaluation and want to argue with the design before compute gets spent on it — that is still the most useful thing anyone could do. Reach me via Medium or GitHub.