One Object: Memory, Navigation, and Reportability in an Operator-Only Language Model

Preprint v1.3 — 2026-09-10. DOI: 10.5281/zenodo.22684392. Versioned; dated revisions supersede (the DOI resolves to the latest version). §12 reports an ongoing experimental program and carries its own dates.

Cite as: Sutherland, G. (2026). One Object: Memory, Navigation, and Reportability in an Operator-Only Language Model. Zenodo. doi:10.5281/zenodo.22684392

Garret Sutherland — MirrorEthic LLC


Abstract

We study cl33-opLM, a language model whose only path from computation to output is a stream of explicitly emitted Cl(3,3) bivector operators: a transformer emitter produces per-block operators, a reversible SO(3,3) rotor scan transports state, attention scores are η-metric products of rotors, and the readout sees exclusively operator-derived features. Because generation, memory, steering, and inspection all act on the same algebraic object, questions that are normally interpretive become measurable. We report four results. (1) Two memories: associative recall (MQAR) and relational navigation are mechanistically distinct in this architecture — recall rides a token-similarity induction shortcut that bypasses the operators entirely (operator-dependence ≈ 0), while navigation is operator-native (operator-dependence ≈ 0.85, scale-invariant); we show the recall shortcut and the bypass are the same operation at matched training budget. (2) Parity where it counts: on permutation-composition navigation — a known transformer weak spot — cl33 matches a learning-rate-tuned transformer baseline exactly (0.433 ± 0.009 vs 0.432 ± 0.030) with 3× lower seed variance and a fully inspectable mechanism. (3) The operator stream is a transcript: a probe trained only on emitted operators decodes the current token at 0.86 top-1 over the 50k vocabulary — ~9.6 of ~11.2 bits of token identity per emission, with 80.7% of running text reconstructed verbatim from the operator record alone, and identity confirmed under substitution. The model's only channel to output doubles as a readable record of what passed through it. (4) An architectural workspace: applying the Jacobian-lens verbalizability estimator across the model's depth shows the emitted operators are the most output-transparent representations in a three-model comparison (final-agreement 0.613, above both GPT-2 controls' best layers; replicated on a second model), while the transported state is completely non-verbalizable (0.000) — a split that holds across two held-out domains. The workspace/ocean dichotomy that emerges statistically in standard transformers is, here, a line in the wiring diagram. We frame what the architecture buys (transparency, reversibility, exact memory, causal steering surface), what it costs (a real, measured tax: +29% bits-per-byte and ~5 points average MCQ against a param-and-context-matched from-scratch control at 5B tokens — though continued prose training moves the same model to exact bits-per-byte parity with token-matched Pythia-160m (1.155 at 6.2B tokens; caveats on ruler alignment in §10), showing the tax is data-bound rather than an architecture floor; we claim no perplexity or reasoning superiority), and what it does not claim.

On the axis the architecture is built for, we report capability results with matched baselines that collapse or cannot produce the artifact: exact retrieval at load (1.000 vs a matched transformer's ≈chance at T=1024/KV=64), editable memory with transaction semantics and bit-exact undo, and causally-faithful provenance receipts (attribution@1 1.000 vs a transformer best-head 0.017). These were measured on a reversible-tape intermediate, and its honest failure mode — attention re-carries recall unless the memory is the sole cross-segment channel (§5.4) — is what taught the sole-channel-by-construction design of the current memory organ (§12). That organ's first belief-install battery (install 0.750, algebraic delete 0.000, receipts 24/24) is reported as a pilot. We close with the governance reading: because the operator record is the mandatory carrier of the computation (ablating it multiplies perplexity by orders of magnitude) and the reversible scan makes that record losslessly replayable (exact to 2e-14), the perplexity tax inverts into the price of an audit guarantee that post-hoc interpretability cannot provide at any cost — with the guarantee scoped to the record's mandatoriness, replayability, editability, and decodability, not correctness, and with the scan's own causal share at LM inference measured (independently) as decorative (§3.2).


1. Introduction

Interpretability of standard transformers is archaeology: the artifact is finished, and we dig. Probes, sparse autoencoders, and lens methods recover correlates of computation, with faithfulness always in question because the object recovered is not the object the model computes with.

cl33-opLM inverts the arrangement. The architecture is chosen so that the model's only route to the output runs through a typed, invertible, algebraically-structured object — a per-token, per-block set of Cl(3,3) bivector operators. There is nothing behind the operators to be unfaithful to: the readout is constructed to see operator-derived features only, and we verify the construction empirically (ablating operators multiplies perplexity 8–42× at 32M, 206× at 163M — a ratio that climbs with capability). Memory is an operation on operators (§5). Navigation is composition of operators (§4). Reportability is a property of operators (§9). Steering is addition of operators (App. C). One object.

The purpose of this paper is not to claim superiority over transformers at language modeling. The transparent model pays a real, measured premium against a param+context- matched from-scratch control at 5B tokens (+29% bits-per-byte, §10 Table 1) — and the same machinery-bearing model, continued on prose to 6.2B tokens, reaches exact bits-per-byte parity with token-matched Pythia-160m (§10 Table 1b). Both numbers are reported: the matched cost of the machinery at a fixed budget is real; the ceiling it was once mistaken for is not. The claim is narrower and, we think, more useful: when the generative object is algebraic and exclusive, capabilities and their mechanisms become the same thing, and questions like "does the model use its history as a trajectory or a bag?" or "which internal content is reportable?" stop being interpretive and acquire experiments with pass/fail readouts. Every positive claim in this paper carries an operator-ablation control (the bypass blade): if a behavior survives with the algebra removed, it is a shortcut and it does not count — including our own memory mechanism, which we first built wrong and report as such (§5.1).

The intuitive form of our results: a transformer does not carry a transcript of words — it carries a trajectory of prediction-relevant events from which the words are mostly recoverable. This is true of standard transformers too (our emitter's hidden state decodes tokens at 0.91), but there the events and their accumulation are entangled in one vector. cl33 factors the process into inspectable parts: the operators are the event record (token identity decodes at 0.86, degrades to near-synonyms because the record stores what a token does to prediction, not its orthography — §7); the transported state is the accumulated path (literally s_t = R_t···R_1·s_0: zero token content, non-verbalizable — §3, §9; its causal share is smaller than this factoring once suggested — the §3.1 audit finds flow rides the operator features, with the state a passenger on the probes run so far); and generation reads both through a mandatory interface (bypass 8–42× at 32M, 206× at scale; emission loss 0.48 bits — §2, §7). Every major result in this paper — the two-memory dichotomy, the tape's reader taxonomy, the order-readout null, the workspace/ocean wall — is a facet of this factorization.

A living artifact. The claims in this paper are demonstrated interactively at cl33.t3atlas.dev: a transparent-chat interface over the trained model in which every response is accompanied by its operator transcript, the reverse readout (the §7 inversion run live — the conversation decoded back out of the operators), rewind, and steering. The demo is not an illustration of the paper; it is the same model, same algebra, same receipts, publicly poke-able — and the bottleneck claim is measured on the serving checkpoint itself: ablating the operators multiplies the demo model's perplexity 314× (the 8–206× ratios elsewhere in this paper are the training-lineage measurements at their own scales). Where a reviewer would ask "does this survive contact with arbitrary input," the demo is our answer.

History, reported as science. This architecture's development includes a falsification worth stating plainly rather than politely omitting, because it explains the design and produced one of its cleanest results. The original cl33-opLM carried its recurrent state as a grade-1 6-vector under the grade-preserving rotor action — and its associative memory failed structurally (MQAR ≈ 1/K at every load, §3): a key→value binding is a grade-2 bivector, and a grade-1 carrier has nowhere to hold one. A four-arm causal ladder built in the successor Genesis program isolated the error exactly — the grade-1 arm reproduces the failure, the grade-2 wedge store binds exactly, a param-matched generic control shows the fix is structure rather than dimension (mqar_left_regular.py, reproducible). For a period the language-model line was formally closed as superseded on the strength of that diagnosis. It reopened when two things became clear: continued training closed the model-class fluency gap (Table 1b), and the representation lesson could be incorporated without abandoning the architecture — superposition belongs in an explicit grade-2 store, not the grade-1 carrier, which is exactly how the memory results in this paper are constructed (§5, App. F) and how the ongoing organ program is built (§12). The vector representation failed for one memory formulation; the failure was localized by experiment; the design absorbed the lesson. We consider this sequence — claim, falsification, diagnosis, reincorporation — the paper's strongest argument that the methodology works.

How to read this paper. The systems description that organizes everything that follows: a neural engine (the transformer emitter — opaque, nonlinear, where the learned computation lives), a geometric bus (the operator stream — mandatory, typed, inspectable, the only path to output), and causal peripherals (memory, sensors, steering — machinery attached to the bus rather than baked into the weights, each attaching under a zero-init discipline that makes it bit-exactly removable). The paper is arranged in three parts. Part I — The Object (§2, §7, §8, §9, §10): the bus exists, transcribes both content and state, is the verbalizability peak while the transported state is the floor, and costs a measured, data-bound tax. Part II — Memory on the object (§3, §4, §5, §6, §12): what memory the bus does not give for free, the falsifications that located the sole-channel requirement, and the register organ that now installs, adopts, and deletes beliefs with receipts. Part III — Governance (§13): what the receipts buy, precisely which reversion claim is being made, and the three doors through which weaker guarantees re-enter. Section numbers are stable identifiers carried from earlier drafts; reading order is the Part structure.

Contributions

  1. The two-memory dichotomy with a mechanistic identity proof: recall-as-induction and routing-around-the-operators are the same operation (§3).
  2. Pre-registered navigation parity with a tuned transformer on permutation composition, with tighter variance and an inspectable mechanism (§4).
  3. The reversible tape as the load-bearing intermediate: exact forward-readable memory native to the rotor scan, with the addressing/content separation principle ("token content may SELECT, only algebra may FLOW"). Its capability results (MQAR 1.000, zero bypass, an attention-free challenger at matched probe scale) hold in sole-channel configurations; its measured failure mode — attention re-carries recall whenever it coexists (§5.4) — is what taught the sole-channel-by-construction design of the memory organ (§5, §12).
  4. The inversion battery: the input words decode back out of the operator stream — 0.86 top-1 over 50k vocab from one emission, +4.7 bits beyond prefix predictability, identity confirmed under fixed-context token substitution, 80.7% verbatim document reconstruction with semantic (near-synonym) degradation (§7).
  5. A claims ladder for self-structure with adjudicated results: transparent recall (closed), structural path-sensitivity (passed), semantic order-readout (null, twice, with a mechanistic explanation), self-modeling (live, carrier localized) (§8).
  6. Jacobian-lens evidence that the workspace/ocean split is architectural in cl33: operators are the verbalizability peak of a three-model study (replicated on two cl33 runs, domain-invariant across two held-out domains); transported state is invisible to the lens (§9) and output-causally inert under direct steering — a residual bias out-steers an operator-field bias ~25× with no positional compounding at either site (§3.1, closed 2026-09-08).
  7. A capability battery on the axis the architecture is for, each against a matched baseline that collapses or cannot produce the artifact: exact retrieval at long sequence length (1.000 vs a matched transformer's ≈chance at T=1024/KV=64 — with the scoping that addressing is short-range: distance-resolved recall collapses with query–pair gap and is a measured limitation, §5.4 — across a pre-registered fairness sweep), editable memory with transaction semantics and bit-exact undo (provable unlearning-with-receipt), and causally-faithful provenance (attribution@1 1.000 vs a transformer best-head 0.017) (§6).
  8. The measured transparency tax and its governance inversion: +29% bits-per-byte and ~5 pp average MCQ against a param-and-context-matched from-scratch control (honestly reported; no superiority claimed, and the tax itself is now shown to be data-bound — continued training reaches class-ruler bpb parity at matched tokens, §10 Table 1b), and the argument that because the reverse trace is the computation, that tax buys an audit guarantee post-hoc interpretability cannot provide at any cost — scoped to invertibility/editability/decodability, not correctness (§10).


Part I — The Object

The bus exists, transcribes content and state, is the verbalizability peak, and costs a measured tax.

2. Architecture (summary)

(Full spec: CL33_OPLM_ARCHITECTURE.md v2.2; tape extension: REVERSIBLE_TAPE_DESIGN.md, CL33_OPLM_V24_TAPE_LM.md.)

Precision discipline: split precision (amp_emit) — fp16 emitter, fp32 algebra. Full fp16 NaNs from η-score/wedge overflow.


7. Reading the words back: the operator stream as transcript

(Full protocol, pre-registration, and adjudication: INVERSION_BATTERY.md. Design: tiered decoders + prefix-only control + fixed-context substitution, document-disjoint splits.)

Sections 3–5 established the forward direction: the operators are causally necessary for the output (ablation multiplies loss 8–206× depending on scale), and by construction nothing else reaches the readout. This section measures the inverse: given only the emitted tuple O_t = concat(B_state, B_q, B_k) (1440-d per position), can the input words be decoded back out? If yes, the operator record is a bidirectional transcript — input decodable backward, next-token decodable forward (§9), trajectory invertible in the middle — and the residual skeptic's position ("the transformer emitter does the real work behind the bottleneck") has nowhere to live: whatever the emitter computes either passes through the operators (measured here) or does not affect output (ablation, §2).

Setup. Frozen scale-C (163M, ~449k steps), 1.23M positions of held-out mix text across 5 domains, GPT-2 BPE (50,257 vocab, unigram floor 0.082 / 11.16 bits), document-disjoint train/eval splits, shuffled-pairing and Gaussian floors.

Results.

Decoder top-1 CE bits
O_t → tok_t, linear 0.725 3.04
O_t → tok_t, MLP 0.862 1.58
prefix-only O_{t−4..t−1} (control) 0.356 6.44
frozen LM's own head (strongest prefix baseline) 0.425 5.12
emitter hidden h_t (ceiling) 0.906 1.10
shuffled pairing 0.082 (= floor) 10.94

Interventions (identity, not expectation). Holding the prefix fixed and substituting the token at position t, the frozen decoder identifies injected mid-frequency tokens — including semantically strange ones in alien contexts — at 0.82 (plausible 0.94, random-vocab 0.32, rare>20k 0.12, MLP probe). Errors land on the contextually-expected token only 2.9% of the time (linear probe). The channel reads actual token identity, not contextual expectation; identification is frequency-graded at the rare tail — a reading we note is the post-hoc adjudication: the pre-registered prediction that random and rare tokens would identify within 15pp of mid-frequency ones failed (the battery's falsification line — errors collapsing onto expected tokens — did not trigger, which is the part that carries the claim). The same-token-across-contexts test quantifies the blend: one token's operator representation moves substantially with context (within-type variance 65% of global) while remaining decodable (83% mean recovery).

Reconstruction. Decoding entire held-out documents from the operator log alone (±2 window) recovers 80.7% of tokens exactly (per-document 0.52–0.98 — narrowly missing the pre-registered ≥0.85 gate, which we report rather than round), and the misses are near-synonyms — "growth"→"innovation", "petroleum products"→"oil products", "says"→"said". The log reads back as a meaning-preserving paraphrase where it is not verbatim: the identity code degrades semantically, not randomly, consistent with a frequency-graded code on a semantic operator geometry (App. C).

What we claim: the operator stream is a bidirectional, frequency-graded transcript — ~9.6 of ~11.2 bits of current-token identity per emission, 4.7 bits beyond predictability, with identity confirmed under substitution. What we do not claim: verbatim recovery of the deep-rare vocabulary; the code is distributed, context-modulated, and frequency-weighted.

8. The claims ladder: what the model can read of itself

We evaluate self-structure with an explicit ladder, most-skeptical-first framing, and the bypass blade at every rung. (Full protocol and adjudications: SELF_STRUCTURE_CLAIMS.md.)

Claim 1 — transparent recall. The model retrieves via a mechanism we can read. CLOSED by §5: recall is a named read (increment mode) of a lossless record, with operator-ablation destroying it.

Claim 2a — structural path-sensitivity. The model's behavior depends on its trajectory as a path, not as a bag of visited states. Test: counterfactual history — block-swap surgery on the tape record preserving the operator-entry multiset while recomputing the displacements it implies (the P1 intervention; the distinct P2 same-visited-state variant probes the state-mode reader and is not conflated here), injected via tape-override hooks. PASSED for displacement mode: outputs shift when path structure shifts, and the shift vanishes under operator ablation (so it is not a token-order artifact). This is exactly what the §5.3 taxonomy predicts: displacement is the global-path reader; only the path reader should be path-sensitive — and increment/state modes serve as built-in negative controls.

Claim 2b — semantic order-readout. The model can report trajectory order in language (narrative-order probes). NULL, twice (full model, then scan-only retest — which came back null/inverted: records get discounted). Mechanistic explanation: order-content is attention-dominated at readout; the trajectory channel influences prediction (2a) without being linearly decodable into tokens. We initially treated this as a failed experiment; §9 reframes it — the trajectory lives in the non-verbalizable part of the model, and the null is the ocean behaving like the ocean. A CE-pure rule (no auxiliary order-pressure losses just to pass our own probe) is in force; the aux variant is deferred as a side experiment, not a rescue.

Claim 3 — self-modeling (the prize; held skeptically). The model distinguishes its own generation history from a compatible foreign one. LIVE, not passed. Current evidence: a generator-identity signal 4× the content signal, carried on the operator record (blade decomposition: operators-only 0.092 vs tokens-only 0.0012 — the carrier is blade-2, the algebra, not the tokens). The phase-0 probe design confounded live-vs-dead records with own-vs-foreign and was redesigned; the honest current statement is that we can name the boundary — ownership vs familiarity — and have localized the carrier, but the discriminating probe (twin-with-same-seed training; ownership under matched familiarity) has not been run. We commit to reporting it either way.

Blade rule (all rungs). Any behavior surviving operator ablation is bypass and does not count — the same standard the tape was held to.


9. The workspace is a wire: Jacobian-lens results

Instrument. The Jacobian lens (Anthropic) asks, mechanically, what an internal activation is disposed to make the model say: lens_l(h) = unembed(J_l h), with J_l the input-averaged Jacobian from layer l to the final representation. No trained probes. On large models it reveals a workspace: a privileged subspace whose contents are verbalizable, surrounded by an ocean of computation that shapes behavior but is never linearly reportable — a structural echo of Global Workspace Theory's broadcast bottleneck.

Method. Faithful estimator port to cl33 (one-hot cotangents at all valid target positions, source-averaged, SKIP_FIRST=16; batch-replication for tractability). Taps: the six emitter layers h1–h6, the emitted operators B (post-clip), the transported state S (differentiable scan copy). Controls: stock GPT-2 124M and the warm-start +5B-mix GPT-2 control. Two cl33 models are reported — scale-C (163M) and scale-B (236M), both at 5B-token final — so the split is tested across two independent runs. Fitting used domain-balanced prompts (300 for the GPT-2 controls, 60 for cl33), 40 disjoint eval prompts. Metrics: future-token MRR and final-agreement per tap.

Results (final-agreement by depth, in-domain).

Model Shape
stock GPT-2 124M .000 × 7 layers, then cliff: .018 / .033 / .138 / .288
warm-start (+5B) no dead zone; smooth ramp .043 → .583
cl33 scale-C 5B emitter .088 → .517; B = .613; S = .000
cl33 scale-B 5B emitter flat .049 → .097; B = .538; S = .000

Five readings:

  1. No workspace at 124M by default. Stock GPT-2's first seven layers are pure ocean — the pre-registered prediction, confirmed. Everything interesting it computes early is linearly invisible to its own output basis until the end.
  2. The operators are the verbalizability peak of their own stack on every domain tested, and above every control layer in-domain — 0.613 (scale-C), above the warm control's final layer, replicated at 0.538 on the independently-trained scale-B. (One out-of-domain exception, reported: on PubMed the warm control's final layer, 0.563, edges scale-C's operator tap, 0.552.) Not mysterious, and that is the point: in a transformer the lens asks whether a layer happens to contain the output; in cl33 the operators are the only path to the output, so training must put the output there. The workspace is not emergent — it is the bottleneck, made.
  3. The state is completely non-verbalizable. S scores 0.000 on both models — floor. This is the paper's ocean, realized as a named architectural component: the workspace/ocean boundary, statistical in standard models, is here the line between "emitted" and "transported" in the wiring diagram. An earlier draft paired this with "causally load-bearing"; the §3.1 audit weakened that pairing, and the 2026-09-08 constant-bias probe now closes it (§3.1, point 3): at the output, a residual bias out-steers an operator-field bias ~25×, and neither effect compounds with position — the state's accumulated history does not reach the readout. S is the non-verbalizable and output-causally-inert residue of the architecture: non-reportable and, for output steering, ≈0 causal share. The workspace/ocean split is thus doubly clean — what is transported is neither readable nor (at the output) load-bearing. The retro- explanation of the Claim-2b null ("order-content lives on the S side") remains demoted to hypothesis.
  4. Upstream emitter organization is model-dependent; the bottleneck is not. scale-C's emitter pre-organizes — a smooth ramp .088 → .517 into B — while scale-B's emitter stays near-flat (.049 → .097) and jumps at B, closer to the original "dead emitter, workspace at B" pre-registration. Whether the bottleneck's gradient reaches back through the emitter varies with training (scale-B saw ~6× fewer optimizer updates); that B is the peak and S is the floor does not.
  5. The split is domain-invariant (confound closed). The fit/eval prompts above share the training-mix domain, which could inflate absolute levels. To rule this out we re-fit and re-evaluated all four models on two held-out domains none trained on — PubMed abstracts (biomedical) and congressional bills (legal). The two load-bearing signatures survive both: cl33's B remains the verbalizability peak (scale-C B .552 / .751, scale-B B .490 / .721 on PubMed / legal) and S remains at floor (≤ .018 everywhere); stock GPT-2 keeps its dead-then-cliff shape (no mid-workspace) and warm-start keeps its bottleneck-free ramp, on both domains. Levels move with domain; shape does not. One honest scoping point falls out: on legal text — so formulaic that even shallow features predict the next token — the emitter lights up too (scale-C .54–.62), so reading #4's "dead emitter" contrast is itself a property of harder-to-predict text. What is architectural, and domain-robust, is B-is-peak, S-is-floor; how far B towers over the emitter is not.

Caveats. 124M–236M-class models only; first-pass estimator with 60 fitting prompts for cl33 (one repetitive-text lens-fit artifact noted: a single mid-emitter spike on scale-B/legal, with B still the clear peak); the emitter-organization effect (reading #4) deserves a controlled comparison — the same emitter trained without the bottleneck — before strong claims. The 5B-final rerun and the OOD domain-invariance control, both pre-registered gates in earlier drafts, are now discharged.

9.6 A regime sensor on the bottleneck (2026-09-08 — first measurement, one control open)

If the operators transcribe the emitter's state as well as its content, the bottleneck is an instrument, not only a record. First measurement: a plain difference-of-means direction in operator space separates text regimes (fiction vs code) at 0.962 held-out; four regime centroids (fiction / code / math / encyclopedic) classify a single token's emission four-ways at 0.802 (chance 0.25), spread across a multi-axis manifold (PC variance 0.52/0.28/0.19) whose axes each use the full compact+noncompact blade geometry (~uniform sector mass — no sector starvation). The load-bearing control: hold the token fixed and vary only the surrounding domain. Classification survives essentially unchanged (pooled 0.811: " the" 0.905, "." 0.877, "," 0.700) — the signal is contextual computational state, not lexical identity: the same symbol is emitted as a measurably different operator depending on the regime the emitter is in. Because the readout consumes only operators, the material the sensor reads is on the causal path by construction; whether the readout uses the particular direction the sensor reads is a separate question — one the closed-loop steering test below partially answers, and the open control decides. A first closed-loop test (steer the residual toward the math centroid, read the sensor) shows a monotone dose-response with the fixed-token invariance intact, and at sufficient gain the generated text itself turns numeric. One control is open and we hold the stronger claim until it closes: at low gain the sensor crosses the regime boundary before sampled generation does, and sensor and actuator directions derive from the same contrastive means — so the sensor-leads-behavior gap must be adjudicated (faithful early warning vs partial injection echo) by classifying sampled generations across the gain range before "state sensor" hardens beyond "readable projection of upstream computation."


10. Cost accounting and limits

The transparency tax, measured (locked, matched, fp32). With the from-scratch matched control complete, the tax is no longer provisional. On a single-session, full-set, right-padded, fp32 lm-eval-harness (v0.4.11), on a common wikitext ruler at 512:

model params tokens wt-bpb ↓ avg-acc (8 MCQ)
cl33 scale-C 163M 5B 1.309 0.435
cl33 scale-B 236M 5B 1.218 0.440
A1 (vanilla, matched control) 166M 4.9B† 1.012 0.483

The matched vanilla transformer wins bits-per-byte by +29% over scale-C (and +20% over the larger scale-B) and 6–7 of 8 downstream tasks. This is the honest cost of forcing every prediction through a reversible, per-token operator bottleneck. We report it plainly and make no claim of perplexity or reasoning competitiveness. An earlier draft guessed a ≈1.2× premium; the measured figure is larger (~1.29× on the loss ratio) and we correct it. Bits-per-byte is the confound-free ruler; word-perplexity dramatizes the same underlying gap through tokenization and chunking differences, and we do not report it as a cross-comparison.

A scaling dissociation, stated plainly. From 163M→236M, cl33's bpb improves (1.309→1.218) while avg-MCQ is flat (0.435→0.440): 73M extra parameters bought zero downstream accuracy. Downstream MCQ is param-saturated / data-bound at 5B and does not track cl33's own capability scaling — at fixed data it is the wrong yardstick for this architecture (the axis it is built for is §6; Table 1b below shows MCQ does move when the data moves). The model is genuinely fluent (coherent multi-sentence English) and its operator index sharpens with scale (nearest-neighbor same-category rate 0.79 at 5B vs 0.68 at 32M).

Table 1b — the tax is data-bound, not an architecture floor (2026-08-03). The scale-B model, continued on a prose-weighted mix to 6.2B tokens at 1024 context, was benchmarked on the same full-set fp32 right-padded harness (new session; the locked Table 1 above is untouched):

model tokens wt-bpb ↓ avg-acc (8 MCQ)
cl33 prose base (scale-B cont.) 6.2B 1.155 0.451
Pythia-160m @6.3B (class ruler) 6.3B 1.155 0.462
(ref) cl33 scale-B 5.0B 1.218 0.440

The bits-per-byte reaches parity with token-matched Pythia-160m to three decimals (1.15501 vs 1.15452) — the model-class fluency gap closes at matched tokens. Against the (now un-matched, still-5B) A1 control the gap narrows from +20% to +14%; the true matched tax at 6.2B is unmeasured — A1's own convergence (1.017 → 1.012 over its final 0.5B) suggests it sits near +14–15% rather than the +29% headline, but that is an inference, not a measurement. Three honest caveats: the Pythia row is carried over from the model-class context session (fp16; its source table marks it "scale intuition," not precision-matched — the parity claim is therefore cross-session), the prose-weighted mix partially aligns with the wikitext ruler (Pythia's Pile is also prose-heavy, so the class comparison remains roughly fair), and A1 was not extended, so Table 1's matched tax remains the only fully controlled number. MCQ also moved for the first time (+1.1pp avg; BoolQ +5.0pp — notable because a bias-immune re-analysis of BoolQ found cl33 the only model in the comparison with significant passage signal: answer-bias- corrected AUC 0.522, p = 0.037, versus Pythia's raw 0.61 collapsing to chance AUC 0.494 once its yes-prior is removed), qualifying the "param-saturated" reading above: MCQ was data-bound, and prose tokens were what it wanted. The claim discipline is unchanged — no superiority is asserted — but "the bottleneck costs a fixed 29%" is now falsified by the architecture's own training curve: the tax is a property of a checkpoint, not of the algebra.

The governance consequences of this cost table — where the tax inverts, and what exactly is guaranteed — are collected in Part III (§13).

Methodological note that gates the numbers. The lm-eval adapter inherited a left-padding step that corrupts the reversible scan (leading pads inject spurious rotors); it drove one binary task below chance until we replaced it with right-padding (batched output then matched batch-1 exactly). Every operator-model eval must right-pad. All cl33 harness numbers are full fp32 (emitter and algebra); the fp16-emitter split-precision path is training-only and never touched the eval numbers (verified: the eval adapter runs amp_emit=False).

† A1 stopped at 4.915B tokens (its GPU was reallocated ~1.7% short of the 5.0B budget). Its bits-per-byte had converged — 1.017 → 1.012 over the final 0.5B — so the token gap does not move the tax; we report it here rather than paper over it.



Part II — Memory on the Object

What the bus does not give for free, the falsifications that located the sole-channel requirement, and the organ that installs, adopts, and deletes with receipts.

3. Two memories, one identity

Question. When the model retrieves, does it use the operator machinery or route around it?

Setup. MQAR (multi-query associative recall; Beck et al. harness) for associative memory; a KGE-style relational task (compose R·head ≈ tail along paths) for relational memory. Both scored with the bypass blade.

Result (the dichotomy).

Capability Mechanism Works? Operator-dependence
Associative recall (MQAR) prev-token key / adjacency (induction) 0.92–0.97 ≈ 0 (91% bypass in the worst configuration)
Relational navigation operator-trajectory composition groks to 0.95 ≈ 0.85, scale-invariant

Result (the identity). The induction shortcut and the bypass are not two problems — they are the same operation. Recall-by-induction is token-similarity matching that never consults the algebra; the wedge-memory design matrix (§5.1) shows every configuration that recalls well before the tape does so through a token-content value channel, and every configuration that keeps values algebraic fails to recall (≈ 1/KV) at matched training budget — Appendix F shows the algebra-value store eventually groks to full recall at roughly 10× the training cost, so the identity is about learning efficiency and shortcut formation, not impossibility. The rotor scan cannot perform outer-product writes: a rotor is orthogonal under η, so "store this key–value pair" has no native expression in the transported state. This is the precise gap the tape (§5) is designed to close, and the reason its design rule is addressing/content separation rather than "more algebra."

Interpretation. Prior to the tape, cl33's honest self-description was: a navigator that borrows a transformer's memory. Navigation — composing a trajectory of operators — is native, transparent, and load-bearing. Recall was borrowed, opaque, and bypassing. The rest of the paper is about making memory native (§5) and then asking what the model can read of its own trajectory (§7–§9).

3.1 Channel-attribution audit (2026-07-15): which operator path is causal?

Operator-dependence throughout this paper is measured by full operator ablation (zeroing B everywhere). That blade is sound for the bypass question — does the capability live outside the algebra? — but it cannot localize within the algebra, because the emitted operators reach the readout by two distinct paths: via the scan (they transform the transported state) and directly as readout features (B plus its grade tower). An audit on a branching-track probe (infer a flow k from a demo, carry it across a filler gap, apply it to a novel start) decomposed the two, with three results.

  1. The flow rides the feature path, not the state. Swapping B-conditioned operators into the scan alone leaves the branch on A (flip 0.000); swapping them into the readout features alone flips it completely (1.000). Replacing the scan outright with identity rotors costs ≈0 accuracy at every distance (0.96–1.00, including gap-1024 extrapolation at 8× the training gap); the state-only channel is weak and distance-limited (0.49→0.19 with distance).
  2. Distance-flatness comes from attention, not transport. The emitter is itself a transformer attending over the whole visible context; the flow-conditioned operator it emits at the query is how the rule crosses the gap. The η-attention can be removed (scan_only) without loss; the emitter's internal attention cannot be — and masking the demo at generation time collapses continuation to chance even when the rule is inferable from tokens in view (with the recorded caveat that PAD-masked demos are out-of-distribution, so part of the collapse may be brittleness rather than dependence). Continuation is demo-conditioned emission, re-read each step, not a rule stored in the transported state. Relatedly, on-manifold state perturbations show no restoring dynamic: the output projection is insensitive to ~1× a class-separation of displacement at emission time, and trajectories knocked off their flow stay off — exactly what reversibility (no contraction) predicts.
  3. The causal share of S at the output — measured (2026-09-08). A constant-bias steering probe on the frozen prose base closes what earlier drafts left open. At matched perturbation scale, a residual-stream bias produces ~25× the output effect (KL to the unperturbed distribution) of an operator-field bias on B_s; and at neither site does the effect grow with position (late/early ratio 1.0–1.1× at every magnitude). The second number is the decisive one: a persistent operator bias provably accumulates in the transported state through the rotor product, so if S's accumulated displacement reached the readout, the effect would compound along the sequence. It does not. Scan-compounding is falsified at the output: the state remains the substrate on which operators act, not a carrier whose history the readout consumes. The KGE navigation task of this section has still not been decomposed path-by-path ("operator-carried," op-dep 0.85, remains the honest label there), but the general hedge — "S's causal share is unresolved" — is now resolved for output steering: it is ≈0.

Design consequence (feeds §5). The tape's recall/receipt properties were measured directly — in configurations where the tape is the sole cross-position channel — and the same audit week recorded the matching caution: in any configuration where attention coexists with the tape, tape-causality must be re-verified, because if attention re-carries recall the tape becomes a passenger and editing it will not move the output. The subsequent v2.4 run measured exactly that failure mode (§5.4), which is why the successor designs make the memory the sole cross-segment channel by construction. The audit also sharpens the division of labor: the scan state alone neither recalls (§3) nor transports rules across distance; every distance-free capability we have measured is carried by an attention-addressed read of the operator record.

3.2 Independent replication and extension (Watson, 2026-09-10): the scan is

decorative at LM inference

An independent, pre-registered, eval-only audit by Nell Watson on the released checkpoints (both models, the shipped frozen slice and WikiText-103, fp32, hash-verified) first reproduced the bottleneck table to one decimal, then extended the §3.1 decomposition to the language model itself. Five conditions; PPL ratio vs native: operators zeroed 270×/115× (the bottleneck, replicated); S pinned to s₀ with attention on 1.19–1.22×; S time-shuffled 1.00×; η-attention zeroed 1.00×; S pinned and η-attention zeroed (readout sees only the local operator window through the grade tower) 1.00×. The time-shuffle control is the decisive one — we replicated it on our own harness (0.996) — and it shows the pinning cost is an artifact of the intervention rather than lost ordered history (Watson's shared-LayerNorm explanation is the candidate mechanism; it has not yet been separately isolated). The conclusion, stronger than §3.1 and measured on the LM proper: at inference on these checkpoints, the reversible scan and the operator-native attention over it are causally decorative; every bit of context reaches the output through the emitter's own attention, expressed in the emitted operators. Two scopings keep this honest in both directions: (a) in sole-channel/tape configurations (§5, §6) the scan's prefix products are the memory and ablating them destroys recall — the decorativeness is a property of the full LM configuration, not of the substrate; (b) the current evidence is consistent with training preferring the grade tower's direct route to B_t, B_{t−1}, B_{t−2} once it became available, but does not yet distinguish "the tower out-competed the scan" from "the scan mattered only as a training scaffold" from "the full-scale model never needed it" — and, as Watson notes, "the reversible path cannot do work" is not shown either. The v0 pre-tower model did use the transported state (the model_v2 docstring's "collapses to unigram" prediction was written for, and was true of, that configuration; it is stale for the shipped one). Half the mechanism is closed by algebra rather than ablation: §5.1's binding impossibility (a grade-1 carrier under the grade-preserving action cannot hold grade-2 content) already establishes the substrate-inadequacy branch for the memory function, with its constructive converse measured in the left-regular successor line. The two controls that would close the residual — whether the state could have carried generic context and the tower out-competed it (a no-tower model), or contributed nothing even during training (a no-scan-from-init model) — are queued as a v1.3 table at probe scale, scoped to that narrower question.

One reading of this result strengthens the transcript claims rather than qualifying them. If the readout consumes only a local operator window and the transported history carries nothing, then each emission must be a contextually complete summary: the emitter's attention integrates the visible history into every B_t before emission. That is the mechanism behind the density of the record — why a single emission decodes its token at 0.86 with +4.7 bits beyond prefix predictability (§7), and why regime state reads out of one token's operators under lexical control (§9.6). The log is information-rich because the decision path is local: everything the context contributes has to be compressed into the emission itself. And the completeness guarantee is unchanged in the only direction governance needs: nothing behaviorally relevant can bypass the record — what the model does not transcribe, it provably cannot act on.


4. Navigation parity (pre-registered)

Task. Permutation composition — apply a composition of permutations read from context; a TC⁰-flavored weak spot for fixed-depth transformers and the capability class closest to cl33's native operation (rotors compose; permutations are rotors at the integer points).

Protocol (pre-registered; deviations logged in the pre-reg memory). Learning-rate sweep for both models; winner LR re-run with 3 seeds; report mean ± sd of the mean-of-3 at the winning LR. (An earlier "cl33 matches-or-beats" framing from a max-of-5 readout was retracted when the mean-of-3 protocol shed the max-inflation — we report the corrected protocol only.)

Result.

Model Accuracy (mean of 3 seeds)
cl33-opLM 0.433 ± 0.009
tuned transformer 0.432 ± 0.030

Statistically indistinguishable. Two secondary observations: cl33's seed variance is ≈3× tighter, and one cl33 seed (43) reached its score with zero measured bypass — an existence proof that the parity-level solution can live entirely in the algebra. Mechanism note: grokking on this task coincides with rotor spectra collapsing toward permutation scale.

What we claim: parity + consistency + inspectability on the task family nearest the architecture's native operation. What we do not claim: superiority. A gauge-throttle optimizer intervention intended to force more-algebraic solutions was falsified (every lever hurt; weight decay worst) and is reported in App. D as a negative result.


5. The reversible tape: exact memory native to the scan

5.1 The failure that shaped the design

We first attempted associative memory as a wedge-product matrix store (C += k∧v read by q⌋C — mLSTM's matrix memory expressed in Cl(3,3)). The design matrix:

Key source Value source MQAR Bypass
operator (context-mixed) any ≈ 1/KV —
state (context-mixed) any ≈ 1/KV —
token (context-free) token projection 0.92 high
token (pure copy) token 0.97 91%
token (context-free) algebra-only untested untested

Context-mixed keys cannot re-form at query time (falsified twice); token values recall but bypass. The untested cell — token addressing with algebra-only content — is the tape.

5.2 Design principle

Token content may SELECT; only algebra may FLOW.

The scan already writes a lossless record: prefix products P_t = R_t·…·R_1 give the exact relative transport R_{t←i} = P_t · ηP_iᵀ η between any two positions, for free, by reversibility. The tape adds only a read head: a context-free grade-1 projection of the token embedding produces match scores α over strictly-past positions (with an induction shift: match position i, read i+1); the returned content is Σ α_i · f(tape_i) where f yields exclusively algebra objects. Token information can influence which positions are read (a low-bandwidth softmax channel) but no token content vector ever reaches the readout. Numerical discipline: raw prefix products diverge by T ≈ 1024 (error 1.3e6); anchored re-orthogonalization every K = 64 steps is mandatory and cheap.

5.3 Value modes and what each one is

Mode Content (dim) Reads the record as MQAR LM PPL (32M/8k) Ablated ratio
increment B_i (15) local adjacency 1.000 (all KV) 22.94 4.88×
displacement R_{t←i} (6) global path ~1.000 20.10 2.17×
state s_i (6) unstructured field no binding 20.84 9.99×
multi all three (21) — 1.000 20.47 —
no tape (v2.2 base, matched 32M) — — ≈ 1/KV 21.30 11.74×

Three findings:

  1. Recall without bypass, both blades. Increment mode solves MQAR perfectly at every key–value load, and ablating the operators destroys recall (ratio 4.88×, far above the ≈1 a bypass would show). The two-memory gap of §3 is closed natively in this configuration — scan-only at probe scale, with the tape the sole cross-position channel; §5.4 records what happened when that condition was relaxed.
  2. The mode taxonomy is a reader taxonomy. The same tape read three ways yields three different capabilities: increments give adjacency (recall), displacements give path integration (helps general prediction: best PPL), states give a cheap high-transparency feature channel. This taxonomy does independent work in §8.
  3. The attention-free challenger. Scan + tape with no attention at all reaches VAL 19.02 at matched 32M/8k — beating the full model with tape (20.10) and the no-tape baseline (21.30). Honesty: this is not sub-quadratic — tape addressing is an O(T²) softmax over 6-d token keys. The win is simplification + quality, not complexity class. (A pre-registered 5B challenger run of this config was planned; it was superseded by the v2.6 redesign before running — the 5B evidence on η-attention's inertness, §3.1, arrived first.)

5.4 Pre-registered v2.4 gates

The 5B retrain (multi-mode tape, anchored displacement) carried pre-registered gates — PPL vs the from-scratch matched control, MQAR at deploy, bypass ratios per mode, J-lens rerun (§9) — recorded in CL33_OPLM_V24_TAPE_LM.md before the runs.

A co-adaptation regression, diagnosed and repaired (reported as a result). The first co-trained tape run (warm-started from a no-tape 5B checkpoint, with an auxiliary MQAR loss to teach addressing) reached the tape gate — recall 0.77 within the first fifth of the token budget — but regressed language modeling: a matched-stream fp32 comparison showed the shared base pulled off-optimum (base-alone val 31→88 across training). The mechanism is not a gradient conflict — the LM and auxiliary gradients are near-orthogonal on the shared emitter (cosine ≈ 0) — but a gradient magnitude imbalance: the auxiliary signal is ~4.6× larger there and additionally reshapes the shared readout toward memory tokens. The fix is a principled coupling, not a loss weight: cap the auxiliary gradient magnitude to a fraction (κ = 0.5) of the LM gradient's on the shared emitter, and route it off the shared readout so memory output is carried by a dedicated head. Under the fix the result exceeds the naive run on both axes: the deployed model returns to the LM baseline (1024-context fp32 val 32.5→31.4 over steps 2k–6k, versus scale-B's 30.97) while recall climbs past the naive run's plateau (0.77→0.80→0.89) and the hollowed base heals monotonically (base-alone 87→62→52). The naive coupling had bought recall 0.77 at a standing +11% LM penalty; the principled coupling buys recall 0.89 at ≈baseline LM (+1.2% vs scale-B, with the caveat that the deployed 1024-context val and the 512-context baseline are not exactly the same ruler). This is the kind of pathology the operator factoring makes visible and addressable: because the memory path and the base path are typed and separable, the imbalance is measurable per-parameter-group and correctable as a gradient operator rather than a blind hyperparameter search.

What the calibrated instruments then showed (the honest coda). Three follow-up measurements, run within days of the repair, scope the 0.89 sharply and are the reason the successor designs exist. First, the 0.89 is a short-range number: the benchmark's query–pair gap is ~40 tokens, and a calibrated distance sweep found recall collapsing 0.94 → 0.05 as the gap grows toward 448 (worse with natural-text filler) — the work log records this as falsifying the tape's distance-independent framing, and we adopt that verdict here. Second, the learned address projection — context-free by design — was co-opted by language-modeling context on natural text (exact-match addressing 0.006 against 0.26 on synthetic), the design property failing to survive co-training. Third, decode-level attribution in the full configuration showed the tape's own vote on recalled tokens at 0.000, with --no_tape recall ≥ tape-on: in the presence of attention, attention re-carried recall and the tape rode as a passenger — precisely the §3.1 caution realized. These three results — distance-limited addressing, address co-option, and passenger collapse — are not appendix caveats; they are the measured boundary of the v2.4 design and the direct motivation for the dual-address revision and the segment-recurrent redesign in which the memory is the sole cross-segment channel by construction (§12).


6. Capability doors: editable, provenance-bearing, load-invariant memory

(Full pre-registration + adjudication: CAPABILITY_DOORS.md. All results here are on the synthetic MQAR task at probe scale, in scan-only configurations where the tape is the sole cross-position channel. The §5.4 coda applies: in configurations where attention coexists with the tape, edit/receipt causality must be re-verified — the v2.4 full config measured the tape as a recall passenger, and editing a passenger does not move the output.)

Once memory is a typed, exact, position-addressed record (§5), three capabilities follow that no opaque KV cache offers. Each carries the bypass blade.

Door 2 — editable memory (transaction semantics). The counterfactual-override path is used as a write API on a trained increment-tape model: swap two records (queries return the swapped value, efficacy 1.00), delete a record (target recall → 0.008 ≈ chance, other pairs unchanged — locality 1.00), transplant a record from a different sequence (foreign value retrieved, 1.00 — records are portable at this scale), and undo (bit-exact restoration, max |Δlogits| = 0). Identity override is also exact (Δ = 0). This is surgical context editing, auditable RAG, and — since delete is targeted forgetting with a receipt and an undo — machine unlearning with provenance.

Door 3 — provenance receipts (faithful, causal attribution). The tape's address weights are a faithful attribution: argmax-α identifies the ground-truth write position with attribution@1 = 1.000, and the receipt is causal — deleting the argmax- addressed record drops that query's recall to 0.000 while deleting a non-source record leaves it at 1.000. A matched transformer's best attention head reaches 0.977 only under oracle head-selection, and that oracle head is not stable across load (it drifts L1H2→L1H3→L0H1 across vocab/length; where the transformer's attribution collapses under load, the diagnosis is that no attribution circuit formed — not that attention "smears"). The tape's address is a single canonical, causally editable interface; the transformer's is an emergent, condition-dependent one.

Door 1 — load-invariance at T = 1024, and a pre-registered fairness sweep. On MQAR the scan+tape holds ceiling as load grows (KV 64/128/256 at T=1024 → 0.997/0.999/0.999, zero bypass), while the same fixed-96-d-state model without the tape sits at chance — the tape rescues exact recall the recurrence structurally cannot do. Validation stops at T=1024: a T=2048/KV=128 run did not grok within the grid's step budget (an honest gap; T≥2048 is unverified). The matched d128 transformer, by contrast, failed to learn T=1024/KV=64 across a pre-registered six-config sweep (three learning rates at depth 2, plus depth-3, depth-4, and 2×-width runs at the winning rate; all 12k steps, all ≈ chance), where the tape reaches 0.99 in under 5k steps. We scope this honestly: MQAR is provably solvable by attention at sufficient scale (Zoology built it to showcase attention), so the claim is not that transformers cannot — it is capacity + learning efficiency: the tape forms exact, zero-bypass, editable, receipt-bearing retrieval at lower capacity and fewer steps than the transformer needs to form any retrieval at all.

12. The memory organ: from deadlock to a certified peripheral (2026-08 → )

The results above treat memory as a channel inside the operator LM. A successor program — reported here as measured intermediates with pre-registrations, not finished claims — grafts the Genesis two-channel memory (validated separately on frozen GPT-2/Gemma/R1 backbones) onto the frozen prose base of Table 1b as an explicit register-bank organ: 0.1–2.5M trainable parameters (per arm) against 236M frozen, wedge-plane records with trained-encoder addressing, zero-init injection, and a removability gate that holds at literally 0.00e+00 ΔCE in every run (organ-off ≡ base).

The program's value so far is a diagnostic ladder — each rung a small set of pre-stated changes with separated instruments, each verdict measured:

  1. Interface deadlock. Next-token CE alone never differentiates addressing — the register softmax stays exactly uniform (9,000 steps with raw-embedding features; 3,000 with emitter-tap features before that arm was cut). Cause: a chicken-and-egg between addressing and decoding, plus retrieval-demand density orders below every precedent that formed recall.
  2. Labels break it. Planted-episode data makes the correct register a free label; a supervised address loss moves address accuracy 0.000 → 0.5 in 100 steps — what CE never did in 9,000.
  3. Context is mandatory for addressing. Raw-embedding queries plateau structurally (the query token is a generic "is"); features tapped from three emitter blocks address unique-vocabulary episodes at 1.000.
  4. Retrieval forms. With episode demand, seam-weighted loss, and per-8-token registers: sustained recall 0.458 with recall-off = 0.000 (zero bypass) — and a clean granularity dose-response (P=8 beats P=16 on both axes). Measured at full episode density: recall never crossed the pre-set 0.5 threshold that would have opened the density anneal, so retrieval-under-realistic-sparsity — the program's pre-registered decision point — remains unmeasured, and clean-text ΔCE stayed negative at this density.
  5. The decode ceiling is the read mixture, not the decoder. At matched read width (H=8), two structurally different decoders — a free logit matrix and a KGE-style codebook whose write and read sides are the same tensor — land on the same conditional recall (0.565 / 0.567 given a correct register), localizing the ceiling to the unsupervised within-register mixture both consume. A 2×-width read measures conditional recall 0.716 (n=102) — width partially dilutes mixture interference — but pays for it in slower addressing, leaving total recall unchanged; the mixture, not the decoder, remains the binding constraint. The pre-registered test that removes the mixture entirely (per-token registers: the register is the binding; predicted conditional decode ≥ 0.85) has now run — rung 6.
  6. Mixture removed: the memory path is exact, and the losses decompose into three named numbers. With per-token registers and the tied codebook, the memory path alone decodes at 1.000 (53/53): given the right register and a pure read, the record is not lossy at all. The remaining gap to headline recall then separates cleanly — soft-α selection conditional 0.463 vs hard top-1 read 0.756 (the α-softness cost, ~0.29), hard read 0.756 vs pure decode 1.000 (the integration cost through the frozen base, ~0.24), and addressing itself at 0.57–0.69. Nothing is mysterious in the residual: an exact record, read through a soft mixer, injected into a base that never trained to consume it. Each named loss is a plumbing target, not a representation problem — the record layer is exact on that evidence; which value substrate feeds it best is adjudicated by the control arm below.

The endpoint has a first measurement. The battery the governance section implies — install → adoption → certificate-bearing forget → revert — was run as a pilot on the rung-6 checkpoint: 24 facts never present in any context, written through the organ's own write path onto virtual segments, queried cold. Install top-1 0.750 (median rank 1, mean Δlogprob +8.80); in generation, 4 of 4 sampled queries adopt the installed fact; algebraic delete → 0.000; organ-off → 0.000 (zero bypass); address receipts 24/24 (the α argmax lands on the installed register in every case). Install efficacy landing on the same number as the hard-read integration ceiling (0.750 ≈ 0.756) says the pilot is at the measured plumbing limit — installs are as good as integration currently allows, no worse. Pilot-scale, single-seed, one config; reported as such.

The installed belief passes through the transcript. A four-cell factorial on the same checkpoint separates the two paths by which a read can reach the output: internal injection only 0.407, direct logit path only 0.017, both 0.712, neither 0.000 — interaction +0.288, internal-carrier fraction 0.571. The dominant carrier is the pre-bottleneck injection: the retrieved memory alters which operators get emitted, and the behavior change rides those operators to the readout. The receipt is direct: on install queries the emitted B_s displaces by 0.777 on hits vs 0.591 on misses. Because the readout consumes only operators (§2), this closes the chain memory → exposed cognition → behavior with a measurement at each arrow — the property an opaque-base memory graft cannot exhibit even in principle.

Two methods results the sweep paid for. First: injection-layer arms L6 and L11 both matched the incumbent layer (L9) on every training-time validation metric, yet measured sharply different install efficacy (L11: 0.083 with a validation profile identical to L9's at the same step). Selection by validation recall would have chosen wrongly; only the install battery measures leverage, and the pre-registered rule — primary metric: install efficacy; validation curves never select — is what caught it. Second, and larger: the incumbent's efficacy number (0.750) came from a checkpoint trained 3× longer than the sweep arms (no intermediate checkpoint had been kept) — precisely the step-mismatched comparison the pre-registration forbade. Re-running L9 at the arms' budget eliminated its apparent advantage entirely: at matched 5k steps, L9 measures 0.208 — identical to L6 (secondaries lean L9: median rank 7 vs 20; secondaries do not select). The "3.6× layer effect" the mismatched comparison suggested was in fact a training-duration effect: install leverage at the same layer grows 0.208 → 0.750 from 5k → 15k steps of interface training. The corrected reading: among tolerant layers, where you inject matters far less than how long the interface trains — and a plausible, val-supported, wrong conclusion was two pre-registered rules away from the paper.

Speed vs ceiling, adjudicated under pre-registration. A control arm with learned operator-derived values (no codebook) was predicted to hit a representational ceiling well below the codebook's exact decode. The prediction landed in its pre-registered middle branch: the control escaped the predicted ceiling family (hard-read conditional 0.871; memory-only 0.914) yet stayed short of exact (codebook 1.000), and it out-integrated the codebook at every in-pipeline mode while acquiring faster — opening the program's density anneal first (dense-plant training relaxed to realistic sparsity with recall retained at 0.44 and clean-text ΔCE healed to ≈0: the dense curriculum is a learning scaffold, not an operating requirement). The codebook's exactness holds at the record layer; its end-to-end advantage is not established, and the two arms differ by a curriculum confound (only the control reached the sparse regime) that the next arm is designed to remove. We report the branch structure rather than a winner because the pre-registration was written before the numbers existed.

The remaining open items are plumbing — addressing sharpness, per-position integration gating, read abstention — and the full battery at scale. Every rung above, including the four nulls and the wrong predictions, is logged with dates and checkpoints in the project work log.


Part III — Governance

What the receipts buy, which reversion claim is being made, and the doors through which weaker guarantees re-enter.

13. Governance: the audit object

Where the tax inverts. Table 1 compares cl33 to an unconstrained transformer on perplexity. In a regulated or high-stakes deployment the relevant comparison is different — cl33 versus transformer-plus-its-audit-stack on guarantee-per-dollar — and it flips. Post-hoc interpretability (SAEs, probes, attention attribution) is itself a compute-and-process tax that buys only an approximation: §6 Door 3 makes this concrete, with the transformer's best attention head attributing the true source at 0.017 (rank 40) and drifting across load, versus the tape's exact, causal 1.000 receipt. cl33's audit trail is the computation's mandatory carrier — the operator record through which all output flows (270× on ablation, independently replicated) — and the reverse trace (R⁻¹ = ηRᵀη, exact to 2e-14) makes that record losslessly replayable, not a learned proxy of it. The refinement §3.2 requires, stated plainly: the record is causal; the scan's invertibility is a property of the log, not of the decision path, whose transported state is measured decorative at LM inference. For deployments that must prove what a model did — machine unlearning with a receipt (Door 2), decision provenance, tamper-evident memory — the ~1.29× perplexity tax becomes the price of a guarantee that post-hoc tooling cannot provide at any cost. We scope this precisely: what is guaranteed is that the computation is invertible, the memory editable with receipts, the emission decodable — not correctness, calibration, or non-hallucination. Transparency is not correctness; a cl33 model can be confidently wrong, but you can prove exactly how it got there and edit it.

What "revert" means — a three-tier hierarchy. "Install a belief, remove it, and the model reverts" is three claims of different strengths, and the architecture supports stating which one is being made:

  1. Computational reversion (strongest, measured). With the base frozen and the organ zero-initialized-and-gated, deleting a register restores the bit-exact forward computation on the same prompt: organ-off ≡ base to 0.00e+00 ΔCE in every run, and the §12 pilot measures delete → 0.000 with the installed fact unreachable. This tier is a property of the frozen-substrate design, not of training success.
  2. Behavioral reversion (fresh prompts). After deletion, novel prompts probing the belief return to baseline behavior. Measured in the pilot at the same 0.000; at scale this tier requires a battery, not an identity argument, and we scope it as such.
  3. Conversational reversion (weakest — inexact by construction, auditable). If the installed belief influenced generated text that re-entered context, deletion does not un-say it: the influence has sedimented into the conversation history, which is ordinary context the model rightly conditions on. This is the operator-model form of the sediment lesson (not-recoverable ≠ not-behaviorally-causal). What the architecture offers here is not exact reversion but an audit trail: the receipts date-stamp exactly which generations occurred under the installed belief, so the contaminated span is identifiable even though it is not erasable.

Where the sediment caveat re-enters (three doors, named so they can be watched). The tier-1 guarantee is conditional, and each condition is a door through which weaker-than-computational reversion re-enters: (a) co-training — any arm that unfreezes the base (§12's v3.1 option) trades bit-exact organ-off for "characterized degradation," a trade that must be made knowingly and labeled; (b) persistent banks — when memories are written by processes that themselves read memory, deleting a record does not delete its causal descendants; a write-provenance cascade (each record carrying the receipts of the reads that produced it) is the pre-registered requirement before any long-lived bank is deployed; (c) context sedimentation — tier 3 above, present in any deployment where outputs re-enter inputs. None of these voids the architecture's guarantees; each converts one guarantee from an identity into a measurement, and the receipts are what make that measurement possible.


14. Related work

Linear-recurrent and matrix memories. mLSTM/xLSTM (Beck et al., 2024), DeltaNet-style fast weights (Schlag et al., 2021; Yang et al., 2024), and state-space models of the Mamba lineage (Gu & Dao, 2023; Dao & Gu, 2024) form the honest competitive set for Appendix F (the linear-transparent store): linearity without receipts. The lineage runs from the original fast-weight programmer (Schmidhuber, 1992) through fast-weight associative memory (Ba et al., 2016) and the linear-attention recurrence (Katharopoulos et al., 2020) to their identification (Schlag et al., 2021). Our wedge store is the Cl(3,3) expression of the mLSTM outer-product memory, and its measured failure modes (91% bypass with token keys; slow grokking with algebra values) delimit ours rather than flatter them. The recall–throughput frontier this family trades on is mapped by Arora et al. (2024); our tape sits off that frontier deliberately — an honest O(T²) for exact, receipted recall. Memory-based test-time adaptation (Sun et al., 2024) is the nearest frame for the tape-as-TTT reading (§6): where TTT layers update a hidden model by gradient, the tape substitutes an exact, undoable memory write. What none of this family provides — and what the tape and register formulations are for — is causal attribution receipts and exact, certificate-bearing deletion.

Induction heads and in-context retrieval. The two-memory identity of §3 (recall-as-induction = routing-around-the-operators) is an architectural corollary of the induction-head literature (Olsson et al., 2022), and our recall instrument is the MQAR task of Arora et al. (2023): we do not dispute that transformers ride token-similarity shortcuts; we make the shortcut visible and ablatable and show what remains without it.

Model editing and unlearning. ROME/MEMIT-style weight editing (Meng et al., 2022; Meng et al., 2023) modifies parameters with no certificate and contested locality (Hase et al., 2023); machine-unlearning methods (Bourtoule et al., 2021) approximate removal with no proof object. Our editable-memory results (§6, Door 2) claim something narrower and stronger where it applies: exact algebraic erasure from the explicit memory state, with a receipt, accompanied by behavioral reversion — scoped explicitly NOT as unlearning of the backbone.

Steering and representation control. Steering by added directions is established for standard transformers — representation engineering (Zou et al., 2023), activation addition (Turner et al., 2023), and persona vectors (Chen et al., 2025). §9.6 and App. C place cl33 in that family with one structural difference: the steered object's downstream effect is readable from a mandatory typed interface (the operator stream), so dose–response, saturation, and regime adoption are measured on the causal path rather than probed beside it.

Reversible architectures. RevNets (Gomez et al., 2017) and reversible transformers (Kitaev et al., 2020) invert computation for activation-memory efficiency — the inversion exists to avoid storing. Here reversibility is the explanation interface: R⁻¹ = ηRᵀη is run backward to produce the audit trail, and the 2e-14 exactness is a property the paper's governance claims rest on.

Geometric and Clifford networks. Clifford/geometric-algebra layers (Brandstetter et al., 2023; Ruhe et al., 2023; Brehmer et al., 2023) are typically used as equivariance-preserving feature maps inside otherwise-standard readouts. cl33's distinction is exclusivity: the algebra is not a feature map but the only path to output, which is what converts interpretability questions into ablation experiments.

Lens methods and verbalizability. Logit-lens (nostalgebraist, 2020) and tuned-lens (Belrose et al., 2023) read representations through trained probes, inheriting the faithfulness question shared by probing classifiers generally (Belinkov, 2022) and attention attribution in particular (Jain & Wallace, 2019); sparse-autoencoder dictionaries (Bricken et al., 2023; Cunningham et al., 2024) buy interpretable features at the same post-hoc remove. The Jacobian-lens instrument of §9 estimates verbalizability without a trained readout, and the three-model comparison places the operator stream above both GPT-2 controls' best layers.

Global Workspace Theory (Baars, 1988; Dehaene et al., 1998) supplies naming only (reportable bottleneck vs non-reportable processing); no cognitive claim is made or needed. TC⁰ expressivity limits (Merrill et al., 2022; Merrill & Sabharwal, 2023) supply the provenance of the permutation-composition task used for the §4 parity study, with the task's transformer-hardness established by Liu et al. (2023) and its state-tracking framing by Merrill et al. (2024). Baselines throughout are GPT-2 (Radford et al., 2019) and the token-matched Pythia suite (Biderman et al., 2023).


References


Appendices (planned)

Appendix R (drafted): reproducibility manifest

Every table maps to an in-tree script, a checkpoint, and an exact command, and every result cites its dated work-log entry (all paths relative to the cl33_oplm project root; run logs preserved verbatim).

Released artifacts. The two central frozen checkpoints are public at huggingface.co/mirrorethic/cl33-oplm — the prose base of Table 1b (236.5M, step 189307) and the serving/demo model of cl33.t3atlas.dev (236.5M, step 13996) — with the load-only model definition, the reverse-readout probe (its held-out evaluation card stored inside the artifact), SHA-256 hashes, and two verified one-command reproductions: repro_bottleneck.py (zeroing the emitted operators multiplies perplexity exactly 270× on a shipped frozen slice of the validation mix — a different draw of the same distribution measured the paper's 314× — and ≈106×/≈112× on public WikiText-103, the off-domain form of §1's bottleneck claim) and repro_reverse_readout.py (text decoded from the operator stream alone, probe top-1 0.860 held-out). The bundle's REPRODUCE.md states its scope plainly: it contains what is necessary to independently test the published claims on frozen artifacts, not the training stack or the §12 program (which the paper itself labels ongoing). Reproducibility surface ≠ complete source disclosure; what a reader needs is the artifact and the measurement, and both are now unconditional.

result script(s) checkpoint / data record
Table 1 (matched tax) cl33_lm_eval.py, bench_table.py (lm-eval 0.4.11, fp32, right-pad, full-set, single session) scale-C / scale-B / A1 ckpts runs_eval/modelclass/table.md; WORK_LOG 2026-07-13
Table 1b (prose parity) cl33_lm_eval.py --ckpt <prose ckpt> scaleB_prose_1024/ckpt.pt (6.2B tok) runs_eval/scaleB_prose6B_lmeval.json; WORK_LOG 2026-08-03
§4 nav parity nav_cl33.py, nav_transformer.py, nav_data.py 5-seed pre-registered pair prereg record + WORK_LOG (pre-fork)
§5/§6/App.A MQAR + doors mqar_cl33.py, door1_transformer.py, door2_edit_probe.py, tape scripts probe-scale ckpts results_mesh/asus_cl33_mqar/ (mirrored)
§7 inversion inversion battery scripts per INVERSION_BATTERY.md scale-C results_mesh/ + WORK_LOG 2026-07-10/11
§9 J-lens code_jacobian_lens/ harness scale-B/C + GPT-2 controls results_mesh/asus_jlens/ (13 eval JSONs + lenses)
§12 organ ladder train_v30.py, v30_organ.py, episodes_v30.py, probe_v30_address.py frozen prose base + organ ckpts per run v30_* run dirs (train.log val lines); WORK_LOG 2026-08-03 → 09-08
§12 install pilot + factorial install_pilot.py, cond probes v30_c3b_p1/ckpt.pt (+ L6/L11 arms) WORK_LOG 2026-09-08 10:20→16:35
§9.6 regime sensor + §3.1 lever probe inline probes (WORK_LOG-referenced) scaleB_prose_1024/ckpt.pt (frozen, eval-only) WORK_LOG 2026-09-08 14:35→15:10; CL33_OPLM_INSTRUMENTABLE_LM.md

Baselines: Pythia-160m rows use public HuggingFace revisions (step-matched); GPT-2 rows use the public 124M release. The eval adapter's right-padding requirement (§10 methodological note) applies to every cl33 row.


Sources of record: WORK_LOG.md (this dir), REVERSIBLE_TAPE_DESIGN.md, SELF_STRUCTURE_CLAIMS.md, CL33_OPLM_V24_TAPE_LM.md, CL33_OPLM_ARCHITECTURE.md, CL33_OPLM_V23_RM.md, nav parity pre-reg memory, J-lens artifacts at /mnt/consciousness-storage/jlens/.

Appendix F (moved from §11): Linear-transparent memory — the trilemma and its break

(Design + pre-registration: LINEAR_TRANSPARENT_MEMORY.md.)

The tape's exactness and receipts come at an O(T²) softmax address (it materializes a distribution over all past records). The field's linear-time memories (mLSTM, DeltaNet, Mamba) compress records into a fixed-size state — O(T) — but have neither receipts nor surgical editability. Can one architecture have all three of {linear, transparent, recall}? We hold both endpoints: the wedge (mLSTM expressed in Cl(3,3): O(T), but its token-content value channel bypasses the operators 91%) and the tape (O(T²), zero-bypass, attribution@1 = 1.0).

A three-corner trilemma, then broken. At a probe configuration (KV=8), three variants: token-value wedge — recall 0.87 but 95% bypass, O(T); tape — recall 1.0, zero bypass, O(T²); and the untested cell, an algebra-value wedge (token key + algebra-only value) — O(T) and zero-bypass, but recall only 0.26 at a short training budget. The apparent trilemma (any two of the three) dissolved on inspection: the algebra-value wedge was still climbing. Given the budget to converge, it grokked to 0.99 recall — linear, zero-bypass, and accurate, all three — via a phase-transition jump (0.39→0.94 between 3k and 5k steps, cross-entropy collapsing), at ~10× the tape's learning time.

Why it groks (a mechanistic account, contributed by the collaboration). Quadratic memory externalizes address formation into an explicit distribution over separately-kept records — nothing to learn. Linear memory must learn an internal geometry that simultaneously prevents write-interference and supports decoding. The token-value wedge is fast only because token embeddings hand it a ready-made near-orthogonal basis (that is the bypass); the algebra-value wedge must build that geometry from the operators, which is a genuine — and slow — representation-learning problem. The grok jump is the signature of finding that geometry.

The honest boundary. This is an existence proof at toy scale: at a realistic 8k vocab the algebra-value wedge did not grok within the same budget — the fixed-size compression's decoding capacity is the binding constraint, and it scales with vocabulary. So linear + transparent + recall is achievable but not yet tractable at scale — a mechanistically understood open problem, not a wall. The design compass it hands us (delta-rule writes to reduce interference by construction, orthogonal initialization, decorrelation penalties) is the next work.