# One Object: Memory, Navigation, and Reportability in an Operator-Only Language Model

**Preprint v1.3 — 2026-09-10.** DOI: [10.5281/zenodo.22684392](https://doi.org/10.5281/zenodo.22684392).
Versioned; dated revisions supersede (the DOI resolves to the latest version). §12
reports an ongoing experimental program and carries its own dates.

*Cite as: Sutherland, G. (2026). One Object: Memory, Navigation, and Reportability in
an Operator-Only Language Model. Zenodo. doi:10.5281/zenodo.22684392*

Garret Sutherland — MirrorEthic LLC

---

## Abstract

We study cl33-opLM, a language model whose only path from computation to output is a
stream of explicitly emitted Cl(3,3) bivector operators: a transformer emitter produces
per-block operators, a reversible SO(3,3) rotor scan transports state, attention scores
are η-metric products of rotors, and the readout sees exclusively operator-derived
features. Because generation, memory, steering, and inspection all act on the same
algebraic object, questions that are normally interpretive become measurable. We report
four results. **(1) Two memories:** associative recall (MQAR) and relational navigation
are mechanistically distinct in this architecture — recall rides a token-similarity
induction shortcut that bypasses the operators entirely (operator-dependence ≈ 0), while
navigation is operator-native (operator-dependence ≈ 0.85, scale-invariant); we show the
recall shortcut and the bypass are the same operation at matched training budget. **(2) Parity where it counts:** on
permutation-composition navigation — a known transformer weak spot — cl33 matches a
learning-rate-tuned transformer baseline exactly (0.433 ± 0.009 vs 0.432 ± 0.030) with 3×
lower seed variance and a fully inspectable mechanism. **(3) The operator stream is a
transcript:** a probe trained only on emitted operators decodes the current token at
0.86 top-1 over the 50k vocabulary — ~9.6 of ~11.2 bits of token identity per emission,
with 80.7% of running text reconstructed verbatim from the operator record alone, and
identity confirmed under substitution. The model's only channel to output doubles as a
readable record of what passed through it. **(4) An architectural workspace:** applying the Jacobian-lens
verbalizability estimator across the model's depth shows the emitted operators are the
most output-transparent representations in a three-model comparison (final-agreement
0.613, above both GPT-2 controls' best layers; replicated on a second model), while the
transported state is completely non-verbalizable (0.000) — a split that holds across two
held-out domains. The workspace/ocean dichotomy that emerges statistically in
standard transformers is, here, a line in the wiring diagram. We frame what the
architecture buys (transparency, reversibility, exact memory, causal steering surface),
what it costs (a real, measured tax: **+29% bits-per-byte and ~5 points average MCQ**
against a param-and-context-matched from-scratch control at 5B tokens — though continued
prose training moves the same model to **exact bits-per-byte parity with token-matched
Pythia-160m** (1.155 at 6.2B tokens; caveats on ruler alignment in §10), showing the tax
is data-bound rather than an
architecture floor; we claim no perplexity or reasoning *superiority*), and what it does
not claim. 

On the axis the architecture is built *for*, we report capability results with matched
baselines that collapse or cannot produce the artifact: **exact retrieval** at load
(1.000 vs a matched transformer's ≈chance at T=1024/KV=64), **editable memory** with
transaction semantics and bit-exact undo, and **causally-faithful provenance receipts**
(attribution@1 1.000 vs a transformer best-head 0.017). These were measured on a
reversible-tape intermediate, and its honest failure mode — attention re-carries recall
unless the memory is the *sole* cross-segment channel (§5.4) — is what taught the
sole-channel-by-construction design of the current memory organ (§12). That organ's
first belief-install battery (install 0.750, algebraic delete 0.000, receipts 24/24) is
reported as a pilot.
We close with the governance reading: because the operator record is the *mandatory
carrier* of the computation (ablating it multiplies perplexity by orders of magnitude)
and the reversible scan makes that record *losslessly replayable* (exact to 2e-14), the
perplexity tax inverts into the price of an audit guarantee that post-hoc
interpretability cannot provide at any cost — with the guarantee scoped to the record's
mandatoriness, replayability, editability, and decodability, not correctness, and with
the scan's own causal share at LM inference measured (independently) as decorative
(§3.2).

---

## 1. Introduction

Interpretability of standard transformers is archaeology: the artifact is finished, and
we dig. Probes, sparse autoencoders, and lens methods recover *correlates* of
computation, with faithfulness always in question because the object recovered is not
the object the model computes with.

cl33-opLM inverts the arrangement. The architecture is chosen so that the model's only
route to the output runs through a typed, invertible, algebraically-structured object —
a per-token, per-block set of Cl(3,3) bivector operators. There is nothing *behind* the
operators to be unfaithful to: the readout is constructed to see operator-derived
features only, and we verify the construction empirically (ablating operators multiplies
perplexity 8–42× at 32M, 206× at 163M — a ratio that *climbs* with capability). Memory
is an operation on
operators (§5). Navigation is composition of operators (§4). Reportability is a property
of operators (§9). Steering is addition of operators (App. C). One object.

The purpose of this paper is not to claim superiority over transformers at language
modeling. The transparent model pays a real, measured premium against a param+context-
matched from-scratch control at 5B tokens (+29% bits-per-byte, §10 Table 1) — and the
same machinery-bearing model, continued on prose to 6.2B tokens, reaches exact
bits-per-byte parity with token-matched Pythia-160m (§10 Table 1b). Both numbers are
reported: the matched cost of the machinery at a fixed budget is real; the ceiling it
was once mistaken for is not. The claim is narrower and, we think, more useful: **when the
generative object is algebraic and exclusive, capabilities and their mechanisms become
the same thing**, and questions like "does the model use its history as a trajectory or
a bag?" or "which internal content is reportable?" stop being interpretive and acquire
experiments with pass/fail readouts. Every positive claim in this paper carries an
operator-ablation control (the *bypass blade*): if a behavior survives with the algebra
removed, it is a shortcut and it does not count — including our own memory mechanism,
which we first built wrong and report as such (§5.1).

The intuitive form of our results: **a transformer does not carry a transcript of
words — it carries a trajectory of prediction-relevant events from which the words
are mostly recoverable.** This is true of standard transformers too (our emitter's
hidden state decodes tokens at 0.91), but there the events and their accumulation
are entangled in one vector. cl33 factors the process into inspectable parts: the
**operators are the event record** (token identity decodes at 0.86, degrades to
near-synonyms because the record stores what a token does to prediction, not its
orthography — §7); the **transported state is the accumulated path** (literally
s_t = R_t···R_1·s_0: zero token content, non-verbalizable — §3, §9; its *causal*
share is smaller than this factoring once suggested — the §3.1 audit finds flow
rides the operator features, with the state a passenger on the probes run so far);
and **generation reads both through a mandatory
interface** (bypass 8–42× at 32M, 206× at scale; emission loss 0.48 bits — §2, §7).
Every major result
in this paper — the two-memory dichotomy, the tape's reader taxonomy, the
order-readout null, the workspace/ocean wall — is a facet of this factorization.

**A living artifact.** The claims in this paper are demonstrated interactively at
**cl33.t3atlas.dev**: a transparent-chat interface over the trained model in which every
response is accompanied by its operator transcript, the reverse readout (the §7 inversion
run live — the conversation decoded back out of the operators), rewind, and steering.
The demo is not an illustration of the paper; it is the same model, same algebra, same
receipts, publicly poke-able — and the bottleneck claim is measured on the *serving
checkpoint itself*: ablating the operators multiplies the demo model's perplexity
**314×** (the 8–206× ratios elsewhere in this paper are the training-lineage
measurements at their own scales). Where a reviewer would ask "does this survive
contact with arbitrary input," the demo is our answer.

**History, reported as science.** This architecture's development includes a
falsification worth stating plainly rather than politely omitting, because it explains
the design and produced one of its cleanest results. The original cl33-opLM carried its
recurrent state as a grade-1 6-vector under the grade-preserving rotor action — and its
associative memory failed structurally (MQAR ≈ 1/K at every load, §3): a key→value
binding is a grade-2 bivector, and a grade-1 carrier has nowhere to hold one. A
four-arm causal ladder built in the successor Genesis program isolated the error
exactly — the grade-1 arm reproduces the failure, the grade-2 wedge store binds exactly,
a param-matched generic control shows the fix is structure rather than dimension
(`mqar_left_regular.py`, reproducible). For a period the language-model line was
formally closed as superseded on the strength of that diagnosis. It reopened when two
things became clear: continued training closed the model-class fluency gap (Table 1b),
and the representation lesson could be incorporated without abandoning the architecture —
superposition belongs in an explicit grade-2 store, not the grade-1 carrier, which is
exactly how the memory results in this paper are constructed (§5, App. F) and how the
ongoing organ program is built (§12). The vector representation failed for one memory
formulation; the failure was localized by experiment; the design absorbed the lesson.
We consider this sequence — claim, falsification, diagnosis, reincorporation — the
paper's strongest argument that the methodology works.

**How to read this paper.** The systems description that organizes everything that
follows: a **neural engine** (the transformer emitter — opaque, nonlinear, where the
learned computation lives), a **geometric bus** (the operator stream — mandatory,
typed, inspectable, the only path to output), and **causal peripherals** (memory,
sensors, steering — machinery attached to the bus rather than baked into the weights,
each attaching under a zero-init discipline that makes it bit-exactly removable). The
paper is arranged in three parts. **Part I — The Object** (§2, §7, §8, §9, §10): the
bus exists, transcribes both content and state, is the verbalizability peak while the
transported state is the floor, and costs a measured, data-bound tax. **Part II —
Memory on the object** (§3, §4, §5, §6, §12): what memory the bus does *not* give for
free, the falsifications that located the sole-channel requirement, and the register
organ that now installs, adopts, and deletes beliefs with receipts. **Part III —
Governance** (§13): what the receipts buy, precisely which reversion claim is being
made, and the three doors through which weaker guarantees re-enter. Section numbers
are stable identifiers carried from earlier drafts; reading order is the Part
structure.

### Contributions

1. The **two-memory dichotomy** with a mechanistic identity proof: recall-as-induction
   and routing-around-the-operators are the same operation (§3).
2. **Pre-registered navigation parity** with a tuned transformer on permutation
   composition, with tighter variance and an inspectable mechanism (§4).
3. The **reversible tape** as the load-bearing intermediate: exact forward-readable
   memory native to the rotor scan, with the addressing/content separation principle
   ("token content may SELECT, only algebra may FLOW"). Its capability results
   (MQAR 1.000, zero bypass, an attention-free challenger at matched probe scale)
   hold in sole-channel configurations; its measured failure mode — attention
   re-carries recall whenever it coexists (§5.4) — is what taught the
   sole-channel-by-construction design of the memory organ (§5, §12).
4. The **inversion battery**: the input words decode back out of the operator
   stream — 0.86 top-1 over 50k vocab from one emission, +4.7 bits beyond prefix
   predictability, identity confirmed under fixed-context token substitution,
   80.7% verbatim document reconstruction with semantic (near-synonym) degradation
   (§7).
5. A **claims ladder for self-structure** with adjudicated results: transparent recall
   (closed), structural path-sensitivity (passed), semantic order-readout (null, twice,
   with a mechanistic explanation), self-modeling (live, carrier localized) (§8).
6. **Jacobian-lens evidence** that the workspace/ocean split is architectural in cl33:
   operators are the verbalizability peak of a three-model study (replicated on two cl33
   runs, domain-invariant across two held-out domains); transported state is invisible to
   the lens (§9) *and* output-causally inert under direct steering — a residual bias
   out-steers an operator-field bias ~25× with no positional compounding at either
   site (§3.1, closed 2026-09-08).
7. A **capability battery on the axis the architecture is for**, each against a matched
   baseline that collapses or cannot produce the artifact: exact retrieval at long
   sequence length (1.000 vs a matched transformer's ≈chance at T=1024/KV=64 — with the
   scoping that addressing is short-range: distance-resolved recall collapses with
   query–pair gap and is a measured limitation, §5.4 — across a pre-registered
   fairness sweep), **editable memory** with transaction semantics and bit-exact undo
   (provable unlearning-with-receipt), and **causally-faithful provenance** (attribution@1
   1.000 vs a transformer best-head 0.017) (§6).
8. The **measured transparency tax and its governance inversion**: +29% bits-per-byte and
   ~5 pp average MCQ against a param-and-context-matched from-scratch control (honestly
   reported; no superiority claimed, and the tax itself is now shown to be data-bound —
   continued training reaches class-ruler bpb parity at matched tokens, §10 Table 1b), and the argument that because the reverse trace
   *is* the computation, that tax buys an audit guarantee post-hoc interpretability cannot
   provide at any cost — scoped to invertibility/editability/decodability, not correctness
   (§10).

---

---

# Part I — The Object

*The bus exists, transcribes content and state, is the verbalizability peak, and costs a measured tax.*

## 2. Architecture (summary)

*(Full spec: CL33_OPLM_ARCHITECTURE.md v2.2; tape extension: REVERSIBLE_TAPE_DESIGN.md,
CL33_OPLM_V24_TAPE_LM.md.)*

- **Geometry.** Cl(3,3) with η = diag(−1,−1,−1,+1,+1,+1). The 15-dimensional bivector
  space spans 6 rotations and 9 boosts; rotors R = exp(B·G) form SO(3,3).
- **Emitter.** A standard transformer (6 layers at probe scale; 8–12 at 5B scale) reads
  tokens and emits, per block and per token, three bivectors: B_state (the operator),
  B_q, B_k (query/key rotors). Coefficients are norm-clipped.
- **Reversible scan.** s_t = R_t · s_{t−1}, fp32 with matrix_exp (bf16 diverges).
  Inverses are free: R⁻¹ = ηRᵀη. The reverse trace *is* the explanation — replaying the
  scan backwards recovers every intermediate state exactly — 2e-14 in fp32 (fp64
  reaches machine ε); anchored long-horizon reads reach 6.1e-5 at T = 1024.
- **Operator-native attention.** Q and K are rotors; the score is their η-metric
  product. No dot-product attention over hidden states exists anywhere.
- **Operator-only readout.** Final features are a grade tower over (attention output,
  state, operators, wedge grades); the readout is linear on these. The emitter's hidden
  states never reach the output. This is the transparency bottleneck, and it is the
  load-bearing design decision.
- **Bypass blade (used throughout).** For any capability, re-evaluate with operators
  ablated (identity rotors / zeroed operator features). Report the ratio of ablated loss
  to the unigram/chance floor. Ratio ≈ 1 ⇒ the capability lived in a bypass. On the LM
  task the base model's ablated-PPL ratio climbs from 8.3× to 42× over training at 32M,
  and reaches 206× at 163M/5B — transparency *strengthens* with capability rather than
  eroding.

**Precision discipline:** split precision (`amp_emit`) — fp16 emitter, fp32 algebra.
Full fp16 NaNs from η-score/wedge overflow.

---

## 7. Reading the words back: the operator stream as transcript

*(Full protocol, pre-registration, and adjudication: INVERSION_BATTERY.md.
Design: tiered decoders + prefix-only control + fixed-context substitution,
document-disjoint splits.)*

Sections 3–5 established the forward direction: the operators are causally necessary
for the output (ablation multiplies loss 8–206× depending on scale), and by
construction nothing else reaches the readout. This section measures the **inverse**:
given only the emitted tuple O_t = concat(B_state, B_q, B_k) (1440-d per position),
can the *input* words be decoded back out? If yes, the operator record is a
bidirectional transcript — input decodable backward, next-token decodable forward
(§9), trajectory invertible in the middle — and the residual skeptic's position
("the transformer emitter does the real work behind the bottleneck") has nowhere
to live: whatever the emitter computes either passes through the operators
(measured here) or does not affect output (ablation, §2).

**Setup.** Frozen scale-C (163M, ~449k steps), 1.23M positions of held-out mix
text across 5 domains, GPT-2 BPE (50,257 vocab, unigram floor 0.082 / 11.16 bits),
document-disjoint train/eval splits, shuffled-pairing and Gaussian floors.

**Results.**

| Decoder | top-1 | CE bits |
|---|---|---|
| O_t → tok_t, linear | 0.725 | 3.04 |
| O_t → tok_t, MLP | **0.862** | **1.58** |
| prefix-only O_{t−4..t−1} (control) | 0.356 | 6.44 |
| frozen LM's own head (strongest prefix baseline) | 0.425 | 5.12 |
| emitter hidden h_t (ceiling) | 0.906 | 1.10 |
| shuffled pairing | 0.082 (= floor) | 10.94 |

- **Bits beyond prefix predictability: 4.7** (3.4 against the LM's own head) —
  the decoding is not a language-redundancy artifact.
- **Emission loss ≈ 0.5 bits** (operators 1.58 vs emitter hidden 1.10): the
  bottleneck transmits nearly everything the emitter knows about the current
  token — and this loss *shrank* from ~1 bit at 32M. Like the bypass ratio, the
  bottleneck's fidelity improves with capability.
- **The code is distributed and role-structured.** No single block exceeds 0.08
  alone (32 blocks ≈ 0.01–0.08 each, from the pilot's B_state-only random-split
  battery; full stream 0.725 linear); families order
  B_state (0.605) > B_k (0.523) > B_q (0.460) — the state operator carries the
  most identity, the key rotor advertises *what is here*, the query rotor encodes
  *what to seek*. (The pre-registration predicted all three families ≈ equal; the
  measured ordering is a pre-reg deviation, reported as such.)
- **One object, three tenses.** The same emission also carries the previous token
  (B_state decodes tok_{t−1} at 0.29 — pilot battery) and the forthcoming
  prediction (O at t−1 decodes tok_t at 0.377 — the forecast channel §9 sees from
  the output side), while the transported state S decodes ≈ nothing (0.02, pilot):
  content rides on the operators from the decode direction too. (An earlier draft added
  "path rides on the state"; that is now closed rather than hedged — S is the
  accumulated path *by construction*, but its causal share at the output is measured
  ≈0: near-zero on flow probes and 25×-dominated by the residual under direct
  steering, with no positional compounding — §3.1.)

**Interventions (identity, not expectation).** Holding the prefix fixed and
substituting the token at position t, the frozen decoder identifies injected
mid-frequency tokens — including semantically strange ones in alien contexts —
at **0.82** (plausible 0.94, random-vocab 0.32, rare>20k 0.12, MLP probe). Errors
land on the contextually-expected token only 2.9% of the time (linear probe). The
channel reads actual token identity, not contextual expectation; identification is
frequency-graded at the rare tail — a reading we note is the post-hoc adjudication:
the pre-registered prediction that random and rare tokens would identify within 15pp
of mid-frequency ones *failed* (the battery's falsification line — errors collapsing
onto expected tokens — did not trigger, which is the part that carries the claim). The same-token-across-contexts test quantifies
the blend: one token's operator representation moves substantially with context
(within-type variance 65% of global) while remaining decodable (83% mean recovery).

**Reconstruction.** Decoding entire held-out documents from the operator log alone
(±2 window) recovers **80.7% of tokens exactly** (per-document 0.52–0.98 — narrowly
missing the pre-registered ≥0.85 gate, which we report rather than round), and the
misses are near-synonyms — "growth"→"innovation", "petroleum products"→"oil
products", "says"→"said". The log reads back as a meaning-preserving paraphrase
where it is not verbatim: the identity code degrades *semantically*, not randomly,
consistent with a frequency-graded code on a semantic operator geometry (App. C).

**What we claim:** the operator stream is a bidirectional, frequency-graded
transcript — ~9.6 of ~11.2 bits of current-token identity per emission, 4.7 bits
beyond predictability, with identity confirmed under substitution. **What we do
not claim:** verbatim recovery of the deep-rare vocabulary; the code is
distributed, context-modulated, and frequency-weighted.

## 8. The claims ladder: what the model can read of itself

We evaluate self-structure with an explicit ladder, most-skeptical-first framing, and
the bypass blade at every rung. (Full protocol and adjudications:
SELF_STRUCTURE_CLAIMS.md.)

**Claim 1 — transparent recall.** The model retrieves via a mechanism we can read.
**CLOSED** by §5: recall is a named read (increment mode) of a lossless record, with
operator-ablation destroying it.

**Claim 2a — structural path-sensitivity.** The model's behavior depends on its
trajectory *as a path*, not as a bag of visited states. Test: counterfactual history —
block-swap surgery on the tape record preserving the operator-entry multiset while
recomputing the displacements it implies (the P1 intervention; the distinct P2
same-visited-state variant probes the state-mode reader and is not conflated here),
injected via tape-override hooks. **PASSED** for displacement mode:
outputs shift when path structure shifts, and the shift vanishes under operator
ablation (so it is not a token-order artifact). This is exactly what the §5.3 taxonomy
predicts: displacement is the global-path reader; only the path reader should be
path-sensitive — and increment/state modes serve as built-in negative controls.

**Claim 2b — semantic order-readout.** The model can *report* trajectory order in
language (narrative-order probes). **NULL, twice** (full model, then scan-only retest —
which came back null/inverted: records get *discounted*). Mechanistic explanation:
order-content is attention-dominated at readout; the trajectory channel influences
prediction (2a) without being linearly decodable into tokens. We initially treated this
as a failed experiment; §9 reframes it — the trajectory lives in the non-verbalizable
part of the model, and the null is the ocean behaving like the ocean. A CE-pure rule
(no auxiliary order-pressure losses just to pass our own probe) is in force; the aux
variant is deferred as a side experiment, not a rescue.

**Claim 3 — self-modeling (the prize; held skeptically).** The model distinguishes
*its own* generation history from a compatible foreign one. **LIVE, not passed.**
Current evidence: a generator-identity signal 4× the content signal, carried on the
operator record (blade decomposition: operators-only 0.092 vs tokens-only 0.0012 — the
carrier is blade-2, the algebra, not the tokens). The phase-0 probe design confounded
live-vs-dead records with own-vs-foreign and was redesigned; the honest current
statement is that we can *name the boundary* — ownership vs familiarity — and have
localized the carrier, but the discriminating probe (twin-with-same-seed training;
ownership under matched familiarity) has not been run. We commit to reporting it either
way.

**Blade rule (all rungs).** Any behavior surviving operator ablation is bypass and does
not count — the same standard the tape was held to.

---

## 9. The workspace is a wire: Jacobian-lens results

**Instrument.** The Jacobian lens (Anthropic) asks, mechanically, what an internal
activation is *disposed to make the model say*: lens_l(h) = unembed(J_l h), with J_l the
input-averaged Jacobian from layer l to the final representation. No trained probes. On
large models it reveals a *workspace*: a privileged subspace whose contents are
verbalizable, surrounded by an *ocean* of computation that shapes behavior but is never
linearly reportable — a structural echo of Global Workspace Theory's broadcast
bottleneck.

**Method.** Faithful estimator port to cl33 (one-hot cotangents at all valid target
positions, source-averaged, SKIP_FIRST=16; batch-replication for tractability). Taps:
the six emitter layers h1–h6, the emitted operators B (post-clip), the transported state
S (differentiable scan copy). Controls: stock GPT-2 124M and the warm-start +5B-mix
GPT-2 control. Two cl33 models are reported — scale-C (163M) and scale-B (236M), both at
5B-token final — so the split is tested across two independent runs. Fitting used
domain-balanced prompts (300 for the GPT-2 controls, 60 for cl33), 40 disjoint eval
prompts. Metrics: future-token MRR and final-agreement per tap.

**Results (final-agreement by depth, in-domain).**

| Model | Shape |
|---|---|
| stock GPT-2 124M | .000 × 7 layers, then cliff: .018 / .033 / .138 / **.288** |
| warm-start (+5B) | no dead zone; smooth ramp .043 → **.583** |
| cl33 scale-C 5B | emitter .088 → .517; **B = .613**; **S = .000** |
| cl33 scale-B 5B | emitter flat .049 → .097; **B = .538**; **S = .000** |

Five readings:

1. **No workspace at 124M by default.** Stock GPT-2's first seven layers are pure
   ocean — the pre-registered prediction, confirmed. Everything interesting it computes
   early is linearly invisible to its own output basis until the end.
2. **The operators are the verbalizability peak of their own stack on every domain
   tested, and above every control layer in-domain** — 0.613 (scale-C), above the warm
   control's final layer, replicated at 0.538 on the independently-trained scale-B.
   (One out-of-domain exception, reported: on PubMed the warm control's final layer,
   0.563, edges scale-C's operator tap, 0.552.) Not mysterious, and that is the point: in a transformer the lens asks whether
   a layer *happens* to contain the output; in cl33 the operators are the *only path to*
   the output, so training must put the output there. The workspace is not emergent — it
   is the bottleneck, made.
3. **The state is completely non-verbalizable.** S scores 0.000 on both models — floor.
   This is the paper's ocean, realized as a *named architectural component*: the
   workspace/ocean boundary, statistical in standard models, is here the line between
   "emitted" and "transported" in the wiring diagram. An earlier draft paired this with
   "causally load-bearing"; the §3.1 audit weakened that pairing, and the 2026-09-08
   constant-bias probe now closes it (§3.1, point 3): at the output, a residual bias
   out-steers an operator-field bias ~25×, and neither effect compounds with position —
   the state's accumulated history does not reach the readout. S is the non-verbalizable
   *and* output-causally-inert residue of the architecture: non-reportable and, for
   output steering, ≈0 causal share. The workspace/ocean split is thus doubly clean —
   what is transported is neither readable nor (at the output) load-bearing. The retro-
   explanation of the Claim-2b null ("order-content lives on the S side") remains
   demoted to hypothesis.
4. **Upstream emitter organization is model-dependent; the bottleneck is not.** scale-C's
   emitter *pre-organizes* — a smooth ramp .088 → .517 into B — while scale-B's emitter
   stays near-flat (.049 → .097) and *jumps* at B, closer to the original "dead emitter,
   workspace at B" pre-registration. Whether the bottleneck's gradient reaches back
   through the emitter varies with training (scale-B saw ~6× fewer optimizer updates);
   that B *is* the peak and S *is* the floor does not.
5. **The split is domain-invariant (confound closed).** The fit/eval prompts above share
   the training-mix domain, which could inflate absolute *levels*. To rule this out we
   re-fit and re-evaluated all four models on two held-out domains none trained on —
   PubMed abstracts (biomedical) and congressional bills (legal). The two load-bearing
   signatures survive both: cl33's **B remains the verbalizability peak** (scale-C B
   .552 / .751, scale-B B .490 / .721 on PubMed / legal) and **S remains at floor**
   (≤ .018 everywhere); stock GPT-2 keeps its dead-then-cliff shape (no mid-workspace)
   and warm-start keeps its bottleneck-free ramp, on both domains. Levels move with
   domain; shape does not. One honest scoping point falls out: on legal text — so
   formulaic that even shallow features predict the next token — the *emitter* lights up
   too (scale-C .54–.62), so reading #4's "dead emitter" contrast is itself a property of
   harder-to-predict text. What is architectural, and domain-robust, is *B-is-peak,
   S-is-floor*; how far B towers over the emitter is not.

**Caveats.** 124M–236M-class models only; first-pass estimator with 60 fitting prompts
for cl33 (one repetitive-text lens-fit artifact noted: a single mid-emitter spike on
scale-B/legal, with B still the clear peak); the emitter-organization effect (reading #4)
deserves a controlled comparison — the same emitter trained without the bottleneck —
before strong claims. The 5B-final rerun and the OOD domain-invariance control, both
pre-registered gates in earlier drafts, are now discharged.

### 9.6 A regime sensor on the bottleneck (2026-09-08 — first measurement, one control open)

If the operators transcribe the emitter's *state* as well as its content, the bottleneck
is an instrument, not only a record. First measurement: a plain difference-of-means
direction in operator space separates text regimes (fiction vs code) at **0.962**
held-out; four regime centroids (fiction / code / math / encyclopedic) classify a
**single token's emission** four-ways at **0.802** (chance 0.25), spread across a
multi-axis manifold (PC variance 0.52/0.28/0.19) whose axes each use the full
compact+noncompact blade geometry (~uniform sector mass — no sector starvation). The
load-bearing control: **hold the token fixed and vary only the surrounding domain.**
Classification survives essentially unchanged (pooled **0.811**: " the" 0.905, "." 0.877,
"," 0.700) — the signal is *contextual computational state, not lexical identity*: the
same symbol is emitted as a measurably different operator depending on the regime the
emitter is in. Because the readout consumes only operators, the material the sensor reads is on the
causal path by construction; whether the readout *uses the particular direction* the
sensor reads is a separate question — one the closed-loop steering test below partially
answers, and the open control decides. A first closed-loop test (steer the residual toward the math centroid,
read the sensor) shows a monotone dose-response with the fixed-token invariance intact,
and at sufficient gain the *generated text* itself turns numeric. **One control is
open and we hold the stronger claim until it closes:** at low gain the sensor crosses
the regime boundary before sampled generation does, and sensor and actuator directions
derive from the same contrastive means — so the sensor-leads-behavior gap must be
adjudicated (faithful early warning vs partial injection echo) by classifying sampled
generations across the gain range before "state sensor" hardens beyond "readable
projection of upstream computation."

---

## 10. Cost accounting and limits

**The transparency tax, measured (locked, matched, fp32).** With the from-scratch matched
control complete, the tax is no longer provisional. On a single-session, full-set,
right-padded, fp32 lm-eval-harness (v0.4.11), on a common wikitext ruler at 512:

| model | params | tokens | wt-bpb ↓ | avg-acc (8 MCQ) |
|---|---|---|---|---|
| cl33 scale-C | 163M | 5B | 1.309 | 0.435 |
| cl33 scale-B | 236M | 5B | 1.218 | 0.440 |
| **A1 (vanilla, matched control)** | 166M | 4.9B† | **1.012** | **0.483** |

The matched vanilla transformer wins bits-per-byte by **+29%** over scale-C (and **+20%**
over the larger scale-B) and **6–7 of 8** downstream tasks. This is the honest cost of
forcing every prediction through a reversible, per-token operator bottleneck. We report it
plainly and make **no** claim of perplexity or reasoning competitiveness. An earlier draft
guessed a ≈1.2× premium; the measured figure is larger (~1.29× on the loss ratio) and we
correct it. Bits-per-byte is the confound-free ruler; word-perplexity dramatizes the
same underlying gap through tokenization and chunking differences, and we do not report
it as a cross-comparison.

**A scaling dissociation, stated plainly.** From 163M→236M, cl33's bpb improves
(1.309→1.218) while avg-MCQ is flat (0.435→0.440): 73M extra parameters bought *zero*
downstream accuracy. Downstream MCQ is param-saturated / data-bound at 5B and does not
track cl33's own capability scaling — at fixed data it is the wrong yardstick for this
architecture (the axis it *is* built for is §6; Table 1b below shows MCQ does move when
the *data* moves). The model is genuinely fluent (coherent multi-sentence
English) and its operator index sharpens with scale (nearest-neighbor same-category rate
0.79 at 5B vs 0.68 at 32M).

**Table 1b — the tax is data-bound, not an architecture floor (2026-08-03).** The
scale-B model, continued on a prose-weighted mix to 6.2B tokens at 1024 context, was
benchmarked on the same full-set fp32 right-padded harness (new session; the locked
Table 1 above is untouched):

| model | tokens | wt-bpb ↓ | avg-acc (8 MCQ) |
|---|---|---|---|
| cl33 prose base (scale-B cont.) | 6.2B | **1.155** | 0.451 |
| Pythia-160m @6.3B (class ruler) | 6.3B | 1.155 | 0.462 |
| (ref) cl33 scale-B | 5.0B | 1.218 | 0.440 |

The bits-per-byte reaches **parity with token-matched Pythia-160m to three decimals**
(1.15501 vs 1.15452) — the model-class fluency gap closes at matched tokens. Against the
(now un-matched, still-5B) A1 control the gap narrows from +20% to +14%; the true
matched tax at 6.2B is unmeasured — A1's own convergence (1.017 → 1.012 over its final
0.5B) suggests it sits near +14–15% rather than the +29% headline, but that is an
inference, not a measurement. Three honest caveats: the Pythia row is carried over from
the model-class context session (fp16; its source table marks it "scale intuition,"
not precision-matched — the parity claim is therefore cross-session), the prose-weighted
mix partially aligns with the wikitext ruler (Pythia's Pile is also prose-heavy, so the
class comparison remains roughly fair), and A1 was not extended, so Table 1's matched
tax remains the only fully controlled number. MCQ also moved for the first time
(+1.1pp avg; BoolQ +5.0pp — notable because a bias-immune re-analysis of BoolQ found
cl33 the only model in the comparison with significant passage signal: answer-bias-
corrected AUC 0.522, p = 0.037, versus Pythia's raw 0.61 collapsing to chance AUC 0.494
once its yes-prior is removed), qualifying
the "param-saturated" reading above: MCQ was *data*-bound, and prose tokens were what it
wanted. The claim discipline is unchanged — no superiority is asserted — but "the
bottleneck costs a fixed 29%" is now falsified by the architecture's own training curve:
the tax is a property of a checkpoint, not of the algebra.

The governance consequences of this cost table — where the tax inverts, and what
exactly is guaranteed — are collected in Part III (§13).

*Methodological note that gates the numbers.* The lm-eval adapter inherited a left-padding
step that **corrupts the reversible scan** (leading pads inject spurious rotors); it drove
one binary task below chance until we replaced it with right-padding (batched output then
matched batch-1 exactly). Every operator-model eval must right-pad. All cl33 harness
numbers are full fp32 (emitter *and* algebra); the fp16-emitter split-precision path is
training-only and never touched the eval numbers (verified: the eval adapter runs
`amp_emit=False`).

† A1 stopped at 4.915B tokens (its GPU was reallocated ~1.7% short of the 5.0B budget). Its
bits-per-byte had converged — 1.017 → 1.012 over the final 0.5B — so the token gap does not
move the tax; we report it here rather than paper over it.

- **Perplexity ruler.** bits-per-byte, matched params/context/tokenizer/corpus, single
  session. word-perplexity against differently-chunked or longer-context baselines is not
  comparable and we do not report it as such.
- **Compute.** fp32 matrix_exp scan is the tax; split precision recovers most emitter
  throughput. Tape adds O(T²) addressing over 6-d keys (cheap constants, honest
  quadratic).
- **Reversibility scope.** Per-step exact inversion with anchoring; boosts create an
  expressivity/reversibility tension (expansion-as-memory) that we surface rather than
  hide — the compact/non-compact decomposition of learned directions is a pre-committed
  analysis (App. C).
- **Not claimed:** LM superiority; sub-quadratic memory; Claim 3; any consciousness
  gloss on the GWT correspondence — the correspondence is structural (reportable
  bottleneck vs non-reportable processing), and we use it only as a naming convention.
- **Negative results reported:** wedge associative memory (bypasses or fails);
  gauge-throttle optimizer (all levers hurt); Claim-2b (twice); sequential falsified
  configs in the lineage (v6-nuclear, Dec-steering, lattice-correlation-only,
  widen64) — each killed one alternative and is part of why the surviving architecture
  is this simple.

---

---

# Part II — Memory on the Object

*What the bus does not give for free, the falsifications that located the sole-channel requirement, and the organ that installs, adopts, and deletes with receipts.*

## 3. Two memories, one identity

**Question.** When the model retrieves, does it use the operator machinery or route
around it?

**Setup.** MQAR (multi-query associative recall; Beck et al. harness) for associative
memory; a KGE-style relational task (compose R·head ≈ tail along paths) for relational
memory. Both scored with the bypass blade.

**Result (the dichotomy).**

| Capability | Mechanism | Works? | Operator-dependence |
|---|---|---|---|
| Associative recall (MQAR) | prev-token key / adjacency (induction) | 0.92–0.97 | ≈ 0 (91% bypass in the worst configuration) |
| Relational navigation | operator-trajectory composition | groks to 0.95 | ≈ 0.85, scale-invariant |

**Result (the identity).** The induction shortcut and the bypass are not two problems —
they are the same operation. Recall-by-induction *is* token-similarity matching that
never consults the algebra; the wedge-memory design matrix (§5.1) shows every
configuration that recalls well before the tape does so through a token-content value
channel, and every configuration that keeps values algebraic fails to recall (≈ 1/KV)
*at matched training budget* — Appendix F shows the algebra-value store eventually groks to
full recall at roughly 10× the training cost, so the identity is about learning
efficiency and shortcut formation, not impossibility.
The rotor scan cannot perform outer-product writes: a rotor is orthogonal under η, so
"store this key–value pair" has no native expression in the transported state. This is
the precise gap the tape (§5) is designed to close, and the reason its design rule is
addressing/content separation rather than "more algebra."

**Interpretation.** Prior to the tape, cl33's honest self-description was: *a navigator
that borrows a transformer's memory*. Navigation — composing a trajectory of operators —
is native, transparent, and load-bearing. Recall was borrowed, opaque, and bypassing.
The rest of the paper is about making memory native (§5) and then asking what the model
can read of its own trajectory (§7–§9).

### 3.1 Channel-attribution audit (2026-07-15): which operator path is causal?

Operator-dependence throughout this paper is measured by *full* operator ablation
(zeroing B everywhere). That blade is sound for the bypass question — does the capability
live outside the algebra? — but it cannot localize *within* the algebra, because the
emitted operators reach the readout by two distinct paths: **via the scan** (they
transform the transported state) and **directly as readout features** (B plus its grade
tower). An audit on a branching-track probe (infer a flow k from a demo, carry it across
a filler gap, apply it to a novel start) decomposed the two, with three results.

1. **The flow rides the feature path, not the state.** Swapping B-conditioned operators
   into the scan alone leaves the branch on A (flip 0.000); swapping them into the
   readout features alone flips it completely (1.000). Replacing the scan outright with
   identity rotors costs ≈0 accuracy at every distance (0.96–1.00, including gap-1024
   extrapolation at 8× the training gap); the state-only channel is weak *and*
   distance-limited (0.49→0.19 with distance).
2. **Distance-flatness comes from attention, not transport.** The emitter is itself a
   transformer attending over the whole visible context; the flow-conditioned operator
   it emits at the query is how the rule crosses the gap. The η-attention can be removed
   (scan_only) without loss; the emitter's internal attention cannot be — and masking
   the demo at generation time collapses continuation to chance even when the rule is
   inferable from tokens in view (with the recorded caveat that PAD-masked demos are
   out-of-distribution, so part of the collapse may be brittleness rather than
   dependence). Continuation is *demo-conditioned emission*, re-read
   each step, not a rule stored in the transported state. Relatedly, on-manifold state
   perturbations show no restoring dynamic: the output projection is insensitive to
   ~1× a class-separation of displacement at emission time, and trajectories knocked
   off their flow *stay* off — exactly what reversibility (no contraction) predicts.
3. **The causal share of S at the output — measured (2026-09-08).** A constant-bias
   steering probe on the frozen prose base closes what earlier drafts left open. At
   matched perturbation scale, a residual-stream bias produces **~25×** the output
   effect (KL to the unperturbed distribution) of an operator-field bias on B_s; and at
   *neither* site does the effect grow with position (late/early ratio 1.0–1.1× at every
   magnitude). The second number is the decisive one: a persistent operator bias
   provably accumulates in the transported state through the rotor product, so if S's
   accumulated displacement reached the readout, the effect would compound along the
   sequence. It does not. Scan-compounding is falsified *at the output*: the state
   remains the substrate on which operators act, not a carrier whose history the readout
   consumes. The KGE navigation task of this section has still not been decomposed
   path-by-path ("operator-carried," op-dep 0.85, remains the honest label there), but
   the general hedge — "S's causal share is unresolved" — is now resolved for output
   steering: it is ≈0.

**Design consequence (feeds §5).** The tape's recall/receipt properties were measured
directly — in configurations where the tape is the sole cross-position channel — and the
same audit week recorded the matching caution: in any configuration where attention
coexists with the tape, tape-causality must be re-verified, because if attention
re-carries recall the tape becomes a passenger and editing it will not move the output.
The subsequent v2.4 run measured exactly that failure mode (§5.4), which is why the
successor designs make the memory the sole cross-segment channel by construction. The
audit also sharpens the division of labor: the scan state alone neither recalls (§3) nor
transports rules across distance; every distance-free capability we have measured is
carried by an attention-addressed read of the operator record.

### 3.2 Independent replication and extension (Watson, 2026-09-10): the scan is
### decorative at LM inference

An independent, pre-registered, eval-only audit by Nell Watson on the *released
checkpoints* (both models, the shipped frozen slice and WikiText-103, fp32,
hash-verified) first reproduced the bottleneck table to one decimal, then extended the
§3.1 decomposition to the language model itself. Five conditions; PPL ratio vs native:
operators zeroed **270×/115×** (the bottleneck, replicated); S pinned to s₀ with
attention on **1.19–1.22×**; S *time-shuffled* **1.00×**; η-attention zeroed **1.00×**;
S pinned *and* η-attention zeroed (readout sees only the local operator window through
the grade tower) **1.00×**. The time-shuffle control is the decisive one — we replicated
it on our own harness (0.996) — and it shows the pinning cost is an artifact of the
intervention rather than lost ordered history (Watson's shared-LayerNorm explanation
is the candidate mechanism; it has not yet been separately isolated). The conclusion, stronger than §3.1 and
measured on the LM proper: **at inference on these checkpoints, the reversible scan and
the operator-native attention over it are causally decorative; every bit of context
reaches the output through the emitter's own attention, expressed in the emitted
operators.** Two scopings keep this honest in both directions: (a) in sole-channel/tape
configurations (§5, §6) the scan's prefix products *are* the memory and ablating them
destroys recall — the decorativeness is a property of the full LM configuration, not of
the substrate; (b) the current evidence is *consistent with* training preferring the grade tower's
direct route to B_t, B_{t−1}, B_{t−2} once it became available, but does not yet
distinguish "the tower out-competed the scan" from "the scan mattered only as a
training scaffold" from "the full-scale model never needed it" — and, as Watson notes,
"the reversible path cannot do work" is not shown either. The v0 pre-tower
model did use the transported state (the model_v2 docstring's "collapses to unigram"
prediction was written for, and was true of, that configuration; it is stale for the
shipped one). Half the mechanism is closed by algebra rather than ablation: §5.1's binding
impossibility (a grade-1 carrier under the grade-preserving action cannot hold grade-2
content) already establishes the substrate-inadequacy branch *for the memory function*,
with its constructive converse measured in the left-regular successor line. The two
controls that would close the residual — whether the state could have carried *generic*
context and the tower out-competed it (a no-tower model), or contributed nothing even
during training (a no-scan-from-init model) — are queued as a v1.3 table at probe
scale, scoped to that narrower question.

One reading of this result strengthens the transcript claims rather than qualifying
them. If the readout consumes only a local operator window and the transported history
carries nothing, then *each emission must be a contextually complete summary*: the
emitter's attention integrates the visible history into every B_t before emission.
That is the mechanism behind the density of the record — why a single emission decodes
its token at 0.86 with +4.7 bits beyond prefix predictability (§7), and why regime
state reads out of one token's operators under lexical control (§9.6). The log is
information-rich *because* the decision path is local: everything the context
contributes has to be compressed into the emission itself. And the completeness
guarantee is unchanged in the only direction governance needs: nothing behaviorally
relevant can bypass the record — what the model does not transcribe, it provably
cannot act on.

---

## 4. Navigation parity (pre-registered)

**Task.** Permutation composition — apply a composition of permutations read from
context; a TC⁰-flavored weak spot for fixed-depth transformers and the capability class
closest to cl33's native operation (rotors compose; permutations are rotors at the
integer points).

**Protocol (pre-registered; deviations logged in the pre-reg memory).** Learning-rate
sweep for *both* models; winner LR re-run with 3 seeds; report mean ± sd of the
mean-of-3 at the winning LR. (An earlier "cl33 matches-or-beats" framing from a
max-of-5 readout was retracted when the mean-of-3 protocol shed the max-inflation —
we report the corrected protocol only.)

**Result.**

| Model | Accuracy (mean of 3 seeds) |
|---|---|
| cl33-opLM | **0.433 ± 0.009** |
| tuned transformer | **0.432 ± 0.030** |

Statistically indistinguishable. Two secondary observations: cl33's seed variance is
≈3× tighter, and one cl33 seed (43) reached its score with *zero measured bypass* — an
existence proof that the parity-level solution can live entirely in the algebra.
Mechanism note: grokking on this task coincides with rotor spectra collapsing toward
permutation scale.

**What we claim:** parity + consistency + inspectability on the task family nearest the
architecture's native operation. **What we do not claim:** superiority. A gauge-throttle
optimizer intervention intended to force more-algebraic solutions was falsified (every
lever hurt; weight decay worst) and is reported in App. D as a negative result.

---

## 5. The reversible tape: exact memory native to the scan

### 5.1 The failure that shaped the design

We first attempted associative memory as a wedge-product matrix store (C += k∧v read by
q⌋C — mLSTM's matrix memory expressed in Cl(3,3)). The design matrix:

| Key source | Value source | MQAR | Bypass |
|---|---|---|---|
| operator (context-mixed) | any | ≈ 1/KV | — |
| state (context-mixed) | any | ≈ 1/KV | — |
| token (context-free) | token projection | 0.92 | high |
| token (pure copy) | token | 0.97 | **91%** |
| **token (context-free)** | **algebra-only** | **untested** | **untested** |

Context-mixed keys cannot re-form at query time (falsified twice); token values recall
but bypass. The untested cell — token *addressing* with algebra-only *content* — is the
tape.

### 5.2 Design principle

> **Token content may SELECT; only algebra may FLOW.**

The scan already writes a lossless record: prefix products P_t = R_t·…·R_1 give the
exact relative transport R_{t←i} = P_t · ηP_iᵀ η between any two positions, for free,
by reversibility. The tape adds only a *read head*: a context-free grade-1 projection of
the token embedding produces match scores α over strictly-past positions (with an
induction shift: match position i, read i+1); the returned content is Σ α_i · f(tape_i)
where f yields exclusively algebra objects. Token information can influence *which*
positions are read (a low-bandwidth softmax channel) but no token content vector ever
reaches the readout. Numerical discipline: raw prefix products diverge by T ≈ 1024
(error 1.3e6); anchored re-orthogonalization every K = 64 steps is mandatory and cheap.

### 5.3 Value modes and what each one is

| Mode | Content (dim) | Reads the record as | MQAR | LM PPL (32M/8k) | Ablated ratio |
|---|---|---|---|---|---|
| increment | B_i (15) | local adjacency | **1.000** (all KV) | 22.94 | **4.88×** |
| displacement | R_{t←i} (6) | **global path** | ~1.000 | **20.10** | 2.17× |
| state | s_i (6) | unstructured field | no binding | 20.84 | **9.99×** |
| multi | all three (21) | — | 1.000 | 20.47 | — |
| no tape (v2.2 base, matched 32M) | — | — | ≈ 1/KV | 21.30 | 11.74× |

Three findings:

1. **Recall without bypass, both blades.** Increment mode solves MQAR *perfectly* at
   every key–value load, and ablating the operators *destroys* recall (ratio 4.88×,
   far above the ≈1 a bypass would show). The two-memory gap of §3 is closed natively
   in this configuration — scan-only at probe scale, with the tape the sole
   cross-position channel; §5.4 records what happened when that condition was relaxed.
2. **The mode taxonomy is a reader taxonomy.** The same tape read three ways yields
   three different capabilities: increments give adjacency (recall), displacements give
   path integration (helps general prediction: best PPL), states give a cheap
   high-transparency feature channel. This taxonomy does independent work in §8.
3. **The attention-free challenger.** Scan + tape with *no attention at all* reaches
   VAL 19.02 at matched 32M/8k — beating the full model with tape (20.10) and the
   no-tape baseline (21.30). Honesty: this is not sub-quadratic — tape addressing is an
   O(T²) softmax over 6-d token keys. The win is *simplification + quality*, not
   complexity class. (A pre-registered 5B challenger run of this config was planned;
   it was superseded by the v2.6 redesign before running — the 5B evidence on
   η-attention's inertness, §3.1, arrived first.)

### 5.4 Pre-registered v2.4 gates

The 5B retrain (multi-mode tape, anchored displacement) carried pre-registered gates —
PPL vs the from-scratch matched control, MQAR at deploy, bypass ratios per mode, J-lens
rerun (§9) — recorded in CL33_OPLM_V24_TAPE_LM.md *before* the runs.

**A co-adaptation regression, diagnosed and repaired (reported as a result).** The first
co-trained tape run (warm-started from a no-tape 5B checkpoint, with an auxiliary MQAR
loss to teach addressing) reached the tape gate — recall 0.77 within the first fifth of
the token budget —
but *regressed language modeling*: a matched-stream fp32 comparison showed the shared
base pulled off-optimum (base-alone val 31→88 across training). The mechanism is not a
gradient *conflict* — the LM and auxiliary gradients are near-orthogonal on the shared
emitter (cosine ≈ 0) — but a gradient *magnitude* imbalance: the auxiliary signal is
~4.6× larger there and additionally reshapes the shared readout toward memory tokens. The
fix is a principled coupling, not a loss weight: **cap the auxiliary gradient magnitude
to a fraction (κ = 0.5) of the LM gradient's on the shared emitter, and route it off the
shared readout** so memory
output is carried by a dedicated head. Under the fix the result exceeds the naive run on
*both* axes: the deployed model returns to the LM baseline (1024-context fp32 val 32.5→31.4
over steps 2k–6k, versus scale-B's 30.97) **while recall climbs past the naive run's
plateau** (0.77→0.80→0.89) and the hollowed base heals monotonically (base-alone 87→62→52).
The naive coupling had bought recall 0.77 at a standing +11% LM penalty; the principled
coupling buys recall 0.89 at ≈baseline LM (+1.2% vs scale-B, with the caveat that the
deployed 1024-context val and the 512-context baseline are not exactly the same ruler). This is the
kind of pathology the operator factoring makes *visible and
addressable*: because the memory path and the base path are typed and separable, the
imbalance is measurable per-parameter-group and correctable as a gradient operator rather
than a blind hyperparameter search.

**What the calibrated instruments then showed (the honest coda).** Three follow-up
measurements, run within days of the repair, scope the 0.89 sharply and are the reason
the successor designs exist. First, the 0.89 is a *short-range* number: the benchmark's
query–pair gap is ~40 tokens, and a calibrated distance sweep found recall collapsing
0.94 → 0.05 as the gap grows toward 448 (worse with natural-text filler) — the
work log records this as falsifying the tape's distance-independent framing, and we
adopt that verdict here. Second, the *learned* address projection — context-free by
design — was co-opted by language-modeling context on natural text (exact-match
addressing 0.006 against 0.26 on synthetic), the design property failing to survive
co-training. Third, decode-level attribution in the full configuration showed the tape's
own vote on recalled tokens at 0.000, with `--no_tape` recall ≥ tape-on: in the presence
of attention, attention re-carried recall and the tape rode as a passenger — precisely
the §3.1 caution realized. These three results — distance-limited addressing, address
co-option, and passenger collapse — are not appendix caveats; they are the measured
boundary of the v2.4 design and the direct motivation for the dual-address revision and
the segment-recurrent redesign in which the memory is the sole cross-segment channel by
construction (§12).

---

## 6. Capability doors: editable, provenance-bearing, load-invariant memory

*(Full pre-registration + adjudication: CAPABILITY_DOORS.md. All results here are on the
synthetic MQAR task at probe scale, in scan-only configurations where the tape is the
sole cross-position channel. The §5.4 coda applies: in configurations where attention
coexists with the tape, edit/receipt causality must be re-verified — the v2.4 full
config measured the tape as a recall passenger, and editing a passenger does not move
the output.)*

Once memory is a typed, exact, position-addressed record (§5), three capabilities follow
that no opaque KV cache offers. Each carries the bypass blade.

**Door 2 — editable memory (transaction semantics).** The counterfactual-override path is
used as a *write* API on a trained increment-tape model: **swap** two records (queries
return the swapped value, efficacy 1.00), **delete** a record (target recall → 0.008 ≈
chance, other pairs unchanged — locality 1.00), **transplant** a record from a *different*
sequence (foreign value retrieved, 1.00 — records are portable at this scale), and **undo**
(bit-exact restoration, max |Δlogits| = 0). Identity override is also exact (Δ = 0). This
is surgical context editing, auditable RAG, and — since delete is targeted forgetting with
a receipt and an undo — **machine unlearning with provenance**.

**Door 3 — provenance receipts (faithful, causal attribution).** The tape's address
weights are a *faithful* attribution: argmax-α identifies the ground-truth write position
with **attribution@1 = 1.000**, and the receipt is **causal** — deleting the argmax-
addressed record drops that query's recall to **0.000** while deleting a non-source record
leaves it at **1.000**. A matched transformer's best attention head reaches 0.977 *only*
under oracle head-selection, and that oracle head is **not stable** across load (it drifts
L1H2→L1H3→L0H1 across vocab/length; where the transformer's attribution collapses under
load, the diagnosis is that no attribution circuit formed — not that attention "smears"). The tape's address is a single canonical, causally
editable interface; the transformer's is an emergent, condition-dependent one.

**Door 1 — load-invariance at T = 1024, and a pre-registered fairness sweep.** On MQAR
the scan+tape holds ceiling as load grows (KV 64/128/256 at T=1024 → 0.997/0.999/0.999,
zero bypass), while the same fixed-96-d-state model *without* the tape sits at chance —
the tape rescues exact recall the recurrence structurally cannot do. Validation stops at
T=1024: a T=2048/KV=128 run did not grok within the grid's step budget (an honest gap;
T≥2048 is unverified). The matched d128 transformer, by contrast, failed to learn
T=1024/KV=64 across a pre-registered six-config sweep (three learning rates at depth 2,
plus depth-3, depth-4, and 2×-width runs at the winning rate; all 12k steps, all ≈
chance), where the tape reaches 0.99 in under 5k steps. **We scope this honestly:** MQAR is provably
solvable by attention at sufficient scale (Zoology built it to showcase attention), so the
claim is *not* that transformers cannot — it is **capacity + learning efficiency**: the
tape forms exact, zero-bypass, editable, receipt-bearing retrieval at lower capacity and
fewer steps than the transformer needs to form any retrieval at all.

## 12. The memory organ: from deadlock to a certified peripheral (2026-08 → )

The results above treat memory as a channel *inside* the operator LM. A successor program
— reported here as measured intermediates with pre-registrations, not finished claims —
grafts the Genesis two-channel memory (validated separately on frozen GPT-2/Gemma/R1
backbones) onto the frozen prose base of Table 1b as an explicit **register-bank organ**:
0.1–2.5M trainable parameters (per arm) against 236M frozen, wedge-plane records with trained-encoder
addressing, zero-init injection, and a removability gate that holds at literally
0.00e+00 ΔCE in every run (organ-off ≡ base).

The program's value so far is a **diagnostic ladder** — each rung a small set of
pre-stated changes with separated instruments, each verdict measured:

1. **Interface deadlock.** Next-token CE alone never differentiates addressing — the
   register softmax stays exactly uniform (9,000 steps with raw-embedding features;
   3,000 with emitter-tap features before that arm was cut). Cause: a
   chicken-and-egg between addressing and decoding, plus retrieval-demand density orders
   below every precedent that formed recall.
2. **Labels break it.** Planted-episode data makes the correct register a *free label*;
   a supervised address loss moves address accuracy 0.000 → 0.5 in 100 steps — what CE
   never did in 9,000.
3. **Context is mandatory for addressing.** Raw-embedding queries plateau structurally
   (the query token is a generic "is"); features tapped from three emitter blocks address
   unique-vocabulary episodes at 1.000.
4. **Retrieval forms.** With episode demand, seam-weighted loss, and per-8-token
   registers: sustained recall **0.458 with recall-off = 0.000** (zero bypass) — and a
   clean granularity dose-response (P=8 beats P=16 on both axes). Measured at full
   episode density: recall never crossed the pre-set 0.5 threshold that would have
   opened the density anneal, so retrieval-under-realistic-sparsity — the program's
   pre-registered decision point — remains unmeasured, and clean-text ΔCE stayed
   negative at this density.
5. **The decode ceiling is the read mixture, not the decoder.** At matched read width
   (H=8), two structurally different decoders — a free logit matrix and a KGE-style
   codebook whose write and read sides are the *same tensor* — land on the same
   conditional recall (0.565 / 0.567 given a correct register), localizing the ceiling
   to the unsupervised within-register mixture both consume. A 2×-width read measures
   conditional recall 0.716 (n=102) — width partially dilutes mixture interference —
   but pays for it in slower addressing, leaving total recall unchanged; the mixture,
   not the decoder, remains the binding constraint. The pre-registered test that
   removes the mixture entirely (per-token registers: the register *is* the binding;
   predicted conditional decode ≥ 0.85) has now run — rung 6.
6. **Mixture removed: the memory path is exact, and the losses decompose into three
   named numbers.** With per-token registers and the tied codebook, the memory path
   *alone* decodes at **1.000** (53/53): given the right register and a pure read, the
   record is not lossy at all. The remaining gap to headline recall then separates
   cleanly — soft-α selection conditional **0.463** vs hard top-1 read **0.756** (the
   α-softness cost, ~0.29), hard read **0.756** vs pure decode **1.000** (the
   integration cost through the frozen base, ~0.24), and addressing itself at
   **0.57–0.69**. Nothing is mysterious in the residual: an exact record, read through
   a soft mixer, injected into a base that never trained to consume it. Each named
   loss is a plumbing target, not a representation problem — the *record layer* is
   exact on that evidence; which value substrate feeds it best is adjudicated by the
   control arm below.

**The endpoint has a first measurement.** The battery the governance section implies —
install → adoption → certificate-bearing forget → revert — was run as a pilot on the
rung-6 checkpoint: 24 facts never present in any context, written through the organ's
own write path onto virtual segments, queried cold. Install top-1 **0.750** (median
rank 1, mean Δlogprob +8.80); in *generation*, 4 of 4 sampled queries adopt the
installed fact; algebraic delete → **0.000**; organ-off → **0.000** (zero bypass);
address receipts **24/24** (the α argmax lands on the installed register in every
case). Install efficacy landing on the same number as the hard-read integration
ceiling (0.750 ≈ 0.756) says the pilot is *at* the measured plumbing limit — installs
are as good as integration currently allows, no worse. Pilot-scale, single-seed,
one config; reported as such.

**The installed belief passes through the transcript.** A four-cell factorial on the
same checkpoint separates the two paths by which a read can reach the output: internal
injection only **0.407**, direct logit path only **0.017**, both **0.712**, neither
**0.000** — interaction **+0.288**, internal-carrier fraction **0.571**. The dominant
carrier is the *pre-bottleneck* injection: the retrieved memory alters which operators
get emitted, and the behavior change rides those operators to the readout. The receipt
is direct: on install queries the emitted B_s displaces by **0.777** on hits vs
**0.591** on misses. Because the readout consumes only operators (§2), this closes the
chain **memory → exposed cognition → behavior** with a measurement at each arrow —
the property an opaque-base memory graft cannot exhibit even in principle.

**Two methods results the sweep paid for.** First: injection-layer arms L6 and L11 both
matched the incumbent layer (L9) on every training-time validation metric, yet measured
sharply different install efficacy (L11: **0.083** with a validation profile *identical*
to L9's at the same step). Selection by validation recall would have chosen wrongly;
only the install battery measures leverage, and the pre-registered rule — primary
metric: install efficacy; validation curves never select — is what caught it. Second,
and larger: the incumbent's efficacy number (0.750) came from a checkpoint trained 3×
longer than the sweep arms (no intermediate checkpoint had been kept) — precisely the
step-mismatched comparison the pre-registration forbade. Re-running L9 at the arms'
budget *eliminated its apparent advantage entirely*: at matched 5k steps, L9 measures
**0.208 — identical to L6** (secondaries lean L9: median rank 7 vs 20; secondaries do
not select). The "3.6× layer effect" the mismatched comparison suggested was in fact a
*training-duration* effect: install leverage at the same layer grows 0.208 → 0.750
from 5k → 15k steps of interface training. The corrected reading: among tolerant
layers, *where* you inject matters far less than *how long the interface trains* —
and a plausible, val-supported, wrong conclusion was two pre-registered rules away
from the paper.

**Speed vs ceiling, adjudicated under pre-registration.** A control arm with learned
operator-derived values (no codebook) was predicted to hit a representational ceiling
well below the codebook's exact decode. The prediction landed in its pre-registered
middle branch: the control *escaped* the predicted ceiling family (hard-read
conditional 0.871; memory-only 0.914) yet stayed short of exact (codebook 1.000), and
it out-integrated the codebook at every in-pipeline mode while acquiring faster —
opening the program's density anneal first (dense-plant training relaxed to realistic
sparsity with recall retained at 0.44 and clean-text ΔCE healed to ≈0: the dense
curriculum is a learning scaffold, not an operating requirement). The codebook's
exactness holds at the record layer; its end-to-end advantage is *not established*,
and the two arms differ by a curriculum confound (only the control reached the sparse
regime) that the next arm is designed to remove. We report the branch structure rather
than a winner because the pre-registration was written before the numbers existed.

The remaining open items are plumbing — addressing sharpness, per-position integration
gating, read abstention — and the full battery at scale. Every rung above, including
the four nulls and the wrong predictions, is logged with dates and checkpoints in the
project work log.

---

# Part III — Governance

*What the receipts buy, which reversion claim is being made, and the doors through which weaker guarantees re-enter.*

## 13. Governance: the audit object

**Where the tax inverts.** Table 1 compares cl33 to an *unconstrained* transformer on
perplexity. In a regulated or high-stakes deployment the relevant comparison is
different — cl33 versus *transformer-plus-its-audit-stack* on guarantee-per-dollar —
and it flips. Post-hoc interpretability (SAEs, probes, attention attribution) is itself
a compute-and-process tax that buys only an *approximation*: §6 Door 3 makes this
concrete, with the transformer's best attention head attributing the true source at
**0.017** (rank 40) and drifting across load, versus the tape's **exact, causal 1.000**
receipt. cl33's audit trail is the computation's *mandatory carrier* — the operator record
through which all output flows (270× on ablation, independently replicated) — and the
reverse trace (R⁻¹ = ηRᵀη, exact to 2e-14) makes that record losslessly replayable, not
a learned proxy of it. The refinement §3.2 requires, stated plainly: the *record* is
causal; the scan's invertibility is a property of the *log*, not of the decision path,
whose transported state is measured decorative at LM inference. For deployments that must *prove*
what a model did — machine unlearning with a receipt (Door 2), decision provenance,
tamper-evident memory — the ~1.29× perplexity tax becomes the **price of a guarantee
that post-hoc tooling cannot provide at any cost**. We scope this precisely: what is
guaranteed is that the computation is *invertible*, the memory *editable with receipts*,
the emission *decodable* — **not** correctness, calibration, or non-hallucination.
Transparency is not correctness; a cl33 model can be confidently wrong, but you can
prove exactly how it got there and edit it.

**What "revert" means — a three-tier hierarchy.** "Install a belief, remove it, and the
model reverts" is three claims of different strengths, and the architecture supports
stating which one is being made:

1. **Computational reversion (strongest, measured).** With the base frozen and the organ
   zero-initialized-and-gated, deleting a register restores the *bit-exact* forward
   computation on the same prompt: organ-off ≡ base to 0.00e+00 ΔCE in every run, and
   the §12 pilot measures delete → 0.000 with the installed fact unreachable. This tier
   is a property of the frozen-substrate design, not of training success.
2. **Behavioral reversion (fresh prompts).** After deletion, novel prompts probing the
   belief return to baseline behavior. Measured in the pilot at the same 0.000; at scale
   this tier requires a battery, not an identity argument, and we scope it as such.
3. **Conversational reversion (weakest — inexact by construction, auditable).** If the
   installed belief influenced *generated text* that re-entered context, deletion does
   not un-say it: the influence has sedimented into the conversation history, which is
   ordinary context the model rightly conditions on. This is the operator-model form of
   the sediment lesson (not-recoverable ≠ not-behaviorally-causal). What the
   architecture offers here is not exact reversion but an *audit trail*: the receipts
   date-stamp exactly which generations occurred under the installed belief, so the
   contaminated span is identifiable even though it is not erasable.

**Where the sediment caveat re-enters (three doors, named so they can be watched).**
The tier-1 guarantee is conditional, and each condition is a door through which
weaker-than-computational reversion re-enters: **(a) co-training** — any arm that
unfreezes the base (§12's v3.1 option) trades bit-exact organ-off for "characterized
degradation," a trade that must be made knowingly and labeled; **(b) persistent banks**
— when memories are written *by processes that themselves read memory*, deleting a
record does not delete its causal descendants; a write-provenance cascade (each record
carrying the receipts of the reads that produced it) is the pre-registered requirement
before any long-lived bank is deployed; **(c) context sedimentation** — tier 3 above,
present in any deployment where outputs re-enter inputs. None of these voids the
architecture's guarantees; each converts one guarantee from an identity into a
measurement, and the receipts are what make that measurement possible.

---

## 14. Related work

**Linear-recurrent and matrix memories.** mLSTM/xLSTM (Beck et al., 2024), DeltaNet-style
fast weights (Schlag et al., 2021; Yang et al., 2024), and state-space models of the
Mamba lineage (Gu & Dao, 2023; Dao & Gu, 2024) form the honest competitive set for
Appendix F (the linear-transparent store): linearity without receipts. The lineage runs
from the original fast-weight programmer (Schmidhuber, 1992) through fast-weight
associative memory (Ba et al., 2016) and the linear-attention recurrence (Katharopoulos
et al., 2020) to their identification (Schlag et al., 2021). Our wedge store is the
Cl(3,3) expression of the mLSTM outer-product memory, and its measured failure modes
(91% bypass with token keys; slow grokking with algebra values) delimit ours rather than
flatter them. The recall–throughput frontier this family trades on is mapped by Arora
et al. (2024); our tape sits off that frontier deliberately — an honest O(T²) for exact,
receipted recall. Memory-based test-time adaptation (Sun et al., 2024) is the nearest
frame for the tape-as-TTT reading (§6): where TTT layers update a hidden model by
gradient, the tape substitutes an exact, undoable memory write. What none of this family
provides — and what the tape and register formulations are for — is causal attribution
receipts and exact, certificate-bearing deletion.

**Induction heads and in-context retrieval.** The two-memory identity of §3
(recall-as-induction = routing-around-the-operators) is an architectural corollary of
the induction-head literature (Olsson et al., 2022), and our recall instrument is the
MQAR task of Arora et al. (2023): we do not dispute that transformers ride
token-similarity shortcuts; we make the shortcut *visible and ablatable* and show what
remains without it.

**Model editing and unlearning.** ROME/MEMIT-style weight editing (Meng et al., 2022;
Meng et al., 2023) modifies parameters with no certificate and contested locality (Hase
et al., 2023); machine-unlearning methods (Bourtoule et al., 2021) approximate removal
with no proof object. Our editable-memory results (§6, Door 2) claim something narrower
and stronger where it applies: exact algebraic erasure *from the explicit memory state*,
with a receipt, accompanied by behavioral reversion — scoped explicitly NOT as
unlearning of the backbone.

**Steering and representation control.** Steering by added directions is established
for standard transformers — representation engineering (Zou et al., 2023), activation
addition (Turner et al., 2023), and persona vectors (Chen et al., 2025). §9.6 and
App. C place cl33 in that family with one structural difference: the steered object's
downstream effect is readable from a mandatory typed interface (the operator stream),
so dose–response, saturation, and regime adoption are measured on the causal path
rather than probed beside it.

**Reversible architectures.** RevNets (Gomez et al., 2017) and reversible transformers
(Kitaev et al., 2020) invert computation for activation-memory efficiency — the
inversion exists to avoid storing. Here reversibility is the *explanation interface*:
R⁻¹ = ηRᵀη is run backward to produce the audit trail, and the 2e-14 exactness is a
property the paper's governance claims rest on.

**Geometric and Clifford networks.** Clifford/geometric-algebra layers (Brandstetter
et al., 2023; Ruhe et al., 2023; Brehmer et al., 2023) are typically used as
equivariance-preserving feature maps inside otherwise-standard readouts. cl33's
distinction is exclusivity: the algebra is not a feature map but the *only* path to
output, which is what converts interpretability questions into ablation experiments.

**Lens methods and verbalizability.** Logit-lens (nostalgebraist, 2020) and tuned-lens
(Belrose et al., 2023) read representations through trained probes, inheriting the
faithfulness question shared by probing classifiers generally (Belinkov, 2022) and
attention attribution in particular (Jain & Wallace, 2019); sparse-autoencoder
dictionaries (Bricken et al., 2023; Cunningham et al., 2024) buy interpretable features
at the same post-hoc remove. The Jacobian-lens instrument of §9 estimates
verbalizability without a trained readout, and the three-model comparison places the
operator stream above both GPT-2 controls' best layers.

**Global Workspace Theory** (Baars, 1988; Dehaene et al., 1998) supplies naming only
(reportable bottleneck vs non-reportable processing); no cognitive claim is made or
needed. **TC⁰ expressivity limits** (Merrill et al., 2022; Merrill & Sabharwal, 2023)
supply the provenance of the permutation-composition task used for the §4 parity study,
with the task's transformer-hardness established by Liu et al. (2023) and its
state-tracking framing by Merrill et al. (2024). **Baselines** throughout are GPT-2
(Radford et al., 2019) and the token-matched Pythia suite (Biderman et al., 2023).

---

## References

- Arora, S., Eyuboglu, S., Timalsina, A., Johnson, I., Poli, M., Zou, J., Rudra, A., & Ré, C. (2023). Zoology: Measuring and improving recall in efficient language models. arXiv:2312.04927.
- Arora, S., Eyuboglu, S., Zhang, M., Timalsina, A., Alberti, S., Zinsley, D., Zou, J., Rudra, A., & Ré, C. (2024). Simple linear attention language models balance the recall-throughput tradeoff. arXiv:2402.18668.
- Ba, J., Hinton, G., Mnih, V., Leibo, J. Z., & Ionescu, C. (2016). Using fast weights to attend to the recent past. *NeurIPS 2016*. arXiv:1610.06258.
- Baars, B. J. (1988). *A Cognitive Theory of Consciousness.* Cambridge University Press.
- Beck, M., Pöppel, K., et al. (2024). xLSTM: Extended Long Short-Term Memory. *NeurIPS 2024*. arXiv:2405.04517.
- Belinkov, Y. (2022). Probing classifiers: Promises, shortcomings, and advances. *Computational Linguistics*, 48(1), 207–219.
- Belrose, N., Ostrovsky, I., McKinney, L., Furman, Z., Smith, L., Halawi, D., Biderman, S., & Steinhardt, J. (2023). Eliciting latent predictions from transformers with the tuned lens. arXiv:2303.08112.
- Biderman, S., Schoelkopf, H., Anthony, Q., et al. (2023). Pythia: A suite for analyzing large language models across training and scaling. *ICML 2023*. arXiv:2304.01373.
- Bourtoule, L., Chandrasekaran, V., Choquette-Choo, C. A., et al. (2021). Machine unlearning. *IEEE S&P 2021*, 141–159. arXiv:1912.03817.
- Brandstetter, J., van den Berg, R., Welling, M., & Gupta, J. K. (2023). Clifford neural layers for PDE modeling. *ICLR 2023*. arXiv:2209.04934.
- Brehmer, J., de Haan, P., Behrends, S., & Cohen, T. (2023). Geometric algebra transformer. *NeurIPS 2023*. arXiv:2305.18415.
- Bricken, T., Templeton, A., Batson, J., et al. (2023). Towards monosemanticity: Decomposing language models with dictionary learning. *Transformer Circuits Thread.* transformer-circuits.pub/2023/monosemantic-features.
- Chen, R., Arditi, A., Sleight, H., Evans, O., & Lindsey, J. (2025). Persona vectors: Monitoring and controlling character traits in language models. arXiv:2507.21509.
- Cunningham, H., Ewart, A., Riggs, L., Huben, R., & Sharkey, L. (2024). Sparse autoencoders find highly interpretable features in language models. *ICLR 2024*. arXiv:2309.08600.
- Dao, T., & Gu, A. (2024). Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. arXiv:2405.21060.
- Dehaene, S., Kerszberg, M., & Changeux, J.-P. (1998). A neuronal model of a global workspace in effortful cognitive tasks. *PNAS*, 95(24), 14529–14534.
- Gomez, A. N., Ren, M., Urtasun, R., & Grosse, R. B. (2017). The reversible residual network: Backpropagation without storing activations. *NeurIPS 2017*. arXiv:1707.04585.
- Gu, A., & Dao, T. (2023). Mamba: Linear-time sequence modeling with selective state spaces. arXiv:2312.00752.
- Hase, P., Bansal, M., Kim, B., & Ghandeharioun, A. (2023). Does localization inform editing? Surprising differences in causality-based localization vs. knowledge editing in language models. *NeurIPS 2023*. arXiv:2301.04213.
- Jain, S., & Wallace, B. C. (2019). Attention is not explanation. *NAACL-HLT 2019*. arXiv:1902.10186.
- Katharopoulos, A., Vyas, A., Pappas, N., & Fleuret, F. (2020). Transformers are RNNs: Fast autoregressive transformers with linear attention. *ICML 2020*. arXiv:2006.16236.
- Kitaev, N., Kaiser, Ł., & Levskaya, A. (2020). Reformer: The efficient transformer. *ICLR 2020*. arXiv:2001.04451.
- Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., & Zhang, C. (2023). Transformers learn shortcuts to automata. *ICLR 2023* (oral). arXiv:2210.10749.
- Meng, K., Bau, D., Andonian, A., & Belinkov, Y. (2022). Locating and editing factual associations in GPT. *NeurIPS 2022*. arXiv:2202.05262.
- Meng, K., Sen Sharma, A., Andonian, A., Belinkov, Y., & Bau, D. (2023). Mass-editing memory in a transformer. *ICLR 2023*. arXiv:2210.07229.
- Merrill, W., Sabharwal, A., & Smith, N. A. (2022). Saturated transformers are constant-depth threshold circuits. *TACL*, 10, 843–856. arXiv:2106.16213.
- Merrill, W., & Sabharwal, A. (2023). The parallelism tradeoff: Limitations of log-precision transformers. *TACL*, 11. arXiv:2207.00729.
- Merrill, W., Petty, J., & Sabharwal, A. (2024). The illusion of state in state-space models. *ICML 2024*. arXiv:2404.08819.
- nostalgebraist (2020). interpreting GPT: the logit lens. *LessWrong.* lesswrong.com/posts/AcKRB8wDpdaN6v6ru.
- Olsson, C., Elhage, N., Nanda, N., et al. (2022). In-context learning and induction heads. *Transformer Circuits Thread.* arXiv:2209.11895.
- Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. *OpenAI technical report.*
- Ruhe, D., Gupta, J. K., de Keninck, S., Welling, M., & Brandstetter, J. (2023). Geometric Clifford algebra networks. *ICML 2023*. arXiv:2302.06594.
- Schlag, I., Irie, K., & Schmidhuber, J. (2021). Linear transformers are secretly fast weight programmers. *ICML 2021*. arXiv:2102.11174.
- Schmidhuber, J. (1992). Learning to control fast-weight memories: An alternative to dynamic recurrent networks. *Neural Computation*, 4(1), 131–139.
- Sun, Y., Li, X., Dalal, K., et al. (2024). Learning to (learn at test time): RNNs with expressive hidden states. arXiv:2407.04620.
- Turner, A. M., Thiergart, L., Udell, D., Leech, G., Mini, U., & MacDiarmid, M. (2023). Activation addition: Steering language models without optimization. arXiv:2308.10248.
- Yang, S., Wang, B., Zhang, Y., Shen, Y., & Kim, Y. (2024). Parallelizing linear transformers with the delta rule over sequence length. *NeurIPS 2024*. arXiv:2406.06484.
- Zou, A., Phan, L., Chen, S., et al. (2023). Representation engineering: A top-down approach to AI transparency. arXiv:2310.01405.

---

## Appendices (planned)

- **A. Full MQAR battery** — all KV loads × modes × η sweep; TTT-free.
- **B. Numerics** — anchored prefix products, T=1024 divergence study, fp32/bf16.
- **C. Steering & the four operations** — per the publication gates: full analogy
  battery with rank distributions (not exemplars), read-vs-write baselines on emitter
  hiddens and matched transformer embeddings, steering statistics (N×M success rate,
  self-PPL coherence, off-target drift), compact/non-compact decomposition.
- **D. Negative results in full** — wedge matrix, gauge throttle, 2b probes, phase-0
  Claim-3 confound and redesign.
- **E. 5B pre-registration** — gates verbatim, with hashes/dates, before unblinding.
- **F. Linear-transparent memory** — the trilemma and its grok-driven break (moved from §11; drafted below).

### Appendix R (drafted): reproducibility manifest

Every table maps to an in-tree script, a checkpoint, and an exact command, and every
result cites its dated work-log entry (all paths relative to the `cl33_oplm` project
root; run logs preserved verbatim).

**Released artifacts.** The two central frozen checkpoints are public at
**huggingface.co/mirrorethic/cl33-oplm** — the prose base of Table 1b (236.5M, step
189307) and the serving/demo model of cl33.t3atlas.dev (236.5M, step 13996) — with the
load-only model definition, the reverse-readout probe (its held-out evaluation card
stored inside the artifact), SHA-256 hashes, and two verified one-command
reproductions: `repro_bottleneck.py` (zeroing the emitted operators multiplies
perplexity **exactly 270×** on a shipped frozen slice of the validation mix — a
different draw of the same distribution measured the paper's 314× — and ≈106×/≈112×
on public WikiText-103, the off-domain form of §1's bottleneck claim) and `repro_reverse_readout.py` (text decoded from the operator
stream alone, probe top-1 0.860 held-out). The bundle's `REPRODUCE.md` states its
scope plainly: it contains what is necessary to independently test the published
claims on frozen artifacts, not the training stack or the §12 program (which the
paper itself labels ongoing). Reproducibility surface ≠ complete source disclosure;
what a reader needs is the artifact and the measurement, and both are now
unconditional.

| result | script(s) | checkpoint / data | record |
|---|---|---|---|
| Table 1 (matched tax) | `cl33_lm_eval.py`, `bench_table.py` (lm-eval 0.4.11, fp32, right-pad, full-set, single session) | scale-C / scale-B / A1 ckpts | `runs_eval/modelclass/table.md`; WORK_LOG 2026-07-13 |
| Table 1b (prose parity) | `cl33_lm_eval.py --ckpt <prose ckpt>` | `scaleB_prose_1024/ckpt.pt` (6.2B tok) | `runs_eval/scaleB_prose6B_lmeval.json`; WORK_LOG 2026-08-03 |
| §4 nav parity | `nav_cl33.py`, `nav_transformer.py`, `nav_data.py` | 5-seed pre-registered pair | prereg record + WORK_LOG (pre-fork) |
| §5/§6/App.A MQAR + doors | `mqar_cl33.py`, `door1_transformer.py`, `door2_edit_probe.py`, tape scripts | probe-scale ckpts | `results_mesh/asus_cl33_mqar/` (mirrored) |
| §7 inversion | inversion battery scripts per `INVERSION_BATTERY.md` | scale-C | `results_mesh/` + WORK_LOG 2026-07-10/11 |
| §9 J-lens | `code_jacobian_lens/` harness | scale-B/C + GPT-2 controls | `results_mesh/asus_jlens/` (13 eval JSONs + lenses) |
| §12 organ ladder | `train_v30.py`, `v30_organ.py`, `episodes_v30.py`, `probe_v30_address.py` | frozen prose base + organ ckpts per run | `v30_*` run dirs (train.log val lines); WORK_LOG 2026-08-03 → 09-08 |
| §12 install pilot + factorial | `install_pilot.py`, cond probes | `v30_c3b_p1/ckpt.pt` (+ L6/L11 arms) | WORK_LOG 2026-09-08 10:20→16:35 |
| §9.6 regime sensor + §3.1 lever probe | inline probes (WORK_LOG-referenced) | `scaleB_prose_1024/ckpt.pt` (frozen, eval-only) | WORK_LOG 2026-09-08 14:35→15:10; CL33_OPLM_INSTRUMENTABLE_LM.md |

Baselines: Pythia-160m rows use public HuggingFace revisions (step-matched); GPT-2 rows
use the public 124M release. The eval adapter's right-padding requirement (§10
methodological note) applies to every cl33 row.

---

*Sources of record: WORK_LOG.md (this dir), REVERSIBLE_TAPE_DESIGN.md,
SELF_STRUCTURE_CLAIMS.md, CL33_OPLM_V24_TAPE_LM.md, CL33_OPLM_ARCHITECTURE.md,
CL33_OPLM_V23_RM.md, nav parity pre-reg memory, J-lens artifacts at
/mnt/consciousness-storage/jlens/.*
### Appendix F (moved from §11): Linear-transparent memory — the trilemma and its break

*(Design + pre-registration: LINEAR_TRANSPARENT_MEMORY.md.)*

The tape's exactness and receipts come at an O(T²) softmax address (it materializes a
distribution over all past records). The field's linear-time memories (mLSTM, DeltaNet,
Mamba) compress records into a fixed-size state — O(T) — but have neither receipts nor
surgical editability. Can one architecture have all three of {linear, transparent,
recall}? We hold both endpoints: the **wedge** (mLSTM expressed in Cl(3,3): O(T), but its
token-content value channel bypasses the operators 91%) and the **tape** (O(T²),
zero-bypass, attribution@1 = 1.0).

**A three-corner trilemma, then broken.** At a probe configuration (KV=8), three variants:
token-value wedge — recall 0.87 but 95% bypass, O(T); tape — recall 1.0, zero bypass,
O(T²); and the untested cell, an **algebra-value wedge** (token key + algebra-only value)
— O(T) and zero-bypass, but recall only 0.26 *at a short training budget*. The apparent
trilemma (any two of the three) dissolved on inspection: the algebra-value wedge was still
climbing. Given the budget to converge, it **grokked to 0.99 recall — linear, zero-bypass,
and accurate, all three** — via a phase-transition jump (0.39→0.94 between 3k and 5k steps,
cross-entropy collapsing), at ~10× the tape's learning time.

**Why it groks (a mechanistic account, contributed by the collaboration).** Quadratic
memory externalizes address formation into an explicit distribution over separately-kept
records — nothing to learn. Linear memory must *learn* an internal geometry that
simultaneously prevents write-interference and supports decoding. The token-value wedge is
fast only because token embeddings hand it a ready-made near-orthogonal basis (that is the
bypass); the algebra-value wedge must build that geometry from the operators, which is a
genuine — and slow — representation-learning problem. The grok jump is the signature of
finding that geometry.

**The honest boundary.** This is an existence proof at toy scale: at a realistic 8k vocab
the algebra-value wedge did not grok within the same budget — the fixed-size compression's
decoding capacity is the binding constraint, and it scales with vocabulary. So linear +
transparent + recall is *achievable but not yet tractable at scale* — a mechanistically
understood open problem, not a wall. The design compass it hands us (delta-rule writes to
reduce interference by construction, orthogonal initialization, decorrelation penalties)
is the next work.

