CAPABILITY · current: v5.0MOUHN · DAS — DURABLE ANCHOR SYSTEM

Zero forgetting, without storing a single example.

Two tools against catastrophic forgetting in language models, now on DAS 5.0 — faster and with negative measured damage compared to the v3 protocol below. Neither one stores an example — no replay buffer, no rehearsal, no access to old data at the moment of protecting.

§1–§6 report the fully audited DAS 3.0 protocol on Qwen3 — the baseline this claim is measured against. §7 traces every version since, including why v5.0 is better (~2× faster than v4.0, negative damage), and directional confirmation on DeepSeek, Mistral and Gemma.

92.0%of forgetting blocked, at zero learning cost
0old examples stored — no replay buffer, no rehearsal
0.07sper knowledge write, with no gradient step

Two tools.

Different regimes, same constraint: neither one keeps the old data around.

DAS_V3_FIN

Knowledge insertion into an already-trained model, with no gradients.

Old data storednone
Cost0.07s / write
Tuningone strength parameter
DAS_V3_MODEL_TRAINING

Continued, sequential training across the whole model.

Old data storednone
Cost~2× the training step
Tuningone intensity parameter

This page reports what was measured. The protocol is described in enough detail to be audited or replicated with another method.

V1 → V3

What changed since the first result we published.

V1
One model
Qwen2.5-3B
→
V3
Three sizes
Qwen3 0.6B · 1.7B · 8B
Regime
Insertion only
→
Regime
Insertion and whole-model training
Stated limitations
3
→
Stated limitations
8, plus a section on a metric that lies

Nothing in V1 was wrong — it's the same exact-insertion guarantee still measured in §3 here, at 0.0256 nats of damage. What's new in V3 is scale and honesty about the harder regime: whole-model continued training, where the guarantee is no longer bit-for-bit and has to be tuned, measured, and reported as a trade-off instead of a single clean number.

View the complete V1 ↔ V3 comparison
DAS V1DAS V3 (this page)
Regime testedknowledge insertion, 3 sequential domainswhole-model continued training (§2) and zero-gradient insertion (§3)
Model(s)Qwen2.5-3B-Instruct — one sizeQwen3 0.6B / 1.7B / 8B — three sizes, same vocabulary
Headline number0.000000 deviation, bit-for-bit92.0% forgetting blocked at zero learning cost
Retention guaranteeexact, in the regime testedexact in §3; tunable and approximate in §2 — see §6
Cost measured137 ms/step, 5.3× faster than full fine-tune0.07s per write; learning cost ≈0 at whole-model scale
Scale testedno — single model sizeyes (§5) — protection gets cheaper as the model grows
Stated limitations38
01

Evaluation protocol

Controlled across three model sizes. Same vocabulary, same corpus, same evaluation.

Metrics. damage = change in held-out log-ppl — what the model already knew and nobody asked to change. learned = ppl on the freshly trained block. retained block 1 = ppl on the first block after training all eight, the standard backward-transfer metric. profile = the ppl of each of the 8 blocks at the end.

Comparisons are always on Pareto: damage and learning together, never one axis alone. §4 explains why that isn't pedantry.

Qwen3 — identical vocabulary, V=151936
0.6B1.7B8B
↓
Portuguese Wikipedia
↓
B1
B2
B3
B4
B5
B6
B7
B8
held-out
8 × 512 · sequential, disjoint1024 · never trained
View the exact protocol
ModelsQwen3-0.6B (d=1024), Qwen3-1.7B (d=2048), Qwen3-8B (d=4096)
Controlall three are Qwen3ForCausalLM with identical V=151936 — size changes, vocabulary doesn't
CorpusPortuguese Wikipedia
Blocks8 × 512 positions, sequential and disjoint from each other
Held-out slice1024 positions, disjoint from every block, never trained
OptimizerAdam
Precisionfp32
02

Continued training — the whole model

Qwen3-0.6B, all 197 linear layers covered = 100% of 596M parameters, 8 sequential blocks.

92%

of forgetting blocked.

Zero measurable learning cost: what the model was meant to learn reads 1.00 at every protection level, identical to running unprotected.

off
+2.4927 nats
DAS — low
+0.2005 nats

damage to the held-out slice — the knowledge nobody asked to change

View the full intensity sweep
Intensityheld-out ppldamage (nats)blockedlearned block 8retained block 1
off148.63+2.49270.0%1.0018.48
low15.02+0.200592.0%1.003.16
medium13.21+0.072597.1%1.005.54
high12.05−0.0193100.8%1.029.31

base model: held-out 12.29 · block 1 11.65

The profile goes flat.

Per-block ppl at the end of training. One curve is a forgetting ramp; the other is a line.

off — classic forgetting rampDAS, low intensity — flat
0 5 10 15 20 1 2 3 4 5 6 7 8

ppl per block · block index on the x-axis

Aggregating held-out plus the eight blocks into one scalar: 1.60 versus 5.55 for off. The protected model ends up good at everything it saw; the unprotected one only at what it saw last.

There is a stopping point, and it's identifiable. Sweeping six intensity levels, what was taught saturates: from the recommended level on, loosening further buys zero extra learning and only costs general ability. The marginal exchange rate drops from 1.32 at the last useful step to 0.40 at the next, then goes negative — the last two levels get worse on both axes at once. Under-protecting is waste, not economy.

View the six-level intensity sweep
Intensitygeneral ability vs. basetaught material vs. base
high−2%−78%
medium+7%−89%
recommended+22%−91%
very low+39%−91%
minimum+52%−91%
near-zero+71%−91%
03

Post-training insertion

Qwen3-8B. New knowledge written into the model with no gradient step at all.

01

10 facts that share a suffix

learned
+0.0256 nats damage
0.07s
02

5 real facts from 2026, in Q/A form

5 / 5
answered under paraphrase
+0.54% ppl
03

An unseen 101-token paragraph

101 / 101
exact, greedy generation
+0.44% ppl

Generalizing to paraphrase is what separates this from memorizing: the knowledge goes in through one wording and answers under another.

The cost is predictable before writing, and depends on the shape of what's taught: content that concentrates load on the same outputs costs ~6× more (+2.60%) than prose, which spreads it out (+0.44%). Prose is the easy case.

Auditable and reversible. What was taught can be read back out of the model, and the write can be undone — exact in fp32, ~5e-4 in bf16 from rounding. Good for governance; bad for privacy, and that's a real limitation.

04

A metric that lies.

Applies to any continual-learning method, not just this one.

blocked %
92.0→97.1→100.8

↑ the metric keeps improving — past perfect, even

retained block 1
3.16→5.54→9.31

↓ the model keeps getting worse over the same rows

The "high" level shows negative forgetting — a score better than perfect — and it's the worst row in the table: it barely learned, so it had nothing left to forget. learned last block is blind for the opposite reason: the last block is the only one that hasn't suffered any interference yet. It reads ~1.00 always, by construction.

Recommendation for anyone measuring this: publish the per-block profile and the held-out ppl, kept separate. A single scalar that sums held-out with trained blocks rewards memorizing the training itself — in our case the unprotected arm "wins" a naive aggregate, with a profile that is a textbook forgetting ramp.

05

Protection gets cheaper on bigger models

Learning cost normalized against the unprotected run of the same model, at low intensity.

Qwen3-0.6B d=1024
+0.308 nats
Qwen3-1.7B d=2048
+0.203 nats
Qwen3-8B d=4096
+0.056 nats

At the same intensity, the bigger model wins on both axes at once — less damage and less cost. This isn't a trade-off sliding along one curve: the whole curve moves outward as the model grows. That's the opposite of what you'd expect if protection and capacity were competing for the same fixed resource.

Important caveat: this section measures a restricted scenario — one layer. In the whole-model regime (§2), the learning cost disappears entirely. Don't use these numbers to size a full-model training run.

View all three intensities
low intensitymediumhigh
d=1024+0.308+0.571+1.115
d=2048+0.203+0.414+0.820
d=4096+0.056+0.135+0.386

bases differ — ppl 12.32 / 8.72 / 5.99 — so absolute nats don't compare across sizes

06

Not a Qwen trick.

Later rounds of testing, run outside the audited v3 protocol above, checked the same anchoring approach on other model families. Forgetting blocked stays in the same band every time.

Qwen separate round, not §2
92% blocked
DeepSeek
92% blocked
Mistral
95% blocked
Gemma
97% blocked

The same anchoring approach was run against Qwen, DeepSeek, Mistral, and Gemma checkpoints, in testing separate from the §1–§6 Qwen3 protocol above. In every case the model retained what it already knew while still learning the new material — no family-specific tuning, no architecture-specific code path.

What's not yet published here: the full protocol for these four runs — block layout, corpus, seed count — lives outside this repository and isn't reproduced with the same level of detail as §5's Qwen3 sweep. The 92% for Qwen here is a different measurement than the 92.0% headline number in §2 — same ballpark, not the same run. Treat this whole section as a directional result confirming the method isn't architecture-specific, not as an audited benchmark on the level of §2–§5.

07

Version history.

What changed release to release, and what each round of testing actually measured. Current: DAS 5.0.

View the version timeline
VersionWhat changedResult
v1 – v2Early post-step protection and projection attemptsseveral sub-approaches failed testing — 0% survival under maximum attack
DAS 3.0Protected subspace + cross-fact coordination~98% of forgetting damage removed
DAS 3.1 / DAS WeightProtecting during training vs. restoring the original state afterwardrestore-after: 100% recovered · protect-during: failed
DAS 3.3Tested on a hybrid model (Qwen3.5)~9× less damage than baseline, no need to unlock the model head
DAS 4.0Post-training repair via regression/KL against a teacher modelerror 10.7% (CE) → 1.7% (KL) · damage −3% (negative = improvement) · QDAS 4.0 quantized: ~2.84× less VRAM
DAS 5.0Closed loop with a holdout sentinel monitoring damage in real time~2× faster than 4.0, negative damage, a floor bug fixed

Testing done across these rounds, without the internal mechanism: retention vs. forgetting measured across multiple domains (one comparison had 63% class overlap between domains — a finding that invalidated an apparently-good 60% ceiling as data leakage); a maximum adversarial attack run against post-step protection to see if it survived; DAS run inside a public benchmark with 4 comparison arms — retention scored exactly 0 damage but finished last on the benchmark's other criteria, meaning "doesn't forget" alone wasn't enough to win overall; cost testing (4 sequential capsules at ~130% cost each, versus fused/cached variants that zeroed the extra inference cost); multiple seeds on every key result; and the baseline's learning rate re-tuned before comparing against DAS, so the reported advantage isn't inflated by an unfair baseline.

Internal mechanism details are withheld deliberately — the numbers above are what changed, not how.

08

Limitations

None of these is resolved, and all of them bear on how to read the numbers above.

01

One seed

No result carries an error bar. Measured noise floor: 0.022 nats — nothing below that is claimed anywhere on this page.

05

Never combined

FIN and Model_Training have never been used together on the same model.

02

Whole-model coverage

Measured only on Qwen3-0.6B. §2 hasn't been repeated on the bigger models; §5 is a restricted scenario.

06

No tested fix

The retention loss described in §4 has no fix we have tested.

03

Cross-model result is directional

§6 shows DeepSeek, Mistral and Gemma all land in the same 92–97% band, but with less protocol detail than the Qwen3 runs — not yet audited to the same standard.

07

Baseline only

No comparison yet against EWC, LwF or buffer rehearsal. Until that exists, §2 says "better than doing nothing", not "better than the state of the art".

04

Eight-block horizon

Saturation was observed at 20 blocks in a separate design.

08

Approximate

In the whole-model regime the guarantee is approximate everywhere and exact nowhere. §3 — insertion — is the one place it is exact.

DAS is integrated with ACE (Adaptive Core Experts) — growing the model's knowledge with a new domain costs a fraction of what it used to.

See ACE →

DAS is the continual-learning piece behind MH-AI — the AI architecture we're building for genuine scientific reasoning.

Meet MH-AI →

Published with the numbers that don't flatter it, too. That's the bar for calling a result a result.