Two tools against catastrophic forgetting in language models, now on DAS 5.0 — faster and with negative measured damage compared to the v3 protocol below. Neither one stores an example — no replay buffer, no rehearsal, no access to old data at the moment of protecting.
§1–§6 report the fully audited DAS 3.0 protocol on Qwen3 — the baseline this claim is measured against. §7 traces every version since, including why v5.0 is better (~2× faster than v4.0, negative damage), and directional confirmation on DeepSeek, Mistral and Gemma.
Different regimes, same constraint: neither one keeps the old data around.
Knowledge insertion into an already-trained model, with no gradients.
Continued, sequential training across the whole model.
This page reports what was measured. The protocol is described in enough detail to be audited or replicated with another method.
What changed since the first result we published.
Nothing in V1 was wrong — it's the same exact-insertion guarantee still measured in §3 here, at 0.0256 nats of damage. What's new in V3 is scale and honesty about the harder regime: whole-model continued training, where the guarantee is no longer bit-for-bit and has to be tuned, measured, and reported as a trade-off instead of a single clean number.
| DAS V1 | DAS V3 (this page) | |
|---|---|---|
| Regime tested | knowledge insertion, 3 sequential domains | whole-model continued training (§2) and zero-gradient insertion (§3) |
| Model(s) | Qwen2.5-3B-Instruct — one size | Qwen3 0.6B / 1.7B / 8B — three sizes, same vocabulary |
| Headline number | 0.000000 deviation, bit-for-bit | 92.0% forgetting blocked at zero learning cost |
| Retention guarantee | exact, in the regime tested | exact in §3; tunable and approximate in §2 — see §6 |
| Cost measured | 137 ms/step, 5.3× faster than full fine-tune | 0.07s per write; learning cost ≈0 at whole-model scale |
| Scale tested | no — single model size | yes (§5) — protection gets cheaper as the model grows |
| Stated limitations | 3 | 8 |
Controlled across three model sizes. Same vocabulary, same corpus, same evaluation.
Metrics. damage = change in held-out log-ppl — what the model already knew and nobody asked to change. learned = ppl on the freshly trained block. retained block 1 = ppl on the first block after training all eight, the standard backward-transfer metric. profile = the ppl of each of the 8 blocks at the end.
Comparisons are always on Pareto: damage and learning together, never one axis alone. §4 explains why that isn't pedantry.
| Models | Qwen3-0.6B (d=1024), Qwen3-1.7B (d=2048), Qwen3-8B (d=4096) |
| Control | all three are Qwen3ForCausalLM with identical V=151936 — size changes, vocabulary doesn't |
| Corpus | Portuguese Wikipedia |
| Blocks | 8 × 512 positions, sequential and disjoint from each other |
| Held-out slice | 1024 positions, disjoint from every block, never trained |
| Optimizer | Adam |
| Precision | fp32 |
Qwen3-0.6B, all 197 linear layers covered = 100% of 596M parameters, 8 sequential blocks.
of forgetting blocked.
Zero measurable learning cost: what the model was meant to learn reads 1.00 at every protection level, identical to running unprotected.
damage to the held-out slice — the knowledge nobody asked to change
| Intensity | held-out ppl | damage (nats) | blocked | learned block 8 | retained block 1 |
|---|---|---|---|---|---|
| off | 148.63 | +2.4927 | 0.0% | 1.00 | 18.48 |
| low | 15.02 | +0.2005 | 92.0% | 1.00 | 3.16 |
| medium | 13.21 | +0.0725 | 97.1% | 1.00 | 5.54 |
| high | 12.05 | −0.0193 | 100.8% | 1.02 | 9.31 |
base model: held-out 12.29 · block 1 11.65
Per-block ppl at the end of training. One curve is a forgetting ramp; the other is a line.
ppl per block · block index on the x-axis
Aggregating held-out plus the eight blocks into one scalar: 1.60 versus 5.55 for off. The protected model ends up good at everything it saw; the unprotected one only at what it saw last.
There is a stopping point, and it's identifiable. Sweeping six intensity levels, what was taught saturates: from the recommended level on, loosening further buys zero extra learning and only costs general ability. The marginal exchange rate drops from 1.32 at the last useful step to 0.40 at the next, then goes negative — the last two levels get worse on both axes at once. Under-protecting is waste, not economy.
| Intensity | general ability vs. base | taught material vs. base |
|---|---|---|
| high | −2% | −78% |
| medium | +7% | −89% |
| recommended | +22% | −91% |
| very low | +39% | −91% |
| minimum | +52% | −91% |
| near-zero | +71% | −91% |
Qwen3-8B. New knowledge written into the model with no gradient step at all.
10 facts that share a suffix
5 real facts from 2026, in Q/A form
An unseen 101-token paragraph
Generalizing to paraphrase is what separates this from memorizing: the knowledge goes in through one wording and answers under another.
The cost is predictable before writing, and depends on the shape of what's taught: content that concentrates load on the same outputs costs ~6× more (+2.60%) than prose, which spreads it out (+0.44%). Prose is the easy case.
Auditable and reversible. What was taught can be read back out of the model, and the write can be undone — exact in fp32, ~5e-4 in bf16 from rounding. Good for governance; bad for privacy, and that's a real limitation.
Applies to any continual-learning method, not just this one.
↑ the metric keeps improving — past perfect, even
↓ the model keeps getting worse over the same rows
The "high" level shows negative forgetting — a score better than perfect — and it's the worst row in the table: it barely learned, so it had nothing left to forget. learned last block is blind for the opposite reason: the last block is the only one that hasn't suffered any interference yet. It reads ~1.00 always, by construction.
Recommendation for anyone measuring this: publish the per-block profile and the held-out ppl, kept separate. A single scalar that sums held-out with trained blocks rewards memorizing the training itself — in our case the unprotected arm "wins" a naive aggregate, with a profile that is a textbook forgetting ramp.
Learning cost normalized against the unprotected run of the same model, at low intensity.
At the same intensity, the bigger model wins on both axes at once — less damage and less cost. This isn't a trade-off sliding along one curve: the whole curve moves outward as the model grows. That's the opposite of what you'd expect if protection and capacity were competing for the same fixed resource.
Important caveat: this section measures a restricted scenario — one layer. In the whole-model regime (§2), the learning cost disappears entirely. Don't use these numbers to size a full-model training run.
| low intensity | medium | high | |
|---|---|---|---|
| d=1024 | +0.308 | +0.571 | +1.115 |
| d=2048 | +0.203 | +0.414 | +0.820 |
| d=4096 | +0.056 | +0.135 | +0.386 |
bases differ — ppl 12.32 / 8.72 / 5.99 — so absolute nats don't compare across sizes
Later rounds of testing, run outside the audited v3 protocol above, checked the same anchoring approach on other model families. Forgetting blocked stays in the same band every time.
The same anchoring approach was run against Qwen, DeepSeek, Mistral, and Gemma checkpoints, in testing separate from the §1–§6 Qwen3 protocol above. In every case the model retained what it already knew while still learning the new material — no family-specific tuning, no architecture-specific code path.
What's not yet published here: the full protocol for these four runs — block layout, corpus, seed count — lives outside this repository and isn't reproduced with the same level of detail as §5's Qwen3 sweep. The 92% for Qwen here is a different measurement than the 92.0% headline number in §2 — same ballpark, not the same run. Treat this whole section as a directional result confirming the method isn't architecture-specific, not as an audited benchmark on the level of §2–§5.
What changed release to release, and what each round of testing actually measured. Current: DAS 5.0.
| Version | What changed | Result |
|---|---|---|
| v1 – v2 | Early post-step protection and projection attempts | several sub-approaches failed testing — 0% survival under maximum attack |
| DAS 3.0 | Protected subspace + cross-fact coordination | ~98% of forgetting damage removed |
| DAS 3.1 / DAS Weight | Protecting during training vs. restoring the original state afterward | restore-after: 100% recovered · protect-during: failed |
| DAS 3.3 | Tested on a hybrid model (Qwen3.5) | ~9× less damage than baseline, no need to unlock the model head |
| DAS 4.0 | Post-training repair via regression/KL against a teacher model | error 10.7% (CE) → 1.7% (KL) · damage −3% (negative = improvement) · QDAS 4.0 quantized: ~2.84× less VRAM |
| DAS 5.0 | Closed loop with a holdout sentinel monitoring damage in real time | ~2× faster than 4.0, negative damage, a floor bug fixed |
Testing done across these rounds, without the internal mechanism: retention vs. forgetting measured across multiple domains (one comparison had 63% class overlap between domains — a finding that invalidated an apparently-good 60% ceiling as data leakage); a maximum adversarial attack run against post-step protection to see if it survived; DAS run inside a public benchmark with 4 comparison arms — retention scored exactly 0 damage but finished last on the benchmark's other criteria, meaning "doesn't forget" alone wasn't enough to win overall; cost testing (4 sequential capsules at ~130% cost each, versus fused/cached variants that zeroed the extra inference cost); multiple seeds on every key result; and the baseline's learning rate re-tuned before comparing against DAS, so the reported advantage isn't inflated by an unfair baseline.
Internal mechanism details are withheld deliberately — the numbers above are what changed, not how.
None of these is resolved, and all of them bear on how to read the numbers above.
No result carries an error bar. Measured noise floor: 0.022 nats — nothing below that is claimed anywhere on this page.
FIN and Model_Training have never been used together on the same model.
Measured only on Qwen3-0.6B. §2 hasn't been repeated on the bigger models; §5 is a restricted scenario.
The retention loss described in §4 has no fix we have tested.
§6 shows DeepSeek, Mistral and Gemma all land in the same 92–97% band, but with less protocol detail than the Qwen3 runs — not yet audited to the same standard.
No comparison yet against EWC, LwF or buffer rehearsal. Until that exists, §2 says "better than doing nothing", not "better than the state of the art".
Saturation was observed at 20 blocks in a separate design.
In the whole-model regime the guarantee is approximate everywhere and exact nowhere. §3 — insertion — is the one place it is exact.
DAS is integrated with ACE (Adaptive Core Experts) — growing the model's knowledge with a new domain costs a fraction of what it used to.
See ACE →DAS is the continual-learning piece behind MH-AI — the AI architecture we're building for genuine scientific reasoning.
Meet MH-AI →Published with the numbers that don't flatter it, too. That's the bar for calling a result a result.