The problem
A language model's context window is finite. Every approach to giving it memory beyond that window pays one of two costs: either the model has to guess its way back to what it once knew — regenerating retrieved text token by token, which means it can misstate a true fact even when the retrieval itself was correct — or the system quietly accepts that some information will drift, degrade, or vanish as conversations grow long.
Most memory layers built for language models today choose the first cost, because it's the one that's easy to build: retrieve some text, put it back in the prompt, let the model talk. It works, and it's why retrieval-augmented approaches are the default. It also means the failure mode is silent — a subtly wrong answer looks exactly like a right one until someone checks.
What MH-X targets
Can a model recall something it was told earlier without ever regenerating it — reading the fact rather than reconstructing it from a guess — while still being addressable the way a person would ask for it, in ordinary language, tolerant of rephrasing? Three properties are being pursued together, deliberately, because any one alone is easy.
Nothing is served back to a conversation unless it has just been checked, on the spot, against what was actually stored. When verification fails, the system says so instead of answering.
The same memory mechanism has been exercised on two independently-built inference stacks and two very different model sizes, without depending on any single vendor's API or serving layer.
Long-term storage cost stays close to the size of the raw text itself; the more expensive representation is only built for the pieces that actually get asked for again.
Measured, not assumed.
Previously-stored content reconstructed exactly, confirmed by restarting the process that holds it — validated independently on two unrelated inference engines and on models roughly 24x apart in parameter count.
Overhead for held content measured at a small, near-constant multiple of the original text size — the basis for keeping a very large accumulated history without a proportional cost in specialized storage.
A dedicated step rejects a mismatched answer outright, measured to catch retrieval errors before they reach a conversation, rather than after.
A gap identified by direct comparison against an existing internal system was closed the same day: memory is now scoped to be forgettable on request, separate from what's meant to persist.
In keeping with how every page on this site is written: a result without its limits isn't a result.
- Recall inside a long, already-growing conversation is measurably less reliable today than recall at the start of a fresh one — an identified regression, not a hidden one, and the current open problem on this line.
- Everything above comes from a single, intensive day of disciplined measurement on small-to-mid-size open models. It has not been run at production scale, nor against the retrieval-augmented systems already in wide use.
- This is a research log, not a shipped capability. Nothing here is a product claim.
MH-X sits alongside DAS and ACE as continual-learning and memory infrastructure feeding into MH-AI — measured the same way, held to the same bar.
See MH-AI →Never present as proven what hasn't been measured — including when the measurement is a limit, not a win.