RESEARCHMOUHN · MH-X — MEMORY INFRASTRUCTURE

Memory a model can trust without re-deriving it.

Most ways of giving a language model memory beyond its context window pay one of two costs: the model regenerates what it once knew and can misstate it even when the retrieval was correct, or the system quietly accepts that information will drift as conversations grow. MH-X is a research line asking whether both costs can be avoided at once.

The problem

A language model's context window is finite. Every approach to giving it memory beyond that window pays one of two costs: either the model has to guess its way back to what it once knew — regenerating retrieved text token by token, which means it can misstate a true fact even when the retrieval itself was correct — or the system quietly accepts that some information will drift, degrade, or vanish as conversations grow long.

Most memory layers built for language models today choose the first cost, because it's the one that's easy to build: retrieve some text, put it back in the prompt, let the model talk. It works, and it's why retrieval-augmented approaches are the default. It also means the failure mode is silent — a subtly wrong answer looks exactly like a right one until someone checks.

What MH-X targets

Can a model recall something it was told earlier without ever regenerating it — reading the fact rather than reconstructing it from a guess — while still being addressable the way a person would ask for it, in ordinary language, tolerant of rephrasing? Three properties are being pursued together, deliberately, because any one alone is easy.

Verified, not assumed

Nothing is served back to a conversation unless it has just been checked, on the spot, against what was actually stored. When verification fails, the system says so instead of answering.

Portable across backends

The same memory mechanism has been exercised on two independently-built inference stacks and two very different model sizes, without depending on any single vendor's API or serving layer.

Cheap to keep

Long-term storage cost stays close to the size of the raw text itself; the more expensive representation is only built for the pieces that actually get asked for again.

WHAT HAS BEEN MEASURED SO FAR

Measured, not assumed.

Exact reconstruction

Previously-stored content reconstructed exactly, confirmed by restarting the process that holds it — validated independently on two unrelated inference engines and on models roughly 24x apart in parameter count.

Storage cost

Overhead for held content measured at a small, near-constant multiple of the original text size — the basis for keeping a very large accumulated history without a proportional cost in specialized storage.

Verification catches errors early

A dedicated step rejects a mismatched answer outright, measured to catch retrieval errors before they reach a conversation, rather than after.

Privacy gap closed same-day

A gap identified by direct comparison against an existing internal system was closed the same day: memory is now scoped to be forgettable on request, separate from what's meant to persist.

What is not yet true

In keeping with how every page on this site is written: a result without its limits isn't a result.

  • Recall inside a long, already-growing conversation is measurably less reliable today than recall at the start of a fresh one — an identified regression, not a hidden one, and the current open problem on this line.
  • Everything above comes from a single, intensive day of disciplined measurement on small-to-mid-size open models. It has not been run at production scale, nor against the retrieval-augmented systems already in wide use.
  • This is a research log, not a shipped capability. Nothing here is a product claim.

Never present as proven what hasn't been measured — including when the measurement is a limit, not a win.