Case Study

When the Memory Wasn’t the Problem

A user-visible memory failure in my AI companion system stayed misdiagnosed for six weeks. This is how a purpose-built RAG debugger eventually found the real cause — and what happened when I turned the same scrutiny on the debugger itself. Numbers are from a live system; conversation content is paraphrased and anonymized throughout.


The system

A long-term memory engine behind several conversational agents. A vector store holds roughly 4,700 memories for the main persona; a relational store holds hard biographical facts. A semantic extractor decides what is worth keeping from each exchange. Retrieval runs a pipeline: source exclusion, a literal-match channel for acronyms, reranking, a temporal cutoff, diversity selection, and a character budget.

One of the rooms hosts three agents at once. They share a conversation but keep separate memories, which is the point — one of them can know something the others don’t.

The failure

July 13. The user asks the three agents whether they remember what he had told them a few minutes earlier, mentioning that he had just moved from his phone to his desktop.

All three answer confidently. All three are wrong. He corrects them — the real topic had been something else entirely — and they confabulate a second time, just as confidently, now around the corrected topic.

Three minutes later he explains the failure to himself: the memory system is young, it doesn’t record everything yet, this is expected for now.

That explanation was wrong. It held for six weeks.

The wrong fix

Four days later a fix shipped: three anti-confabulation rules added to the agents’ prompts. Don’t claim to remember what you don’t have. Say you don’t know.

That is a prompt fix for an architectural failure. But there was no way to know that at the time, because there was no way to see what the model had actually received.

The deeper problem was diagnostic, not technical: the failure was shaped exactly like the failure we were expecting. We were actively working on memory quality, so a memory-shaped symptom went into the memory drawer. Nothing about it looked like a routing bug, so nobody looked for one.

The debugger

The tool that eventually resolved this is a read-only inspector over the whole memory path. Two views:

One design decision mattered more than the rest: the debugger does not simulate production, it calls the same context-assembly function the live endpoint calls. There is no second implementation to drift out of sync. Where that principle was violated, it cost me — twice, as it turns out.

Auditing the debugger

By late August the debugger had become load-bearing. A relevance harness ran on top of it, and parameter decisions ran on top of that harness. It had never itself been tested.

So I audited it, with four invariants over the read trace:

  1. No entry may appear in the final stage without existing in some earlier stage.
  2. The budget-trimmed stage must be a subset of the stage before it.
  3. Every entry in the last stage must appear in the assembled prompt.
  4. The number of memory lines in the prompt must equal the last stage’s count exactly.

Invariants 2, 3 and 4 passed on every probe. Invariant 1 failed on every probe.

Two entries in the final stage existed in no earlier stage. They were behavioural rules — a separate retrieval channel that had never been given a trace stage at all. Roughly a quarter of the prompt had no provenance the debugger could show. A second channel for imported reference material had the same gap.

This also explained a complaint I had failed to reproduce. The user had reported that a behavioural rule kept “coming back into memory”, and I could not trace where it entered. It wasn’t traceable. The trace never contained it. The debugger had not been lying — it had been silent, which is harder to notice.

Two more findings from the same pass:

All three are now fixed: the missing channels have stages, the growing stage is split into before and after the injection, and both history builders call one shared function. The invariant suite runs as a script and exits non-zero on violation.

When the measurement lies

The audit forced a second, more uncomfortable review: of the measurements themselves. I found eight instances across one month, in two families. (The catalogue has since grown to ten.)

Measuring the wrong quantity. The regression harness counted how many memories reached the prompt. The question we actually cared about was whether they were the right memories. So the harness reported “26/26, no change” when a new retrieval channel had taken a specific query from zero relevant results to three, and it reported success when a diversity-parameter change increased volume by 37% while the target query still returned nothing on topic.

The instrument silently measuring nothing. An offline comparison returned identical numbers for four different configurations. The conclusion “this parameter has no effect” was one keystroke from being written down. The real cause was a missing environment variable: a user-scoping filter silently excluded the entire semantic channel, so only an unfiltered fallback channel was being measured. A second script pointed its store path at a directory that didn’t exist, quietly created an empty database, and reported zero results for everything.

The rule I now apply before believing any measurement: prove the instrument measures anything at all before trusting what it says. Every harness prints its record count on startup, and identical numbers across configurations are treated as instrument failure until proven otherwise.

Once the harness measured relevance instead of volume, a parameter that had been raised a week earlier turned out to buy nothing:

Memory slotsRecallOn-topic shareEntries returned
570%34.8%46
680%38.8%49
8 (shipped)80%38.5%52
1280%37.0%54

The entire gain happened at six. Going to eight added three more entries and zero additional answers.

Widening the diversity selector was worse than useless — it actively destroyed relevance, which is obvious in hindsight and was not obvious at all beforehand:

Diversity slotsRecallOn-topic share
370%43.8%
5 (current)80%38.8%
860%27.1%
1250%28.8%

The selector carries a similarity penalty. The more slots it is given, the harder it works to find something different — so it pushes out the on-topic entries that answer the question and pulls in unrelated ones. At twelve slots, half the questions stop being answered.

The most useful number from this pass was the one that didn’t move. On-topic share sits near 38% regardless of how the retrieval parameters are tuned, and every off-topic entry traces back to the extractor rather than to retrieval. Retrieval was faithfully serving what it had been given. That reframed the roadmap: the ceiling is on the write path, not the read path.

The actual root cause

The July failure was resolved six weeks later, by an unrelated report: history wasn’t loading in one browser.

The server was fine. The conversation the user had asked about was on disk the whole time, and still is. The identity of the conversation lived in browser storage, and it governed both directions:

The room had one continuous thread of 735 messages. Next to it sat eleven other conversation identifiers, five of them real conversations holding 201, 174, 114, 41 and 30 messages: 560 messages of history cut off from view. The interface has never had a “new conversation” button. Every fork was an accident.

Which means the agents in July were not failing to recall. They had been handed an empty context and behaved exactly as any language model behaves with an empty context. We had spent four days hardening prompts against confabulation, when the model had nothing to confabulate from.

The fix moves conversation identity to the server in both directions: the client’s stored identifier is a hint, never an authority, and a request without one joins the current thread instead of creating a new one. The orphaned conversations were left untouched — they are intact data, and merging them is a separate decision.

A prediction that failed

Three days before, I had fixed a different bug in the same room. Conversation history was being sent to the model with speaker labels stripped and consecutive turns from different agents merged into a single block. Since every one of those turns is tagged as the assistant’s own, each agent read the others’ lines as its own past speech — and would occasionally claim a sentence another agent had said.

I predicted that this would also leak verbal mannerisms: one agent has a signature grunt, and I expected to find it spreading. It was a falsifiable claim, so I measured it across 711 agent turns from before the fix.

AgentTurns containing the ticShare
A (owner of the tic)172 / 24570.2%
B0 / 2000.0%
C2 / 2660.8%

The prediction was wrong. The personas held their voices perfectly through a month of ambiguous history. The bug corrupted attribution of events without touching style — the persona instruction turned out to be stronger than the conversational evidence. Those are two independent layers of character identity, and I would not have guessed they were separable.

What’s next

Takeaways

  1. A failure that matches your current worry gets filed under your current worry. Six weeks of wrong diagnosis cost nothing in code and everything in direction. The correction did not come from thinking harder about memory; it came from an unrelated bug report that made the routing visible.
  2. An instrument that is silent is more dangerous than one that is wrong. The debugger never made a false claim. It simply had nothing to say about a quarter of the prompt, and that gap looked exactly like completeness.
  3. Validate the instrument before the system. Everything downstream of a debugger inherits its blind spots. The invariants that caught this are four assertions and a hundred lines — cheaper than any of the decisions they now protect.
  4. Prefer measurements that can embarrass you. The most valuable results here were a parameter change that bought nothing, a tuning direction that made things worse, and a prediction that data refused to support.

Numbers are from a single production system over one measurement window. Conversation content is paraphrased; agent names, stored memories and personal details are omitted. The system is a personal project, not a commercial product.