Home

From Compression to Reasoning: Learning Hindsight After Mastra

I have been exploring memory management in agents, and along the way I came across reflect in Hindsight. It is a call the agent makes when a user asks a question: it reads the memory bank and returns prose for that user.

The name immediately reminded me of Mastra’s observational memory, where the Reflector is a background agent that rewrites the observation log so the log stops growing.

In Hindsight it answers the user. In Mastra it compresses history. I had assumed they were the same operation under two names, and reading the code showed they are not.

I wrote about Mastra’s memory before, in this post. Coming back to it with Hindsight open in another tab is what made the mismatch worth writing down.

Compression is Mastra’s Reflector’s whole job

Observational memory exists because context windows fill up, and models perform worse when they are full. Mastra compresses old messages into observations when history crosses 30k tokens, then compresses the observations into reflections when they cross 40k.

Reflections do not stack. Each one rewrites the entire log. The output becomes the new log, and Mastra appends new observations after it, so the log stays close to the reflection threshold no matter how long you run.

messages:      0 ──────────────────────────► 30k ──┐
                                                  │
observations:   └────────────► 40k ──► reflect ────┘
                             (rewrites the whole log)

The Reflector call:

packages/memory/src/processors/observational-memory/reflector-runner.ts L301–310
const agent = new Agent({
  id: "observational-memory-reflector",
  name: "Reflector",
  instructions: buildReflectorSystemPrompt(this.reflectionConfig.instruction, extractors),
  model: agentModel,
  maxRetries: 0,
  ...(memory ? { memory } : {}),
});

The last line of its prompt:

packages/memory/src/processors/observational-memory/reflector-agent.ts L74
IMPORTANT: your reflections are THE ENTIRETY of the assistants memory. Any
information you do not add to your reflections will be immediately forgotten.
Make sure you do not leave out anything.

So the Reflector overwrites the log. The new reflection becomes the memory, and anything it leaves out is gone.

Compression lands around 5x to 40x on tool-heavy workloads. The read path stays simple: recent messages plus the whole log, every turn, with nothing to select.

Hindsight puts four types in one bank

Hindsight is an agent memory system from Vectorize. The pitch is that agents should learn over time instead of replaying conversation history. It keeps a persistent bank per agent: you send it text, an LLM extracts facts and files them, and later you ask for the relevant ones back.

The bank holds four kinds of memory. Where Mastra keeps one observation log, Hindsight splits it:

  • World: objective facts about the outside world
  • Experience: the agent’s own experiences, written in first person
  • Opinion: judgements with a confidence score and a timestamp
  • Observation: summaries synthesized from facts underneath

Three operations move data in and out: retain, recall, reflect. The paper covers all three.

Storage is Postgres with pgvector, MIT licensed, one Docker container. If you want memory running with the least setup, there is a two-line LiteLLM wrapper that recalls before the call and retains after it.

Hindsight’s retain does not reframe the user

Mastra’s Observer turns a user assertion into an observed fact. User says “I have two kids”, the log stores “User stated has two kids”, and the prompt has to add “USER ASSERTIONS ARE AUTHORITATIVE” so the model does not start trusting its own summary over the human.

Hindsight’s extraction prompt classifies instead:

hindsight-api-slim/hindsight_api/engine/retain/fact_extraction.py L1091–1092
fact_type:
- "world": Objective/external facts, including the user's preferences, rules,
  corrections, constraints, plans, traits, or context. These stay "world" even
  when the user states them during an assistant interaction.
- "assistant": Actions, experiences, or observations the assistant/agent actually
  performed. Use this for the assistant/agent doing, trying, learning, deciding,
  recommending, or responding — not merely for user facts mentioned in
  conversation.

A user’s preference gets filed as a world fact. The agent’s own actions and lessons go somewhere else, in the experience network.

It also normalizes time, which decides whether temporal queries work later:

hindsight-api-slim/hindsight_api/engine/retain/fact_extraction.py L1099–1104
- CRITICAL: Convert ALL relative temporal expressions to absolute dates in the
  fact text itself. "yesterday" → write the resolved date (e.g. "on November 12,
  2024"), NOT the word "yesterday"
- Coarse dates (only a year, or only a month, is stated): span the WHOLE period —
  "in 2015" → 2015-01-01 to 2015-12-31

Every fact carries occurred_start, occurred_end, and mentioned_at, which track when the thing happened and when the agent heard about it. That looks fussy until you ask what they decided in March and the only timestamp you kept was ingestion time.

Hindsight’s consolidation is Mastra’s opposite

Hindsight consolidates facts into observations in the background. One canonical observation per facet, accumulating evidence instead of replacing it, with a proof_count and the source facts kept around.

The processing rules are a numbered list. Rules 7 and 8 carry the argument:

hindsight-api-slim/hindsight_api/engine/consolidation/prompts.py L51–53
7. PRESERVE HISTORY: observations that record significant events (sold, died,
   moved, changed) are important history — never DELETE them. Only delete an
   observation when it is restated identically or truly meaningless. Be very
   conservative with deletes.

8. NO COMPUTATION: you do not have the full picture — never calculate, derive,
   or adjust numeric values. If the user says "I have 2 dogs" and then "I have a
   dog named Rex", do NOT update the count to 3 — you don't know if Rex is one
   of the 2 or a new one.

Compare that to “any information you do not add will be immediately forgotten.”

Both systems consolidate, with opposite instructions. Mastra’s consolidator compresses hard because a token budget forces it. Hindsight’s gets told to stay conservative with deletes and never do arithmetic, because a memory system that invents a count is worse than one that forgot something.

Rule 1 goes after the same problem: PREFER UPDATE OVER CREATE. Both systems are trying to stop the log from turning into fifty restatements of one thing.

Hindsight’s recall thinks in tokens, not top-k

Observational memory has no recall in its default config. The whole log sits in context, which is the point.

Hindsight does:

results = client.recall(
    bank_id="my-bank",
    query="What do I know about Alice?",
    max_tokens=4096,
)

You give it a token budget, it returns as much relevant memory as fits. Underneath, four arms run in parallel:

hindsight-api-slim/hindsight_api/engine/search/retrieval.py L181–209
from ..memories import get_memories

unified = await get_memories().recall_unified(
    conn=pool,
    bank_id=bank_id,
    fact_types=fact_types,
    query_embedding=query_embedding_str,
    query_text=query_text,
    limit=thinking_budget,
    temporal_window=temporal_constraint,
    min_semantic=min_semantic,
    min_keyword=min_keyword,
    enable_text_search=enable_text_search,
    enable_graph=enable_graph_retrieval,
)

Dense vectors, BM25, graph traversal over entity/temporal/causal links, and interval filtering. Reciprocal rank fusion merges them, a cross-encoder reranks, then it trims to budget.

One correction to my earlier post: Mastra added a retrieval mode since then. retrieval: true gives the agent a recall tool for paging back to raw messages, and { vector: true } adds semantic search. The target still differs. Mastra’s recall recovers source text that was already summarized away. Hindsight’s queries a graph of facts that were never summarized.

Hindsight’s reflect is a tool loop

Hindsight’s reflect is an agentic loop. The agent decides what to look up until it has enough:

hindsight-api-slim/hindsight_api/engine/reflect/tools.py L1–8
Implements hierarchical retrieval:
1. search_mental_models - User-curated summaries (highest quality)
2. search_observations - Consolidated knowledge with freshness
3. recall - Raw facts as ground truth

The order runs from the most distilled thing to the rawest. The prompt is explicit about the fallback:

hindsight-api-slim/hindsight_api/engine/reflect/prompts.py L336–356
- MANDATORY: If search_mental_models and search_observations both return 0
  results, you MUST call recall() before giving up
- This is the source of truth that other levels are built from

Mastra’s Reflector has to write the memory in one pass. Hindsight’s checks the raw facts before committing to an answer.

Banks also carry three disposition traits, skepticism, literalism, and empathy, each scored 1 to 5, that change how reflect reasons rather than what it retrieves. The Hindsight docs explain the presets.

Side by side

Mastra: ReflectorHindsight: reflect
TriggerObservations cross a token thresholdA user asks a question
RunsBackground, off the critical pathForeground, on the request path
InputThe observation logFour networks, retrieved per query
OutputA smaller observation logProse for the user, plus opinion state
Optimizes forBounded contextAnswer quality
Told toOverwrite. Omit and it’s forgottenVerify against raw facts first
Side effectsNoneOpinions gain confidence-weighted judgements

Neither approach is wrong. Mastra forgets things you wanted to keep. Hindsight accumulates things you did not.

What’s new here

Facts and opinions stay apart. You can ask what the agent knows separately from what it thinks, and opinions carry confidence scores. Mastra’s log cannot tell you whether “Alice is on the West Coast” came from Alice or from a guess, because it is one flat text field.

The experience network gives the agent its own memories. They are written in first person, so the agent can recall what it tried and what failed, not just facts about the user.

You can trace an answer back. Observations keep the source facts with exact quotes, and when facts conflict, both versions survive with timestamps. The Reflector would have dropped one of them.

What to keep, and when you decide that

The LLM call count bothered me enough that I went looking at their own examples. The chat memory app in the cookbook is 55 lines and does the obvious thing:

packages/memory/src/index.ts L2264–2275
export async function storeConversation(
  userId: string,
  userMessage: string,
  assistantMessage: string,
) {
  const conversation = `User: ${userMessage}\nAssistant: ${assistantMessage}`;

  await hindsightClient.retain(userId, conversation, {
    context: "conversation",
    metadata: { timestamp: new Date().toISOString() },
  });
}

It retains every turn with no filter, one LLM extraction call per message. A user with a few hundred conversations has paid for a few hundred extractions, most of which are “user said hello”, and the bank holds all of them.

The escape hatch is a retain mission: plain language you write that goes into the extraction prompt:

Always include technical decisions, API design choices, and architectural trade-offs. Ignore meeting logistics, greetings, and social exchanges.

You set it on the bank or through HINDSIGHT_API_RETAIN_MISSION. The env var lands on config.retain_mission, and one small function turns it into a preamble:

hindsight-api-slim/hindsight_api/engine/retain/fact_extraction.py L1770–1787

The code keeps it out of the cached system prompt on purpose, since that prompt has to stay identical for every bank, and prepends it to the user message instead. It works like a Mastra system prompt, except it steers what gets written rather than how the agent behaves.

The mission has a sharp edge. If it excludes everything in a document, that document produces zero memories, and since recall and reflect search memories rather than documents, the source becomes unreachable. Nothing errors, and the retain reports success.

You find out from the retain.completed webhook carrying memory_unit_count: 0, or from the hindsight.retain.documents.total{outcome="no_facts"} metric. Extraction is not deterministic either, so a borderline document can yield facts on one run and none on the next. A zero means “run it again with a wider mission”, not “this document has nothing in it”.

Mastra took a different route. Working memory is a typed template, and leaving the rest out is the design:

public defaultWorkingMemoryTemplate = `
# User Information
- **First Name**:
- **Last Name**:
- **Location**:
- **Occupation**:
- **Interests**:
- **Goals**:
- **Events**:
- **Facts**:
- **Projects**:
`;

Mastra filters on write too. Turn on manageWorkingMemory and the observer gets a WorkingMemoryExtractor added to its extraction step, which sends it this:

packages/memory/src/processors/observational-memory/working-memory-extractor.ts L65–86

Update working memory with durable facts from the observations you made. Return null when no working memory update is needed.

“Durable facts” is doing the same job a mission does, except it is hardcoded in the extractor instead of written by you. The default for manageWorkingMemory is false, which is why it is easy to miss. You set workingMemory.agentManaged: true to keep the main agent’s own tools enabled alongside it.

The two filter differently when they are wrong. If a Hindsight mission excludes everything, you get zero memories and a successful response. If the Mastra extractor returns something that does not match the configured schema, it throws and nothing is saved, which is louder but loses the update too. Neither one tells you at the time that a real conversation just went into a hole.

Where I would use each

Observational memory for one long thread doing tool-heavy work. A coding agent driving Playwright MCP, where one page snapshot is 50,000 tokens and the agent has to keep going. Turn it on, it works, and nothing in the user’s turn waits on the Reflector.

Hindsight when memory has to span sessions and you need to know when. Per-user assistants, support agents, anything where “what did they tell me last month” is a real question.

Hindsight when the agent is supposed to have opinions. A PM agent weighing risk, a sales agent working out why one email landed. The opinion network is the point, and observational memory has nowhere to put it.

Neither, probably, for an n8n workflow. Four networks, a graph, a consolidator and a tool loop is a lot of infrastructure to remember a preference.


The Mastra docs are good, and the Hindsight paper is short for what it covers. Worth reading both if you are picking one.