From Compression to Reasoning: Learning Hindsight After Mastra
I have been exploring memory management in agents, and along the way I came across reflect in
Hindsight. It is a call the agent makes when a user
asks a question: it reads the memory bank and returns prose for that user.
The name immediately reminded me of Mastra’s observational memory, where the Reflector is a background agent that rewrites the observation log so the log stops growing.
In Hindsight it answers the user. In Mastra it compresses history. I had assumed they were the same operation under two names, and reading the code showed they are not.
I wrote about Mastra’s memory before, in this post. Coming back to it with Hindsight open in another tab is what made the mismatch worth writing down.
Compression is Mastra’s Reflector’s whole job
Observational memory exists because context windows fill up, and models perform worse when they are full. Mastra compresses old messages into observations when history crosses 30k tokens, then compresses the observations into reflections when they cross 40k.
Reflections do not stack. Each one rewrites the entire log. The output becomes the new log, and Mastra appends new observations after it, so the log stays close to the reflection threshold no matter how long you run.
messages: 0 ──────────────────────────► 30k ──┐
│
observations: └────────────► 40k ──► reflect ────┘
(rewrites the whole log)
The Reflector call:
packages/memory/src/processors/observational-memory/reflector-runner.ts L301–310const agent = new Agent({
id: "observational-memory-reflector",
name: "Reflector",
instructions: buildReflectorSystemPrompt(this.reflectionConfig.instruction, extractors),
model: agentModel,
maxRetries: 0,
...(memory ? { memory } : {}),
});
The last line of its prompt:
packages/memory/src/processors/observational-memory/reflector-agent.ts L74IMPORTANT: your reflections are THE ENTIRETY of the assistants memory. Any
information you do not add to your reflections will be immediately forgotten.
Make sure you do not leave out anything.
So the Reflector overwrites the log. The new reflection becomes the memory, and anything it leaves out is gone.
Compression lands around 5x to 40x on tool-heavy workloads. The read path stays simple: recent messages plus the whole log, every turn, with nothing to select.
Hindsight puts four types in one bank
Hindsight is an agent memory system from Vectorize. The pitch is that agents should learn over time instead of replaying conversation history. It keeps a persistent bank per agent: you send it text, an LLM extracts facts and files them, and later you ask for the relevant ones back.
The bank holds four kinds of memory. Where Mastra keeps one observation log, Hindsight splits it:
- World: objective facts about the outside world
- Experience: the agent’s own experiences, written in first person
- Opinion: judgements with a confidence score and a timestamp
- Observation: summaries synthesized from facts underneath
Three operations move data in and out: retain, recall, reflect. The
paper covers all three.
Storage is Postgres with pgvector, MIT licensed, one Docker container. If you want memory running with the least setup, there is a two-line LiteLLM wrapper that recalls before the call and retains after it.
Hindsight’s retain does not reframe the user
Mastra’s Observer turns a user assertion into an observed fact. User says “I have two kids”, the log stores “User stated has two kids”, and the prompt has to add “USER ASSERTIONS ARE AUTHORITATIVE” so the model does not start trusting its own summary over the human.
Hindsight’s extraction prompt classifies instead:
hindsight-api-slim/hindsight_api/engine/retain/fact_extraction.py L1091–1092fact_type:
- "world": Objective/external facts, including the user's preferences, rules,
corrections, constraints, plans, traits, or context. These stay "world" even
when the user states them during an assistant interaction.
- "assistant": Actions, experiences, or observations the assistant/agent actually
performed. Use this for the assistant/agent doing, trying, learning, deciding,
recommending, or responding — not merely for user facts mentioned in
conversation.
A user’s preference gets filed as a world fact. The agent’s own actions and lessons go somewhere else, in the experience network.
It also normalizes time, which decides whether temporal queries work later:
hindsight-api-slim/hindsight_api/engine/retain/fact_extraction.py L1099–1104- CRITICAL: Convert ALL relative temporal expressions to absolute dates in the
fact text itself. "yesterday" → write the resolved date (e.g. "on November 12,
2024"), NOT the word "yesterday"
- Coarse dates (only a year, or only a month, is stated): span the WHOLE period —
"in 2015" → 2015-01-01 to 2015-12-31
Every fact carries occurred_start, occurred_end, and mentioned_at, which track when the thing
happened and when the agent heard about it. That looks fussy until you ask what they decided in
March and the only timestamp you kept was ingestion time.
Hindsight’s consolidation is Mastra’s opposite
Hindsight consolidates facts into observations in the background. One canonical observation per
facet, accumulating evidence instead of replacing it, with a proof_count and the source facts kept
around.
The processing rules are a numbered list. Rules 7 and 8 carry the argument:
hindsight-api-slim/hindsight_api/engine/consolidation/prompts.py L51–537. PRESERVE HISTORY: observations that record significant events (sold, died,
moved, changed) are important history — never DELETE them. Only delete an
observation when it is restated identically or truly meaningless. Be very
conservative with deletes.
8. NO COMPUTATION: you do not have the full picture — never calculate, derive,
or adjust numeric values. If the user says "I have 2 dogs" and then "I have a
dog named Rex", do NOT update the count to 3 — you don't know if Rex is one
of the 2 or a new one.
Compare that to “any information you do not add will be immediately forgotten.”
Both systems consolidate, with opposite instructions. Mastra’s consolidator compresses hard because a token budget forces it. Hindsight’s gets told to stay conservative with deletes and never do arithmetic, because a memory system that invents a count is worse than one that forgot something.
Rule 1 goes after the same problem: PREFER UPDATE OVER CREATE. Both systems are trying to stop the
log from turning into fifty restatements of one thing.
Hindsight’s recall thinks in tokens, not top-k
Observational memory has no recall in its default config. The whole log sits in context, which is
the point.
Hindsight does:
results = client.recall(
bank_id="my-bank",
query="What do I know about Alice?",
max_tokens=4096,
)
You give it a token budget, it returns as much relevant memory as fits. Underneath, four arms run in parallel:
hindsight-api-slim/hindsight_api/engine/search/retrieval.py L181–209from ..memories import get_memories
unified = await get_memories().recall_unified(
conn=pool,
bank_id=bank_id,
fact_types=fact_types,
query_embedding=query_embedding_str,
query_text=query_text,
limit=thinking_budget,
temporal_window=temporal_constraint,
min_semantic=min_semantic,
min_keyword=min_keyword,
enable_text_search=enable_text_search,
enable_graph=enable_graph_retrieval,
)
Dense vectors, BM25, graph traversal over entity/temporal/causal links, and interval filtering. Reciprocal rank fusion merges them, a cross-encoder reranks, then it trims to budget.
One correction to my earlier post: Mastra added a retrieval mode since then. retrieval: true gives
the agent a recall tool for paging back to raw messages, and { vector: true } adds semantic
search. The target still differs. Mastra’s recall recovers source text that was already summarized
away. Hindsight’s queries a graph of facts that were never summarized.
Hindsight’s reflect is a tool loop
Hindsight’s reflect is an agentic loop. The agent decides what to look up until it has enough:
Implements hierarchical retrieval:
1. search_mental_models - User-curated summaries (highest quality)
2. search_observations - Consolidated knowledge with freshness
3. recall - Raw facts as ground truth
The order runs from the most distilled thing to the rawest. The prompt is explicit about the fallback:
hindsight-api-slim/hindsight_api/engine/reflect/prompts.py L336–356- MANDATORY: If search_mental_models and search_observations both return 0
results, you MUST call recall() before giving up
- This is the source of truth that other levels are built from
Mastra’s Reflector has to write the memory in one pass. Hindsight’s checks the raw facts before committing to an answer.
Banks also carry three disposition traits, skepticism, literalism, and empathy, each scored 1
to 5, that change how reflect reasons rather than what it retrieves. The
Hindsight docs explain the presets.
Side by side
| Mastra: Reflector | Hindsight: reflect | |
|---|---|---|
| Trigger | Observations cross a token threshold | A user asks a question |
| Runs | Background, off the critical path | Foreground, on the request path |
| Input | The observation log | Four networks, retrieved per query |
| Output | A smaller observation log | Prose for the user, plus opinion state |
| Optimizes for | Bounded context | Answer quality |
| Told to | Overwrite. Omit and it’s forgotten | Verify against raw facts first |
| Side effects | None | Opinions gain confidence-weighted judgements |
Neither approach is wrong. Mastra forgets things you wanted to keep. Hindsight accumulates things you did not.
What’s new here
Facts and opinions stay apart. You can ask what the agent knows separately from what it thinks, and opinions carry confidence scores. Mastra’s log cannot tell you whether “Alice is on the West Coast” came from Alice or from a guess, because it is one flat text field.
The experience network gives the agent its own memories. They are written in first person, so the agent can recall what it tried and what failed, not just facts about the user.
You can trace an answer back. Observations keep the source facts with exact quotes, and when facts conflict, both versions survive with timestamps. The Reflector would have dropped one of them.
What to keep, and when you decide that
The LLM call count bothered me enough that I went looking at their own examples. The chat memory app in the cookbook is 55 lines and does the obvious thing:
packages/memory/src/index.ts L2264–2275export async function storeConversation(
userId: string,
userMessage: string,
assistantMessage: string,
) {
const conversation = `User: ${userMessage}\nAssistant: ${assistantMessage}`;
await hindsightClient.retain(userId, conversation, {
context: "conversation",
metadata: { timestamp: new Date().toISOString() },
});
}
It retains every turn with no filter, one LLM extraction call per message. A user with a few hundred conversations has paid for a few hundred extractions, most of which are “user said hello”, and the bank holds all of them.
The escape hatch is a retain mission: plain language you write that goes into the extraction prompt:
Always include technical decisions, API design choices, and architectural trade-offs. Ignore meeting logistics, greetings, and social exchanges.
You set it on the bank or through HINDSIGHT_API_RETAIN_MISSION. The env var lands on
config.retain_mission, and one small function turns it into a preamble:
The code keeps it out of the cached system prompt on purpose, since that prompt has to stay identical for every bank, and prepends it to the user message instead. It works like a Mastra system prompt, except it steers what gets written rather than how the agent behaves.
The mission has a sharp edge. If it excludes everything in a document, that document produces zero
memories, and since recall and reflect search memories rather than documents, the source becomes
unreachable. Nothing errors, and the retain reports success.
You find out from the retain.completed webhook carrying memory_unit_count: 0, or from the
hindsight.retain.documents.total{outcome="no_facts"} metric. Extraction is not deterministic
either, so a borderline document can yield facts on one run and none on the next. A zero means “run
it again with a wider mission”, not “this document has nothing in it”.
Mastra took a different route. Working memory is a typed template, and leaving the rest out is the design:
public defaultWorkingMemoryTemplate = `
# User Information
- **First Name**:
- **Last Name**:
- **Location**:
- **Occupation**:
- **Interests**:
- **Goals**:
- **Events**:
- **Facts**:
- **Projects**:
`;
Mastra filters on write too. Turn on manageWorkingMemory and the observer gets a
WorkingMemoryExtractor added to its extraction step, which sends it this:
Update working memory with durable facts from the observations you made. Return null when no working memory update is needed.
“Durable facts” is doing the same job a mission does, except it is hardcoded in the extractor
instead of written by you. The default for manageWorkingMemory is false, which is why it is easy
to miss. You set workingMemory.agentManaged: true to keep the main agent’s own tools enabled
alongside it.
The two filter differently when they are wrong. If a Hindsight mission excludes everything, you get zero memories and a successful response. If the Mastra extractor returns something that does not match the configured schema, it throws and nothing is saved, which is louder but loses the update too. Neither one tells you at the time that a real conversation just went into a hole.
Where I would use each
Observational memory for one long thread doing tool-heavy work. A coding agent driving Playwright MCP, where one page snapshot is 50,000 tokens and the agent has to keep going. Turn it on, it works, and nothing in the user’s turn waits on the Reflector.
Hindsight when memory has to span sessions and you need to know when. Per-user assistants, support agents, anything where “what did they tell me last month” is a real question.
Hindsight when the agent is supposed to have opinions. A PM agent weighing risk, a sales agent working out why one email landed. The opinion network is the point, and observational memory has nowhere to put it.
Neither, probably, for an n8n workflow. Four networks, a graph, a consolidator and a tool loop is a lot of infrastructure to remember a preference.
The Mastra docs are good, and the Hindsight paper is short for what it covers. Worth reading both if you are picking one.