Friday, July 24, 2026
Tech

How to Build AI Agents With Memory Using Weaviate Engram

How to Build AI Agents With Memory Using Weaviate Engram

LLMs are stateless. Every API call starts cold. That works for one-shot answers, and fails for agents that must remember preferences, past decisions, and lessons across sessions.

Weaviate Engram is a managed memory service built on Weaviate for exactly that problem. You send raw conversations or events. Engram extracts structured memories, reconciles them with what it already knows, and stores them for semantic search. Your agent stays fast because memory work runs asynchronously, while recall stays precise because retrieval is backed by Weaviate’s vector index.

This guide shows how to wire Engram into a real agent loop.

Why agents need Engram (not just a bigger context window)

Stuffing full chat history into every request looks simple. It does not scale.

  • Long context raises cost and latency on every turn.
  • Models still get lost in the middle.
  • Raw transcripts are noisy, contradictory, and outdated.
  • Multi-agent workflows split one task across multiple windows, so “one transcript” is not enough.

Engram’s model is different: actively maintain memories. Extract facts. Deduplicate. Update when preferences change. Retrieve only what is relevant for the next decision.

What Engram is

Engram is a memory server for LLM agents and apps. It exposes a REST API (https://api.engram.weaviate.io) and a Python SDK (weaviate-engram).

Core capabilities:

Core concepts (keep these straight)

  • Memories — discrete facts, embedded as vectors for search
  • Topics — categories that guide extraction (e.g. UserKnowledge, experience)
  • Groups — bundles of topics + a pipeline for one use case (often default)
  • Scopes — who a memory belongs to:
  • project-wide (shared learning)
  • user-scoped (hard isolation via multi-tenancy)
  • property-scoped (e.g. one summary per conversation_id)
  • Pipelines — async graphs that extract, reconcile, and commit

Templates like Personalization get you started without designing pipelines from scratch.

Setup

  1. Create an Engram project in Weaviate Cloud (Personalization template is a good start).
  2. Create an API key and save it immediately.
  3. Install the client:

pip install weaviate-engram anthropic

# or: uv add weaviate-engram

export ENGRAM_API_KEY=”eng_…”

export ANTHROPIC_API_KEY=”sk-ant-…”

import os

from engram import EngramClient

client = EngramClient(api_key=os.environ[“ENGRAM_API_KEY”])

The agent memory loop

A practical agent loop with Engram has three steps each turn:

  1. Recall — search memories for the current user message
  2. Act — call the LLM with recent turns + recalled context
  3. Remember — fire-and-forget the new exchange into Engram

1) Store conversations (async)

run = client.memories.add(

[

{“role”: “user”, “content”: “I just moved to Berlin and prefer specialty coffee, not chains.”},

{“role”: “assistant”, “content”: “Got it — I’ll keep specialty spots in Berlin in mind.”},

],

user_id=”alice”,

group=”default”,

)

print(run.run_id, run.status)

Engram returns a run_id immediately. The pipeline:

  1. Extract — pull topic-matching facts
  2. Transform — dedupe / merge with existing memories
  3. Commit — persist to Weaviate

You can poll with client.runs.wait(run.run_id) when you need consistency before the next search. In most chat UIs, fire-and-forget is fine because the latest turn is already in short-term context.

Other input types:

  • String — app events (“User viewed pricing page”)
  • Pre-extracted — agent decides what to remember via tool calls

2) Recall before the model responds

from engram import HybridRetrieval

results = client.memories.search(

query=”What kind of coffee does the user like?”,

user_id=”alice”,

group=”default”,

retrieval_config=HybridRetrieval(limit=5),

)

memory_context = “\n”.join(f”- {m.content}” for m in results)

Retrieval options:

Minimal memory-enabled agent

import os

import anthropic

from engram import EngramClient, HybridRetrieval

engram = EngramClient(api_key=os.environ[“ENGRAM_API_KEY”])

llm = anthropic.Anthropic()

user_id = “alice”

recent = [] # short-term: last few turns only

def agent_turn(user_input: str) -> str:

# 1) Recall long-term memory

results = engram.memories.search(

query=user_input,

user_id=user_id,

group=”default”,

retrieval_config=HybridRetrieval(limit=5),

)

memory_context = “\n”.join(f”- {m.content}” for m in results) or “- (none yet)”

system = f”””You are a helpful agent with persistent memory.

What you remember about this user:

{memory_context}

Use memories when relevant. Do not invent facts not present here or in the chat.”””

recent.append({“role”: “user”, “content”: user_input})

# 2) Act with short-term context + recalled memory

response = llm.messages.create(

model=”claude-sonnet-4-5-20250929″,

max_tokens=1024,

system=system,

messages=recent[-6:], # last ~3 exchanges

)

assistant = response.content[0].text

recent.append({“role”: “assistant”, “content”: assistant})

# 3) Remember asynchronously

engram.memories.add(

[recent[-2], recent[-1]],

user_id=user_id,

group=”default”,

)

return assistant

This pattern replaces growing history with search + a small recent window, which cuts tokens while keeping personalization.

Give the agent control with tools

Automatic recall before every turn is simple. Tool-based recall is more powerful for multi-step agents.

Expose Engram as tools:

This matches the Hermes Agent plugin model (engram_search, engram_store, engram_fetch).

Sketch:

tools = [

{

“name”: “search_memory”,

“description”: “Search long-term memories about the current user.”,

“input_schema”: {

“type”: “object”,

“properties”: {“query”: {“type”: “string”}},

“required”: [“query”],

},

},

{

“name”: “store_memory”,

“description”: “Store or correct a fact about the user.”,

“input_schema”: {

“type”: “object”,

“properties”: {“content”: {“type”: “string”}},

“required”: [“content”],

},

},

]

def handle_tool(name: str, args: dict, user_id: str):

if name == “search_memory”:

return [

m.content

for m in engram.memories.search(

query=args[“query”],

user_id=user_id,

retrieval_config=HybridRetrieval(limit=5),

)

]

if name == “store_memory”:

run = engram.memories.add(args[“content”], user_id=user_id)

return {“run_id”: run.run_id, “status”: run.status}

When the agent “forgets,” it stores a correcting memory. Engram’s reconcile pipeline supersedes the old one instead of leaving contradictions in the store.

Continual learning for agents (not only users)

Engram is not limited to user preferences. Configure topics like experience or feedback so agents learn workflows over time:

  • User says genre filtering should use a genres property, not near-text search.
  • Engram extracts feedback, transforms it into an experience memory, and commits it.
  • Next task, the agent searches experience memories and avoids the same mistake.

Scope choices matter:

  • Project-wide experience — team agents improve together
  • User-scoped experience — personal agents that never leak learning across users

Design patterns that work in production

  1. Always pass user_id for user-scoped topics — Engram enforces isolation; do not invent a shared memory bag.
  2. Use hybrid search by default — best balance of meaning and exact terms.
  3. Keep short-term history short — last 2–3 exchanges + recalled memories.
  4. Fire-and-forget adds; wait only when needed — e.g. before a critical next-step search.
  5. Use bounded topics for profiles — one UserProfile per user, fetched into the system prompt every turn.
  6. Let agents store corrections — do not delete as the primary “forget”; reconcile instead.
  7. Separate groups by use case — personalization vs continual learning stay clean.

REST fallback (any language)

curl -X POST “https://api.engram.weaviate.io/v1/memories” \

-H “Authorization: Bearer $ENGRAM_API_KEY” \

-H “Content-Type: application/json” \

-d ‘{

“input”: {“string”: {“content”: [“The user prefers dark mode.”]}},

“user_id”: “alice”

}’

curl -X POST “https://api.engram.weaviate.io/v1/memories/search” \

-H “Authorization: Bearer $ENGRAM_API_KEY” \

-H “Content-Type: application/json” \

-d ‘{

“query”: “What UI preferences does the user have?”,

“user_id”: “alice”,

“retrieval_config”: {“retrieval_type”: “hybrid”, “limit”: 5}

}’

Summary

Building agents with memory is not “save the transcript.” It is extract, reconcile, scope, and retrieve.

With Weaviate Engram you get:

  1. A low-latency write path (memories.add) that pipelines extraction in the background
  2. Weaviate-backed search (vector / bm25 / hybrid) for relevant recall
  3. Hard multi-tenant isolation by user and soft isolation by properties
  4. Two integration styles: auto-recall into the prompt, or agent-controlled tools

Start with the Personalization template, wire the search → respond → store loop, then add tool-based recall and experience topics as your agent grows.

More in Tech