Chapter 1
Agent memory starts with clear boundaries
Agent memory is application-managed information retained so later work can use it under defined rules. A model call does not create durable application memory by itself, although a provider or application may retain state and pass earlier content into later calls. The useful boundary is operational: who stores the information, for how long, for which identity, and with what authority.
A context window is the bounded input available to a model for one inference. A history is an ordered record of earlier messages, events, or operations. State holds current task or process values. A profile holds structured attributes for an entity. Each can be durable, but each has a different selection and authority contract.
Retrieval-augmented generation, or RAG, fetches external items for a query and conditions generation on them. A vector database stores or indexes vectors so similar items can be found, often with filters or hybrid search. Both can support memory, but neither defines identity, formation, updates, retention, or action policy. See the agent memory and RAG comparison.
How the memory boundaries connect
- History records events, while formation selects memories.
- State and profiles expose known fields directly.
- RAG and indexes retrieve candidate context.
- Assembly places selected evidence in model input.
- The model returns text or a tool call. The call remains a proposal until host validation.
- The system of record supplies authoritative state.
| Mechanism | Primary job | Typical lifetime | Authority | Common failure |
|---|---|---|---|---|
| Context window | Supply input for one inference | One call, then application refill | No business authority | Relevant material is omitted or poorly placed |
| History | Preserve an ordered trace | Session or durable archive | Evidence of recorded events | Volume, ambiguity, repetition |
| State | Represent a current process value | Task or business lifetime | May be authoritative when declared | Stale replicas or unclear ownership |
| Profile | Address known entity attributes | Across sessions until changed or erased | Depends on field source and owner | Display name mistaken for identity |
| RAG | Retrieve items for a query | Corpus dependent | Relevance does not grant permission | Plausible passage treated as current truth |
| Agent memory | Manage retained information over time | Policy defined | Advisory unless another system owns the fact | Unclear formation, updates, isolation, deletion |
| System of record | Own declared business state | Business and legal lifetime | Authoritative within its stated scope | Memory overrides live state |
A database row is enough for known, directly queried fields with clear ownership. RAG may be enough for documents. Memory is a system responsibility, not a database category. Postgres, vector or graph indexes, key-value stores, and object storage can each support it when the application tests lifecycle semantics.
A larger input does not define source quality, identity, update semantics, or erasure. Those remain system design responsibilities.
Lost in the Middle reported position sensitivity in its published multi-document question answering and key-value retrieval tasks. The result is bounded to those experiments, so teams should test their own prompts. External research S15
Chapter 2
Memory types overlap across three axes
Agent memory has no single settled hierarchy. A useful taxonomy keeps representation form, functional role, and lifecycle operation on separate axes. One record can occupy several functional categories and pass through many operations. These labels organize design choices. They do not imply that software thinks or remembers in the biological way a person does.
How to read the three axes
| Axis | Available labels | AC example |
|---|---|---|
| Representation | Token-level, parametric, or latent | A structured textual event record is token-level memory |
| Function | Working, episodic, factual, entity, and other roles | The August 11 visit is episodic and entity-linked. The August 14 noise statement is factual. |
| Operation | Formation through forgetting | The visit is formed and retrieved. A preference may later be superseded, retained, or erased. |
Axis one: representation form
Token-level memory keeps information in discrete text, record, graph, or key-value form. Parametric memory places information in model parameters. Latent memory retains internal continuous representations. The forms differ in inspection, update, and deletion properties.
Axis two: cognitive or functional role
In cognitive theory, working memory is limited-capacity temporary storage and manipulation or attentional control. This guide uses it by analogy for task-active agent state. Short-term memory is temporarily accessible, capacity-limited information. Its relation to working memory and its decay are theory-dependent.
In cognitive theory, episodic memory is recollection of experienced events in subjective time. Agent systems use the term by analogy for event or trajectory records. Semantic memory holds generalized knowledge not tied to one episode. Factual memory expresses a proposition about an entity or world state. Its label does not make it true.
In cognitive theory, procedural memory concerns nondeclarative skills and habits. CoALA extends the agent mapping to implicit model knowledge and explicit agent code External research S02. Experiential memory retains cases, strategies, and lessons from prior outcomes. It can combine episodes with procedures.
Prospective memory is a retained intention to perform an action at a future cue while other work continues. The stated Thursday preference is factual memory. A delayed intention to resolve its date or check the calendar at a known cue is prospective. PM-Bench studies delayed intentions and cues, not general memory quality. External research S13
Entity memory organizes records around a stable entity and its relations. It depends on identity resolution. Alex is not a unique key, so the application must define how account and unit identifiers relate.
Axis three: lifecycle operation
Formation selects and labels candidates from source material. Consolidation reorganizes retained items. Retrieval selects items for a task. Reflection derives an inference or lesson. Evolution changes status or representation. Forgetting reduces future use.
Generative Agents, Reflexion, ExpeL, and Voyager combine these operations differently in their reported simulations and tasks. They illustrate designs rather than one winning taxonomy. External research S04 S06 S07 S08
Chapter 3
How agent memory moves from source to deletion
A memory lifecycle moves from source through formation, consolidation, storage, retrieval, context assembly, reflection, then evolution or forgetting. It is iterative. Retrieval failures can expose poor formation, contradictions can trigger evolution, and policy can remove records. Derived items should remain linked to their evidence.
How memory moves through the lifecycle
- Capture source, actor, time, scope, and record ID.
- Form candidates through extraction, filtering, and a write decision.
- Consolidate duplicates while preserving evidence links.
- Store records and build suitable indexes.
- Route and rank retrieval for each task.
- Assemble bounded context with uncertainty labels.
- Store reflections as derived records.
- Evolve or forget under explicit rules.
- Feed failures back into earlier stages.
Formation chooses what can become memory
Formation segments sources, extracts candidates, binds supporting spans, filters disallowed content, then accepts, rejects, or holds each write. Retain material only when a later task needs it and identity, purpose, source, update, retention, and erasure rules are defined. Reject filler, unsupported inference, duplicates, secrets, and purposeless personal data.
SeCom reported gains from segment-level construction and compression-based denoising in its evaluated conversation benchmarks. This supports testing those methods, not a universal claim. External research S18
Consolidation changes organization, not evidence ownership
Consolidation can deduplicate, merge, summarize, link, or abstract while retaining source pointers. Conflicts may preserve both statements, contest one, or select a scoped current value.
Storage and indexing are separate duties
Storage retains canonical records and lifecycle fields. Rebuildable vector, keyword, temporal, entity, graph, or structured indexes support lookup without owning provenance, retention, or authoritative values.
Retrieval routes questions before ranking answers
Semantic retrieval finds similar meaning. Keyword retrieval finds exact terms. Other methods handle time, entities, graphs, or fields. Methods can combine, and no index type has a universal advantage.
Retrieval router
Alex + unit 4B
Resolve partition and entity, then read the attribute or episode.
what happened after August 11
Order by valid time while preserving transaction time.
cases like unresolved compressor noise
Use similarity with entity, source, status, and time constraints.
may the host book a follow-up
Apply evidence policy, then check live state outside memory.
A-MEM, Zep and Graphiti, MAGMA, and Hindsight describe linked, temporal, graph, or reflective designs. Zep and Graphiti is vendor-authored, and Hindsight is vendor-led. Results remain setup-specific. External research S09 S10 S11 S12
Context assembly makes retrieval usable
Context assembly selects, orders, labels, and compresses retrieved items under a budget while exposing source, time, status, and uncertainty. Instructions, task state, and memories may carry different authority.
Reflection and evolution must preserve the distinction between statement and inference
Reflection derives an evidence-linked lesson or inference. Evolution applies add, update, supersede, contest, merge, no-op, or delete. A current attestation need not be objective truth.
How valid and transaction time differ
| Record | Valid time | Transaction time | Later state |
|---|---|---|---|
| Technician visit | August 11 | Recorded August 14 | Historical episode remains |
| Unresolved compressor noise | From August 14 until changed | Recorded August 14 | Separate factual record |
| Thursday preference | From August 14 until superseded | Recorded August 14 | Target date and timezone may need resolution |
| Friday possibility | From August 14 until superseded | Recorded August 14 | Tentative, not confirmed |
| Booking | No valid time until the calendar accepts a slot | Not recorded | The calendar owns any accepted slot |
Forgetting names the actual effect
Expiry ends eligibility, suppression hides, compaction reduces detail, archival moves, and deletion removes covered data. Erasure checks records, indexes, links, caches, and derivatives, then states what remains.
Chapter 4
A reference architecture for agent memory
A sound memory architecture separates sources, identity, formation, records, indexes, retrieval, policy, execution, and business truth. Memory can recommend evidence for an answer or action. The trusted host must still authenticate the actor, check permission and live state, enforce business rules, and ask the system of record to create any side effect.
How the architecture handles an action
- Sources arrive with record identifiers.
- Application logic resolves identity and partitions.
- Formation and write policy create records.
- Canonical data retains lifecycle fields.
- Indexes support task-specific retrieval.
- Retrieval and context assembly prepare task evidence.
- The runtime returns an answer or tool call.
- Memory policy evaluates evidence for the proposed action.
- The tool call remains a proposal until host validation.
- The host checks identity, permission, state, and rules.
- The system of record accepts or rejects the mutation.
- Audit and evaluation records preserve review data.
A memory record is more than content
- Identity
- Memory, tenant, user, and entity keys.
- Content
- A bounded proposition, event, method, or intention.
- Type
- Functional role and representation form.
- Source
- Source kind, record ID, actor, and span.
- Time
- Valid time and transaction time.
- Trust
- Confidence, attestation, contest, and corroboration state.
- Lifecycle
- Status, supersession, retention, and schema version.
- Derivation
- Transformation method and evidence links.
Provenance records the entities, activities, and actors that produced or changed an item. It names the claim source, transformation, and later derivations. W3C PROV supplies a storage-neutral model. External research S20
Records and indexes have different recovery plans
A memory store owns canonical content and lifecycle metadata. An index accelerates a query pattern. Vector, keyword, temporal, entity, and graph indexes can point to one record. A rebuild should not change source, attestation, or retention state.
Read consistency is a named contract
A claim of consistency is incomplete unless it names the partition, store, time, version, and replica behavior. Read-after-write, monotonic reads, a version floor, and bounded replica lag are distinct guarantees with different costs.
RFC 9110 defines an idempotent method as one where multiple identical requests have the same intended effect as one such request. That method property does not define an application retry-key contract. External research S27
This guide's operational recommendation is to bind each retry key to a tenant and operation scope, require payload equality, set an expiry, define concurrent-request behavior, and specify how callers resolve an ambiguous outcome.
Chapter 5
Evaluate the whole memory loop
Agent memory evaluation needs four layers: formation quality, retrieval and answer quality, lifecycle correctness, and action plus operational safety. A single recall score cannot show whether records were sourced, updates replaced stale facts, erasure reached derivatives, or a retry duplicated a booking. Test with costs and permissions from the intended workload.
How evaluation proceeds
- Define inputs, identities, permissions, times, and expected outcomes.
- Check supported formation and rejected content.
- Check retrieval and evidence-grounded answers.
- Apply updates, expiry, erasure, and retries.
- Run actions through the real host path.
- Classify each failure by its earliest cause.
- Rerun fixtures after controlled changes.
Layer one: formation quality
Candidate precision is the share of formed candidates that should be retained. Unsupported candidate rate counts records without source support. Source-span coverage checks evidence links. Duplicate rate, compression ratio, and personally identifiable information (PII) handling add operational context. A lower memory count is not a quality result by itself.
Layer two: retrieval and answer quality
Recall at k asks what share of required items appears in the first k results. Precision at k asks how many returned items are relevant. Mean reciprocal rank rewards the first correct result near the top. Normalized discounted cumulative gain rewards graded relevance near the top. These metrics score ranking, not answer fidelity.
Score answer accuracy, evidence alignment, temporal order, multi-hop reasoning, and abstention separately. Record dataset version, model, prompt, retrieval budget, and judge for comparable runs.
Layer three: lifecycle correctness
Introduce corrections, contradictions, late events, expiry, erasure, and retries. Inspect records, indexes, assembled context, derivatives, caches, and audit scope against expected behavior written before the run.
Layer four: action and operational safety
A false act permits the host to proceed when evidence policy should not. A false abstain withholds valid support. An unnecessary ask repeats a settled question. Test host bypass, duplicate side effects, receipt accuracy, isolation, failures, and whole-path latency. Weight errors by workload cost.
| Layer | Expected result | Trap | Useful measure |
|---|---|---|---|
| Formation | Keep episode, unresolved-noise fact, stated preference, fallback, and task separate | Mark conditional Friday as confirmed | Unsupported rate and source coverage |
| Retrieval | Return August 11, the August 14 noise fact, and Thursday | Rank an old closed repair first | Recall, order, and answer accuracy |
| Lifecycle | In a branch, confirm Friday and supersede Thursday while preserving history | Delete the visit with the preference | Update, expiry, and erasure checks |
| Action and operations | Resolve date and timezone, check availability, and prevent retry duplicates | Treat preference as calendar proof | False act, duplicates, receipts, latency |
In their published tasks, LoCoMo evaluates long conversational histories with question answering, summarization, and temporal questions, LongMemEval evaluates extraction, multi-session reasoning, updates, time, and abstention, and PM-Bench evaluates delayed intentions. Public benchmarks omit your permissions, systems of record, latency path, and action costs. External research S16 S17 S13
Compare vendor results only when dataset, model, prompt, retrieval budget, judge, and costs match. Run a release gate from your own permissions, failures, action costs, and latency budget.
Chapter 6
Relevant memory is not action authority
Agent memory does not authorize action. It can supply policy evidence, but the trusted host still validates every tool call, verifies identity and permission, checks live state, enforces rules, and controls side effects. Security must cover writes, retrieval, use, retries, receipts, retention, erasure, and audit.
Define the action boundary before assigning controls
A tool call is a structured model or runtime request. It remains a proposal until the trusted host validates it. An action is the intended operation. A side effect is the resulting external change.
Evidence supports a memory, answer, or decision. Its source names the origin. Confidence is a system score, not a calibrated probability unless measured, and cannot authorize. Corroboration requires an independent source. Copies or transformations of one source do not qualify.
Confirmation or attestation records that an authenticated actor affirmed a scoped claim. It records an assertion, not objective truth. A policy decision applies rules to evidence, actor, action, and risk. Act continues host checks, ask requires more information, and abstain means memory supplies no support for the action.
How a proposed action reaches execution
- Retrieve sourced, timed evidence.
- Evaluate versioned action policy.
-
The policy opens three parallel branches.
- Act continues only to host enforcement.
- Ask collects missing information and re-evaluates policy without a side effect.
- Abstain stops the action path because memory supplies no support.
- On the act branch, the host verifies identity, permission, live state, and business rules.
- The system of record accepts or rejects the side effect.
- The host reports a receipt, while downstream checks can strengthen execution proof.
Policy decision and enforcement are separate functions
NIST attribute-based access control guidance and the XACML standard distinguish a policy decision function from a policy enforcement function. XACML uses Permit, Deny, Indeterminate, and NotApplicable. Those outcomes differ from the memory-evidence triad above. External research S22 S24
Identity must bind every write and read
Partition keys scope storage and retrieval. They are not identity assertions. A project credential authenticates a project, not the end user. The customer host authenticates people and resolves partitions. Cross-user leakage tests use reused names, missing tenant keys, and cache collisions.
OAuth 2.0 supplies delegated authorization and scoped access concepts. It proves neither a remembered claim nor Alex's identity. Authorization, attestation, and truth remain separate. External research S23
Persistent context creates durable attack paths
Prompt injection becomes more dangerous when malicious instructions are formed into durable memory. Memory and context poisoning changes later behavior through inserted content. Controls inspect writes, label untrusted sources, constrain retrieval, preserve instruction priority, and isolate privileged tools.
OWASP's 2026 agentic guidance names tool misuse, identity and privilege abuse, and memory poisoning. It is guidance, not certification. MEXTRA demonstrated black-box memory extraction in two evaluated agent settings, without establishing rates for other systems. External research S25 S19
Retries and receipts require bounded claims
Consequential operations need an idempotency contract because network failures create ambiguity. The host might book and lose the response. Define retry-key scope, changed-payload conflicts, retention, and resolution of an unknown result.
An execution receipt links a prior decision to what the host reports. It can carry status, an opaque reference, policy alignment, and time. Without an independent calendar check, the receipt is host-reported evidence, not proof of an appointment.
PII, retention, erasure, and audit form one review
PII controls start before storage and external model calls. Data minimization limits fields, retention sets a lifetime, and erasure verifies records, indexes, links, caches, and derivatives. Suppression, expiry, backups, and business records need separate treatment.
An audit is a reviewable record of selected operations, actors, times, decisions, and outcomes. Claims of complete, immutable, signed, or tamper-evident coverage need direct proof and scope. NIST guidance leaves event and retention controls to each organization. External research S26
Chapter 7
When to use a database, RAG, built memory, or a vendor
Use ordinary database and profile tables when fields are known, authoritative, and directly queried. Use RAG when the main job is retrieving passages from a document corpus. Add memory lifecycle components when user-specific or experience-specific knowledge must be formed, updated, retrieved, and removed over time. Add explicit action policy when remembered evidence can affect a side effect.
| Need | Likely starting point | Reason |
|---|---|---|
| Known current fields | Database or profile service | Known schema and authoritative ownership |
| Document questions | Keyword, vector, or hybrid RAG | The job is finding passages |
| Cross-session personal or case knowledge | Built or purchased memory components | Lifecycle operations need shared rules |
| Consequential action | Memory policy and host enforcement | Evidence cannot bypass host checks |
Build versus buy scorecard
| Control area | Question | Required proof artifact |
|---|---|---|
| Isolation and identity | How are identity boundaries resolved? | Negative cross-partition tests and identity diagram |
| Provenance and time | Are sources, derivations, and both times inspectable? | Record export and late-update trace |
| Update semantics | How do corrections and conflicts evolve? | Versioned API responses and failure cases |
| Retention and erasure | What does deletion cover? | Deletion test and residue report |
| Action policy and host boundary | Can policy and enforcement be traced? | Policy trace, host path, and bypass test |
| Receipts and retries | How are duplicate effects prevented? | Retry replay, conflict, timeout, and downstream check |
| Audit coverage | Which operations enter each audit scope? | Signed export, event inventory, and crash gap |
| Read consistency | Which write-to-read relation is guaranteed? | OpenAPI contract, token test, and topology |
| Evaluation and latency | Does the test match your workload? | Reproducible suite, costs, and latency trace |
| Cost and portability | What does the priced workload cost, and how does exit work? | Priced workload trace, export/import test, and exit procedure |
| Deployment and status | What exists now? | Dated status ledger and known limitations |
Demand artifacts, not logos or maturity labels. Build when your team can own lifecycle incidents. Buy when inspectable vendor behavior reduces that burden.
Chapter 8
Build memory in ten testable steps
Start with actions, authority, identity, and evidence. Define a neutral sourced record, retrieve for tested questions, specify updates and deletion, let the runtime propose tool calls, evaluate action evidence, validate in the host, record bounded receipts, and turn failures into regression cases.
-
Step 1
Inventory actions and systems of record
List each action, risk owner, authorization source, live-state dependency, and authoritative system.
-
Step 2
Define partition keys and identity boundaries
Name tenant, project, user, entity, and case keys, plus who resolves each value.
-
Step 3
Define records and evidence links
Require type, source, attestation, both times, status, supersession, retention, and schema version.
-
Step 4
Build bounded formation with source spans
Segment events, extract candidates, bind supporting spans, apply PII rules, and reject unsupported detail.
-
Step 5
Add retrieval for tested query types
Route tested queries under source, identity, time, status, and budget rules.
-
Step 6
Add consolidation, evolution, and forgetting
Specify merge, update, contest, no-op, expiry, suppression, archival, and erasure. Preserve derivation links.
-
Step 7
Evaluate consequential tool calls before host execution
Keep the call proposed. Record action, evidence, policy, outcome, reason, time, and responses to uncertain evidence.
-
Step 8
Enforce decisions in the customer host
The server host authenticates, authorizes, validates, checks live state and rules, and controls effects.
-
Step 9
Record host-reported execution receipts
Attach a terminal host report, retry key, downstream reference, and independent-check status.
-
Step 10
Build regression, privacy, retry, and erasure tests
Version fixtures for identity, updates, PII, injection, retries, receipts, and erasure.
Illustration
Neutral memory record
The record separates the preference, repair episode, and noise fact. Missing target date and timezone remain explicit.
{
"memory_id": "mem_example_01",
"partition_key": "tenant_demo:user_042",
"content": "Thursday afternoon works for the current AC follow-up",
"memory_type": ["factual", "entity"],
"source": "user_statement",
"source_record_id": "conversation_2026_08_14",
"confidence": 0.92,
"confirmation_state": "stated",
"entity": "unit_4B",
"attribute": "follow_up_preference",
"target_date": null,
"timezone": null,
"valid_time": {
"from": "2026-08-14T15:10:00Z",
"to": null
},
"system_time": "2026-08-14T15:10:04Z",
"status": "current",
"supersedes": null,
"retention_class": "service_preference",
"schema_version": 1
}
Illustration
Neutral branched decision and receipt
This branch starts after Alex confirms Friday and the missing details are resolved. The host validates the call, and the receipt reports calendar acceptance.
{
"branch": "later_confirmed_friday",
"resolved_target": {
"date": "2026-08-21",
"timezone": "America/New_York"
},
"decision": {
"decision_id": "decision_example_01",
"action_name": "book_ac_follow_up",
"evidence_ids": ["mem_visit_01", "mem_noise_01", "mem_friday_confirmed"],
"policy_version": "booking_policy_3",
"outcome": "act",
"reason_code": "confirmed_preference_requires_host_checks",
"created_at": "2026-08-15T10:00:00Z"
},
"receipt": {
"decision_id": "decision_example_01",
"host_reported_status": "calendar_accepted",
"downstream_reference": "calendar_ref_opaque",
"policy_alignment": "aligned",
"receipt_time": "2026-08-15T10:00:03Z"
}
}
Keep server credentials out of clients, public snippets, and logs. Derive partition keys after host authentication.
Current ContextDB behavior
How ContextDB maps to the implementation plan
Statuses below reflect repository commit 73beb047. They describe tested code, not outcomes.
| Capability | Exact status | Current mapping and limit |
|---|---|---|
| Apache SDK pycontextdb==0.4.4 | Available | Tested PostgreSQL code has user scope, policy, consistency tokens, and factual ADD, UPDATE, DELETE, and NOOP. Available is not a Cloud status. |
| Cloud gateway and core routes | Cloud |
Remember, batch remember, recall, action recall, direct confirm,
scoped forget, and pending confirmations are part of the current
allowlist. Later rows cover evolve, Formation, and receipts. The
linked OpenAPI lists MCP and every route. The single-process
hosted gateway has no SLA and is not production-ready.
user_id is a partition key, not an identity assertion.
|
| Factual evolution and read consistency | Available in the pinned SDK and exposed in Cloud | ADD, UPDATE, DELETE, and NOOP use the factual state machine. Real mutations return a project memory version and primary WAL position. Token reads use the primary, with no read replica. |
| Action evaluation | Cloud |
recall_for_action returns a durable decision ID,
an act, ask, or abstain outcome, trusted memories for act, and
pending confirmation IDs for ask. The durable decision records
evidence IDs. The result is advisory. The host owns
authentication, live state, and enforcement.
|
| Confirmation | Cloud for direct HTTP /v1/confirm and the MCP confirm tool, Planned for Cloud tickets |
Direct HTTP /v1/confirm and the MCP
confirm tool record authenticated
project-credential context against one memory, then policy must
be evaluated again. The host authenticates the end user.
Confirmation does not prove objective truth. Cloud tickets
remain Planned and do not call SDK confirmation.
|
| Execution receipts | Cloud | One structured terminal receipt can attach to an action decision with a required idempotency key. Policy misalignment is retained. The receipt is host-reported. ContextDB does not query the downstream system to prove the outcome. |
| Formation | Alpha for direct extraction, Hosted Alpha for asynchronous jobs | Direct extraction accepts bounded structured text turns, processes PII, and returns sourced factual candidates in propose or commit mode. It does not form experiential, working, prospective, procedural, multimodal, or graph memory. One single-instance Formation process runs on the gateway VM, with no fleet, audio, cancellation, or failover and no availability SLO. |
| Forgetting, erasure, and retention | Alpha for on-request forgetting and erasure, Planned for scheduled retention | Forget covers one memory, slot, or verified partition. Partition erasure requires confirmation and idempotency. It is synchronous, capped at 10,000 active memories, and not one atomic transaction with the Cloud audit event. Scheduled retention is Planned. |
| Audit and PII boundaries | Alpha with bounded coverage and a known SDK limitation | Cloud control events, read access, and SDK operations use separate chains. No single export covers every operation. Coverage starts at deployment, is capped, and has crash gaps. Gateway tests remove plaintext email before provider calls. SDK default annotations retain originals. |
The current status ledger covers adjacent surfaces outside this implementation map. Recheck the pinned known limitations, OpenAPI contract, and changelog before evaluation.
Chapter 9
Memory requirements across four agent types
Voice, support, scheduling, and workflow agents need different latency and tool paths, but they share the same duties: bind memory to identity and sources, retrieve current evidence, keep live state in its system of record, enforce consequential actions in the host, and test updates, retries, and erasure.
| Setting | Memory job | Consequential boundary | Evaluation focus |
|---|---|---|---|
| Voice | Carry sourced customer context across calls while fitting retrieval, identity resolution, and context assembly into a tight response path. | The host controls booking, rescheduling, refunds, messages, and account updates after checking live state. | Measure the whole time budget, identity binding, interruption recovery, stale context, and false act outcomes. |
| Support | Retain case history, preferences, prior outcomes, promises, and unresolved tasks across chat, email, and voice without merging customers who share names. | The host authorizes credits, entitlements, closure, escalation, and account mutation in the relevant business system. | Test stale facts, cross-channel identity, source quality, escalation decisions, PII, retention, and erasure residue. |
| Scheduling | Track stated preferences, hard constraints, tentative options, unresolved tasks, and prospective reminders without treating any preference as inventory. | The calendar owns availability and mutation. The host handles confirmation, permission, conflict rules, and idempotent booking. | Test live availability, confirmation scope, timezone changes, retries, ambiguous failures, and duplicate appointments. |
| Workflow agents | Reuse prior cases, procedures, lessons, and pending intentions while preserving evidence and separating reflected guidance from source facts. | A trusted host validates each proposed tool call and controls each system mutation under current permission, policy, and business rules. | Test privilege scope, prospective cues, poisoning, versioned procedures, host bypass, receipt accuracy, and rollback plans. |
Reference
Agent memory glossary
Definitions here are operational, and cognitive terms are analogies. Hu et al. and CoALA propose mappings. Entity memory, the expanded lifecycle, execution receipt, and read-consistency contract are guide-specific. The act, ask, and abstain triad is ContextDB-specific.
- Context window
- The bounded model input for one inference, without inherent durable memory.
- History
- An ordered record of prior messages, events, or operations.
- State
- Current values needed to continue a task or represent a process.
- Profile
- A structured set of attributes associated with a person, account, device, or other entity.
- Retrieval-augmented generation
- RAG retrieves external items and conditions generation on them.
- Vector database
- A store or index for vectors used in similarity and hybrid search.
- Working memory
- Limited-capacity temporary storage and manipulation or attentional control in cognitive theory, used by analogy for task-active agent state.
- Short-term memory
- Temporarily accessible, capacity-limited information whose relation to working memory and decay is theory-dependent.
- Episodic memory
- Cognitive recollection of experienced events in subjective time, used by analogy for agent event or trajectory records.
- Semantic memory
- Generalized knowledge not tied to one remembered episode.
- Factual memory
- A sourced, time-scoped proposition whose label does not establish truth.
- Procedural memory
- Cognitive nondeclarative skills and habits. CoALA extends the agent mapping to implicit model knowledge and explicit agent code.
- Experiential memory
- Cases, lessons, or skills retained from prior attempts and outcomes.
- Prospective memory
- A retained intention triggered by a future time, event, or state.
- Entity memory
- A guide-specific category organizing memory around a resolved entity, its attributes, and relations.
- Formation
- Candidate selection through segmentation, extraction, filtering, source binding, and write decisions.
- Consolidation
- Reorganization by deduplication, merging, summarization, linking, or abstraction.
- Storage
- Retention of canonical records and lifecycle metadata.
- Indexing
- Creation of structures that support named retrieval methods.
- Retrieval
- Selection of retained items for a query, task, or action check.
- Context assembly
- Selection and ordering of instructions, state, and evidence for one call.
- Reflection
- A derived inference or lesson kept distinct from source statements.
- Evolution
- Change through add, update, supersede, contest, merge, no-op, or delete.
- Forgetting
- Reduced future use through expiry, suppression, compaction, archival, or deletion.
- Provenance
- Entities, activities, and actors that produced or changed a record.
- Bitemporality
- Valid time in the world plus transaction time in the system.
- Tool call
- A structured request that remains a proposal until the trusted host validates it.
- Action
- An intended operation that may be evaluated without execution.
- Side effect
- An externally observable state change caused by execution.
- Evidence
- A scoped, sourced record supporting a memory, answer, or decision.
- Source
- The person, system, event, or record where a claim originated.
- Confidence
- A system score that is not truth, permission, or assumed calibration.
- Corroboration
- Support from an independent source, excluding copies and transformations.
- Confirmation or attestation
- A scoped actor affirmation that does not prove objective truth.
- Policy decision
- A versioned evaluation of evidence, actor, action, and risk attributes.
- Act, ask, and abstain
- ContextDB-specific illustrative outcomes to continue checks, request information, or state that memory supplies no support for the action.
- Host enforcement
- Customer code that applies checks and controls the side effect.
- System of record
- The declared authoritative system for a business object or state.
- Execution receipt
- A guide-specific decision-linked record reporting what the host says happened.
- Audit
- A scoped record of selected operations, actors, decisions, and outcomes.
- Read-consistency contract
- A guide-specific write-to-read relation naming partition, store, version, time, and replica behavior.
- Idempotency
- Repeated logical execution has the same intended effect as one attempt.
- Retention
- Policy and mechanism governing record and derivative lifetimes.
- Erasure
- Scoped deletion verified across records, indexes, links, caches, and derivatives.
- Poisoning
- Content insertion that changes later behavior through memory or context.
Reference
Frequently asked questions
Answers preserve the guide's operational boundaries.
What is memory management for AI agents?
Memory management for AI agents controls how information is formed, stored, indexed, retrieved, changed, used, retained, and deleted across time. Storage is one stage. Applications must define identity, provenance, conflict handling, context assembly, and whether remembered evidence may support an answer or action without granting execution authority.
What is the difference between an AI agent's context window and memory?
A context window is the bounded input for one model inference. Memory is information retained under rules for later work. Selected memories may enter each new window. A larger window changes input capacity, while memory management governs selection, source, identity, updates, retention, and deletion across calls.
Do larger context windows remove the need for agent memory?
A larger window can hold more history for one call, but it does not define persistence, source ownership, identity, contradiction handling, or derivative erasure. Lost in the Middle reported position sensitivity in its published multi-document question answering and key-value retrieval tasks. Test full context and selected retrieval on your workload because results, cost, and latency vary.
How is agent memory different from RAG?
Retrieval-augmented generation retrieves external items for a query and conditions generation on them. It can be enough when the main job is finding document passages. Agent memory may also govern formation, identity, provenance, updates, retention, and forgetting. The needed boundary depends on whether task-specific records must change across time.
What types of memory should an AI agent have?
Choose types from the work. Working memory supports the current task. Episodic memory records events. Factual and entity memory hold sourced propositions. Procedural and experiential memory retain methods. Prospective memory carries delayed intentions. Keep representation, functional role, and lifecycle operation as separate axes because categories overlap.
What should an AI agent store in memory?
Store information that later tasks need and that the system can source, scope, update, retain, and erase. Candidates include preferences, case events, unresolved tasks, verified constraints, and reviewed lessons. Keep transcripts in an source store. Avoid filler, unsupported inference, secrets, or PII without a purpose and policy.
How should agent memory handle updates and contradictions?
Preserve source and time, then apply update, supersede, contest, merge, no-op, or delete. A newer statement can become the current attestation without becoming objective truth. Keep historical events separate from current values. Consequential actions should ask or abstain when unresolved conflicts make the evidence unsafe.
How do you evaluate an AI agent memory system?
Evaluate formation quality, retrieval and answer quality, lifecycle correctness, and action plus operational safety. Public benchmarks support conclusions only in their published tasks. Add fixtures for your identities, permissions, systems of record, latency path, and action costs. Test stale facts, erasure, cross-user leakage, retries, false act, duplicate effects, and receipts.
What are the main security risks in agent memory?
Risks include durable prompt injection, poisoning, cross-user retrieval, private memory extraction, stale facts, excessive tool permissions, host bypass, duplicate effects, incomplete erasure, PII, and audit gaps. Apply controls at write, retrieval, context assembly, policy, host execution, and deletion boundaries. Benchmarks do not prove those controls exist.
When should recalled memory require confirmation?
Require confirmation when policy needs an actor attestation that evidence lacks, especially before costly or hard-to-reverse actions. Name the actor, claim, scope, and time. Confirmation proves neither objective truth nor live availability. Independent corroboration and host authorization may remain necessary. Policy may allow low-risk actions without confirmation.
How should an AI agent forget or erase memory?
Name the effect. Expiry stops later use. Suppression hides an item. Compaction reduces detail. Archival moves it. Erasure deletes a scope and verifies records, indexes, links, caches, summaries, and reflections. Document remaining lineage, backups, receipts, audit data, and separate business records instead of calling each mechanism deletion.
Should a team build or buy an agent memory layer?
Build when team-owned lifecycle code meets the workload and your team can operate identity, privacy, evaluation, and incidents. Buy when inspectable vendor behavior reduces that burden. First test whether a profile table or RAG is enough. Demand artifacts for isolation, provenance, updates, erasure, action policy, retries, consistency, latency, cost, portability, and current status.
Evidence
Primary sources and maturity
Sources state maturity and narrow claim boundaries. S28 and S29 are one self-authored work.
- External research Hu et al., 2026. External preprint. Memory survey frame.
- External research CoALA, Sumers et al., 2024. Peer-reviewed TMLR paper. Cognitive architecture.
- External research Zhang et al., 2025. Peer-reviewed ACM TOIS memory survey.
- External research Generative Agents, 2023. Peer-reviewed UIST paper.
- External research MemGPT, 2023. External preprint. Virtual context management.
- External research Reflexion, 2023. Peer-reviewed NeurIPS paper.
- External research ExpeL, 2024. Peer-reviewed AAAI paper.
- External research Voyager, Wang et al., 2024. Peer-reviewed TMLR paper. Executable skill library.
- External research A-MEM, 2025. Peer-reviewed NeurIPS paper. Linked notes.
- External research Zep and Graphiti, 2025. Vendor-authored external preprint. Temporal graph.
- External research MAGMA, 2026. Peer-reviewed ACL paper. Multi-graph architecture.
- External research Hindsight, 2026. Vendor-led peer-reviewed ACL system demonstration.
- External research PM-Bench, 2026. Accepted at COLM 2026. Linked arXiv version. Prospective cues.
- External research Lewis et al. on RAG, 2020. Peer-reviewed NeurIPS paper.
- External research Lost in the Middle, 2024. Peer-reviewed TACL paper. Position sensitivity.
- External research LoCoMo, 2024. Peer-reviewed ACL paper.
- External research LongMemEval, 2025. Peer-reviewed ICLR paper.
- External research SeCom, 2025. Peer-reviewed ICLR paper.
- External research MEXTRA, 2025. Peer-reviewed ACL paper.
- External research W3C PROV-DM. W3C Recommendation.
- External research Snodgrass and Ahn, 1986. Peer-reviewed IEEE Computer paper. Bitemporality.
- External research NIST SP 800-162. NIST Special Publication guidance. Attribute-based access control.
- External research OAuth 2.0, RFC 6749. IETF Proposed Standard on the Standards Track. Delegated authorization.
- External research XACML 3.0. OASIS Standard. Policy decision and enforcement.
- External research OWASP Top 10 for Agentic Applications 2026. Security guidance. Agent risks.
- External research NIST SP 800-53 Revision 5. NIST Special Publication guidance. Audit controls.
- External research HTTP Semantics, RFC 9110 section 9.2.2. Internet standard. Identical-request method idempotency.
- Self-authored preprint Gaurav Sharma (2026), ContextDB: A Unified Context Layer for AI Agents - Replacing the Patchwork with a Memory Operating System. SSRN preprint. Directional observations, not controlled customer evidence.
- Self-authored preprint Zenodo mirror of the same Gaurav Sharma (2026) preprint. Self-authored preprint mirror. Directional observations, not controlled customer evidence or independent corroboration.