Enterprise AI agent memory evaluation guide
How to evaluate memory for an AI agent that can take action.
A practical security, architecture, and procurement checklist for teams moving voice, support, and workflow agents from pilot to production.
conversation → memory → policy → host → receipt
Updated August 27, 2026 · 18-minute read.
Short answer
Do not evaluate agent memory as a better vector search.
Evaluate the whole path from untrusted conversation to persistent memory to tool execution. The system should show who owns the memory, where it came from, its validity state, whether policy allows it to support an action, and what the host reports afterward.
user partition + source + confidence + validity + policy version + action decision + evidence IDs + request ID + execution receipt
API and MCP confirmation record the authenticated project credential against one scoped memory, not an end-user identity. The customer host authenticates and retains any end-user attestation. Console confirmation records operator context. None proves objective truth. ContextDB advises. The customer host enforces.
1. Buying trigger
Persistent memory needs controls when agents act.
A support chatbot remembering a preferred language is one problem. An agent using remembered details to refund a payment, book a service, alter a plan, dispatch a worker, or update an account is a different problem.
| Evaluation path | Conditions | Boundary |
|---|---|---|
| Governed memory layer |
|
Use source, validity, action policy, host enforcement, and receipts. |
| Generic memory API may be enough |
|
Keep application-level logging and deletion limits explicit. |
2. Reference architecture
Put memory between the agent and your database.
The model interprets language. The memory layer stores durable context and evaluates whether it may support an action. The host enforces the result. The CRM, booking system, payment system, or operational database remains authoritative.
-
01 / Channels
Voice, chat, and workflow
Untrusted language and events enter through customer channels.
conversation + event -
02 / Runtime
Agent runtime
The model, tools, and orchestration interpret the request.
model + tools + orchestration -
03 / Memory
Memory control
Scope, provenance, validity, and policy govern durable context.
remembered evidence × policy -
04 / Enforce
Customer tool host
The host enforces the advisory decision under its own authorization.
act | ask | abstain -
05 / Authority
Systems of record
CRM, payments, scheduling, and operational data remain authoritative.
mutation + external reference
source → candidate → PII handling → commit
Form memory deliberately. Extract bounded candidates, preserve exact source evidence, process PII first, and separate proposals from committed writes.
recall → recall_for_action
Separate conversation context from action evidence. Broad recall can support a natural conversation. Consequential tools should use a stricter action gate.
decision → host → receipt
Keep enforcement in your host. The host executes only after an allowed decision, then reports succeeded, failed, or skipped with an idempotent receipt.
3. Threat model
Test failures that appear after the original conversation is over.
Persistent memory changes the time horizon of an attack. A bad write can influence an unrelated session days later, when the original malicious input is no longer visible in the active prompt.
| Threat | Failure | Test |
|---|---|---|
| Cross-customer memory leakage | An observed memory ID or user identifier retrieves another customer or project. | Replay real identifiers with a different project key and user partition. Expect 404, never an empty-but-successful cross-scope response. |
| Memory and context poisoning | Attacker-controlled or weakly sourced content becomes durable trusted context. | Inject an instruction through a transcript, third-party source, or tool result. Inspect whether it is rejected, marked suspect, or held for confirmation. A later API or MCP confirmation records project-credential context. The customer host retains any end-user attestation. |
| Provenance collapse | User-stated, agent-inferred, and imported data become indistinguishable after extraction. | Submit the same claim through three source types and compare confidence, confirmation state, and action eligibility. |
| Stale or conflicting memory | An old preference or superseded instruction continues to drive actions. | Write a conflicting fact to the same stable slot. Verify which generation is current and whether the conflict blocks action. |
| Tool execution after ask or abstain | The model or host ignores the gate and calls a consequential tool anyway. | Force an ask or abstain result, simulate execution, and verify the receipt is retained as a policy violation. |
| Ambiguous retries | A network timeout creates duplicate memories, confirmations, or downstream actions. | Repeat each mutation with the same idempotency key and payload, then with the same key and a changed payload. |
OWASP identifies Memory and Context Poisoning as ASI06 in its Top 10 for Agentic Applications 2026. Use that framework alongside your existing application and identity threat models.
4. Control checklist
Questions to ask before approving an enterprise memory service.
| Area | Question | Evidence to request |
|---|---|---|
| Tenancy | Is organization and project scope derived from the server credential? | Cross-org and cross-project replay tests. |
| User partition | Can a caller override or wildcard another customer's partition? | Identifier validation and cross-user 404 tests. |
| Credentials | Are API keys shown once, hashed at rest, scoped, and revocable? | Key lifecycle API and revocation test. |
| PII | Does detection and encryption happen before embeddings or model formation calls? | Data-flow diagram and stored annotation inspection. |
| Provenance | Does every fact preserve source type, confidence, and evidence? | Memory schema and conflicting-source test. |
| Formation | Can extraction propose candidates without committing them? | Propose/commit modes and bounded provider failures. |
| Regression tests | Can the same memory behavior run before and after a model, policy, prompt, or source change? | Deterministic cases, explicit baseline, evidence/outcome diff, and privacy-safe export. |
| Action policy | Can conversational context be excluded from consequential tool decisions? | Separate recall and pre-action gate APIs. |
| Confirmation | Is human attestation scoped to the exact memory and customer? | Pending queue, confirm API, actor record, and foreign-ID tests. |
| Execution | Can the host report what it actually attempted? | Decision ID, terminal receipt, idempotency, and policy-alignment field. |
| Audit | Can an authorized customer independently verify an export? | Real entries, hash chain, asymmetric signature, published public key. |
| Logging | Are memory content, queries, credentials, and webhook URLs excluded? | Forbidden-field tests and sample structured logs. |
| Failure posture | What happens when metadata, memory, encryption, or audit storage is unavailable? | Fail-closed tests and readiness behavior. |
| Deletion | Can the vendor prove scoped deletion and retention behavior? | Deletion API, verification job, retention policy, and current limitations. |
| Identity | Does the service support your IdP, service identity, and provisioning model? | SSO, SCIM, workload identity, and roadmap status. |
| Encryption keys | Who controls the master key and its rotation? | KMS architecture, per-project data keys, BYOK policy, and revocation test. |
| Operations | What availability, backup, recovery, and data-location commitments exist? | Published commitments, status history, backup tests, and contract language. |
5. Evaluation plan
Run a narrow evaluation that can fail for a named reason.
Pick one workflow with an observable action. Do not start with a broad “remember everything” requirement.
- Define scope. One project, one user identifier scheme, one action, one operational system.
- Define allowed memory. List durable facts, forbidden data, source types, confirmation rules, and validity windows.
- Integrate the client action flow. Recall, remember, evaluate action, confirm when required, re-evaluate, run the host action, and report execution.
- Run boundary attacks. Cross-user replay, foreign project key, poisoned transcript, conflicting fact, stale fact, duplicate retry.
- Trace one action. From conversation turn to memory candidate to decision to receipt to external reference.
- Pin a memory baseline. Save deterministic recall and action cases, pass one suite, then prove a deliberate memory change is reported as a regression.
- Verify operations. Revoke the key, stop a dependency, inspect readiness, and verify a signed audit export.
Pass criteria
Suggested pass criteria
foreign project + user
Foreign project and user identifiers never return a scoped resource.
contested | policy-untrusted ≠ act
Confirmation is one trust signal, not the only path to act. Contested
or policy-untrusted evidence cannot produce act. The host
branches on act,
ask, or abstain.
decision ID + terminal receipt
Every tested consequential action has a decision ID and terminal receipt.
same key + changed payload → conflict
Changing an idempotent retry payload returns a conflict instead of another write.
pinned baseline → evidence + outcome diff
A pinned Memory CI suite catches changed evidence IDs or action outcomes.
content-free export → offline verify
Audit export entries contain no customer content and verify offline.
missing control → named blocker
The team can name which missing controls block production approval.
6. ContextDB mapping
How the current hosted alpha maps to the checklist.
| Maturity | Capability | Current scope | Qualification |
|---|---|---|---|
| Available in Alpha | Memory and action boundaries |
|
Hosted behavior remains Alpha and requires host enforcement. |
| Enterprise Alpha | Operator and audit controls |
|
An export verifies its signed record, not objective truth or downstream state. |
| Hosted and Private Alphas | Formation, Memory CI, and Managed Sources |
|
One instance of each runs on one VM. This does not establish scale, high availability, failover, or a worker fleet. |
| Planned | Controls not available today |
|
No delivery date or production commitment is stated. |
ContextDB has no compliance certifications or uptime commitments today. See the security page and current control table before beginning an evaluation.
Sources
Research and external frameworks used in this guide.
- ContextDB: A Unified Context Layer for AI Agents (Sharma, 2026): a synthesis of 200+ agent-memory papers and directional internal observations reported in a self-authored preprint, not pilot findings. Also archived on Zenodo.
- NIST AI 600-1, Generative AI Profile: governance, provenance, pre-deployment testing, and incident disclosure.
- OWASP Top 10 for Agentic Applications 2026: agentic application threat categories, including Memory and Context Poisoning.
- Okta / AlphaSights enterprise buyer survey, January 2026: identity, least privilege, revocation, and audit requirements.
- PwC AI agent survey, May 2025: adoption, business functions, and trust in high-stakes agent tasks.
- CCW Digital / NiCE AI memory research: the customer effort created by cross-channel systems that forget.
Use the checklist on one real workflow.
Bring the architecture, action, data boundaries, and controls your review requires.