Enterprise AI agent memory evaluation guide

How to evaluate memory for an AI agent that can take action.

A practical security, architecture, and procurement checklist for teams moving voice, support, and workflow agents from pilot to production.

conversation → memory → policy → host → receipt

Updated August 27, 2026 · 18-minute read.

Short answer

Do not evaluate agent memory as a better vector search.

Evaluate the whole path from untrusted conversation to persistent memory to tool execution. The system should show who owns the memory, where it came from, its validity state, whether policy allows it to support an action, and what the host reports afterward.

Minimum production record

user partition + source + confidence + validity + policy version + action decision + evidence IDs + request ID + execution receipt

API and MCP confirmation record the authenticated project credential against one scoped memory, not an end-user identity. The customer host authenticates and retains any end-user attestation. Console confirmation records operator context. None proves objective truth. ContextDB advises. The customer host enforces.

1. Buying trigger

Persistent memory needs controls when agents act.

A support chatbot remembering a preferred language is one problem. An agent using remembered details to refund a payment, book a service, alter a plan, dispatch a worker, or update an account is a different problem.

When persistent memory requires a governed action boundary
Evaluation path Conditions Boundary
Governed memory layer
  • The same customer returns across calls, chats, or workflows.
  • More than one agent or channel consumes the same context.
  • Remembered details influence writes to a business system.
  • A tentative statement must not become a permanent instruction.
  • Security review asks who, what, when, why, and what happened next.
Use source, validity, action policy, host enforcement, and receipts.
Generic memory API may be enough
  • The agent only personalizes low-risk responses.
  • Memory never leaves one application session.
  • No tool changes money, access, schedules, plans, or records.
  • Your team can accept application-level logs and manual deletion.
Keep application-level logging and deletion limits explicit.

2. Reference architecture

Put memory between the agent and your database.

The model interprets language. The memory layer stores durable context and evaluates whether it may support an action. The host enforces the result. The CRM, booking system, payment system, or operational database remains authoritative.

  1. 01 / Channels

    Voice, chat, and workflow

    Untrusted language and events enter through customer channels.

    conversation + event
  2. 02 / Runtime

    Agent runtime

    The model, tools, and orchestration interpret the request.

    model + tools + orchestration
  3. 03 / Memory

    Memory control

    Scope, provenance, validity, and policy govern durable context.

    remembered evidence × policy
  4. 04 / Enforce

    Customer tool host

    The host enforces the advisory decision under its own authorization.

    act | ask | abstain
  5. 05 / Authority

    Systems of record

    CRM, payments, scheduling, and operational data remain authoritative.

    mutation + external reference
Write boundary source → candidate → PII handling → commit

Form memory deliberately. Extract bounded candidates, preserve exact source evidence, process PII first, and separate proposals from committed writes.

Read boundary recall → recall_for_action

Separate conversation context from action evidence. Broad recall can support a natural conversation. Consequential tools should use a stricter action gate.

Execution boundary decision → host → receipt

Keep enforcement in your host. The host executes only after an allowed decision, then reports succeeded, failed, or skipped with an idempotent receipt.

3. Threat model

Test failures that appear after the original conversation is over.

Persistent memory changes the time horizon of an attack. A bad write can influence an unrelated session days later, when the original malicious input is no longer visible in the active prompt.

Persistent-memory failure modes and concrete review tests
Threat Failure Test
Cross-customer memory leakage An observed memory ID or user identifier retrieves another customer or project. Replay real identifiers with a different project key and user partition. Expect 404, never an empty-but-successful cross-scope response.
Memory and context poisoning Attacker-controlled or weakly sourced content becomes durable trusted context. Inject an instruction through a transcript, third-party source, or tool result. Inspect whether it is rejected, marked suspect, or held for confirmation. A later API or MCP confirmation records project-credential context. The customer host retains any end-user attestation.
Provenance collapse User-stated, agent-inferred, and imported data become indistinguishable after extraction. Submit the same claim through three source types and compare confidence, confirmation state, and action eligibility.
Stale or conflicting memory An old preference or superseded instruction continues to drive actions. Write a conflicting fact to the same stable slot. Verify which generation is current and whether the conflict blocks action.
Tool execution after ask or abstain The model or host ignores the gate and calls a consequential tool anyway. Force an ask or abstain result, simulate execution, and verify the receipt is retained as a policy violation.
Ambiguous retries A network timeout creates duplicate memories, confirmations, or downstream actions. Repeat each mutation with the same idempotency key and payload, then with the same key and a changed payload.

OWASP identifies Memory and Context Poisoning as ASI06 in its Top 10 for Agentic Applications 2026. Use that framework alongside your existing application and identity threat models.

4. Control checklist

Questions to ask before approving an enterprise memory service.

Control questions, evidence, and current limitations to review
AreaQuestionEvidence to request
TenancyIs organization and project scope derived from the server credential?Cross-org and cross-project replay tests.
User partitionCan a caller override or wildcard another customer's partition?Identifier validation and cross-user 404 tests.
CredentialsAre API keys shown once, hashed at rest, scoped, and revocable?Key lifecycle API and revocation test.
PIIDoes detection and encryption happen before embeddings or model formation calls?Data-flow diagram and stored annotation inspection.
ProvenanceDoes every fact preserve source type, confidence, and evidence?Memory schema and conflicting-source test.
FormationCan extraction propose candidates without committing them?Propose/commit modes and bounded provider failures.
Regression testsCan the same memory behavior run before and after a model, policy, prompt, or source change?Deterministic cases, explicit baseline, evidence/outcome diff, and privacy-safe export.
Action policyCan conversational context be excluded from consequential tool decisions?Separate recall and pre-action gate APIs.
ConfirmationIs human attestation scoped to the exact memory and customer?Pending queue, confirm API, actor record, and foreign-ID tests.
ExecutionCan the host report what it actually attempted?Decision ID, terminal receipt, idempotency, and policy-alignment field.
AuditCan an authorized customer independently verify an export?Real entries, hash chain, asymmetric signature, published public key.
LoggingAre memory content, queries, credentials, and webhook URLs excluded?Forbidden-field tests and sample structured logs.
Failure postureWhat happens when metadata, memory, encryption, or audit storage is unavailable?Fail-closed tests and readiness behavior.
DeletionCan the vendor prove scoped deletion and retention behavior?Deletion API, verification job, retention policy, and current limitations.
IdentityDoes the service support your IdP, service identity, and provisioning model?SSO, SCIM, workload identity, and roadmap status.
Encryption keysWho controls the master key and its rotation?KMS architecture, per-project data keys, BYOK policy, and revocation test.
OperationsWhat availability, backup, recovery, and data-location commitments exist?Published commitments, status history, backup tests, and contract language.

5. Evaluation plan

Run a narrow evaluation that can fail for a named reason.

Pick one workflow with an observable action. Do not start with a broad “remember everything” requirement.

  1. Define scope. One project, one user identifier scheme, one action, one operational system.
  2. Define allowed memory. List durable facts, forbidden data, source types, confirmation rules, and validity windows.
  3. Integrate the client action flow. Recall, remember, evaluate action, confirm when required, re-evaluate, run the host action, and report execution.
  4. Run boundary attacks. Cross-user replay, foreign project key, poisoned transcript, conflicting fact, stale fact, duplicate retry.
  5. Trace one action. From conversation turn to memory candidate to decision to receipt to external reference.
  6. Pin a memory baseline. Save deterministic recall and action cases, pass one suite, then prove a deliberate memory change is reported as a regression.
  7. Verify operations. Revoke the key, stop a dependency, inspect readiness, and verify a signed audit export.

Pass criteria

Suggested pass criteria

Isolation foreign project + user

Foreign project and user identifiers never return a scoped resource.

Policy contested | policy-untrusted ≠ act

Confirmation is one trust signal, not the only path to act. Contested or policy-untrusted evidence cannot produce act. The host branches on act, ask, or abstain.

Execution decision ID + terminal receipt

Every tested consequential action has a decision ID and terminal receipt.

Idempotency same key + changed payload → conflict

Changing an idempotent retry payload returns a conflict instead of another write.

Regression pinned baseline → evidence + outcome diff

A pinned Memory CI suite catches changed evidence IDs or action outcomes.

Audit content-free export → offline verify

Audit export entries contain no customer content and verify offline.

Approval missing control → named blocker

The team can name which missing controls block production approval.

6. ContextDB mapping

How the current hosted alpha maps to the checklist.

Current capability maturity and explicit limitations
Maturity Capability Current scope Qualification
Available in Alpha Memory and action boundaries
  • Org, project, and user partitions
  • Source, confidence, confirmation, and stable slots
  • Recall and separate pre-action evaluation
  • Decision IDs and execution receipts
  • Idempotent protected mutations
  • Scoped memory, slot, and verified partition erasure
Hosted behavior remains Alpha and requires host enforcement.
Enterprise Alpha Operator and audit controls
  • Owner, operator, and viewer roles
  • Invites and scoped API-key lifecycle
  • Metadata-only signed webhooks
  • Entitlement-gated organization audit log
  • Ed25519-signed export with real entries
An export verifies its signed record, not objective truth or downstream state.
Hosted and Private Alphas Formation, Memory CI, and Managed Sources
  • Durable asynchronous Formation submit and poll
  • Encrypted propose and idempotent commit results
  • One dedicated single-instance Formation process beside the gateway on the same app VM
  • Hosted Alpha Memory CI with a PostgreSQL queue, leases, cancellation, and pinned baselines
  • One dedicated single-instance Evals process beside the Console API on that same VM
  • Available public Memory CI package, public contract, CLI source, and pinned reusable Action
  • Postgres source checkpoints, retries, and DLQ
One instance of each runs on one VM. This does not establish scale, high availability, failover, or a worker fleet.
Planned Controls not available today
  • SSO and SCIM
  • KMS, key rotation, and BYOK
  • Private connectivity and IP allowlists
  • Residency and contractual retention
  • Availability and recovery commitments
No delivery date or production commitment is stated.

ContextDB has no compliance certifications or uptime commitments today. See the security page and current control table before beginning an evaluation.

Sources

Research and external frameworks used in this guide.

Use the checklist on one real workflow.

Bring the architecture, action, data boundaries, and controls your review requires.