Technical evaluation and procurement review
Review an agent before it touches production systems.
Start with one consequential action. Inspect the memory partition, evidence, policy result, host enforcement, receipt, erasure path, and operating limits before deciding whether the workflow may advance.
remembered evidence × policy → act | ask | abstain → host-reported receipt
- Hosted service status
- Cloud The hosted path is available for bounded evaluation. Production readiness is not claimed.
- Enforcement boundary
- ContextDB records an advisory decision. The customer host must authenticate, authorize, enforce the outcome, and execute or stop the external action.
- Evaluation unit
- One named workflow, one action owner, one system of record, and explicit pass, condition, and rejection criteria.
Enterprise rollout questions
Start with the questions that can stop the rollout.
Each question links to a review object, a current capability boundary, or a gate that can return yes, no, or proceed with conditions.
Decision review
One path from memory to a host-enforced action.
An act result records what policy concluded from the
evidence. It does not authorize or execute the downstream action. The
customer host separately applies business authorization and remains the
enforcement point.
-
01 / Authenticate
Resolve the actor and current business state
The host authenticates the user, authorizes the request, and reads the system of record.
actor + authorization + current state -
02 / Partition
Select one bounded memory space
The server maps the authenticated user to an organization, project, and user partition.
org_id · project_id · user_id -
03 / Evaluate
Apply policy to action-relevant evidence
The decision stores an outcome, reason, policy version, evidence IDs, and request ID.
recall_for_action → act | ask | abstain -
04 / Enforce
Stop or continue in the customer host
The host may continue on
actand must pause or stop onaskorabstain.host check → execute | pause | stop -
05 / Record
Attach the host-reported result
A receipt records succeeded, failed, or skipped. It does not independently inspect downstream state.
POST /v1/receipts -
06 / Review
Reconstruct the boundary
Review the request, decision, evidence, policy alignment, actor, and external reference together.
decision_id → evidence → receipt
{
"decision_id": "synthetic-decision",
"outcome": "ask",
"reason": "action-relevant memory awaits confirmation",
"policy_version": "default",
"evidence_ids": [],
"pending_confirmation_ids": ["synthetic-memory"],
"request_id": "synthetic-booking-review",
"receipt": {
"action_name": "appointment.book",
"status": "skipped",
"policy_alignment": "aligned",
"external_ref": null
}
}
Pilot operations
Inspect activity before expanding the test boundary.
The aggregate view shows project activity and organization-scoped active profiles. It does not establish decision quality, availability, or readiness for a broader rollout.
Incident review
Reconstruct the decision, then verify the host report separately.
The review joins a decision with its evidence snapshot, policy result, actor, and host-reported receipt. A receipt is an attestation from the integration. It is not independent evidence from the downstream system.
Current capability boundary
A reviewer should not have to reverse-engineer what is live.
The rows below identify what is currently implemented, its maturity, and the limit a reviewer should carry into a pilot decision.
| Capability | Status metadata | Implemented inventory | Review boundary |
|---|---|---|---|
| Memory and data | Cloud |
|
One deployment encryption key is used for memory content today. The system of record remains authoritative. |
| Action controls | Cloud |
|
ContextDB advises. The customer host enforces the decision and remains responsible for the external action. |
| Human and server access | Cloud |
|
Project keys are server credentials. A caller-supplied
user_id is a partition key, not authentication.
|
| Audit and service operations |
Cloud Enterprise alpha audit export |
|
The live audit chain is not WORM storage, an external transparency log, or a certification. |
| Developer and evaluation surface |
SDK available Hosted Alpha Memory CI |
|
The bounded Memory CI proof passed on 2026-08-23. Hosted execution has no availability commitment. |
| Formation | Hosted Alpha |
|
The ADD-only operated proof is dated 2026-08-21. The same-VM worker restart proof is dated 2026-08-24. |
| Managed Sources |
Private Alpha Postgres private-alpha sync and a Supabase protocol path |
|
Postgres proof is dated 2026-08-20. BigQuery synthetic provider proof is dated 2026-08-26. No provider has customer warehouse proof. |
| Integration and account foundation | Reference and sandbox paths |
|
Reference flows do not indicate a vendor partnership. Billing is shadow or sandbox infrastructure, not live charging. |
Open control register
Treat roadmap requirements as blockers until they exist.
Every status in this table belongs to its row. None is an enabled control, contractual promise, or current service property.
| Control | Status metadata | Requested capability | Current distinction |
|---|---|---|---|
| SSO and SCIM | Planned · Not implemented | SAML and OIDC single sign-on, SCIM user and group provisioning, service accounts, workload identity, and project-scoped roles. | Current human access uses email verification, password plus emailed OTP, team invites, and three organization roles. |
| KMS, key rotation, and BYOK | Planned · Not implemented | KMS-backed envelope encryption, unique data-encryption keys per project, customer-managed keys and BYOK, key rotation, revocation, and crypto-shredding. | Memory content uses one deployment key. The backup bucket uses an operator-managed GCP KMS CMEK. Neither provides customer-managed project key custody or BYOK. |
| Private network and data location | Planned · Not implemented | Private connectivity, customer IP allowlists, data-location options, and contracted retention and deletion policies. | No private connection, customer allowlist, region list, or residency contract is offered. |
| Memory lifecycle automation | Planned · Not implemented | Explicit expiry and retention-policy APIs, scheduled retention and deletion verification, consolidation and stale-memory review jobs, and explicit lifecycle events in audit exports. | Current deletion is request-driven and bounded. Larger partitions require an operator workflow. |
| Policy promotion and large audit export | Planned · Not implemented | Real policy canary routing and promotion gates, signing-key rotation and historical verification keys, asynchronous large audit exports, and SIEM delivery and export history. | Current audit export is synchronous, owner-only, entitlement gated, and bounded. |
| Control-plane and gateway scale | Planned · Not implemented | Postgres metadata storage and managed migrations, shared rate-limit state, multiple gateway workers, and a multi-VM topology. | Dedicated Formation and Evals processes run on the same VM as the current gateway. The bounded restart proof is not a worker fleet. |
| Availability and recovery commitments | Planned · Not implemented | High-availability topology, multi-VM failover, published recovery objectives, contracted support, and availability commitments. | No high-availability or multi-region topology is claimed. No uptime SLA, RPO, RTO, or recovery commitment is offered. |
| Expanded data platforms | Planned · Not implemented | Real customer Snowflake, BigQuery, and Databricks warehouse proof, automated source-secret rotation and deletion, a MongoDB source adapter, decision and receipt export, and continuous hard-delete reconciliation. | BigQuery has bounded synthetic provider evidence. Snowflake and Databricks have contract tests only. Hard-deleted source rows are not detected. |
| Commercial account operations | Planned · Not implemented | Reconciled live billing, enterprise contract and privacy workflows, custom support terms, and account-level usage and export controls. | Current Stripe work is sandbox-only, and the billing ledger is shadow infrastructure. |
| Expanded evaluation reporting | Planned · Not implemented | Extraction precision and conflict reporting, repeated-question and stale-memory indicators, scheduled and durable CI suite execution, and policy-violation and confirmation-resolution reporting. | Current Memory CI uses deterministic assertions against a pinned baseline and content-free result exports. |
| Contracts and certifications | Not available | Certification reports, regulated-data addenda, and contracted assurances. | No compliance certifications are offered today. No BAA is offered today. |
Pilot gates
Run one workflow through architecture, security, and incident review.
A gate needs an owner, evidence, and a recorded result. A demo or screenshot cannot substitute for a failed control requirement.
-
Gate 01 / Entry
Name the action and owner
Identify the booking, refund, update, dispatch, or record write and the person accountable for errors.
pass: action + owner + failure cost -
Gate 02 / Boundary
Define the memory contract
List durable facts, allowed sources, confirmation rules, user partition mapping, and prohibited data.
pass: scope + source + retention decision -
Gate 03 / Enforcement
Make the host prove it stops
Require a decision ID before the tool runs. Exercise
askandabstain, then attach a receipt.pass: no side effect after ask | abstain -
Gate 04 / Adversarial
Attack scope, freshness, and deletion
Replay foreign IDs, use stale or poisoned context, omit confirmation, retry writes, revoke keys, and verify erasure.
pass: fail closed + no cross-scope mutation -
Gate 05 / Operations
Review worker and backup limits
Compare the dated same-VM restart and content-restore evidence with the workflow's actual recovery requirements.
pass: every accepted gap has an owner -
Gate 06 / Decision
Record yes, no, or conditions
Approve only the bounded workflow and environment reviewed. Do not generalize one pass to the service or another action.
result: proceed | condition | reject
Evidence register
Match each claim to its narrowest proof.
Functional tests, synthetic UI fixtures, and bounded operated runs answer different questions. This evidence does not establish scale, high availability, a certification, or a service-level commitment.
v0.4.4 · trust_policy.md · test_trust_model.py
The public policy and executable evals cover evidence classes, confirmation, contest state, injection state, PII-before-embed, and verifiable forgetting. The customer host still enforces the result. [C1]
POST /v1/receipts · screenshot release 20260827
Receipt persistence requires idempotency and records status, policy alignment, action name, and external reference. The screenshots show a synthetic review object, not independent downstream verification. [C2]
POST /v1/forget · memory | slot | user partition
Whole-partition erasure requires typed confirmation, idempotency, a synchronous cap, and immediate residue verification. Scheduled retention and large asynchronous erasure remain open. [C2]
SIGKILL → new PID → synthetic lease advance → terminal result
Systemd restarted the dedicated Formation and Evals processes under
new PIDs. Evals reached worker_lost → passed for 50
cases. This does not establish scale, high availability, natural
lease timing, or a worker fleet.
[C4]
signed PostgreSQL + SQLite bundle → disposable PG16 drill
One exact bundle passed signature, hash, row-count, and supported audit-head checks. Daily backup and weekly drill timers are enabled. The path excludes roles, secret payloads, PITR, and failover. No uptime SLA, HA topology, region list, RPO/RTO, or latency numbers are claimed or published. [C5]
Named responsibilities
Approval belongs to people with different evidence.
Put names beside these roles for the selected workflow. ContextDB cannot assume the customer's authentication, authorization, risk acceptance, or downstream verification duties.
| Role | Owns | Evidence to bring | Decision |
|---|---|---|---|
| Head of AI Platform | Reference architecture, partition mapping, credentials, host enforcement, and integration tests. | Data flow, threat boundaries, failure tests, key-revocation test, and versioned deployment record. | Whether the shared pattern is technically acceptable for this workflow. |
| Security, Privacy, and AI Governance | Prohibited data and actions, identity requirements, retention rules, evidence review, and accepted gaps. | Control requirements, threat model, deletion test, audit sample, and written exceptions. | Approve, reject, or condition the evaluated boundary. |
| Workflow owner | Allowed action, current system of record, customer impact, manual fallback, and business acceptance criteria. | Representative fixtures, expected act/ask/abstain outcomes, error cost, and escalation path. | Whether the workflow result is useful and safe enough under the approved conditions. |
| Service owner or SRE | Deployment, monitoring, webhook deduplication, incident response, backup review, and recovery runbooks. | Readiness checks, alert routes, receipt reconciliation, dependency inventory, and recovery requirements. | Whether the operating boundary meets the selected environment's requirements. |
Sources and citations
Check the source before carrying a claim into procurement.
Product sources establish current code, route, fixture, or bounded operating evidence. Industry sources explain why buyers ask the questions. They do not validate ContextDB.
- C1 / Public SDK v0.4.4
- Trust policy and trust-model evals. These are the public definitions and executable checks for the trust boundary.
- C2 / Product and API boundary
- Product system, trust review, and API reference. These sources describe routes, decision fields, receipt semantics, host enforcement, and current non-claims.
- C3 / Managed Sources proof
- Managed Sources evidence and limits. Postgres bounded proof is dated 2026-08-20. BigQuery bounded synthetic provider proof is dated 2026-08-26.
- C4 / Memory CI proof
- Memory CI evidence and limits. The bounded run is dated 2026-08-23. Dedicated worker restart evidence is dated 2026-08-24.
- C5 / Service operations
- Current service status and security boundary. These pages state the 2026-08-24 worker and content-restore evidence, plus the current limits. No uptime SLA, HA topology, region list, RPO/RTO, or latency numbers are claimed or published. PITR and complete disaster recovery are out of scope.
- C6 / Synthetic UI release
-
Screenshot release
20260827. The analytics and reports figures above are synthetic fixtures. They show review surfaces, not customer outcomes or an availability result. - R1 / Okta-commissioned survey / 2026
- Enterprise buyer survey. AlphaSights surveyed 150 IT and security decision-makers. The report says 69% cited security concerns as slowing adoption and 49% called audit trails critical for production readiness.
- R2 / PwC survey / 2025
- AI agent survey. PwC surveyed 308 US executives. Twenty percent selected financial transactions among tasks they trusted agents to handle.
- R3 / OWASP / 2026
- OWASP Top 10 for Agentic Applications 2026. ASI06 covers memory and context poisoning that can affect later reasoning, planning, and tool use.
Review one bounded workflow.
Bring the action, system of record, control requirements, and named owners. Leave with a pilot decision and an explicit gap register.