This page documents current ContextDB Memory CI behavior. For
vendor-neutral methodology across formation, retrieval, lifecycle, actions,
and release gates, read the
AI agent memory evaluation guide.
Memory CI availability
Distribution / public
CLI and Action
Available
Public distribution: Available (Apache-2.0).contextdb-memory-ci==0.1.0a1
was published through PyPI trusted publishing with attestations, and
its exact registry artifact passed a fresh isolated install, import,
version, and --help check. The reusable Action is public at
memory-ci-action-v0.1.0.
Memory CI uses deterministic assertions. It does not ask an LLM to judge
its own output.
01 / Author
Write deterministic cases
Choose recall or action mode and add at least one text, evidence-ID, or outcome assertion.
recall · action · assertions
02 / Baseline
Pin one passing run
The suite pin is explicit and an enqueued run snapshots the selected baseline.
passing run → pinned baseline
03 / Enqueue
Create one queued run
A required idempotency key returns the same run for an identical request and 409 for conflicting reuse.
POST /evals/v1/.../runs → 202
04 / Execute
Claim, lease, and terminalize
The dedicated Evals process records attempts, heartbeats, bounded case results, and a terminal row for every snapshot.
queued → running → terminal
05 / Gate
Return the release result
The CLI maps deterministic behavior and operational state to a closed exit contract.
unchanged 0 · regressed 1 · error 2
Deterministic cases
State what must appear, what must not, and whether the agent may act.
Each case has a stable user partition, a recall or action query, result
bounds, and deterministic text, evidence-ID, or act, ask, or abstain
assertions.
Trace-to-eval creates a draft with no assertions. An operator must add
at least one deterministic assertion before the case can join a suite.
Catch the old memory ID if it comes back after a correction.
An UPDATE returns the new memory ID and the IDs it superseded. Put the
new ID in expected_evidence_ids and the old ID in
forbidden_evidence_ids. A DELETE regression test can
forbid the deleted ID.
This checks which memory grounded the result, not whether a model
happened to paraphrase the expected words.
after UPDATE
required: mem_new
forbidden: mem_old
recall evidence: [mem_new]
passed
recall evidence: [mem_old]
failed · stale memory returned
See how Memory Evolution returns prior and
deleted IDs for this test.
Hosted execution
A PostgreSQL queue sits behind the public API.
One dedicated single-instance Evals process executes beside the Console
on the same app VM. Queue, snapshot, attempt, lease, case-terminal, and
baseline state remain durable in PostgreSQL.
Idempotent enqueuesuite + baseline + key → one run
Submit with an Idempotency-Key. The same request returns
the same run with 202. Conflicting reuse returns
409.
One named Console worker, now separatedqueued → running → terminal
One separate worker process claims a durable row, records an attempt,
renews its lease, and obeys case, suite, result, and job deadlines.
Bounded recoveryworker_lost → retry_wait | error
An expired lease retries only while attempt and deadline budgets
remain. Exhausted work becomes an explicit error.
Synthetic demoMemory CI, Hosted Alpha synthetic fixture. This is not an availability claim.
Cancellation
Cancellation is cooperative and visible.
A queued run can reach cancelled before dispatch. An
active run can remain cancelling until in-flight work
reaches a safe stopping point.
The public API returns safe status for enqueue, poll, and cancel. It
never returns suite names, case names, queries, assertions, raw
failures, or memory content.
Project-boundone project · one eval_automation purpose
A cbe_ token belongs to one project, is shown once, and is
stored hashed. Keep it in CI secret storage.
Four operations onlyenqueue · poll · cancel · export
The token can enqueue, poll, cancel, and export under
/evals/v1. It cannot call memory routes or author
Testbench cases and suites.
Pinned baselines
Omit the baseline to use the suite pin.
A baseline must be a passing run of the same suite. Enqueue snapshots
the selected baseline. Changing the suite pin later does not rewrite an
existing run.
Comparison checks case status, evidence IDs, and action outcome. The
closed states are no_baseline, unchanged,
and regressed.
request body
{} # use pinned baseline
{"baseline_run_id": "run_0"}
# use this passing run
{"baseline_run_id": null}
# run without baseline
Content-free exports
Export behavior without copying customer conversations.
Keep the cbe_ evaluation token in GitHub Secrets. Project
and suite IDs can use GitHub Variables. The Action installs exactly
contextdb-memory-ci==0.1.0a1. See the
public Action release.
Exit 0passed + unchanged
Passed and unchanged, or an intentional no-baseline pass.
Exit 1failed | regressed
Failed behavior or a regression.
Exit 2operational terminal
Operational error, cancellation, timeout, malformed response, output failure, or no baseline by default.
Omit baseline-run-id to snapshot the suite's pinned passing
baseline. The Action uploads content-free JSON and JUnit artifacts by
default. They omit names, queries, assertions, raw failures, credentials,
and memory content. Shell and metadata tests cover the Action wrapper, and
the packaged CLI has live registry proof. No Action run inside GitHub
against production is claimed.
The actual packaged CLI returned exit 0 for unchanged behavior.
A deliberate memory regression returned exit 1.
Queued cancellation reached cancelled.
Trace-to-eval created no assertions and preserved PII processing.
JSON, JUnit, and Evals control-state scans contained neither the synthetic marker nor the proof token.
Cleanup removed every Memory CI proof resource, including PostgreSQL rows, credentials, secrets, keys, memories, decisions, users, organizations, and projects.
The proof killed the Evals process with SIGKILL and observed systemd start a new PID.
After advancing only the synthetic lease expiry, the new owner recorded worker_lost → passed across all 50 cases.
No attempt or case remained active. JSON and JUnit stayed content-free.
The first attempts exposed gateway throttling. Bounded Retry-After handling and project-key caching were added before the passing run.
Cleanup removed the GCP secret and every mutable synthetic row.
The process proof advances a synthetic lease rather than waiting for its
natural wall-clock expiry. It does not establish scale, high availability,
multi-VM failover, crash-resume availability, customer-data backup/restore,
scheduling, a worker fleet, an SLO, or human usability.
Memory CI FAQ
Memory CI questions.
Does Memory CI use an LLM judge?
No. It evaluates expected text, forbidden text, and action outcomes with deterministic assertions.
Does an action case execute the business action?
No. It evaluates memory policy through the real gateway but does not call your booking, payment, refund, or account tool. ContextDB advises. The customer host enforces. The evaluation stops before the business action.
Can I install the CLI from PyPI?
Yes. Install contextdb-memory-ci==0.1.0a1 from PyPI. It is an alpha client for a Hosted Alpha service with no availability SLA.
Is the reusable GitHub Action available?
Yes. Pin atomsai/contextdb-clients@memory-ci-action-v0.1.0. The public prerelease installs the exact alpha package. Hosted execution remains Hosted Alpha with no availability SLA.
Does a worker failure prove the service resumes?
A bounded same-VM proof killed the dedicated process and observed a new PID finish all 50 cases after the synthetic lease was advanced. That proves this process-restart path, not natural lease timing, multi-VM failover, or an availability commitment.
Pin one passing run before your next memory change.
Create the case in the Console, then use the public API contract from CI.