Current scope: Postgres has operated validation,
preview, backfill, incremental sync, checkpoint, retry, and soft-delete
behavior. Supabase uses the same Postgres-protocol adapter. No separate
production proof is claimed. BigQuery passed a bounded synthetic proof
against a real temporary dataset on 2026-08-26. Snowflake and
Databricks have extensive no-network contract tests only. None has
customer warehouse proof or public access.
A source sync should resume after failure instead of starting over.
01 / Validate
Prove the connection and mapping
The job verifies TLS or provider routing, pinned credentials, one relation, and every mapped column.
credentials · relation · mapping
02 / Preview
Inspect without writing memory
Preview returns bounded PII-processed mapped records and leaves the memory store unchanged.
source rows → PII-safe preview
03 / Backfill
Traverse stable tuple pages
The frozen job config bounds records, query work, response size, wall clock, and concurrency.
updated_at + primary_key
04 / Write
Keep one logical ingestion identity
Each source key and version maps to an idempotent memory write or soft deletion.
source key + version → memory result
05 / Checkpoint
Persist only after record tasks settle
A mutating job cannot succeed while a record remains pending or running, or without a checkpoint cursor.
settled tasks → durable cursor
06 / Increment
Resume from the last checkpoint
Manual or scheduled runs read later updates, propagate mapped soft deletions, and route exhausted records to review.
checkpoint → update | delete | DLQ
Control state
Credentials, jobs, and writes keep separate identities.
Credentials stay sealedpinned Secret Manager version
Source and memory credentials are written to GCP Secret Manager as
pinned versions. The console returns source metadata, never plaintext
credentials.
Every job has a named statequeued | running | retry_wait | terminal
PostgreSQL stores sources, jobs, attempts, records, checkpoints,
ingestion identity, and dead-letter entries. Workers claim bounded
leases and synthesize worker-loss attempts after expiry.
Memory writes are idempotentsource key + version → one logical ingestion
A source record carries a stable source key and version. Restarting a
backfill does not create another logical memory for the same version.
Provider status
What each provider has proved.
Provider maturity and bounded evidence as of 2026-08-27
Source
Current state
Evidence
Postgres
Private Alpha
Bounded operated proof on 2026-08-20: live validation, preview, backfill, update, soft-delete, checkpoint, and cleanup
Supabase
Private Alpha
Uses the Postgres protocol and durable worker contract. No separate production proof is claimed.
Snowflake
Private Alpha
OAuth/PAT SQL API, strict account routing, polling, paging, and partition tests. There is no real warehouse proof.
BigQuery
Private Alpha
Bounded synthetic proof on 2026-08-26 from the production app VM against a real temporary BigQuery table. There is no customer warehouse proof or public access.
Databricks
Private Alpha
Production-bounded SQL Warehouse reader and no-network contract tests. There is no real warehouse proof or public access.
MongoDB
Planned
No source reader is implemented
Warehouse reader contracts
Each sync job has an explicit query budget.
Frozen warehouse query and result boundaries
Provider
Relation and credential
Per-job cap
Fail-closed handling
Snowflake
Strict account routing through the SQL API with OAuth or a programmatic access token.
Polling, paging, response, and wall-clock budgets are frozen with the job.
Only no-network contract evidence exists. No real warehouse proof is claimed.
BigQuery
One strict project.dataset.table. Hosted multi-tenant projects accept a service-account key only.
One dry-run and one actual query under the configured maximum_bytes_billed, capped at 10,000 records.
The billed-byte ceiling, response, page, poll, and wall-clock budgets fail before checkpoint advance.
Databricks
One catalog.schema.table through the SQL Warehouse Statement Execution API. OAuth M2M is preferred and a PAT is the fallback.
One bounded INLINE JSON statement, 10,000 records, and an 8 MiB aggregate result budget.
Databricks rejects external result links and truncated results. An over-budget result fails the job.
Polling and mutations
Your timestamp is the change contract.
Every provider advances with a stable timestamp-and-key cursor. BigQuery
and Databricks stop each read at a five-minute safety lag. Polling is
manual by default. Optional schedules cannot run more often than every
15 minutes. Private-alpha BigQuery and Databricks jobs are capped at
10,000 records.
Every source mutation, including a soft deletion, must advance
updated_at at commit and become visible within the safety lag.
Map a soft-delete timestamp when rows can be removed. Mutations committed
later than that contract can fall behind the cursor. Hard-deleted rows are
not detected.
Mapping source records
Map business records without pretending they are conversations.
A source maps one relation, primary key, user partition, content field,
update timestamp, and optional deletion timestamp. Business records
remain in the source. ContextDB stores the customer context selected for
agent memory.
Selected source records remain evidence, not authorization. If that memory
informs a booking, refund, or account change, ContextDB advises and the
customer host enforces the pre-action result.
Synthetic demoManaged Sources, Private Alpha. Postgres shows the operated path. BigQuery shows a bounded synthetic path. This is not customer warehouse proof.
Failure handling
Inspect a failed record instead of losing an entire sync.
Bounded retriesretryable → retry_wait | failed
Transient source and gateway failures retry within a frozen attempt and wall-clock budget.
Dead-letter reviewexhausted record → project-scoped DLQ
Permanent or exhausted record failures enter a project-scoped DLQ and can be requeued explicitly.
Pause without deletinglive → paused → resume from checkpoint
Operators can pause future synchronization while preserving source metadata, history, and checkpoints.
On 2026-08-26, the production app VM read two synthetic rows from a temporary BigQuery dataset and table in the ContextDB GCP project. The reader passed its dry-run byte gate, validation, actual query, key and timestamp mapping, and one soft-deletion row. The dataset and proof artifacts were deleted.
The BigQuery run is bounded synthetic provider evidence. It is not customer warehouse, scale, worker-crash, scheduled-sync, availability, or SLO evidence.
AvailabilityPrivate Alpha
Contact us to evaluate Managed Sources with your warehouse schema.
Managed Sources FAQ
Managed Sources questions, answered directly.
Does ContextDB replace my database or warehouse?
ContextDB does not replace your system of record. Managed Sources selects mapped context for agent memory and keeps provenance back to the source.
Can I connect Snowflake today?
Not as a publicly available connector. The contract-tested implementation can be evaluated with a Private Alpha design partner, but no real customer warehouse proof has passed.
Can I connect BigQuery now?
Not through a public or self-serve connector. The reader has a bounded synthetic proof against a real temporary BigQuery table, but it can be configured only for a Private Alpha design-partner evaluation and has no customer warehouse proof.
How do I set up BigQuery, and what are the limits?
Provide one strict project.dataset.table mapping, a service-account key for hosted multi-tenant use, mapped primary key, user, content, updated_at, and optional deleted_at columns, plus a maximum bytes billed value. Each job runs one dry-run and one actual query, uses a timestamp-and-key cursor with a five-minute lag, is capped at 10,000 records, and cannot detect hard deletes.
Can I connect Databricks now?
Not through a public or self-serve connector. The SQL Warehouse reader can be configured only for a Private Alpha design-partner evaluation, and it has no real customer warehouse proof.
How do I set up Databricks, and what are the limits?
Provide the workspace host, SQL Warehouse ID, one catalog.schema.table mapping, OAuth M2M credentials or a PAT, and mapped primary key, user, content, updated_at, and optional deleted_at columns. Each job uses one bounded INLINE JSON statement, is capped at 10,000 records, fails closed above an 8 MiB result budget, rejects external links, applies a five-minute lag, and cannot detect hard deletes.
How often do Managed Sources poll?
Runs are manual by default. An optional schedule can run no more often than every 15 minutes. Every source mutation must advance updated_at so the timestamp-and-key cursor can observe it.
Are credentials visible in the console?
No. Credentials are written to Secret Manager and are not returned after source creation.
Does hard-delete reconciliation exist?
No. Map a deletion timestamp for soft-delete propagation. A row removed from the source without that timestamp is not detected.
Does ContextDB automate warehouse secret rotation or deletion?
No. Secret rotation and deletion automation remain open. Real-provider proof also remains open for Snowflake and Databricks. BigQuery has bounded synthetic provider proof, not customer warehouse proof.
Bring one customer-context table into a private-alpha project.
Start with a PII-safe preview, a bounded backfill, and an explicit checkpoint.