Sentinel Signal

MCP Verify — Verify Trust Data Feed, Release 1 (schema/export foundation)

Source: docs/mcp-verify-trust-data-feed-release1.md

Document Content

MCP Verify — Verify Trust Data Feed, Release 1 (schema/export foundation)

New commercial surface: Verify Trust Data Feed, a bulk-dataset-licensing product distinct from the per-request Verify Intelligence API. Instead of querying one server at a time, qualified customers (registries, agent platforms, security vendors, enterprise data teams) will eventually download a daily current-state snapshot and a daily change/delta feed as Parquet/JSONL.gz artifacts.

This is Release 1 of a 4-release plan (see the architecture plan for the full design and the three remaining releases). Release 1 has no customer-visible surface — the projection, schema, export job, and storage all exist and are exercised by tests and a production dry run, but no authenticated route yet exposes any of it. Releases 2-4 (scheduling already shipped alongside Release 1 here; auth/entitlement download routes and the marketing page remain) are tracked separately.

13 new tests added (6 projection, 7 export). Full suite green at the usual ~20 pre-existing/unrelated failures, no new ones.

What shipped

Schema (intelligence/data/schemas.py)

Two versioned, extra="forbid" Pydantic record types:

  • VerifyTrustRecordV1 (verify.trust_data_feed.snapshot.v1) — one row per
  • exportable server: identity/registry fields, trust score/risk/confidence, evidence counts, freshness state, deployment readiness, capability flags, fingerprints, and snapshot_as_of.

  • VerifyTrustChangeRecordV1 (verify.trust_data_feed.change.v1) —
  • near-1:1 with the already-persisted M2 change taxonomy (IntelligenceTrustChange): change_type, severity, field, previous_value/current_value, material, evidence_reference.

Three deliberate naming deviations from the original commercial spec's literal suggested field names (see projection.py's module docstring for the full reasoning):

  1. first_observed_at → **verify_first_tracked_at** (sourced from
  2. Server.created_at) — no column tracks "first seen live on the internet," only "first ingested by Verify." The spec's literal name would have overclaimed.

  3. capability_fingerprint → omitted. No such field exists in the
  4. domain model; the two real fingerprint columns (remote_endpoint_fingerprint, server_card_fingerprint) are exported verbatim instead.

  5. verify_scoring_version → **aliased from Server.public_score_version**
  6. — no standalone scoring-engine version constant exists anywhere in this codebase; this is the closest real "which scoring logic produced this" marker already attached to each server.

canonical_snapshot_id was not carried over from the spec's suggested field list — it's an internal bookkeeping id never exposed to real Intelligence API customers either (confirmed: intelligence/monitor.py's own HTTP-response parser always reads it as ""), so a commercial dataset shouldn't invent exposing it.

Known, documented gap (pre-existing, not introduced by this feature): IntelligenceTrustChange.verify_build has never actually been populated at its one call site (intelligence/service.py's _classify_and_persist_trust_changes). Every change-feed row will show verify_build: null until that separate bug is fixed — documented in the schema's own docstring so it isn't reported as a new defect in the feed.

A small, additive API enrichment (intelligence/schemas.py, intelligence/service.py)

Added deployment_readiness: str | None = None to the existing, live IntelligenceServerTrust model (part of IntelligenceServerResponse, GET /v1/intelligence/servers/{namespace}/{name}). Backward-compatible (optional, defaulted, no schema-version bump needed per the feed's own versioning rule) — sourced from the same canonical verdict-gate code already shown on the public server-detail page (build_production_readiness's code).

This exists because the export job (below) fetches trust facts via HTTP from this same endpoint, and deployment_readiness wasn't previously part of its response. All 51 existing Intelligence tests re-verified passing after this change.

Projection (intelligence/data/projection.py)

project_server_to_trust_record / project_change_to_record — explicit allowlist functions, one field at a time, never **dict/**model_dump() passthrough. This is the single place commercial exportability is defined; an unrelated new column added to Server later can never silently leak into a paid, versioned dataset. is_exportable_server reuses the same canonical/opted-out/hidden-server predicates ServerRepository already applies elsewhere (_is_canonical_server, _hidden_server_filters) rather than inventing a new "exportable" flag.

Export job (intelligence/data/export.py, workers/export_dataset.py)

  • Trust facts are fetched over HTTP, one call per server, to this same
  • service's own authenticated Intelligence Trust API using the internal service token — the same architectural choice intelligence/monitor.py's tick already makes, and for the same reason: a bare python -m worker process has no access to the redaction/suppression closures inside create_app(), so calling the real endpoint is the only way to guarantee this can never diverge from what a real customer's read would show.

  • Bounded batch iteration over the full exportable corpus (offset-loop,
  • ordered by id, same shape as ServerRepository.list_ranked_servers but walking every match instead of stopping at a target rank count) — never one unbounded SELECT *.

  • Formats: JSONL.gz (stdlib gzip) and Parquet (new pyarrow
  • dependency) for the current snapshot; JSONL.gz only for the change feed.

  • Idempotent publish flow: generate → validate → SHA-256 checksum →
  • upload → persist as status="validated" → flip to "published" in the same final commit. A retry for an already-published (dataset, schema_version, format, snapshot_as_of) is a no-op, not a second publish. An empty current-snapshot result raises (never publishes zero rows — that means every fetch failed, not that Verify tracks zero servers); an empty change-delta window is valid and published as-is.

  • **No is_latest pointer column.** "Latest" is derived at read time as
  • MAX(snapshot_as_of) among status="published" rows for a given (dataset, format) — collapses the generate→publish→update-pointer flow into one step, removing a whole class of "forgot to update the pointer" bug.

  • Change-feed catch-up: the delta window's period_start is bounded by
  • the previous successful artifact's snapshot_as_of, not a hardcoded "yesterday" — a missed day self-heals instead of silently dropping changes.

Storage (intelligence/data/storage.py)

Narrow ObjectStore protocol (put/exists/signed_download_url/ metadata); nothing outside this module imports boto3 directly. S3ObjectStore (real, boto3-backed) and InMemoryObjectStore (test double, no real MinIO/boto3 needed — matches the rest of the suite's in-memory-only posture). Backend is self-hosted MinIO, added as a new minio service in docker-compose.yml — no new vendor relationship; the client code is plain S3-compatible so it would work against real AWS S3/R2 later without changes if adoption ever justifies moving off MinIO.

One real bug found by testing against an actual local MinIO container rather than assuming: ServerSideEncryption: "AES256" fails on MinIO without a configured KMS ("KMS not configured for a server side encrypted objects"), unlike real AWS S3 where SSE-S3/AES256 needs no KMS. Removed — the real access controls for this MVP are the private bucket, scoped IAM credentials, and short-lived signed URLs, not encryption-at-rest.

Database

New table data_export_artifacts (migration 0038_data_export_artifact, 25-char revision id, well under the 32-char alembic_version.version_num guard from the A26 incident class). One row per physical file — a logical daily export produces two rows (Parquet + JSONL.gz). No organization_id — artifacts are one shared corpus-wide object per day; entitlement will be enforced at the route layer (Release 3), never by row ownership. File contents live in MinIO, never in Postgres.

Scheduling

Two independent systemd timers (snapshot failures never block delta, and vice versa), following the exact docker compose run --rm --no-deps ... verify-web python -m mcp_verify.workers.<module> pattern the existing analytics-partition-cutover step already uses in deploy.sh — not the bare-host-Python coherence-monitor pattern, which has zero DB/ORM access and can't run this job:

  • sentinel-signal-data-snapshot.{service,timer} — daily at 03:15 UTC.
  • sentinel-signal-data-delta.{service,timer} — daily at 03:45 UTC (after
  • the snapshot, generous gap).

Both installed by deploy.sh the same way the existing coherence-monitor and proof-strip-snapshot timers are (install -m 0644 + systemctl daemon-reload && systemctl enable --now).

Also fixed: a .gitignore bug

The repo's .gitignore had an unanchored data/ rule meant to ignore a top-level runtime-data directory (/data/synthetic-100k, etc.) but which also matched any directory literally named data anywhere in the tree — including this feature's own intelligence/data/ source package, which would have silently never been committed. Fixed by anchoring it to /data/.

Production hardening after the first real-corpus test (v1.0.741+)

The first manual production run of export-verify-dataset.sh snapshot against the real corpus (93,917 exportable servers at the time) exposed problems no test against a synthetic corpus could have caught:

  • Fully serial fetch does not scale. The original implementation made
  • one synchronous HTTP call per server. Against 90k+ servers this ran for 9+ minutes with zero completion and no way to tell if it was healthy or stuck (no per-server progress signal). Fixed by materializing the exportable-server list on the main thread (SQLAlchemy sessions aren't thread-safe), then fanning fetch_trust_facts out across a bounded ThreadPoolExecutor (Settings.data_feed_export_concurrency, default 10 — deliberately conservative, not maximized, because verify-web runs as a single unscaled process that also serves real customer traffic).

  • A full-corpus run legitimately takes on the order of an hour or more,
  • not minutes — the bottleneck is per-server HTTP round-trip latency (measured ~0.5s per warm-cache request against the live Trust API) times ~94k servers divided by the bounded concurrency, and most servers are cold on first touch during a nightly run (a risk flagged in the original architecture plan before this was ever run for real). This is fine for a once-nightly job with no downstream time pressure, but means a smoke test needs a long observation window, not a short one.

  • No per-server progress output made it impossible to distinguish
  • "slow but healthy" from "stalled" during that long window. Fixed by logging one JSON progress line every 1,000 completed servers (and a final one), each with completed/total/fetch_errors/elapsed_seconds/ estimated_seconds_remaining — enough to see the run advancing without 90k+ lines of per-server noise.

  • **timeout did not actually stop the container it was bounding.**
  • export-verify-dataset.sh ran docker compose run as a plain child process of the wrapper bash script. When an operator wrapped an interactive invocation in timeout 900 ... and it fired, timeout's SIGTERM landed on bash, which does not forward signals to a foreground child it's waiting on by default — so the underlying container kept running, orphaned, well past the timeout that was supposed to bound it. Separately, an interrupted local SSH session was found to leave the same kind of orphaned container running remotely, undetected, concurrently with a later retry — doubling load on the single-process verify-web until found and killed manually. Fixed the root cause: the script now execs into the final docker compose run invocation, so the script's own process is the process Compose uses to manage the container's lifecycle, and Compose's own (correct) SIGTERM/SIGINT-forwarding-to- container behavior actually applies.

Production hardening after the first full-corpus completion (v1.0.742+)

A first full run was allowed to complete unattended (~2–2.5 hours). It published zero rows, and surfaced a second, more consequential problem than the runtime itself:

  • **The export job's own concurrent fetches exhausted verify-web's DB
  • connection pool.** Each of the job's bounded HTTP fetches (data_feed_export_concurrency, default 10) calls back into verify-web's own Intelligence API, and each of those calls needs a pool connection for its DB query. Layered on top of normal traffic, this repeatedly hit the process's 15-connection ceiling (pool_size=10 + max_overflow=5), producing sqlalchemy.exc.TimeoutError — and, more importantly, a handful of real customer-facing 500s/503s on unrelated routes sharing the same pool. Fixed by raising verify-web's pool to 20 + 15 (35 total) in docker-compose.yml — Postgres itself was never close to its own limit (max_connections=120, real usage never observed above ~30), so this is pure headroom, not a database capacity change.

  • No log survived to explain the zero-row outcome. The run was
  • triggered by hand over raw SSH (timeout 900 export-verify-dataset.sh snapshot, bypassing systemd), and the container ran with --rm; once it exited, its logs were gone before they could be inspected. This is not a gap in the real production path — the systemd service units (sentinel-signal-data-{snapshot,delta}.service) have no StandardOutput= override, so systemd's default journald capture applies and persists independently of the container's own removal. The lesson is procedural, not code: trigger test/manual runs via systemctl start sentinel-signal-data-<dataset>.service (and journalctl -u ... -f to watch), never a raw docker compose run over SSH, so a failure's actual cause is never lost again.

  • **Nothing prevented two instances of the same export from running at
  • once.** The nightly timer fired mid-run during that same test (a manual run and the real scheduled run overlapped), and would have again 30 minutes later when the delta timer fired. export-verify-dataset.sh now runs each dataset's container under a fixed, predictable name (verify-data-feed-{snapshot,delta}-export) and checks for an already-running instance under that name before starting — skipping (exit 0, not a failure) rather than piling on. This covers every trigger path (timer, manual, or both at once), not just systemd's own single-instance-per-unit behavior, which a raw out-of-band invocation bypasses entirely.

The actual root cause of the zero-row outcome (v1.0.743+)

Redeploying the pool-size/overlap-guard fix re-enabled both timers as a side effect of deploy.sh's (unconditional, idempotent) install step, which immediately fired the delta job to catch up a missed window (Persistent=true). It failed instantly with a clean, fully-captured traceback — proof the journald-capture reasoning above was correct — and revealed the real bug: boto3.exceptions.S3UploadFailedError: ... NoSuchBucket. The MinIO bucket was never created. docker-compose.yml stood up the minio service and the application-side S3ObjectStore client, but nothing anywhere ever provisioned the bucket itself. This is almost certainly what actually killed the first full-corpus run: it likely fetched real data successfully for ~2 hours, then hit this same unhandled S3UploadFailedError at the final publish step — an exception publish_generated_files doesn't catch (it isn't an ExportError), so nothing was published and, at the time, nothing was logged either (the logging gap above). The DB pool exhaustion was real and worth fixing, but it was not what caused zero rows.

Fixed with a minio-init service (docker-compose.yml, mirroring the existing postgres-init pattern exactly): a one-shot minio/mc container, gated on minio's healthcheck, running deploy/ionos/minio-init/ create-data-feed-bucket.sh (mc mb --ignore-existing, idempotent). It runs as part of every normal docker compose up during a deploy, so the bucket is guaranteed to exist before any export job — scheduled or manual — could ever run.

A third, independent bug found by the first fully-supervised run (v1.0.744+)

With the bucket fixed, a real supervised run (systemctl start sentinel-signal-data-snapshot.service, journald-watched throughout) got all the way through fetching ~86,000 servers over ~2h47m — and then still failed at the final publish step: sqlalchemy.exc.InternalError (psycopg.errors.IdleInTransactionSessionTimeout).

Root cause: run_current_snapshot_export opens session once at the top to materialize the exportable-server list (iter_exportable_servers), then never touches it again until the very end, at the final publish step — but the fetch loop in between takes hours against the full corpus. Production's idle_in_transaction_session_timeout is 60 seconds (MCP_VERIFY_DB_IDLE_TRANSACTION_TIMEOUT_MS); Postgres kills the connection long before the loop finishes, and the session doesn't notice until it's used again — discarding every successfully-fetched record in the process.

Fixed by ending the transaction immediately after materializing the server list: session.expire_on_commit = False (so the already-loaded columns project_server_to_trust_record depends on survive the commit instead of being marked stale — the default would raise DetachedInstanceError on the next line instead, since there'd be no live transaction left to reload from), then session.commit() and session.expunge_all(). The final publish step opens a fresh transaction on a healthy pooled connection, only seconds before it's used. A regression test (test_run_current_snapshot_export_ends_transaction_before_fetch_loop) asserts session.in_transaction() is False from inside the fetch call itself, proving the property directly rather than depending on a real 60-second wait.

This was the third of what turned out to be four independent bugs found in sequence by real production testing, each one only reachable after the previous one was fixed. No amount of unit testing against synthetic, fast-completing data would have surfaced any of them — they only exist at real corpus scale and real multi-hour duration.

A fourth bug, right at the finish line (v1.0.745+)

With all three fixes live, a fully-supervised run got all the way through: 86,236 servers fetched, both files (Parquet + JSONL.gz) generated and successfully uploaded to MinIO — real proof the whole pipeline works end to end. It still failed on the very last statement: sqlalchemy.exc.DataError (psycopg.errors.StringDataRightTruncation): value too long for type character varying(32).

DataExportArtifact.schema_version was defined as String(32), but VERIFY_TRUST_RECORD_SCHEMA_VERSION = "verify.trust_data_feed.snapshot.v1" is 34 characters — a plain off-by-a-few-characters mismatch that had been there since Release 1 was first built. Every previous run had failed even earlier in the pipeline (pool exhaustion, missing bucket, idle transaction), so this exact INSERT had simply never executed for real before. SQLite's permissive typing (no VARCHAR length enforcement at all) meant the test suite could never have caught it either — every test insert of this exact value had silently succeeded against the in-memory SQLite fixtures.

Fixed with migration 0039_data_export_schema_widen (String(32) → String(64), verified for real against a throwaway local Postgres container, both upgrade and downgrade) plus a new regression test (test_schema_version_constants_fit_the_db_column) that checks the actual constant lengths against the real column width directly, independent of ever inserting a row — closing the exact blind spot that let this ship in the first place.

The two files from the failed run (real, valid, checksummed) are still sitting in MinIO with no corresponding data_export_artifacts row — never visible to any reader (nothing queries storage directly, only published DB rows), so not a correctness risk, just harmless orphaned storage from a run that failed one statement short of done.

First fully clean production run (v1.0.745)

With all four fixes live, a fully-supervised run against the real production corpus completed end to end: 86,538 of 87,086 servers (99.4% coverage) processed in 2h28m48s, 548 fetch errors (0.63%, individually-unreachable third-party endpoints, not systemic — verified zero customer-facing 5xx and a healthy DB pool throughout). Both formats published: Parquet (9,032,125 bytes) and JSONL.gz (6,196,400 bytes), 86,538 rows each, valid SHA-256 checksums. The delta job was also manually triggered the same way and published cleanly (see below for why it returned zero rows).

A fifth bug: change detection never actually ran (v1.0.746+)

The change_delta dataset published successfully but always with record_count: 0 — not a bug in the export job itself, but in what feeds it. IntelligenceTrustChange (the M2 table this dataset reads from) is written by _classify_and_persist_trust_changes, called from exactly two places: the /v1/intelligence/servers/{ns}/{name}/history endpoint, and the webhook monitor tick — which only evaluates webhook-watched servers. Production had one watch subscription and one snapshot ever recorded, corpus-wide. The main Intelligence read endpoint (GET /v1/intelligence/servers/{ns}/{name}) — the one the Trust Data Feed export job already calls once per canonical server, every day — never called capture_trust_snapshot_if_changed at all, so nothing was driving change detection across the corpus independent of real customer traffic, which barely exists yet for this pre-launch product.

The fix adds capture_trust_snapshot_if_changed directly into the export job's own per-server fetch loop (run_current_snapshot_export), right after fetching each server's trust facts — reusing traffic the job already generates rather than adding a new corpus-wide scan. This was tried first as a change to the read route itself (so any caller would drive capture), but that broke test_monitor_tick_detects_tool_count_only_change_via_http_round_trip: the monitor tick fetches facts via the same route and does its own capture call afterward, so both capturing at once meant whichever ran first "claimed" the baseline and the other saw no change — a real race, not a test artifact. Capturing from the export job's loop instead avoids that overlap entirely.

A second, subtler issue surfaced while testing this fix: fetch_trust_facts (the export job's HTTP parser) never parsed tool_count/transport_type/ has_oauth/has_dcr out of the response, because project_server_to_trust_record reads those straight off the local Server row instead. But compute_trust_fingerprint — the function that decides whether anything changed — does hash those four fields, so tooling/auth/transport changes would have been silently invisible to this new capture path even though score/risk changes worked. Fixed by extending fetch_trust_facts to also parse the response's capabilities block, matching what monitor.py's own _facts_from_trust_api_response already does for the same reason.

New regression test: test_current_snapshot_export_drives_change_detection_across_the_corpus runs the export twice against one seeded server, changes its tool_count between runs, and asserts a second IntelligenceTrustSnapshot and a tooling_changed IntelligenceTrustChange row both get created — proving the export loop now actually drives change detection, not just that it doesn't crash. Capture failures are caught per-server (change_capture_errors, surfaced in progress logs and the final result) so one bad row can never abort a multi-hour corpus export.

This closes the loop for the change_delta dataset, which will start reporting real content starting with the next scheduled run rather than being permanently, silently empty.

Deferred to later releases (per the architecture plan)

  • Release 3 (shipped, see mcp-verify-trust-data-feed-release3.md):
  • entitlement scopes (intelligence:data_snapshot, intelligence:data_delta, intelligence:data_redistribution) and the three authenticated routes (catalog, artifact list, signed-URL download).

  • Release 4 (shipped, see mcp-verify-trust-data-feed-release4.md):
  • /verify-data marketing page, admin visibility endpoint, and a customer-facing data dictionary.

Configuration reference

New environment variables (deploy/ionos/prod.env.example): DATA_FEED_S3_ENDPOINT_URL, DATA_FEED_S3_BUCKET, DATA_FEED_S3_ACCESS_KEY_ID, DATA_FEED_S3_SECRET_ACCESS_KEY, DATA_FEED_S3_REGION, DATA_FEED_SIGNED_URL_TTL_SECONDS, MINIO_DATA_DIR. The access key/secret pair doubles as MinIO's own MINIO_ROOT_USER/ MINIO_ROOT_PASSWORD.

Manual invocation for operators: deploy/ionos/export-verify-dataset.sh <snapshot|delta> (same env-file-loading + docker compose run pattern as backup-postgres.sh).