Sentinel Signal

MCP Verify Round 3 implementation notes

Source: docs/mcp-verify-round-3-implementation.md

Document Content

MCP Verify Round 3 implementation notes

Date: 2026-08-04

This documents the implementation of the Round 3 corrections for MCP Verify. Full diagnostic detail for the two P0s lives in docs/verify-build-scoped-route-cache.md (TASK-08b) and inline in verify/src/mcp_verify/validation/service.py / verify/src/mcp_verify/db/repository.py (TASK-31). This note summarizes what changed and, where the plan's assumptions turned out to be wrong, what the data actually showed.

Implemented

  • TASK-08b: root-path cache. Root-caused to a manually-maintained MCP_VERIFY_BUILD_SHA env var with no script enforcing update order, plus a confirmed dead-code regression (route-cache prewarm writes using pre-build-scoping key formats since the build-scoping commit, so every prewarm write was silently orphaned). Fixed the prewarm key-scoping bug and replaced the manual SSH/sed runbook with deploy/ionos/deploy.sh, which derives the build SHA from git automatically and fails the deploy if the post-deploy route-version smoke check doesn't pass. See docs/verify-build-scoped-route-cache.md for the full diagnosis.
  • TASK-31: status/score/verdict coherence.
  • Diagnosed Degraded against production data (74,804 scored servers, 2026-08-04) instead of assuming the spec's "fires on ~everything" framing: Degraded is actually rare (22 servers, 0.03% of the corpus), and among those, the OAuth-metadata-mismatch check dominates (19/22 corpus-wide; 8/8 of the degraded servers visible in the top 25 by score) -- not auth_required (3/22, 0/25). compute_summary_status's branches were not redefined as a result; they were already a meaningful, rare signal.
  • Found and fixed the actual bug: two independent verdict implementations (db/repository.infer_server_verdict, used for ?verdict= filtering, and insights.build_production_readiness, used for the verdict text on listing/shortlist rows) had diverged -- the latter had no tool_count guard and let status="degraded" straight into safe_for_evaluation unconditionally. build_production_readiness now delegates its verdict code to infer_server_verdict; the two can no longer disagree.
  • Fixed the score floor: ~8 of the ~50 scoring dimensions awarded mid-to-high "safe" credit specifically when there was nothing to inspect (zero tools), conflating "inspected and found safe" with "nothing was inspected." These now distinguish the two cases (using the existing tool_inventory each function already receives) and award low credit for the uninspected case.
  • Added an explicit status-derived score penalty (Degraded -8, Failing -20, applied after the component average, not diluted as 1-of-50 dimensions) and a 0.6x ceiling multiplier for zero-tool servers, applied in compute_current_score.
  • Extended /methodology with a "Status, score, and verdict" section documenting the model.
  • Added scripts/report_status_distribution.py as a standing diagnostic for re-checking the status-branch distribution as the corpus evolves.
  • TASK-29b: metadata-only exclusion. The true default-home path (default_home_servers) already excluded tool_count=0 servers; the gap was the general/filtered listing path (ServerRepository.list_servers / _list_servers_fast) and the JSON /v1/servers API, which had no such default. Added an include_metadata_only parameter (default True at the repository level, so existing API/internal callers are unaffected) and wired the HTML index() route to pass include_metadata_only=False by default, with a new ?show_metadata_only=1 opt-in (the existing ?all_servers=1 escape hatch also implies it, preserving its prior "show everything" meaning). Added a per-namespace cap (max 3 consecutive rows) applied after ordering/dedup in the default, filtered, and commerce-filtered listing branches.
  • TASK-28b: result-count reconciliation. Replaced the two different header strings ("N default candidates out of M indexed" vs. "N servers in current result set") with one consistent "Showing N of M indexed servers" -- N and M are different measurements by design now that metadata-only exclusion and namespace capping apply broadly, so the header states that explicitly instead of implying equality.
  • TASK-13 (partial): Removed the trustsnap_ snapshot id from list-view rows (still present on server-detail/report/policy/badge/compare surfaces). Added an "Evidence age" column. Not done this round: a persisted Risk column (destructive/exec/high-risk tool counts) -- this needs a new Server migration and was judged out of scope for this pass given the schema-change risk on the production catalog; the counting logic already exists inside the relevant score_* functions and just needs extracting into a shared, persisted helper.
  • Quick win: relabeled the nav "Sign in" link (which pointed at /claim, an ownership-claim flow, not authentication) to "Claim/Manage".

Not done this round

  • TASK-06 (percentile): needs a cached score-distribution helper (keyed by the existing catalog_fingerprint invalidation signal) to avoid an N+1 query pattern in the listing route; scoped out to avoid rushing new caching plumbing immediately before a production deploy.
  • TASK-12, TASK-11, TASK-19, TASK-18/20/22-25: unchanged from Round 1/2 backlog, per the spec's own "Everything else is cleanup" framing for this round.

Verification

PYTHONPATH=verify/src pytest verify/tests   # 233 passed
PYTHONPATH=verify/src python3 -c "import mcp_verify.main"

Production status-distribution query used to ground TASK-31 (run via docker compose exec postgres psql, results included above) is reproducible going forward with scripts/report_status_distribution.py.

Deploy verification (post-implementation)

Deployed via deploy/ionos/deploy.sh on 2026-08-04. First deploy (commit 3722dff, v1.0.499): /, /pricing, /trust-index, a server-detail page, /methodology, and /status all reported matching v1.0.499 / build 3722dff6da20ee3d7d93b8ccd4d0ec947d2e7f8f for both body and X-MCP-Verify-Build header -- confirmed independently outside the smoke script too, comparing / against /?x=1 directly. This is the first time the bare root path has matched query-string variants across a deploy in this project's history per the TASK-08b writeup.