MCP Verify — Round 15 Implementation Report (post-1.0.516 verification)
Implements the full "MCP Verify — Requirements, Round 9" doc, found against two fixtures at v1.0.516: awesome-gpitrella/memxus-remote-mcp (write surface, action_safety warning, alias endpoint split, no server card) and flevy-com/kpi-depot-mcp (read-only surface, action_safety ok, real server card without inputSchema). Preceded by a planning-only pass (see verify_round15_architecture_plan in memory), which flagged an evidence discrepancy in the doc's R50 citations — investigated first, resolved with hard data from direct production DB access, documented below. All items shipped (R48, R49, R50-diagnosed, R51, R39). 403 tests passing, up from 397 at the start of this round.
Standing order (§0): classifier freeze
No new capability-classifier rules landed this round.
R50 — Diagnosed first, per the doc's own instruction
Direct production DB queries (not just the public API) settled this conclusively:
- **The doc's specific citations for memxus/kpi-depot don't exist in the
- The underlying phenomenon was real — for a different server,
- Round 13's R28 fix is confirmed working: that same server's next
- **
validate_server_recordis structurally atomic**, confirmed both by
database.** Full validation history for both servers (10 runs, 7 runs — confirmed complete, not truncated) shows every single run at the full tool count. No zero-tool row has ever existed for either server.
github-depixapp/depix-mcp: 3 confirmed zero-tool "healthy" runs, timestamped 2026-08-04T12:16, 2026-08-05T19:32, 2026-08-06T07:32 — all before round 13's deploy (17:46:38Z the same day). The doc's cited kpi-depot Aug 06 07:32:52 is essentially the same timestamp as depix-mcp's last bad run (07:32:34) — almost certainly a fixture mixup in whatever produced the doc's evidence table, not a live gap in either cited server.
validation, 2026-08-06T19:33:42 (after the deploy), correctly shows summary_status="unknown" instead of "healthy" — the exact behavior R28 was built to produce. Corpus-wide, zero zero-tool-"healthy" rows exist with a completed_at after the fix deployed.
code reading (one transaction, stub-insert-then-final-update, single commit) and by direct inspection (no status='running' or orphaned summary_status='unknown' rows exist anywhere in the table). R50.1's literal "reader sees a partial write" hypothesis does not hold.
R50.2: not implemented — there was no partial-write bug to fix.
R50.4: implemented regardless of the above (the plan's own recommendation) — build_evidence_confidence now excludes summary_status == "unknown" runs from history_depth's count, so a history padded with zero-observation runs no longer inflates confidence.
Also found, not fixed, flagged for a future round: while running the diagnostic queries, found a genuinely live pattern of sequential (not concurrent) re-validations of a failing server ~2.5 seconds apart, happening in production right now, apparently bypassing the 4-hour failing-server backoff floor (compute_failing_server_backoff_hours). Confirmed the round-14 partial unique index is correctly in place and enforced (ix_jobs_validate_server_active_unique exists exactly as migrated) — this is a different mechanism than either R27 race already fixed. Out of this round's authorized scope; recorded here rather than investigated further.
R48 — R41/R48 regressed; the masked branch, fixed and tested
Round 14 consolidated the threshold across safe_to_publish, the unsafe_for_write_actions verdict, and write_safe mode, but never broadened the input set to include active_alerts — confirmed live on kpi-depot (action_safety_probe: ok, 2 active high-severity alerts): safe_to_publish/unsafe_for_write_actions said "safe"/"allowed" while write_safe mode alone said "blocked."
R48.1 — build_write_action_governance now accepts active_alerts (already computed earlier in build_server_insights, no reordering needed) and folds a has_blocking_active_alert check (severity in {critical, high}, matching write_safe mode's existing round-13 threshold) into safe_to_publish.
R48.2 — added a parametrized cross-product test ({action_safety: ok, warning, error} × {alerts: absent, present}, 6 cells) asserting safe_to_publish, the unsafe_for_write_actions verdict, and write_safe mode agree in every cell — the test that would have caught round 14's gap immediately, since round 13/14's own fixtures never exercised the alerts-present cell.
R48.3 — write_action_governance now returns blocking_reasons: list[str], generated once alongside safe_to_publish itself; unsafe_for_write_actions's reason field now cites it instead of a generic two-sentence toggle.
R49 — Split "card omits schema" from "card contradicts live"
Root cause in build_schema_divergence_probe: card_tools[name].get("inputSchema") if isinstance(...) else {} made "the card never declared a schema" indistinguishable from "the card declared an empty schema" — the subsequent set-difference comparisons always "found" a divergence, and has_breaking_divergence's check tripped on key presence ("parameters_only_in_card" in item), not value truthiness, so an empty list still counted as a breaking finding. Confirmed live on kpi-depot exactly as the doc describes, and — found while fixing this — the same bug was already baked into an existing test (test_r31_schema_divergence_probe_against_a_third_independent_fixture, the ai.dreamlit/mcp fixture) as an asserted, expected result; updated to assert the corrected behavior instead.
R49.1/R49.2 — per-tool: if the card tool object has no inputSchema key at all, the parameter/required/type dimensions are recorded as an omission (tools_with_omitted_schema), never compared against an implicit empty schema. output_schema_presence (a boolean declared/not-declared check, not a properties diff) is unaffected — still a genuine, always-comparable dimension. Severity gained a "low" tier for omission-only findings, distinct from "medium"/"high"/"critical" contradictions; build_active_alerts now emits a separate, low-severity card_omits_tool_schemas alert instead of the client-breaking server_card_schema_drifted text when that's the only finding.
R49.3 — the deeper, more consequential bug: build_remediations's generic loop fired REMEDIATION_RULES[check_name]'s fixed text for any non-ok status, including missing/not_assessed — confirmed live on memxus (server_card: error, schema_divergence_probe: missing, reason no_server_card_tools), which still carried the full "disagree... construct invalid calls" remediation. REMEDIATION_RULES entries now carry a fires_on_absence: bool per entry (presence-type — absence itself is the finding — vs. comparison-type — absence means nothing was compared). Classified all 25 entries directly (10 presence-type, unchanged; 12 comparison-type, now correctly suppressed on missing/not_assessed — schema_divergence_probe, provenance_divergence_probe, connector_replay_probe, tool_snapshot_probe, prompts_list, resources_list, session_resume_probe, request_association_probe, interactive_flow_probe, action_safety_probe, determinism_probe, probe_noise_resilience; reasoning for each recorded inline in the code).
Standing instruction: re-ran the "unknown reads as pass" sweep (round 6's original ask) specifically for the inverted shape this round introduced ("unknown reads as conflict"). One additional speculative candidate found (tool-snapshot-diff functions comparing inputSchema presence across two time-separated validation runs of the same live server) — flagged as lower-confidence/different bug class, not fixed this round; recorded below.
R51 — Two more "gates that ignore the verdict"
Hosted runtime readiness (build_runtime_readiness): confirmed zero awareness of active_alerts — memxus showed allowed_to_activate: true under a Block-for-Production verdict with 2 active high-severity alerts. Now accepts active_alerts and blocks on the same severity threshold R48.1 established, threaded through all three call sites (a new _active_alerts_for_server helper in main.py, since two of the three call sites are POST endpoints that provision/release a deployment, not server-detail rendering with insights already computed).
**official_registry_probe unrelated peers**: root cause was a SQLAlchemy footgun, not a logic bug — Server.server_card_url == server.server_card_url compiles to IS NULL when the right-hand Python value is None (confirmed: kpi-depot's server_card_url is None), silently matching any other official-registry server that also never populated that field. Each OR clause in the peer-matching query now only applies when the server's own value is actually non-empty.
R39 — The fifth request, delivered as a genuine empirical matrix
Re-verified live first: neither of the doc's cited "absent from breakdown" examples held up against current code/fresh data — schema_divergence_probe deliberately has no score component at all (round 10's original design, unchanged), and provenance_divergence_score's absence on memxus is because its live status is not_assessed (not error as the doc states), which is already-correct, already-established exclusion behavior. The underlying deliverable — a generated, CI-enforced mapping table — was still genuinely incomplete relative to what's been asked five times, independent of whether today's specific fixtures demonstrate a live bug.
New ml/registry/generate_score_status_matrix.py: unlike round 13's narrower AST-derived audit (which only asks "can this component exclude itself, and only for not_assessed"), this one actually executes each in-scope score_* function once per status, with a synthetic checks dict setting that component's own AST-derived source check(s) to each of the seven statuses in turn, and classifies the real return value (excluded / zero / credited:<value>). Scoped deliberately to the 27 score_* functions whose parameters are check-status-driven (a subset of {checks, server, tools, tool_inventory, metadata_documents}) — excludes functions needing session (live DB queries, 2 components) and time-series/statistical functions taking historical_runs (8 components, where "status" isn't a single-check axis). This scope boundary matches every example the doc has ever cited across all five requests, all of which are check-status-driven. verify/tests/test_score_status_matrix.py asserts regeneration matches (the same stale-detector pattern as round 13's registry) plus pinned assertions for the doc's named components.
Found while building this, not fixed, flagged for a future round: transport_compliance_score — one of the three genuinely not_assessed-only components — credits 1.0 (not a full zero) on missing/error/skipped/auth_required status, via elif details.get("bad_protocol_status_code") is None: score += 1.0 inside its post-ok/not_assessed fallthrough branch. Not a violation of R39.1's exclusion rule (it correctly never returns None outside not_assessed), but a small, genuine "absence read as a mildly positive signal" instance — worth a look in a future round, not named in any of the five prior requirements docs.
Test suite
403 tests passing (PYTHONPATH=verify/src pytest verify/tests), up from 397 at the start of this round.
Deferred / carried forward
The newly-discovered live-in-production sequential-revalidation pattern (R50, above) and the transport_compliance_score minor credit-on-absence finding (R39, above) are both flagged, not fixed, out of this round's authorized scope. The one speculative "unknown reads as conflict" candidate from the re-run sweep (tool-snapshot-diff functions comparing schema presence across two time-separated live runs) is recorded but not investigated further.