Sentinel Signal

MCP Verify — Audit Round 16 Remediation: Implementation

Source: docs/mcp-verify-audit-round-16-remediation-implementation.md

Document Content

MCP Verify — Audit Round 16 Remediation: Implementation

Shipped 2026-08-11 (v1.0.547). Implements the plan in docs/mcp-verify-audit-round-16-architecture-plan.md against the "MCP Verify — remediation requirements (audit round 16)" doc (11 Aug 2026 audit, target build v1.0.546). A1-A2/A4-A13 implemented; A14/A15 remain blocked on R11 (labeled dataset) per the audit's own scope guard and are untouched. 518 tests passing (up from 495 at the start of this round).

A13 — version audit

Confirmed not a caching/deploy bug: VERIFY_SITE_VERSION was already the single source of truth, applied uniformly across every page. The actual finding was two rendered prose sentences on /docs/scoring-specification naming a historical planning effort "the post-1.0.503 remediation plan"/"post-1.0.503 review" — reworded to drop the version-number-shaped token (main.py, SCORE_COMPONENT_ZERO_POINT entry + render_scoring_specification_page's "Subscore zero points" section). Added test_no_public_page_renders_a_stray_version_looking_token, a regression sweep over 9 pages asserting no v?1.0.\d+-shaped token appears other than the current version. That test's own first run caught two more instances (post-1.0.508) that manual review missed — fixed the same way.

A1 — schema_divergence_probe

Card-shape detection already existed (name-only vs. object-shaped card tools); the bug was narrower than the audit assumed — only the output-schema comparison (service.py) was unguarded by the same shape check the parameter dimensions already used, so a name-only card entry (card_tool is None) unconditionally read as "declares no output schema," flagging any tool whose live schema did declare one as diverging. Fixed by gating the output-schema comparison on card_tool is not None, same as the parameter dimensions. Alert/remediation wording (insights.py) now names the dimension(s) that actually diverged instead of always saying "parameters," and reserves "clients will construct invalid calls" language for genuine input-schema divergence.

A2 + A3 — one root cause, three separable fixes

  • A2 §3a (ingestion): ingest/registry_client.py's _extract_mcp_endpoint_urls now excludes glama.ai + /mcp/servers/ path shape outright (a directory listing page, not a transport endpoint) — the same exclusion mechanism already used for three other known-non-endpoint hosts.
  • A2 §3b (status/verdict propagation): compute_summary_status (validation/service.py) now checks for initialize's non_mcp_content_type marker before the generic initialize-failure branch, returning unverified instead of failing. Reused the existing unverified status rather than inventing a new one — infer_server_verdict (db/repository.py) already routes any non-"failing" zero-tool server to metadata_only, so no verdict-routing change was needed, only the status-computation change. Verified live-code fan-out: profile headline, /report, incident feed (reads summary_status directly), badge SVG, and both /trust-summary paths all auto-corrected because they read current_status/summary_status directly; client-compatibility verdicts and benchmark task labels were left as-is (a client genuinely cannot use an unreachable server, regardless of whose fault that is — not an adverse claim against the publisher). Caught and fixed a self-introduced regression: build_active_alerts's registry_endpoint_invalid alert originally required summary_status == "failing" — after the status fix, that condition can never be true again for the case the alert exists to describe, making it permanently unreachable. Relaxed to gate on non_mcp_content_type alone.
  • A2 §3c (rendering conflation): build_provenance_divergence_detail (insights.py) no longer overwrites the probe's own status/drift_fields with cross-registry alias-consolidation data. The alias disagreement is now a separately-attributed field (alias_source_disagreements/alias_disagreement_severity), rendered as its own labeled row in the same section — preserving the original R43.2 goal (never show "OK" directly above a real identity split) without one computation impersonating another's result.

A4-A8 — three distinct "unknown reads as pass" bug shapes

Not one shared root cause, per research: (1) an accepted-status set that wrongly included "missing" (A4: compatibility fixtures; A5: connector_publishability_probe's six criteria, which also had an is None-as-ready loophole); (2) a bare unfiltered count (A6: live_check_count = len(checks) → now filters to {"ok","warning"}); (3) an existing fix pattern (validation_has_observed_tool_surface) applied to one function but not its sibling (A7: build_public_server_reputation's 7d/30d ratios now use the same filter build_history_summary already used). A8 combined two fixes: build_evidence_confidence now hard-caps the score below the medium threshold when history_depth == 0, and the "Recent validation runs" table (main.py) now marks rows that don't count toward the "N recent validations" figure, rather than showing an uncorrelated larger row count with no explanation.

A9 — client-compatibility status/blocker consistency

build_client_readiness_verdicts (insights.py) computed its status badge from only the narrow criteria checklist while its own reason text (fixed in R77.1) already read from the full blocker list (criteria ∪ policy gates ∪ request-association/transport-compliance extras). Now computes the full blocker list once and derives both status and reason from it — finishing the R77.1 fix, which explicitly left status alone (confirmed via an existing test's own comment, "R76 is the status fix, not this test," now updated).

A10a/A10b — remediation addressee and severity

Added a structured addressee: "publisher" | "registry" | "verify" to every remediation-row construction site in build_remediations and to every alert in build_active_alerts (insights.py); the alias_endpoint_split and registry_endpoint_invalid rows/alerts are the only registry-addressed cases today. extract_next_actions and render_remediations (main.py) now filter out non-publisher rows, so a registry-addressed finding no longer appears as an instruction the publisher cannot act on — it stays fully visible, correctly attributed, in the Active Alerts section. For severity: rather than touching production_readiness's critical_alerts count (which gates the published verdict — Safe for production/evaluation — and was explicitly out of scope to change per the audit's own "don't change classifier conclusions" guard), the displayed "High/critical-severity alerts" stat (render_production_readiness) now takes max(alert count, remediation-table critical/high count), so the number shown can never disagree with the table shown beside it, without altering what gates a verdict.

A11 — score suppression

build_validation_timeline and build_server_incident_feed (insights.py) both called compute_current_score directly, bypassing public_display_score's suppression gate (no score for failing status or zero-tool servers). The incident feed's latest-run entry now calls public_display_score(server) directly (available in scope); the timeline, which only has per-run data (not the Server object), now applies the same two suppression conditions (summary_status == "failing" or unknown/zero tool count) per-run rather than off current server state — correct for a historical timeline, where a past run's own outcome shouldn't be hidden or revealed by what the server is doing today.

A12 — ranking population parity

Extracted score_percentile's cohort query (db/repository.py: fresh public_display_score recompute over non-failing, tool-bearing, validated servers) into _evidence_comparable_fresh_scores/evidence_comparable_scored_server_count, and pointed the homepage's scored_servers coverage stat (main.py) at the same query instead of a raw, unfiltered current_score IS NOT NULL column count. The two numbers can no longer disagree by construction. Two existing test fixtures had to be corrected (current_score set with empty current_score_components, which real production data never has since one is always derived from the other) — not a fixture-compatibility shim, a fix to unrealistic test data.

Left open

  • A14, A15: blocked on R11 (labeled dataset) per the audit's explicit scope guard. Current-state grounding recorded in the architecture plan for whoever picks up R11.
  • §1 standing constraints: could not locate a verbatim, checked-in copy of "round 15 §8" anywhere in the repo — proceeded on the audit doc's own 6 reconstructed rules. Flagged in the architecture plan as needing the owner to supply the authoritative text.
  • Open questions 1-4 (A2 new-state-vs-metadata_only, A10b severity direction, A12 population correctness, A1 catalog-wide retraction count): resolved via the recommendations already stated in the architecture plan (reuse unverified; display-only union rather than changing verdict gating; ranking pool's stricter definition wins; catalog retraction count to be computed and reported once this ships and the corrected probe reruns across the live catalog).