MCP Verify — Product architecture gap-assessment, Track 1 implementation
Target release: 1.0.536.
Source: the section-55 gap assessment against "Verify — Product Architecture & Development Plan" (a 56-section product/UX doc, not part of the numbered R-round requirements series). That assessment found two live correctness bugs matching this engagement's recurring bug classes ("unknown reads as pass", "verdict divergence across surfaces") plus real ownership-status duplication, and recommended fixing them as a small, contained "Track 1" before any of the larger homepage/IA/pricing work in the doc's own Phase 1-10 sequence. This document covers Track 1 only.
Implemented changes
- Write-action governance "not_assessed" gap:
action_statuscan be
literally "not_assessed" (build_action_safety_probe, validation/service.py, fires when tool_inventory is empty and tools_list never actually succeeded) — a server whose write-action safety was never checked, not a server confirmed clean. write_action_governance.safe_to_publish (insights.py:902) excluded error/warning but not not_assessed, so it computed True anyway, rendering "Safe to publish" beside a visibly "Not_Assessed" governance status. Separately, the high-risk-tools table (render_write_action_governance, main.py) rendered "No high-risk tools were detected on the latest run" for the exact same unassessed case — an empty list reads identically whether the probe ran and found nothing or never ran at all. Fixed both: safe_to_publish now excludes not_assessed, blocking_reasons names it explicitly, and the table now shows "High-risk tool assessment unavailable -- no tools were analyzed." specifically for the unassessed case, never conflating it with a genuinely clean result.
- **Verdict divergence —
build_executive_verdict**: this function's own
prior comment already named it "a third, independent verdict/decision computation from build_production_readiness/infer_server_verdict" and only patched the owner_opted_out special case to agree with the other two computations. The general case still recomputed its own blocker threshold (stale evidence, current_status == "failing", score < 50, a substring match against production_code that a "failing" code wouldn't even trip, since "failing" contains none of "block"/"remediation"/"metadata_only"). decision (Block for production / Allow with approval / Allow for production) is now derived directly from the one canonical production_readiness.code (already present on the same payload, computed by build_production_readiness, which already folds in status/score/freshness/alert-downgrade logic); write-like/high-risk/unauthenticated factors may now only ever tighten an otherwise-allowed canonical state from "Allow for production" to "Allow with approval," never independently override the canonical Block boundary.
- **Verdict divergence —
_server_summary_dict**: feeds the
/compare/{slug} comparison-template page (build_comparison_template_payload → "servers": detail_payloads, rendered directly by render_comparison_template_page/render_compare_page). It computed production_readiness.code via the same canonical infer_server_verdict gate (correct — and deliberately alert-blind, matching the same batch-listing-cost tradeoff already established and documented for infer_server_verdict itself: no per-row active_alerts/evidence_confidence fetch at catalog-listing scale), but derived the label independently via code.replace("_", " ").title() — producing "Safe For Production" (title case) on the comparison-template page for the same server whose profile page, badge, and rankings all correctly show "Safe for production" (sentence case, the one canonical string). Extracted PRODUCTION_READINESS_CODE_LABELS (insights.py, alongside build_production_readiness, which now also reads from it instead of restating each label string inline) as the single source for this text; _server_summary_dict now reads from it too.
- Ownership/claim-status duplication: claim status was independently
re-derived in several places with different fallback rules. One raw-SQL site (main.py, a listing-item builder) reimplemented the claimed/verified/status logic from scratch instead of using the existing _claim_verified_from_value helper. A different site (render_owner_funnel) fell back to "unknown" for a present-but- status-less claim record, where every other surface's established fallback (compare_service._claim_status) is "claimed" — there IS a claim, its status field just isn't populated; a genuine, reader-visible inconsistency where the same claim record could read "claimed" on one surface and "unknown" on another. Added compare_service.claim_status_from_latest_claim(latest_claim) -> (claimed, verified, status_label) as the one canonical derivation, accepting a ServerClaim ORM object, a serialized dict, or None. compare_service._claim_status and main.py's _claim_verified_from_value now both delegate to it; the two divergent sites above now call it directly.
Deliberately not changed
infer_server_verdict(the repository-layer canonical gate) staying_claim_status's exact-caseclaim.status == "verified"comparison was- The 14 existing call sites of
_claim_verified_from_valuewere left
alert-blind for batch-listing endpoints — this is an intentional, already-documented performance tradeoff (R35.2), not a bug. Wiring per-row active_alerts into every listing endpoint would be a real per-row query cost at catalog scale, exactly the kind of regression the source doc's own performance section (§35-38) warns against. Only the label text divergence was a genuine bug; the code-value tradeoff is now explicitly documented at its call site instead of silently differing.
widened to .strip().lower() inside the new shared helper — confirmed every real write site (main.py:15487, claim.status = "verified") already sets this field lowercase, so this is a no-op for real data, not a behavior change; done for robustness now that this is the one canonical comparison every caller relies on.
untouched (it's still correct, just now a one-line delegation to the new shared function) — only the two sites with genuinely divergent behavior were rewritten.
Validation
PYTHONPATH=verify/src pytest verify/tests -q— 487 passed (up frompython3 scripts/export_openapi.py --check— fresh; no response-schema- Root-repo
pytest(app/token_service, run by the pre-commit hook) —
475): 12 new/extended tests, including a property test parametrized over all 5 canonical production_readiness codes proving build_executive_verdict's decision can no longer disagree with the canonical code even under deliberately adversarial score/status inputs that would have tripped the old, now-removed independent thresholds in the opposite direction; an end-to-end build_action_safety_probe test (not a hand-written status string) proving the not_assessed fix holds for the real producer of that status; and a live-route test hitting /servers/{ns}/{name}, /v1/servers/{ns}/{name}, and /compare/oauth-required-vs-no-auth for the same server, asserting "Safe for production" renders identically (and "Safe For Production" renders nowhere) across all three.
changes (only field values, not shapes, changed).
188 passed, unaffected; no files outside verify/ touched.