MCP Verify — Consistency, Polish, and Final UX Architecture Plan
Baseline: v1.0.548 (deployed, live-verified — see docs/mcp-verify-audit-round-16-remediation-implementation.md and memory verify_audit_round16_architecture_plan). This plan responds to the "Verify v1.0.548 — Consistency, Polish, and Final UX Requirements" spec. Status: PLANNED, not yet implemented.
This is deliberately not a redesign. Per the spec's own §17/§20, homepage hero, lifecycle model, coverage funnel, server-report top section, contextual claiming, Observed Attention, pricing packaging, TrustOps architecture, methodology split, and rankings philosophy are preserved as-is. Everything below is grounded in direct reads of v1.0.548 code (four parallel research passes over main.py, insights.py, validation/service.py, db/repository.py, materialization.py, workers/scheduler.py, trustops.py), not the spec's own assumptions — several of which turned out to be more specific or less severe than framed. See the citation lists in each section below.
---
A. Trust Index freshness
Current data source
Single canonical freshness window setting: settings.validation_freshness_alert_hours (config.py:209-211, env MCP_VERIFY_VALIDATION_FRESHNESS_ALERT_HOURS, default 24h) — not duplicated as a magic number at any of its real call sites (db/repository.py:277,286,1283; materialization.py:150,156,...; main.py:7510,7642,8477,16684,5398). Both the homepage and /trust-index call the same function, build_public_coverage_stats() (main.py:19714-19803), against the same field, Server.last_validated_at (set unconditionally on every completed validation attempt, pass or fail — validation/service.py:1031). There is no registry/discovery-timestamp confusion; the audit's "another timestamp" theory does not hold.
Root cause
Not a data-source bug — a label collision on two legitimately different stats that share the bare word "Fresh":
- Homepage's "Fresh" card (
main.py:23661-23666) renderscoverage["healthy_and_fresh"], computed atmain.py:19747-19751ascurrent_status == "healthy" AND last_validated_at >= cutoff. Its own hint text says "Healthy and validated within the current {sla}h freshness window." - Trust Index's "Fresh" card (
main.py:28616) renderscoverage["fresh_validations"], computed atmain.py:19752-19755aslast_validated_at >= cutoffwith no status filter — includesdegraded/failingattempts. Its hint text says "Validated inside the current {sla}h freshness window."
Both cards are labeled the bare word "Fresh"; the only differentiator is one word ("Healthy and…") buried in hover-hint prose most visitors will never read. Under a scheduler that keeps the whole catalog inside the SLA window, fresh_validations trends toward "attempted recently regardless of outcome" ≈ the whole indexed inventory, while healthy_and_fresh is a much smaller, stricter subset. That fully explains the audit's observed symptom — Trust Index "Fresh" reading close to total inventory while the homepage's "Fresh" reads small — without any query bug.
A second, minor, genuinely inconsistent literal was also found while auditing this surface: Trust Index's "Strong live" hint hardcodes "validated within 48h" (main.py:28621) while every other freshness-window reference on both pages templates the real configured value ({sla}h, resolving to 24h by default). This is a stray literal, not a semantic dispute — should be fixed alongside.
A secondary, smaller factor: the homepage route is served through a catalog_fingerprint-keyed cache with a 900s TTL (main.py:7761-7768, config.py:260), while /trust-index computes fresh on every request (no cache). This can produce up to ~15-minute skew between the two pages' numbers even after the label fix — legitimate, and exactly what F4 asks to make transparent rather than eliminate.
A related but independent duplication: materialization.py:74-117's _server_freshness_snapshot() — used for server-detail/shortlist age labels — buckets age into fixed calendar labels ("Verified in last 24h / 7d / 30d / Stale", lines 87-96) that are not driven by the configured SLA parameter (only the separate freshness_sla_status: met/breached field correctly reads the passed-in freshness_sla_hours). The calendar buckets are a defensible, distinct UX concept (age-in-human-terms vs. SLA-compliance), but 720h/168h/24h happening to numerically match TRUSTOPS_TIERS' Community/Pro/Enterprise SLA hours (see §B) is a coincidence worth a one-line code comment so a future reader doesn't conflate the two.
Proposed fix (F1-F4)
- Keep both stats — they measure genuinely different, useful things (attempted-recently vs. healthy-and-recent). Do not collapse them per the spec's own "don't overload Fresh" instruction — the fix is disambiguating labels, not deleting a stat.
- Relabel: Trust Index's card keeps
fresh_validationsbut its visible label changes from bare "Fresh" to "Fresh (any outcome)" or "Recently validated" — the metric literally matches methodology's stated definition ("validated within window"), so keep the number and the value it feeds into/v1/trust-index/latest, just make the label say what it counts, not what it implies. Homepage's card keepshealthy_and_freshand its label changes from bare "Fresh" to "Healthy & Fresh" (already the internal key name — the fix is entirely presentational). - Extract one small shared helper (e.g.
db/repository.py::freshness_stat_definitions()or a new lightweightfreshness.py) that bothbuild_public_coverage_statsandmaterialization.py::_server_freshness_snapshotimport their cutoff/threshold from, removing the two ad hoc inline SQLAlchemy filters currently duplicated inbuild_public_coverage_stats(F3). This does not change either stat's semantics, only removes the duplication risk. - Fix the stray
"48h"literal atmain.py:28621to template the real configured threshold, matching the homepage's equivalent card. - F4 snapshot transparency: add a small "as of" timestamp/note near the Trust Index stat block (it already computes live, so this is just displaying
datetime.now(timezone.utc)at render time) and, on the homepage, surface the cache TTL context already implicit incatalog_fingerprint(e.g. "figures refresh approximately every 15 minutes").
Affected surfaces
main.py:23661-23666 (homepage card), main.py:28600-28625 (render_trust_index_page), main.py:19714-19803 (build_public_coverage_stats), materialization.py:74-117, plus a new test (TEST1 below).
---
B. Plan SLA consistency
Current values and where they live
Three independent tier taxonomies exist, none of which the actual scheduler reads:
| Structure | Location | Shape | |---|---|---| | TRUSTOPS_TIERS | main.py:1123-1126 | 3 tiers (community/pro/enterprise), freshness_sla_hours: 720/168/24 | | VERIFY_PLAN_DEFINITIONS | main.py:1128-1179 | 6 plans (community/pro/publisher/trustops/cloud/docs_compiler/enterprise), free-text "freshness" strings — **disagrees with TRUSTOPS_TIERS**: assigns "24 hour target" to a trustops plan while TRUSTOPS_TIERS assigns 24h to enterprise; its own enterprise entry says "Custom SLA" (no number at all) | | TRUSTOPS_TIER_ORDER | trustops.py:23 | 4-tier ordering (community/publisher/trustops/enterprise) used only for feature gating (FEATURE_MINIMUM_TIER, trustops.py:24-34) — no SLA field at all |
/trustops's own page (main.py:30678-30699) renders both TRUSTOPS_TIERS (tier table) and VERIFY_PLAN_DEFINITIONS (plan cards) on the same page — they disagree with each other in the same render. /pricing's comparison table (main.py:29715) is a fully independent inline literal tuple — ("Best effort", "24 hours", "4-12 hours") — that reads none of the above and contradicts TRUSTOPS_TIERS outright (Pro shown as 24h there vs. 168h in TRUSTOPS_TIERS; Enterprise shown as a "4-12 hours" range that appears nowhere else in the codebase). /docs/scoring-specification (main.py:28183) hand-copies TRUSTOPS_TIERS' three numbers as a separate literal string — currently correct by coincidence, silently driftable. The server-detail "Policy SLA" badge (main.py:27185) reads freshness_sla_hours from a caller-supplied query parameter that defaults to 168 regardless of the server's actual plan/tier (main.py:3747 and three other call sites) — it is not derived from BillingAccount.tier at all today.
Ground truth: what the backend actually does
workers/scheduler.py:431-550 (should_validate_server, compute_validation_priority) drives real revalidation cadence off a single global MCP_VERIFY_STALE_AFTER_HOURS (default 24, config.py:187), modulated only by three boolean flags — claimed (age ≥ stale/2), watched (age ≥ stale/4), high-scoring trust-index candidate — never by BillingAccount.tier. No code path anywhere reads a plan/tier value to pick a differentiated cadence. Every 720h/168h/24h/"4-12h" number currently on the site is a display-only claim with zero corresponding backend enforcement.
Proposed canonical model
This is the one place in the spec where copy cannot simply be corrected without a product decision — see Open Question 1 below. Recommended default (pending sign-off): treat current numbers as aspirational targets, not guarantees, and label them honestly rather than build real tiered scheduling in this pass (that would be a scheduler/infra change, outside "consistency pass" scope per §20's non-goals).
- Introduce one small
PlanEntitlementsconfig object (module-level dict or lightweight dataclass, e.g. newverify/src/mcp_verify/entitlements.py) mergingTRUSTOPS_TIERS' numeric fields with the parts ofVERIFY_PLAN_DEFINITIONSandTRUSTOPS_TIER_ORDERthat are genuinely about entitlement (not feature-gate ordering, which can keep its own concern). Fields per plan:validation_target_hours,freshness_sla_hours,guaranteed: bool,monitoring_frequency,api_access,history_retention_days(only include fields that already have a real value somewhere — don't invent new entitlement dimensions this round). - Rewire
/pricing's comparison table,/trustops's tier table AND plan cards, and/docs/scoring-specification's SLA line to all read from this one object — eliminating the/pricingvsTRUSTOPS_TIERScontradiction and theTRUSTOPS_TIERSvsVERIFY_PLAN_DEFINITIONSself-contradiction on the same/trustopspage. - Default the server-detail "Policy SLA" badge from the server's actual
BillingAccount.tiervia the new object, while preserving the existing query-param override (used for what-if/preview scenarios) as an explicit override rather than the only source. - Per P3: explicitly separate copy into three labeled concepts wherever SLA-adjacent text appears — "Trust freshness definition" (the 24h evidence-age window, product-wide, already correct), "Validation cadence" (how often we attempt to revalidate a plan's servers — today: not actually differentiated by plan), and "Commercial SLA" (what's contractually promised — today: Enterprise only, and only as prose, not enforced). Do not let all three keep using the word "SLA" interchangeably.
Migration approach
Land the PlanEntitlements object and the read-site rewiring as one self-contained commit (values match today's TRUSTOPS_TIERS, i.e. no user-visible number changes yet beyond honesty qualifiers). A second, separate commit corrects the actually-contradictory numbers (/pricing's literal tuple) to match the canonical object, since that's a visible copy change worth its own review/changelog entry.
---
C. Decision taxonomy
What already exists (largely already reconciled)
Two of the three concepts the spec worries about are already the same thing as of a 2026-08-10 fix: build_executive_verdict (main.py:17985) derives its decision field directly from production_readiness.code (main.py:18075-18083) — other factors can only tighten (e.g. safe_for_production → "Allow with approval"), never contradict the readiness boundary. This is enforced by an existing property test, test_arch_track1_executive_verdict_decision_cannot_disagree_with_canonical_production_code (tests/test_score_integrity.py:1161). Production readiness and executive verdict are not the contradiction the spec describes.
The genuinely distinct, genuinely unreconciled concept is build_policy_export_payload (main.py:18175, GET /v1/servers/{ns}/{name}/policy) — allow/requires_human_approval/per-tool decisions, computed from an entirely independent set of inputs (tool risk levels, write-safety, freshness-SLA breach, OAuth posture, client-readiness) that **does not read production_readiness.code or executive_verdict.decision at all**. This is legitimately "recommended runtime policy" as the spec describes it — but today it is not rendered anywhere on the page as a labeled value; the server detail page only links to it ("Policy export" link, main.py:24942). The specific UI juxtaposition the spec describes — "Production readiness: Safe for evaluation" next to "Recommended policy: Allow with approval" on the same page — does not exist yet; it needs to be built, not just relabeled.
One genuine field-naming ambiguity (D3) was found: production_trust_decision is used for two different shapes across two API builders — build_trust_snapshot (main.py:17705-17722) uses it for the full executive_verdict dict; trust_index_item_from_server (main.py:19448-19457) uses the identical field name for the raw infer_server_verdict string code. Same field name, different shape, different endpoints.
Proposed fix (D1-D3)
- D1/D2: Add one new, clearly-labeled line to the server detail page — "Recommended runtime policy: {Allow / Allow with approval / Block}" — sourced from
build_policy_export_payload'sallow/requires_human_approval, positioned near but visually distinct from the existing readiness/decision block (own heading, own hint text explaining "governs what an agent may do at runtime" vs. readiness's "how suitable for evaluation/production adoption"). Add the corresponding domain-definition paragraph to/methodology. - D3: Do not rename
production_trust_decisionanywhere (backward-compat instruction, andbuild_trust_snapshot's usage is not actually ambiguous on its own). Add one new, unambiguous sibling field totrust_index_item_from_server's output only — e.g.production_readiness_code— equal in value to whatproduction_trust_decisionalready returns there, documented as the preferred field going forward; leaveproduction_trust_decisionpresent and unchanged for existing consumers.
Affected surfaces
main.py:24906-24960 (server detail page assembly), main.py:19448-19457 (trust-index item shape), /methodology render function, new TEST3.
---
D. Evidence semantics
What's already fixed (checked, no action needed)
The bulk of "unknown reads as pass" work shipped in audit round 16 (v1.0.548) holds up under a fresh audit: build_write_safe_badge_reason, summarize_write_action_governance_detail, summarize_probe_details's action-safety branch, _readiness_summary_from_blockers, and the SCORE_COMPONENT_ZERO_POINT justification prose were all re-checked and are correctly gated.
One real, still-open gap
render_write_action_governance (main.py:26415-26433)'s no_high_risk_tools_message only special-cases governance_status == "not_assessed":
no_high_risk_tools_message = (
"High-risk tool assessment unavailable -- no tools were analyzed."
if governance_status == "not_assessed"
else "No high-risk tools were detected on the latest run."
)
But build_write_action_governance (insights.py:814-848) defaults status to "missing" — not "not_assessed" — whenever a server has never been validated at all (main.py:14347 _write_action_governance_for_server with latest_validation=None → checks={}). For that never-validated case, this panel currently renders the affirmative "No high-risk tools were detected on the latest run." directly alongside its own sibling "Governance status" badge, which correctly says evidence is unavailable — a same-page contradiction, exactly the pattern already fixed for "not_assessed" but not extended to "missing".
Proposed fix (E1-E2)
Extend the existing special-case condition from governance_status == "not_assessed" to governance_status in {"not_assessed", "missing"}. One-line change, one new regression test.
---
E. Testing gaps
What already exists
Freshness classification, verdict/readiness classification, and unknown-evidence guards are all well covered (test_score_integrity.py, test_insights_goldens.py, test_rank_independence.py — see full inventory in research notes). One genuine cross-surface pattern already exists: test_r3_score_is_consistent_across_surfaces_for_a_healthy_server and test_arch_track1_production_readiness_label_is_consistent_across_surfaces_for_a_healthy_server (test_score_integrity.py:452,496) spin up a real build_test_client() HTTP client and assert one fixture server produces matching values across the HTML server page, raw JSON API, badge JSON, and compare page — but only for current_score and production_readiness.label. This is the pattern to extend, not a new pattern to invent.
Gaps
- No SLA-copy consistency test (nothing currently asserts
/pricing,/trustops, and/docs/scoring-specificationagree on plan numbers). - No test asserting Trust Index's and homepage's two "Fresh" cards use their (post-fix) distinct, documented labels rather than colliding on the bare word.
- No test for the policy-export vs. production-readiness relationship (today true because they're independent by design — TEST3 should assert coexistence without one overwriting the other, not equality).
- No impossible-state/invariant checks anywhere in the codebase today.
New tests (mapped to spec §12)
- TEST1 (freshness): extend the cross-surface
build_test_client()pattern with a fixture server at a knownlast_validated_at; assert homepage "Healthy & Fresh", Trust Index "Fresh (any outcome)", rankings/search?freshness=freshfilter, and the JSON API all agree on the underlying fresh/stale boolean for that server (not that the two aggregate counts are equal — they're intentionally different populations post-fix). - TEST2 (readiness): largely already covered by existing
test_arch_track1_*/test_production_readiness_verdict_matches_canonical_verdict_for_*tests; add the two example combinations from the spec (healthy+production-threshold+fresh; healthy+evaluation-score+fresh) if not already parametrized. - TEST3 (policy vs readiness): new — fixture where
production_readiness.code == "safe_for_evaluation"andpolicy_export.requires_human_approval is Truesimultaneously; assert both values render on the page with their own distinct labels and neither field is silently overwritten by the other's computation. - TEST4 (plan entitlements): new — assert
/pricing,/trustops, and/docs/scoring-specificationall derive their SLA numbers from the samePlanEntitlementsobject (e.g. by monkeypatching the object and asserting all three surfaces reflect the change), guarding against a future surface reintroducing an independent literal. - TEST5 (unknown evidence): extend existing
test_a4_audit16_missing_probes_are_not_assessed_not_passes-style fixture to cover therender_write_action_governance"missing" case fixed in §D. - TEST6 (billing neutrality): already covered (
test_rank_independence.py:54,95) — no new work, just keep running.
Invariant checks (§13)
Implement as a small pure function (e.g. db/repository.py::coverage_stat_invariants(coverage: dict) -> list[str]) returning a list of violated-invariant descriptions, called from both a unit test (fixture-driven) and, cheaply, from the existing /v1/build or an admin diagnostics route so violations are observable in production without a new scheduled job. Only implement invariants that are actually true of this domain model post-fix: production_ready_count <= scored_count, production_ready_count <= fresh_validations, healthy_and_fresh <= healthy_servers, healthy_and_fresh <= fresh_validations. Do not implement fresh_count > validated_count as stated verbatim in the spec — post-fix there are two fresh counts with different, correct relationships to "validated," so the invariant needs to name which one.
---
F. Implementation plan — sequenced commits
Small, independently revertible commits, in the priority order the spec itself specifies (§19).
P0 — data consistency
- Relabel homepage/Trust Index "Fresh" cards (§A.2-3), fix the stray
"48h"literal, extract shared freshness-window helper used bybuild_public_coverage_statsandmaterialization.py::_server_freshness_snapshot. Ship with TEST1. - Introduce
PlanEntitlementsconfig object (§B, new small module), rewire/trustopstier table + plan cards to read from it (no visible number change yet — this commit is pure de-duplication). - Rewire
/pricingcomparison table and/docs/scoring-specification's SLA line to the same object — this commit does change visible/pricingnumbers (the actual contradiction fix) and should get its own CHANGELOG entry. Ship with TEST4. - Default server-detail "Policy SLA" badge from
BillingAccount.tierviaPlanEntitlements, preserving query-param override.
P1 — domain language
- Add labeled "Recommended runtime policy" line to server detail page from
build_policy_export_payload(§C.1); add methodology copy distinguishing the three concepts (§C, §D2 of spec). - Add
production_readiness_codesibling field totrust_index_item_from_server's JSON shape (§C.2, D3). - Fix
render_write_action_governance's"missing"gap (§D). Ship with TEST5. - Grep-and-remove remaining unsubstantiated "first public MCP Trust Index"-style priority claims (§18) — small, standalone copy commit.
P2 — guardrails
- TEST2/TEST3 additions extending the existing
build_test_client()cross-surface pattern. coverage_stat_invariants()helper + unit test + wiring into an existing diagnostics surface.
P3 — polish
- Server-report lower-section progressive disclosure refinement (§S3/S4) — small, only if a concrete over-expanded section is identified; no broad restructuring.
/pricingPro-tier scaling-factor bullets (§PRICE1) — list only factors already confirmed to exist (API volume, monitored servers, history retention — do not invent new ones).
Each numbered item ships as its own commit with its own focused tests; run the full suite (PYTHONPATH=verify/src python3 -m pytest verify/tests -q --no-cov) after each, matching this project's established discipline. Live-verify P0 items against the tracepass/playwright fixtures post-deploy, same as prior rounds — freshness/SLA changes are exactly the kind of thing that can look right in tests and still disagree in production between cached and uncached routes.
---
Open questions (recommend, do not decide unilaterally)
- Are per-plan validation SLAs meant to become real differentiated backend behavior this round, or stay display-only targets, honestly labeled as such? The scheduler today has zero tier-awareness (§B). Recommendation: keep display-only this round (implementing real tiered scheduling is an infra change beyond "consistency pass" scope per the spec's own §20 non-goals) but change qualifying language so Pro/Enterprise numbers read as "target" rather than implying an enforced guarantee, until/unless a future round implements real enforcement.
- Exact final label text for the two "Fresh" cards (§A.2) — recommended "Fresh (any outcome)" / "Healthy & Fresh," but any pair that disambiguates without implying one is more "real" than the other works equally well; this is a copy decision, not an architectural one.
- **Whether
/v1/trust-index/latest's newproduction_readiness_codefield (§C.2) should also be added tobuild_trust_snapshot's output** for symmetry, even though that endpoint's existingproduction_trust_decisionfield isn't ambiguous on its own — low-cost either way, deferred to whoever reviews the diff.
---
Non-goals reaffirmed (per spec §20)
No scoring algorithm changes, no threshold changes, no dimension removal, no homepage/pricing/TrustOps redesign, no analytics reclassification, no raw telemetry exposure, no methodology simplification. Every item above is additive labeling, one shared config/helper extraction, or a one-line gating fix — no route is removed or restructured, and no existing public API field is renamed or removed.