Sentinel Signal

MCP Verify 1.0.628 remediation and R11 groundwork

Source: docs/mcp-verify-v1.0.628-remediation-and-r11-groundwork.md

Document Content

MCP Verify 1.0.628 remediation and R11 groundwork

Date: 2026-08-13

Release scope

This release fixes endpoint/provenance identity, canonical trust fields on cold routes, registry-directed recommendations, and technical compatibility evidence handling. It also adds the non-scoring instruction-reference probe, default-off observed-behavior and action-enum groundwork, release safeguards, and a reproducible R11 evaluation package.

Shipped semantics

  • A shared deterministic URL classifier distinguishes MCP endpoints, directory
  • listings, and invalid URLs. Glama /mcp/servers/{id} URLs remain listing provenance but cannot become validation targets, install endpoints, freshness evidence, alias endpoints, or endpoint-split findings.

  • Alias responses expose valid remote_urls separately from additive
  • listing_urls and remote_url_records provenance.

  • Full detail, policy/report/page fallbacks, compact compare, trust summary,
  • rankings, and shortlists use the same trust-core computation for active alerts, evidence confidence, and production readiness. Cold compare remains visibly partial/warming only for omitted deeper detail.

  • The canonical history loader uses a bounded 20-run window with two full
  • payloads. Lightweight rows retain only the affirmative tools/list status needed for evidence-depth credit, and batched Compare/materialization loads use the same window as single-server routes.

  • Top-level recommendations retain formal remediation priorities. Publisher
  • actions lead only when publisher evidence exists; listing-only and non-MCP records lead with registry correction.

  • Technical compatibility is criteria-based and remains independent of client
  • readiness/publishability. ok and completed are affirmative. A designated auth_required result may prove an auth-challenge criterion. A warning earns credit only when its structured evidence affirmatively proves that specific capability; warning alone never earns credit. Missing, skipped, not-assessed, unknown, and absent evidence fail closed.

  • instruction_tool_reference_probe compares only explicit code-formatted or
  • tightly phrased instruction references with discovered tool identifiers. It is non-scoring and produces a low-severity finding when assessable.

Remediation severity remains a prioritization system independent from active alert severity. A legitimate high-priority remediation does not require an active alert, while the displayed high/critical alert counter reads only the active-alert list.

Default-off groundwork

MCP_VERIFY_ACTION_ENUM_CLASSIFICATION_ENABLED=false keeps public tool risk, badges, scores, and persisted semantics unchanged. The shadow implementation recursively reads action, operation, and op enums through composed JSON Schema nodes, classifies compound values such as archive_by_serial, records per-action variants, and aggregates maximum reachable risk. Run scripts/report_action_enum_shadow.py against a catalog database to record delta counts without writes.

MCP_VERIFY_OBSERVED_BEHAVIOR_CONFLICTS_ENABLED=false keeps observation-based publisher findings disabled. Its strict Pydantic contract accepts only trusted, secret-free runtime evidence. Static inference disagreement cannot create a publisher accusation.

R11 status and gate

The supplied R11 attachment is methodology, not a completed label set. This release therefore does not fabricate reviews or enable either public semantic change. The evaluation package now fixes the sample at 200 servers with seed verify-r11-v1, expands stratification, preserves frozen scores/components, and reports weighted agreement, Spearman correlations, discrimination, per-subscore results, and bootstrap 95% confidence intervals.

Public action-enum semantics require both a completed R11 report no older than 90 days and at least 0.95 precision for accusatory classifications. Any subscore without a confidence interval establishing positive relationship to the adjudicated ordering becomes a removal/rework candidate.

Release safeguards

  • The reference GitHub Action publishes a decision-reason output and its
  • forced minimum-score fixture asserts that exact reason.

  • CI runs the complete Alembic chain on PostgreSQL and checks revision width
  • plus expected production tables.

  • Deployment requires explicit immutable RELEASE_SHA and RELEASE_VERSION;
  • host-side Git fallback is removed.

  • Deployment runs a read-only semantic invariant audit before restarting the
  • service. The restart invalidates build-scoped and in-memory caches.

Verification record

Before release:

  • PYTHONPATH=verify/src:verify/tests python3 -m pytest verify/tests -q
  • python3 -m pytest -m unit -q
  • make test-suites
  • PYTHONPATH=verify/src python3 scripts/verify_predeploy_invariants.py

Deployment identity and live Tracepass/Playwright cold/warm observations are recorded in the final deployment handoff rather than pre-filled here.