Verify: observed-attention cache stampede + export back-off
Targeted production reliability fix for the 2026-09-28 06:02–06:04 UTC stall. Not a feature release.
Incident
- The nightly Trust Data Feed snapshot export (
sentinel-signal-data-snapshot.timer, - Every Intelligence call builds a policy payload →
build_observed_attention→ - The cache was check-then-compute with no single-flight. At each expiry every
- The exporter amplified it: a 10 s client timeout abandoned requests that kept
03:15 UTC) fetches /v1/intelligence/servers/{ns}/{name} for ~97k servers, 10 threads, 10 s client timeout, through the public URL.
_cached_session_sets, a per-process cache of three raw 30-day analytics_events rollups (MCP_VERIFY_OBSERVED_ATTENTION_SESSION_CACHE_SECONDS=1800).
in-flight request recomputed the rollups concurrently → Postgres temp spill → verify-web pool (20 + 15 overflow) exhausted → every DB-backed route stalled (policy 41 s, compare 12 s, /mcp 13 s, 500/503s), not just Intelligence. It recurred at ~30-minute boundaries during every export; 06:02 was the worst.
running server-side and immediately sent the next one. On 09-28 the exporter counted 637 fetch errors while verify-web logged only 52 non-200 responses — ~585 were abandoned requests that later completed 200.
Change
observed_attention._cached_session_sets: single-flight + stale-while-refresh.
- fresh entry → returned;
- expired entry → exactly one caller recomputes; concurrent callers get the stale value;
- failed refresh → previous value kept and served; retry after
- cold start → one caller computes; others wait at most
MCP_VERIFY_OBSERVED_ATTENTION_SESSION_REFRESH_RETRY_SECONDS (default 60);
MCP_VERIFY_OBSERVED_ATTENTION_SESSION_COLD_WAIT_SECONDS (default 60), then ObservedAttentionSessionSetsUnavailable. Attention results are unchanged; only who computes them and when.
- Exporter (
intelligence/data/export.py):
DATA_FEED_EXPORT_CONCURRENCYdefault 10 → 4;DATA_FEED_EXPORT_TIMEOUT_SECONDS60 (was a hard-coded 10);- bounded retries (
DATA_FEED_EXPORT_MAX_ATTEMPTS=3) on timeout / transport ExportFetchBackoffGate: a retryable failure pauses all export threads- result/progress JSON gains
fetch_backoff_pauses.
error / 429 / 5xx only; 4xx and bad JSON fail immediately;
for base * 2**(n-1) s (base 5, cap 120), so the exporter stops adding load to a struggling server instead of refilling its queue;
Known limits
- The single-flight lock is process-local. N web processes can still mean N
- The refreshing request still pays the full recompute latency itself.
- Not changed here (separate decisions): skipping observed attention for the
- Lower concurrency may lengthen the export (baseline 2h32m–2h53m at 10); watch
concurrent recomputes per expiry. Production verify-web runs one uvicorn process, so this collapses to one. If verify-web is ever scaled out, move the session sets to a maintenance-worker-precomputed shared store.
internal export / a lighter export endpoint; Loki health.
run duration alongside fetch_errors.
Validation
verify/tests/test_observed_attention_single_flight.py(12 concurrent callers atverify/tests/test_intelligence_data_export_backoff.py(defaults, exponential +- Production: no DB-wide latency spike at the ≥2 30-minute expiry boundaries
expiry, cold start, bounded wait, failed refresh, retry window, cold-start failure handoff, unchanged results) — 6 of 8 fail on the pre-fix code.
capped backoff, bounded retries, non-retryable errors, shared pause across threads).
after deploy, and during the next 03:15 export; compare export fetch_errors (baseline 469–767/run) and duration.