Pre-registered episode detection
Not yet computed, run python -m src.validate hit-rate after the backfill.
| Result | Channel | Date | Episode |
|---|
What the detector cannot see
Base rates, published beside the headline because a hit rate without its chance rate is a number wearing a costume (added 2026-08-07 after an adversarial review): under the registered ±3-day, any-episode-day criterion the rate is 24 of 29; requiring the episode start within ±3 days gives 19 of 29; randomly-dated events would expect ≈6.8 of 29 by chance; and a naive detector that counts an episode in any channel scores 26 of 29 — so the channel attribution, not the detection alone, is where the five-dictionary apparatus earns its keep. All four numbers recompute nightly from the published files: data/detection_baselines.json.
The hit-rate above is measured against episodes chosen in advance. This is the other half: what the detector structurally misses. Detection fires when a channel's share clears its trailing 90-day mean + 2σ, so a crisis enters the very window its own future threshold is computed from — the mean rises, σ rises faster, and the bar climbs for ninety days. It is not a tuning choice that can simply be improved: every alternative is worse (excluding spike days presumes what a spike is; a fixed threshold abandons the per-channel adaptivity that makes five channels comparable; a longer window delays first detection, which is the point). The registered rule stands and its cost is measured instead.
| Channel | Episode | bar before | bar after | × | days suppressed |
|---|
“Days suppressed” are days after an episode whose coverage cleared the threshold in force the day before that episode began, but not the live one — days a pre-crisis detector would have flagged and this one did not. Read the largest episodes, not the median: a one-day blip barely moves the bar, so the blindness is concentrated exactly where the first crisis was most serious.
Placebo channels
Channels with no India-geopolitics content (IPL cricket, Bollywood) run through the identical pipeline. If they spike alongside real episodes, the pipeline measures general news volume, not channel salience.
Not yet computed.
Dictionary robustness
The full index recomputed under a registered narrower and a registered broader
term set (dictionaries_alt.json), correlated with the primary over the 2022–present
overlap. Reading conventions, stated ex ante: ≥ 0.90 the series does not depend on the
term list; 0.70–0.90 moderately term-dependent — the large moves survive the change
while day-to-day levels shift; < 0.70 materially term-dependent — the variant
produces a different series, and the divergence is discussed individually below, never averaged away.
| Channel | vs narrow | vs broad |
|---|
The weak spot, stated plainly. Gulf & energy vs the narrow variant is 0.527 — the lowest robustness correlation, and it is informative. The narrow variant keeps only four of the channel's eleven registered phrases (Strait of Hormuz, Israel–Iran, Gulf tensions, and one supply-disruption phrasing) and drops the energy-security vocabulary entirely (crude oil supply, energy security, Iran sanctions, Middle East conflict, Gulf of Oman). At 0.527, those two halves of the channel move differently: this channel is a deliberate composite of conflict-proximity coverage and energy-supply coverage, and cutting it to the conflict core yields a different series. Practical reading: treat single-day moves in this channel with more caution than the others and read its receipts before quoting it. The composite is insulated from this choice (vs narrow 0.796, vs broad 0.880).
Primary vs variants, weekly means
Cross-source agreement
GDELT (supply-side: what editors publish) vs Wikipedia pageviews (demand-side: what readers look up). High agreement means the signal is not a GDELT artifact; persistent divergence is a documented finding, not noise.
| Channel | correlation |
|---|
Coverage drift
GDELT's monitored corpus is not constant. Mean daily corpus size by year, and whether channel shares trend with it.
Not yet computed.
Cross-language validity: Hindi Wikipedia
The corpus is English, the matcher is English, and the cross-source check above uses English Wikipedia — every leg Anglophone, so none of it could tell “Indian salience” apart from “what the international English-language press covers about India.” This test can. The same registered articles read on Hindi Wikipedia, with Hindi titles resolved through Wikipedia's own interlanguage links rather than chosen by hand, correlated against the same daily shares. Day-to-day changes is the load-bearing column: levels can agree through a shared trend alone.
| Channel | English, changes | Hindi, changes | English, levels | Hindi, levels | Hindi articles |
|---|
English-language coverage comparison (V5)
The production instrument reads English coverage. This comparison asks how its weekly percentile differs from selected registered-language corpora, each ranked against its own history. Divergence is English minus the mean of the non-English series available for that channel: positive means the English percentile was higher; negative means it was lower. This does not establish which language is correct, represent the full Indian news ecosystem, validate the index, or change any score.
Loading coverage…
| Channel | Languages compared | Latest week | English pctl | non-English mean | English − non-English | mean |gap|, last 26 wks | weeks |
|---|---|---|---|---|---|---|---|
| Loading language comparison… | |||||||
Registered languages load from the payload.