Evidence

Validation

What is currently tested: detection against a frozen episode list, placebo channels, sensitivity across alternative dictionaries, and a cross-source comparison. These diagnostics test instrument behavior; they do not convert salience into risk, establish matched-item precision or population recall, or show superiority over another index.

Data day
See latest published payload
Freshness
Source-by-source status

Pre-registered episode detection

Not yet computed, run python -m src.validate hit-rate after the backfill.

ResultChannelDateEpisode

Placebo channels

Channels with no India-geopolitics content (IPL cricket, Bollywood) run through the identical pipeline. If they spike alongside real episodes, the pipeline measures general news volume, not channel salience.

Not yet computed.

Dictionary robustness

The full index recomputed under a registered narrower and a registered broader term set (dictionaries_alt.json), correlated with the primary over the 2022–present overlap. Reading conventions, stated ex ante: ≥ 0.90 the series does not depend on the term list; 0.70–0.90 moderately term-dependent — the large moves survive the change while day-to-day levels shift; < 0.70 materially term-dependent — the variant produces a different series, and the divergence is discussed individually below, never averaged away.

Channelvs narrowvs broad

The weak spot, stated plainly. Gulf & energy vs the narrow variant is 0.527 — the lowest robustness correlation, and it is informative. The narrow variant keeps only four of the channel's eleven registered phrases (Strait of Hormuz, Israel–Iran, Gulf tensions, and one supply-disruption phrasing) and drops the energy-security vocabulary entirely (crude oil supply, energy security, Iran sanctions, Middle East conflict, Gulf of Oman). At 0.527, those two halves of the channel move differently: this channel is a deliberate composite of conflict-proximity coverage and energy-supply coverage, and cutting it to the conflict core yields a different series. Practical reading: treat single-day moves in this channel with more caution than the others and read its receipts before quoting it. The composite is insulated from this choice (vs narrow 0.796, vs broad 0.880).

Primary vs variants, weekly means

Cross-source agreement

GDELT (supply-side: what editors publish) vs Wikipedia pageviews (demand-side: what readers look up). High agreement means the signal is not a GDELT artifact; persistent divergence is a documented finding, not noise.

Channelcorrelation

Coverage drift

GDELT's monitored corpus is not constant. Mean daily corpus size by year, and whether channel shares trend with it.

Not yet computed.

English-language coverage comparison (V5)

The production instrument reads English coverage. This comparison asks how its weekly percentile differs from selected registered-language corpora, each ranked against its own history. Divergence is English minus the mean of the non-English series available for that channel: positive means the English percentile was higher; negative means it was lower. This does not establish which language is correct, represent the full Indian news ecosystem, validate the index, or change any score.

Loading coverage…

ChannelLanguages comparedLatest week English pctlnon-English mean English − non-English mean |gap|, last 26 wksweeks
Loading language comparison…

Registered languages load from the payload.