How we know the index measures anything real: detection of a pre-registered episode list the query terms never name, placebo channels that must stay quiet, robustness across alternative dictionaries, and agreement across independent sources. What validation cannot do: convert salience into risk.
Not yet computed, run python -m src.validate hit-rate after the backfill.
| Result | Channel | Date | Episode |
|---|
Channels with no India-geopolitics content (IPL cricket, Bollywood) run through the identical pipeline. If they spike alongside real episodes, the pipeline measures general news volume, not channel salience.
Not yet computed.
The full index recomputed under broader and narrower term sets. Correlations above 0.9 mean the result is not an artifact of one term list.
| Channel | vs narrow | vs broad |
|---|
GDELT (supply-side: what editors publish) vs Wikipedia pageviews (demand-side: what readers look up). High agreement means the signal is not a GDELT artifact; persistent divergence is a documented finding, not noise.
| Channel | correlation |
|---|
GDELT's monitored corpus is not constant. Mean daily corpus size by year, and whether channel shares trend with it.
Not yet computed.