Why a second construct instead of a longer index
The live instrument measures the share of GDELT-monitored articles matching a channel's query. That corpus does not exist before 2017 in any comparable form, so extending the index backwards is not a matter of running the same code on older data — the data is not there. What does reach back is GDELT's event database, which codes actor pairs from news text starting in 1979.
So this measures something adjacent and honestly different: the monthly share of global event mentions involving a registered actor pair. India–Pakistan. India–China. It is a proxy for how much of the world's coded news traffic concerned that relationship, month by month, for forty-one years.
The filters, the anchor events, the overlap thresholds and the publication rule were all written
down and signed in analysis/back_extension_memo.md before any query ran. That ordering is
verifiable in the git history rather than asserted here: the commit that froze the filters predates
the commit that first held BigQuery credentials.
Which channels earned a historical series
The memo set the rule in advance: a channel publishes a back-extension only if its historical proxy correlates with the live instrument at r ≥ 0.6 over the 36 months where both exist (2017–2019). Below 0.4 it does not publish at all. Two channels cleared it and two did not, and the two that did not are shown here because a pre-registered test that fails is a result, not an embarrassment.
| Channel | Overlap r | Months | Verdict |
|---|
The border channels track because a border confrontation is an actor-pair event almost by definition — the thing the event coder sees and the thing the article counter sees are the same thing. US & Trade Policy and Gulf & Energy Security do not, and the reason is not a bug: a tariff schedule, a refinery contract or a shipping-lane premium is coverage that mostly does not resolve to a coded bilateral event. Those two have no historical series on this site, and it would be easy and wrong to publish them anyway.
Forty-one years, two channels
Monthly, ranked against each channel's own trailing ten years, so the vertical axis means the same thing in 1984 as in 2014. Marked points are the anchor events named in the memo before anything was computed — eight of the nine, because the ninth belongs to a channel that did not earn a series.
Two breaks in the line, both structural. The series begins December 1983, not January 1979: a rank needs something to rank against, and the window requires sixty months of history before it will report. And there is a 28-month hole from February 2015 to May 2017, which is the GDELT 1.0 → 2.0 transition — the old event stream ends there and the new one is not comparable across the seam at monthly resolution. Both are drawn as gaps rather than bridged, because interpolating across a source change is how a chart starts asserting things nobody measured.
Anchor events: named first, graded after
Nine events were listed in the memo as things a working attention measure ought to notice, with a pass mark of top decile of the trailing ten years. Grading happened after the series existed. Six cleared it; three did not, and the three misses are the informative part.
| Event | Month | Channel | Percentile | Top decile |
|---|
What the three misses mean
Operation Blue Star, June 1984 (52.3). The anchor was filed under the Pakistan channel, and Blue Star was an internal Indian security operation. The India–Pakistan actor pair is the wrong instrument for it. This is a mis-specified anchor rather than a missed event — and it was mis-specified in advance, in writing, which is why it is still on the list.
Sumdorong Chu, October 1986 (54.8). A months-long standoff that Western wire coverage largely did not carry, in the era before either country's press had a global distribution footprint. The honest reading is that the proxy is thin on early China-channel events, which is a real limitation of the series and not a property of the standoff.
Mumbai 26/11, November 2008 (89.1). Missed the top decile by roughly one percentile point. Worth stating plainly rather than rounding into the pass column.
Six of nine is the number. It is reported as six of nine because the alternative — dropping the mis-specified anchor and reporting six of eight — would be choosing the denominator after seeing the result.
What this series cannot do
- It is not the index extended. Do not join it to
history.csv, do not compute a growth rate across 2017, do not put them on one axis. They are different measurements that happen to share a name for the thing they point at. - Levels are not comparable across the 1979/2015 GDELT boundary without care. GDELT 1.0 runs to February 2015 and 2.0 begins there, with a much larger and differently sourced corpus. Ranking against a trailing ten-year window absorbs most of that, which is exactly why the published series is a rank and not a level.
- Shipping has no historical series at all, and never will from this source. Chokepoint salience has no actor-pair coding — there is no country that the Strait of Hormuz is in a bilateral relationship with.
- It says nothing about what comes next. Same as everything else here.