Power and Influence Index · Briefing for our funder

How we're using AI to build this index — and why it changes what your funding buys

A plain-language self-assessment, written by the project's AI system at the principal investigator's request, of what artificial intelligence actually does in this project, what it would have cost to do this work before AI existed, and the safeguards that keep the numbers honest.

August 2026 · Prepared for internal funder discussion — please don't circulate further

1.The short version

This project measures the political landscape around climate and the energy transition — across dozens of countries, year by year. To do that credibly, you need an enormous evidence base: what politicians actually say in parliament, what parties promise, how courts rule, how media covers the issue, what regulators do, and what informed experts judge.

Historically, evidence bases like this have been built by international consortia over decades, at a cost of millions of dollars per evidence stream. AI hasn't just made that old approach faster. It has made a fundamentally different approach possible, with five features we'd highlight:

2.The headline: estimates first, spending second

This is the most important innovation in the project, and it's worth two minutes to understand properly.

The problem

Take South Africa. To place South Africa on our index, we need answers to dozens of specific questions: where do its parties really stand on the energy transition? How do its courts treat climate cases? How is elite rhetoric shifting? For most of those questions, no dataset exists. The traditional answer is to build one — recruit country experts, train coders, wait years. The best-known project of that kind, an expert-scored database of political institutions worldwide, has taken a standing network of thousands of country experts and tens of millions of dollars in cumulative funding. Until that kind of investment completes, a country like South Africa simply can't be scored. You'd have a blank row, for years.

What we do instead: AI stand-in scores

Where hard data doesn't exist yet, we ask several leading AI systems — deliberately from different companies, so they don't share blind spots — to each independently score the question, using the same detailed written scoring instructions a human expert panel would use. Then we treat those AI answers the way a careful scientist treats any imperfect measurement:

The common yardstick

One more piece makes cross-country comparison honest. Every country in our system — over a hundred — is measured on a shared backbone of well-established global datasets that already exist for everyone. That common yardstick is what makes it meaningful to say one country scores higher than another. On top of that backbone, a set of about two dozen focus countries gets the deep, custom measurement. So we never pretend to know every country equally well: depth where we've invested, a common yardstick everywhere, and published margins of error that say which is which.

Why this matters for funding: the logic of investment inverts. Instead of "fund data collection for years and hope the index eventually justifies it," the index exists now, with honest uncertainty — and each new investment in real data visibly shrinks the uncertainty, starting exactly where the model says uncertainty matters most. You can watch your funding tighten the picture, stream by stream, country by country.

3.What would this have cost before AI?

Two respected academic projects give us honest benchmarks.

The Manifesto Project is political science's flagship hand-coding effort. Since 1979, its international network of trained coders has classified about 5,300 party platforms — roughly 3.3 million coded sentences — supported most recently by a 15-year grant from the German national science foundation, in a program reserved for decade-scale research infrastructure. Total investment over its life: conservatively several million euros, for one document type.

ParlaMint is a recent European consortium — more than 90 collaborators across 20+ countries, funded for four years — whose job was just to collect and standardize parliamentary debate transcripts. Not to analyze what anyone said; simply to assemble the text. Even that step, historically, took a multinational team.

Now the comparison. Our parliamentary-speech evidence stream alone contains roughly seven million utterances across the focus countries, in many languages — about twice the total volume the Manifesto Project's human coders produced in four decades. Building it involved two kinds of work:

Doing it the pre-AI wayRough cost (estimate)
First pass: is each of ~7M utterances even about climate/energy? (trained human coders)≈ $1.0–1.4M
Detailed coding of the relevant material (stance, framing)≈ $0.6–1.2M
Quality control (double-coding), training, management, recruiting coders in 10+ languages≈ $0.5–1.3M
This one evidence stream, done by hand≈ $2–4M and 5+ years — if multilingual coders could be recruited at all

And the index has several such streams — speech, party platforms, media, courts, corporate conduct, regulation. The pre-AI version of this evidence base is a $10–20M, decade-long consortium project. The actual classification spend is a rounding error on that. The binding costs are now human judgment and expert validation — which is exactly where research money should bind.

4.The bright line: AI supplies evidence, never verdicts

The question any thoughtful funder should ask: "Isn't this index just AI opinions?" Our answer is built into the architecture, not just our intentions.

The index's actual scores come from a conventional statistical model — the same family of methods used to score standardized tests, where many imperfect questions are combined into one estimate with a margin of error. That model is ordinary, fully inspectable statistical code. AI never touches it. AI-produced classifications and stand-in scores enter it exactly the way human-coded data would: as evidence, each piece carrying a declared level of uncertainty.

5.We test the AI before we trust it

We treat "which AI model should do this task?" as a measurable question, never a vibe. Before any AI classifier is allowed to run at scale, it takes a test: a sample of material that human experts have already judged, scored language by language.

This has produced findings that matter well beyond our project. The sharpest: an inexpensive AI model matched the premium model on English — and its rare mistakes came with low confidence, so a simple "when unsure, escalate to the stronger model" rule would catch them. But on Japanese, the same cheap model failed badly and was highly confident while being wrong — more confident in its errors than in its correct answers. That's the dangerous kind of failure, because it silently defeats the "escalate when unsure" safety net. The finding is now a standing rule: minimum model strength is set per language, based on measurement, and re-tested whenever models change.

We apply the same skepticism up the chain. In one blind test, mid-tier AI reviewers checked every number in an analysis, found them all correct — and still endorsed a wrong explanation of what the numbers meant. Lesson, now standing policy: checking arithmetic is not the same as checking reasoning, and only the strongest available models (or humans) get to be the final reviewer before results move forward.

6.AI checking AI — a quality layer that never existed before

Some of the most valuable AI work in this project has no pre-AI equivalent at any price:

7.A big lab, run by two people

The project runs about eight parallel workstreams, coordinated through machine-maintained decision queues, memos between workstreams, and a permanent decision ledger — so decisions made once stay made, and apparent contradictions get traced to the ruling that governs them instead of being re-argued.

The practical effect on timelines is dramatic and worth stating plainly: work that would read as "six to eight weeks" in a conventional research plan is typically a few focused working sessions of AI execution, plus human review time, plus whatever genuinely external waiting is involved (an expert's schedule, a consultant's deliverable). The AI is essentially never the bottleneck. The human hours in this project go almost entirely to judgment, sign-off, and expert relationships — not to production.

8.What we're honest about

AI classifications are imperfect measurements, and we treat them that way. Error rates are estimated task by task and language by language, and those errors feed into the published margins of error. Where AI providers disagree with each other, the published uncertainty gets wider — never quietly averaged.

Cheap models do the routine work; the strongest models review; humans decide. No change to the production scoring database happens without an explicit human authorization naming exactly what's being changed — and every change is archived and reversible.

Some of our best evidence of reliability is caught failure. Unreproducible figures, confidently-wrong classifications, rules on paper that nothing enforced — all found by checking machinery we deliberately built. The honest claim is not "AI is reliable." It's: we have built the testing and verification layer that makes AI reliable enough, measurably, for this specific scientific use — and that layer is itself one of the project's contributions.

9.Three takeaways

1.AI produces our evidence, never our scores — the index itself is conventional, transparent statistics, and every AI-derived input is labeled, uncertainty-rated, and traceable, down to a published count of real data versus stand-ins behind every score.

2.The stand-in system inverts the funding logic of comparative measurement: instead of spending millions and waiting years for a first estimate, the estimates exist now — and each new investment in real data visibly shrinks the uncertainty, starting where it matters most.

3.Our parliamentary-speech evidence alone is roughly twice what political science's flagship hand-coding project produced in four decades — built by two people, with classification costs in the hundreds of dollars, and every AI classifier tested language by language against human experts before we let it run.