Are We More Certain Than the Evidence Allows?
Are We More Certain Than the Evidence Allows? Why the next frontier in evidence synthesis is communicating uncertainty
Over the past two decades, the education sector has undergone a quiet revolution.
Increasingly, policy decisions are informed by rigorous impact evaluations, systematic reviews, and evidence syntheses rather than anecdote or intuition. Governments, donors, and advocacy organisations now routinely ask not simply whether an intervention works, but which interventions work best and which provide the greatest value for money.
This is a major step forward, but it also raises a new challenge: the more we compare interventions across studies, the more we rely on summary statistics such as standardised effect sizes, meta-analyses, cost-effectiveness estimates and intervention rankings. These summaries are enormously valuable, but they can also create an illusion of certainty that exceeds what the underlying evidence can support.
How comparable are our comparisons?
A recent paper by researchers at the Center for Global Development, The Illusion of Comparability Among Standardised Effect Sizes, highlights one important part of this problem.
The paper shows that standardised effect sizes — the common currency of many education evidence syntheses — are often much less comparable than they appear. The same improvement in children's reading can produce very different effect sizes depending on the assessment, the distribution of scores, and the reference population used for standardisation.
Two programmes may produce very similar improvements in children's learning but yield very different standardised effect sizes, simply because those gains are measured on different assessments or standardised using different score distributions. The authors argue persuasively that researchers should routinely report raw learning gains alongside standardised effect sizes, so that readers can understand what children actually learned rather than relying solely on standardised metrics. It is a reminder that the numbers we compare are not as directly comparable as we often assume.
The challenge doesn't end with effect sizes
Our recent review of the evidence on Teaching at the Right Level (TaRL) led us to many of the same questions from a different perspective. The challenge is not simply that standardised effect sizes can be misleading — it is that every step we take away from the original evidence introduces additional assumptions, and additional opportunities for comparison to become more difficult. Cost-effectiveness provides perhaps the clearest example.
Cost-effectiveness inherits — and compounds — the problem
Cost-effectiveness has become one of the most influential ways of comparing education interventions. Rather than asking simply which programmes improve learning, it asks which produce the greatest learning gains for every dollar invested. This is an important question: governments and donors need to consider costs alongside impacts when deciding how to allocate scarce resources.
However, cost-effectiveness estimates inherit many of the same comparability challenges identified by the CGD paper, and then add several more.
The estimated learning impacts that underpin cost-effectiveness calculations are often based on standardised effect sizes, with all the accompanying challenges of differences in assessment instruments, score distributions and reference populations. Converting those standardised effects into Learning-Adjusted Years of Schooling (LAYS) does not resolve these issues — standard deviation gains are simply translated into LAYS using a constant conversion factor, regardless of the assessment instrument, grade level or educational context (Angrist et al., 2025). The underlying comparability problem is carried forward, not resolved.
Cost-effectiveness estimates then layer two further assumptions on top, and both are rarely stated, let alone tested.
The first is that programme costs and impacts can be meaningfully compared across settings. Programme costs are themselves noisy estimates that vary sharply with local context, and rarely extrapolate cleanly from one setting to another (Popova and Evans, 2016). Costs expressed in PPP-adjusted dollars compound this: PPP conversion factors are themselves estimated with real uncertainty, and nominal exchange rates — particularly in the emerging-market contexts where most of these programmes run — can move sharply within the time it takes to design, cost and publish an evaluation. A cost-effectiveness estimate quoted as "100 USD per child" can mean something quite different in real terms by the time it is used to justify a scale-up decision, especially once a context has since experienced a significant devaluation.
The second is that the relationship between spending and learning is linear and scalable — in two separate senses that both need to hold, not one. It assumes that spending more money buys a proportionally larger dose of the intervention: more coaching visits, more contact hours, more intensive supervision. In practice, doubling a budget rarely doubles what actually reaches a classroom. It then assumes that a larger dose produces a proportionally larger learning gain, and this is where the assumption breaks down most visibly at either end of the cost distribution.
At the cheap end, a small, one-off intervention such as a single round of deworming tablets or a text-message information campaign can generate a spectacular cost-effectiveness ratio simply because dividing a modest effect by a tiny cost, then rescaling that ratio to a standard 100 USD budget, mechanically produces a huge number. Almost nobody actually spends 100 USD deworming one child, and there is a good reason for that: a child who has already been treated gains nothing from a second dose that term, so there is no way to spend the rest of that hypothetical 100 USD on more of the same intervention and get more learning for it. The ratio is being extrapolated far outside the range in which it was ever actually estimated.
At the expensive end, the same assumption fails in the opposite direction: programmes piloted under close, well-resourced supervision routinely show smaller effects once implemented by government systems at scale, even where per-child spending is comparable or higher. Either way — no more dose worth buying, or no more effect from the dose you can buy — a single cost-effectiveness ratio has no way of flagging that it has left solid ground.
Then the uncertainty disappears
Perhaps the most surprising feature of cost-effectiveness analysis is what happens next.
Individual impact evaluations routinely report confidence intervals or standard errors — researchers rightly acknowledge that programme impacts are estimates subject to sampling variation. Yet when those same studies are transformed into cost-effectiveness estimates, the uncertainty is almost always stripped away. Every component is uncertain: the estimated learning impacts have confidence intervals, programme costs are themselves estimates, and each further conversion — to LAYS, to a common currency, to a scaled-up projection — introduces additional uncertainty of its own. But by the time the final cost-effectiveness estimate appears, all of that uncertainty has disappeared, and readers are presented with a single point estimate that looks remarkably precise (see J-PAL’s Cost-Effectiveness page and Angrist et al. 2025)
Imagine if impact evaluations were reported in the same way. Suppose one programme was estimated to improve learning by 0.20 standard deviations and another by 0.25 — no researcher would conclude that the second programme was superior without first considering the confidence intervals. Yet we routinely compare cost-effectiveness estimates without any equivalent indication of uncertainty, and the result is that programmes are often ranked as though the differences between them are known with confidence, when in reality those differences may be well within the uncertainty of the underlying estimates.
Ironically, some of the most influential numbers in global education are accompanied by less information about uncertainty than the impact evaluations from which they were derived.
The next frontier for evidence synthesis
This is not an argument against systematic reviews, meta-analysis or cost-effectiveness analysis. On the contrary: education needs rigorous evidence synthesis more than ever. Few policymakers have the time to read hundreds of individual evaluations, and evidence syntheses play an essential role in helping decision-makers navigate an increasingly complex literature.
But perhaps the next frontier is not simply producing better summary measures. It is producing better summaries of uncertainty.
The recent CGD paper makes an important contribution by encouraging researchers to report raw learning gains alongside standardised effect sizes. We believe the same principle should apply more broadly. Evidence synthesis should become more transparent about the assumptions that underpin cross-study comparisons. Cost-effectiveness analyses should make explicit the assumptions required to compare interventions across contexts and scales — including how sensitive a ranking is to uncertainty on the cost side, not only the impact side. And where possible, summary measures should be accompanied by appropriate measures of uncertainty, just as individual impact evaluations routinely report confidence intervals.
There is a natural extension of that same idea on the cost side. Just as CGD argues that researchers should report the raw learning gain alongside the standardised effect size, cost-effectiveness analyses could report a measure of affordability alongside the $/LAYS or $/SD ratio: what the intervention costs in local currency, relative to what a government already spends per child on education. A programme that adds 5 percent to per-pupil spending and one that adds 60 percent are not comparable propositions for a ministry of finance, even when their headline cost-effectiveness ratios look similar once both have been converted to PPP dollars and divided by an assumed learning gain. Affordability is not a footnote to cost-effectiveness. For the official who actually has to find the money, it may be the more decision-relevant number of the two.
The goal is not to make evidence synthesis more complicated. It is to make it more honest, because policymakers do not need the illusion of precision — they need the best possible representation of what the evidence can and cannot tell us.
Further reading
The Illusion of Comparability Among Standardised Effect Sizes: Why Education Evaluations Should Report Raw Effects (Center for Global Development)
Cost-Effectiveness Analysis in Development: Accounting for Local Costs and Noisy Impacts (World Bank)
The Limits of a "Great Buy": What Rigorous Evidence Reveals about Teaching at the Right Level (AFLEARN)