What Do Average Scores Hide?

28 Sep 2026
Children counting on their fingers
28 Sep 2026

What Do Average Scores Hide?  

Why understanding children's learning means looking beyond headline indicators

Learning assessments are often summarised using a single statistic. Sometimes it is an average score. Increasingly, it is the percentage of children reaching a minimum proficiency benchmark.

These indicators are indispensable. They allow governments to monitor progress, compare education systems and communicate results in a way that is readily understood by policymakers and the public. Without summary indicators, large-scale learning assessments would be difficult to interpret and almost impossible to use for accountability.

Yet every summary comes at a cost. Reducing an assessment to a single number inevitably conceals much of what the assessment has measured. Children do not learn in averages, and education systems rarely face a single learning challenge. Understanding how children learn — and why learning differs between countries — requires looking beneath the headline indicators.

One of the greatest advantages of publicly available assessment microdata is that they make this possible.

Average scores are only the beginning of the story

Consider Comoros and the Democratic Republic of the Congo. Both countries post almost identical average mathematics scores — around 8 out of 21 items correct. At first sight, they appear to have achieved similar learning outcomes.

Now look beneath those averages. Comoros has the fourth-highest proportion of children reaching the programme-specific proficiency benchmark in the entire sample, at around 9%. In the Democratic Republic of the Congo, virtually no child reaches that same benchmark. Although the averages are identical, the underlying distributions — and the educational challenges each country faces — are fundamentally different.

Comoros may need targeted support for the substantial group of children scoring at the very bottom of the distribution, while continuing to stretch those already close to proficiency. The Democratic Republic of the Congo faces a more fundamental challenge: almost no child, at any point in the distribution, has yet acquired foundational numeracy.

The average score cannot distinguish between these situations. Assessment microdata can.
One of the first analyses we undertook in our report was therefore not to compare average scores but to examine the full distribution of children's mathematics performance. The resulting figures immediately revealed patterns that would have remained invisible had we focused only on national means.

The lesson is straightforward: averages are valuable summaries, but they should be viewed as the starting point of analysis rather than its conclusion.

Minimum proficiency tells a different — but still incomplete — story

The increasing emphasis on minimum proficiency reflects an important shift in international education. Rather than focusing solely on average achievement, attention has moved towards whether children have acquired the essential skills needed to participate effectively in school and society.

This is a welcome development. Minimum proficiency benchmarks provide an intuitive measure of whether education systems are ensuring that children acquire foundational competencies. They are particularly valuable for monitoring progress towards Sustainable Development Goal 4.1.1.

However, proficiency benchmarks are also summary statistics. They divide a continuous distribution of learning into two groups: children who have reached the benchmark and those who have not.

This distinction is useful for monitoring, but it inevitably conceals important variation below — and above — the threshold. Children who narrowly miss the benchmark are treated in exactly the same way as children who cannot answer even the simplest questions. Conversely, children who comfortably exceed the benchmark are grouped together with those who only just reach it.

Like average scores, proficiency measures simplify a much richer picture of learning.

Similar headline indicators can conceal very different learning distributions

Tunisia and Nigeria illustrate this vividly. Both countries post almost identical proficiency rates — around 16 to 17% of children reach the programme-specific benchmark. Yet the distributions beneath that number could hardly be more different. In Nigeria, a large share of children score at or near zero. In Tunisia, very few children score zero; instead, most cluster just below the benchmark, close to acquiring it. Two countries, the same headline number, and two entirely different educational challenges.

Looking at both measures changes how we understand inequality

Our report also illustrates this using household assessment data to compare wealth inequalities in learning: the same underlying gap can look negligible or severe depending on whether it is measured using proficiency benchmarks or average scores. We explore this in full — including just how large some of these wealth gaps are — in the next AFLEARN Insight. 

Looking inside the assessment

Perhaps the greatest contribution of publicly available microdata is that they allow researchers to move beyond total scores altogether.
Instead of asking only how well children performed, we can ask which skills they have mastered.

In mathematics, for example, assessments typically include tasks covering number recognition, quantity comparison, addition and pattern recognition. Reading assessments distinguish between skills such as decoding, oral reading and comprehension.

Examining performance on these individual components transforms the assessment from a monitoring instrument into a diagnostic tool.

Mathematics competencies
Looking inside the assessment reveals where children's numeracy begins to break down

Two countries with similar average scores may display very different patterns of strengths and weaknesses. One may struggle primarily with basic number concepts, while another performs well on simple operations but encounters difficulties with more advanced reasoning. Similarly, children may decode words accurately yet struggle to understand connected text.

Reading competencies

The same is true for reading. In Uganda, for example, 76% of children score zero on listening comprehension — the ability to understand simple spoken language — yet around half can correctly recognise every letter sound they are shown. Decoding, in other words, is running well ahead of oral language comprehension. Performance falls further still on connected text: 83% of children score zero on oral reading comprehension.

No single reading score could reveal this pattern. Only by examining individual components can we see precisely where a child's reading development is, and is not, breaking down. These distinctions have important implications for curriculum design, teacher support and classroom instruction. They also illustrate why preserving detailed assessment data matters: without access to item-level information, many of these insights would simply disappear once the national report had been published.

Better diagnosis leads to better policy

Learning assessments are often viewed primarily as accountability tools. They identify whether learning levels are improving and whether systems are meeting national or international targets.

But they can also serve another purpose. When analysed in sufficient depth, they become diagnostic tools that help explain why learning differs across countries and identify where children begin to struggle.

This distinction is important. Monitoring tells us whether there is a problem. Diagnosis helps us understand what kind of problem we are trying to solve.

Published reports necessarily focus on the former. Publicly available microdata make the latter possible.

Looking beyond the headline indicators

The widespread use of average scores and minimum proficiency benchmarks has undoubtedly improved the monitoring of foundational learning. These indicators remain essential for tracking progress and communicating results.

However, they should not be mistaken for complete descriptions of children's learning.

One of the central arguments of our report is that publicly available microdata allow researchers to move beyond these summaries and recover much of the richness that is otherwise hidden within assessment datasets. Looking at score distributions, comparing different summary measures and examining individual assessment components all contribute to a deeper understanding of learning.

Headline indicators will always remain important. But the most interesting — and often the most policy-relevant — stories are usually found beneath them.

AFLEARN Insight #1       AFLEARN Insight #2       AFLEARN Insight #3      Read the report