Guides / How an average lies

Why does the average hide the truth?

4 min read

Fig. 21. A chart of two groups, each rising on its own across five checkpoints. Merge them into one line and the same data heads downward, because the merge quietly changes which group carries the most weight.

In brief

Because an average erases the very groups it was computed from. Combine two groups moving in different directions and the single number can point opposite to both of them, or hide two shapes that look nothing alike behind one tidy figure. One number can only hold one shape, and a crowd rarely lives inside just one. Berkeley’s 1973 admissions data looked biased against women in aggregate and mildly biased toward them once split by department, because women applied to the harder departments. The “average experience” a business reports is often a number no customer in the underlying pile of accounts actually had.

The pile of folders#

Picture a records floor at Berkeley in the fall of 1973: folders by the thousand, one per graduate applicant, sorted by department. Sweep every folder into a single pile and count who got in, and a stark number falls out. About 44% of the men who applied were admitted. About 35% of the women were. 8,442 men applied that year, 4,321 women, 12,763 folders in all, and the nine-point gap between those two percentages looked exactly like the thing a nervous administrator would dread finding.

Two statisticians, Peter Bickel and Eugene Hammel, working with a colleague named William O’Connell, ran the real analysis and published it in Science on 7 February 1975. Berkeley never actually got sued over the gap. Administrators feared they might be, and commissioned the study to see how deep the trouble ran before any lawyer arrived. What the three of them found, once they stopped looking at one giant pile and started looking at each department on its own, upended the story the aggregate number told. The number everyone reached for first was the pile-wide number. The number that actually explained what happened lived one level down, inside six subject departments nobody thought to check until the fear of a lawsuit forced the question.

The department that flips it#

Sorted by department, the picture reversed itself. Bickel and his co-authors found almost no individual department where men were admitted at a meaningfully higher rate than women. Add the departments back together the honest way, weighting each by its own size, and women came out very slightly ahead: a small but statistically significant bias in favor of women, in the paper’s own words. The university-wide gap everyone quoted first turned out to describe a shape none of the actual departments shared.

The reason sits in where people applied, and has nothing to do with who reviewed their files. Women disproportionately applied to the university’s more selective departments, English among them, where even a strong applicant of either sex faced long odds. Men disproportionately applied to departments such as engineering and chemistry, which accepted a much larger share of everyone who applied. Sort the six departments researchers cite most often and four of the six lean toward women once you look inside them. Take the single largest department on that list: men were admitted at about 62% there, women at about 82%, the exact reverse of the university-wide shape. Nobody’s file got read differently because of a name. The applicant pool sorted itself unevenly across rooms with very different odds attached, and the one number covering the whole building never saw the rooms at all. A dean glancing only at the pile-wide figure would have spent an entire budget correcting a bias that actually lived somewhere else entirely: inside admissions committees that, department by department, were largely behaving themselves.

Watch the reversal#

Here are two groups, watched over five checkpoints instead of six departments. Each one is climbing on its own terms, the whole way through. Flip the control and merge them into a single line, and watch what happens to two honest, rising trends the moment they share one number.

12345share of pool0%100%one average, fallinggroup agroup bcombinedNOLEMY GUIDES · PLATE 21After Bickel et al. 1975
Fig. 21. A chart of two groups, each rising on its own across five checkpoints. Merge them into one line and the same data heads downward, because the merge quietly changes which group carries the most weight.

Watch the size of the marks while you flip back and forth. One group starts big inside the pool and shrinks with every checkpoint. The other starts small and grows to take over. Neither group’s own line ever bends downward. The single merged line falls anyway, because the merge is quietly changing whose numbers carry the most weight from one checkpoint to the next. That is the entire mechanism behind the Berkeley numbers, stripped of the departments and the decade, down to the arithmetic underneath it. This is the whole Berkeley story, replayed at the scale of a single figure: no individual group ever moved the wrong way, and the pooled number still did.

Every average hides a shape#

Berkeley is one demonstration of a wider fact: a single summary number can be honest about the arithmetic and still lie about the shape underneath it. A statistician named Francis Anscombe proved this with almost cruel elegance in 1973. He built four different sets of paired data and arranged them so every one carried a mean x of 9, a mean y of 7.50, a correlation of 0.816, the same regression line, matching exactly. Plotted one at a time, the four datasets would look like four copies of the same experiment. Plotted side by side, they look nothing alike, even though every summary statistic describing them stays identical.

Statisticians of Anscombe’s day treated a plotted graph as the rough cousin of a clean calculation, useful for a classroom wall and little else. He built the quartet to answer them directly, and the line still quoted back at students is his own: numerical calculations are exact, but graphs are rough.

Anscombe’s trick found new legs in 2017, when Autodesk researchers Justin Matejka and George Fitzmaurice pushed the same idea further and built the Datasaurus Dozen: thirteen datasets, one shaped unmistakably like a dinosaur, engineered so their headline statistics matched to the same tight tolerance. Look at the summary numbers alone and every one of the thirteen reads as an ordinary, unremarkable scatter. Look at the actual plot and one of them draws a dinosaur. Forty-four years apart, the same lesson landed twice: the number describing a group and the shape of the group itself are two separate facts, and only one of them shows up inside a spreadsheet cell.

Berkeley’s aggregate rate and Anscombe’s four identical regression lines are the same warning, stated twice. An average, a star rating, a headline rate: each one collapses a pile of different individual experiences into a single figure, and that figure can be arithmetically correct while describing an experience no single member of the group actually had. Whenever a single number speaks for a group you have never looked inside, that is the same instinct behind asking who actually wrote the reviews sitting behind it: go find the departments hiding underneath the number.

The objection#

Bickel, Hammel & O’Connell, “Sex Bias in Graduate Admissions: Data from Berkeley,” Science 187 (1975), pp. 398-404, cross-checked against the Stanford Encyclopedia of Philosophy’s entry on Simpson’s paradox. The no-lawsuit correction comes from a researched account of the Berkeley retelling. Francis Anscombe, “Graphs in Statistical Analysis,” The American Statistician 27 (1973), pp. 17-21. All fetched 18 July 2026.