In September 2020, two respected institutions measured the exact same thing — the price of Seoul apartments — and one reported a monthly rise 7.3 times larger than the other. Same city, same month, same homes. The gap was not fraud. It came from choices: which apartments to sample, how to average them, where to start the axis. That single number is a clean lesson in how statistics lie — not by inventing figures, but by choosing which figures to show you. This is part four of a series about not being fooled — the previous part took apart expert authority, why a credential is not the same as being right — and it puts one stubborn question to every chart you meet: but — is this real?
We trust numbers because they look objective. A figure feels like it fell out of the sky, clean and neutral, with no fingerprints on it. But a statistic is a human choice wearing the costume of objectivity. Someone decided what to count, what to leave out, and how to frame it — and those decisions can quietly steer you to the opposite conclusion.
The problem just got worse. A chart used to take effort to build, which slowed down the people making them. Now anyone can prompt an AI to generate a slick, confident, professional-looking chart in about three seconds. Smooth fakes are cheaper than ever, and a smooth fake outruns a rough truth every time.
Key Takeaways
- The same honest data can be drawn two ways — a gentle slope or a cliff — depending only on where the axis starts and which window you crop.
- Most damage isn’t fraud: survivorship bias, Simpson’s paradox, and correlation-as-causation fool honest analysts too.
- You don’t need a statistics degree — just six questions, and one killer question: “could this same number prove the opposite?”
The Oldest Trick: How Statistics Lie Without Inventing Anything
In 1954, a writer named Darrell Huff published a small book with a shameless title: How to Lie with Statistics. Seventy years later it is still the standard primer, because the tricks never went out of style.
His most famous one is the “gee-whiz graph.” Take a boring, gentle rise — say a number creeping from 3% to 4%. Start the vertical axis at zero and it looks like a flat road. Chop the axis so it starts at 3% instead, and the exact same data suddenly looks like a cliff. Nothing about the numbers changed. Only the frame did.
Think of it like a camera. Point it up at someone and they look towering; crouch and shoot down and they look small. The person is the same height. The photographer chose the impression. A chart’s axis is that camera angle.
FIG. 01 — SAME DATA, TWO AXES
One dataset, opposite impressions
Axis starting at zero
Axis chopped off
3% to 4% rise
3% to 4% rise
A gentle uphill
A steep cliff
"Barely moved"
"It exploded"
Calm
Fear or excitement
SOURCE: Darrell Huff, How to Lie with Statistics (1954)
This is the gentlest form of how statistics lie, because no one technically lied. Every number on the page is true. The deception lives entirely in what got emphasized and what got hidden — and that is exactly why it is so hard to catch.
Cherry-Picking the Window
The second everyday trick is choosing when the story starts. Pick the window that flatters your point and drop the rest.
US robbery figures are a tidy example. Look only at 2014 to 2016 and the line shoots up — a crime wave. Zoom out to 2014 through 2019 and the later years fall back down. Same dataset, opposite headline, decided entirely by where you put the scissors.
You have seen the investing version. A fund shows you its return “since 2023” and quietly skips the crash of 2022. The window isn’t neutral. It was selected precisely because it tells the story someone wants told.

When the Numbers Themselves Fool You
Axis tricks are the easy case, because someone chose to mislead. The scarier category is the one where honest, careful analysts fool themselves — and then fool you. This is the second way to understand how statistics lie: not through anyone’s malice, but through a blind spot no one noticed. Here nobody is a villain, and that is the whole point.
Survivorship Bias: The Data That Didn’t Come Back
In World War II, the US military studied bombers returning from missions and mapped where the bullet holes clustered — the fuselage and wings. The obvious fix: add armor where the holes are. The statistician Abraham Wald stopped them cold.
The holes on the survivors, he pointed out, mark the places a plane can be shot and still fly home. The engines had almost no holes — not because they never got hit, but because the planes hit there never came back to be counted. Armor the engines. The most important data was the data that was missing.
The modern version is everywhere. Read ten interviews with founders who made it and try to extract “the formula for success,” and you are studying only the planes that returned. The startups that did the exact same things and died were never invited to the interview.
Simpson’s Paradox: When the Whole and the Parts Disagree
In 1973, graduate admissions at UC Berkeley looked damning: men were admitted at 44% and women at 35%. It reads like open discrimination.
Then someone broke it down by department. Within almost every individual department, women were admitted at equal or higher rates. The overall gap appeared because women applied in larger numbers to the most competitive departments, where everyone’s odds were low. The aggregate said one thing; the pieces said the opposite. (Worth flagging: the famous “lawsuit” over this is an urban legend that grew up around the case — the statistical reversal is real, the courtroom drama is not.)
Correlation Is Not Cause
Ice cream sales and drowning deaths rise together, tightly. Ban ice cream to save lives? Of course not — a hidden third factor, summer heat, drives both. The two are holding hands, but neither is steering.
Children’s shoe sizes and reading ability climb together too; the hidden driver is simply age. And the number of films Nicolas Cage released in a year once tracked the number of pool drownings — pure coincidence, a reminder that two lines moving in step prove nothing on their own.
FIG. 02 — HOW A TRUE NUMBER BECOMES A LIE
Data to distortion in four steps
RAW DATA
An honest original
A dataset where every single number is true and nothing at all has been faked.
THE EDIT
Selective framing
Someone picks the axis, crops the time window, or swaps the denominator to favor one story.
THE CHART
A polished picture
The choice is rendered as a clean, confident, professional-looking chart in seconds.
THE READ
You absorb it
The viewer takes in the conclusion as fact, never seeing the choices baked in behind it.
SOURCE: ThoughtSpot; Sigma Computing
The Average That Describes No One
Nine employees each earn 30 million won a year. The CEO earns 3 billion. The “average” salary at that company is about 327 million won — a figure that describes not one single human in the building. One extreme value dragged the average into fantasy. The honest number here is the median — the person standing in the exact middle — which is 30 million. Whenever a headline shouts an “average,” ask where the middle actually sits.
Korea’s Number Wars: The Illusions in Your Own Life
None of this is a museum piece. The same parts run through the numbers that decide your rent, your vote, and your savings.
Start with the housing fight that opened this piece. For 2020 and 2021, the Korea Real Estate Board reported Seoul apartment gains of 3.01% and 8.02%. KB Kookmin Bank’s index reported 13.06% and 16.4% for the same years — and in September 2020, KB’s monthly figure ran about 7.3 times the Board’s. The difference traces back to method: the Board leaned on a smaller sampled survey of roughly 32,000 units, while KB’s index drew on about 62,000 entries reported through brokerages.
| Seoul apartment price index | 2020 rise | 2021 rise | Sample basis |
|---|---|---|---|
| Korea Real Estate Board | 3.01% | 8.02% | ~32,000 units (sampled survey) |
| KB Kookmin Bank | 13.06% | 16.4% | ~62,000 entries (brokerage input) |
Neither index is “the lie.” They measure with different nets. But which index a politician quotes decides whether a housing policy gets called a triumph or a disaster — and the citizen paying the rent is left holding two truths that point in opposite directions.

The Poll That Says Less Than It Seems
Poll season brings its own illusions. At the standard 95% confidence level, a survey of 1,000 people carries a margin of error of about ±3.1 percentage points; 500 people widens it to ±4.4; 2,000 tightens it to ±2.2. So a candidate “leading” 46 to 44 in a 1,000-person poll is not leading at all — both numbers sit inside the same fog.
There’s a deeper crack too. When only about one in ten people answer a phone interview — and far fewer an automated one — the ten percent who pick up may not resemble the ninety who didn’t. “Within the margin of error” quietly gets sold to you as “surging ahead.”
The Backtest That Fits the Past and Breaks in the Future
Investing ads love a flawless past. A strategy shows a perfect upward curve with no losing stretch, and it feels like proof. It is usually the opposite of proof.
Fit a model tightly enough to old data and it starts memorizing noise instead of signal — a trap called overfitting, or curve-fitting. It looks immaculate on the past it was built from and falls apart on data it has never seen. There is a plain way to hold this: the fact that it worked on history is not the same as it working on your future. A backtest is a plan. Reality is the punch that lands afterward.
The Six Questions That Show How Statistics Lie
Here is the good news the cynics miss. You don’t need a statistics degree to defend yourself. You need a short habit — a handful of questions you run before you let a number into your head.
FIG. 03 — THE SIX-QUESTION HABIT
A field guide to how statistics lie
| # | The question to ask | What it catches |
|---|---|---|
| 1 | Does the axis start at zero, with even gridlines? | Axis truncation, the gee-whiz graph |
| 2 | Why this exact time window? | Cherry-picked periods |
| 3 | What is the denominator: absolute vs rate, base rate? | Ratio illusions |
| 4 | Is a correlation being sold as a cause? | Hidden third variable, spurious links |
| 5 | Who was sampled, from where, and how many? | Representativeness, margin of error |
| 6 | Could this same number prove the opposite? | The killer question |
SOURCE: Huff; Ellenberg; Gallup Korea
These six checks are the whole field guide to how statistics lie. Run any chart or claim past them and most manipulation surfaces on its own. Is the axis honest? Why this particular window? What’s the denominator? Is a correlation being sold as a cause? Who was sampled, and how many? And the last one carries the most weight.
The Killer Question
The single most powerful move in statistical self-defense is to ask: could this same number be used to argue the opposite? If one dataset can be dressed to prove a boom and dressed again to prove a bust, then the number was never the argument — the framing was. Once you can see the framing, the spell breaks.
And notice what this habit is not. It is not “all statistics are manipulation,” and it is not the reflex to sneer at every figure. Good statistics are how we understand a world too big to see by eye. The goal isn’t to throw the good ones out — it’s to be able to tell them apart from the dressed-up ones. This is the same instinct the naturalization series keeps circling: power’s oldest move is dressing a choice up as nature, and a number is just its newest costume.
Bottom Line. A statistic rarely lies by inventing a figure; it lies by choosing which figure to show — so the defense isn’t distrust, it’s the reflex to ask “compared to what, over what window, out of what total?”
Everyday Takeaway. Next time a chart makes you feel something fast — panic, relief, certainty — pause before you share it, and ask the killer question: could this exact number be drawn to prove the opposite? Verification isn’t suspicion. It’s a habit.
Frequently Asked Questions (FAQ)
Q. What is the fastest way to see how statistics lie in a chart? A. Check the vertical axis first. If it does not start at zero, or the gridlines are not evenly spaced, the chart may be exaggerating a small change into a dramatic one. That single glance catches the most common visual trick, the “gee-whiz graph,” before the rest of the image can work on you.
Q. If numbers can mislead this easily, should I just distrust all statistics? A. No — that is the opposite mistake. Most misleading numbers come from selective framing, not invented data, and good statistics remain the best tool we have for understanding a world too large to eyeball. The aim is not blanket suspicion but a habit of asking a few questions so you can tell a solid figure from a dressed-up one.
Q. Does “within the margin of error” mean a poll is useless? A. Not useless, just more modest than the headline suggests. A gap smaller than the margin of error — for a 1,000-person poll, about ±3.1 percentage points — means the two figures are statistically tied, not that one is winning. Treat a lead inside that range as “too close to call,” regardless of how the coverage frames it.
References
- How to Lie with Statistics — Wikipedia
- How to Lie With Statistics — Shortform Summary
- The Legend of Abraham Wald — AMS Feature Column
- Abraham Wald and the Missing Bullet Holes — Ellenberg excerpt
- Simpson’s Paradox: Gender Bias at Berkeley — refsmmat
- Gender Bias, Simpson’s Paradox — Berkeley Math Circle (PDF)
- How to Identify Misleading Graphs — ThoughtSpot
- When Good Graphs Go Bad — Sigma Computing
- Correlation versus Causation — Statistics LibreTexts/03%3A_Relationships/14%3A_Correlations/14.03%3A_Correlation_versus_Causation)
- Spurious relationship — Wikipedia
- KB vs Korea Real Estate Board apartment index gap — Hankyung
- Board “narrowing” vs KB “widening” — Bizwatch
- Mean, median and mode — JMP Statistics Knowledge Portal
- Polling FAQ — Gallup Korea
- 3 Simple Ways To Reduce Curve-fitting — Build Alpha
- Statistical Overfitting and Backtest Performance — Bailey et al. (PDF, LBL)
The Verification Age — Series
- Part 1. Spotting Deepfakes: When Seeing Stopped Being Proof
- Part 2. Verifying AI Content: Why Fluent Never Means True
- Part 3. The Quiet Skill of Vetting Expertise: A Title Is Not the Truth
- Part 4. How Statistics Lie: Six Questions to Catch a Chart in the Act (this article)
- Part 5. Verifying Your Own Judgment: The Easiest Person to Fool Is You
Action sequel to The Anatomy of Naturalization series — from spotting the arbitrary to testing what is real.
