More data is not always better data
Why sample quality beats sample size.
More data is not always better data because a systematic flaw in how a sample was gathered decides whether the numbers describe the thing they claim to, and collecting far more of a biased sample only produces a far more confident-looking version of the same wrong answer. A pile of data earns its trust from where it came from, and the height of the pile has nothing to say about that.
Random scatter shrinks, a steady push does not
A sample is meant to stand in for a larger population it was drawn from, and the point of collecting more of it is to shrink the random noise sitting around the true value, since random errors in individual measurements tend to cancel out as more of them are averaged together. This only works when the errors are truly random, scattered evenly above and below the true value with no consistent direction to them.
A systematic bias pushes results the same way every single time, and that changes everything about how it responds to volume. Collecting more data from a biased process does shrink the random scatter around whatever number the process happens to be producing, but it leaves that number exactly as far from the truth as it was, since the bias is present in every measurement equally and has no tendency to cancel out. Random scatter around an average falls roughly with the square root of the sample size, so a hundred times more data cuts it to a tenth, while a bias stays the same size at a hundred readings or a million. Worse, a large biased sample tends to look more trustworthy than a small one purely by virtue of its size, tightening the reported scatter around a wrong number and making it appear more precisely known than it is.
A smudge on the photocopier glass
Photocopying a document with a faint smudge across one corner, then photocopying that copy again, produces a copy with the same smudge in the same place. Running off a hundred more copies from the same smudged master leaves the smudge undiluted on every one of them, because a photocopier reproduces whatever is on the page in front of it and has no access to a cleaner original hiding underneath.
A biased sampling method works the same way on data. If the method itself favours a certain kind of result (asking only people outside one particular building, timing measurements only during one part of a process's cycle, using an instrument with a fixed calibration error), every additional data point collected through that method reproduces the same bias faithfully, adding volume without adding accuracy. The only way to fix a smudged page is to clean the master before copying again, and the only way to fix a biased sample is to change how it is gathered.
Two million ballots that called it wrong
The best-known case is a 1936 magazine poll of a national presidential election, which collected more than two million returned ballots from lists that leaned heavily towards wealthier households, and predicted the loser would win comfortably. Surveys a small fraction of that size, drawn deliberately to match the voting population, called the result correctly. The smaller surveys' careful method mattered more to their accuracy than the larger poll's sheer size ever could, because size shrinks random noise, and only a representative sampling method fixes bias.
The same trap catches modern data sets gathered from whatever source happens to be convenient: an app's own users, a website's own visitors, a factory's day-shift measurements. Each of these can quietly exclude or over-represent some part of the wider population the conclusion is meant to describe, and the exclusion survives any increase in volume.
Questions to ask before trusting a big sample
Before trusting a result because it was drawn from a very large data set, the more useful question is almost always how that data set was gathered. Did every relevant case have a fair chance of being included? Did the collection method favour certain outcomes over others? Do the conditions under which the data was collected match the conditions the conclusion is being applied to? None of these questions require statistical training to ask, only a willingness to look past an impressive-sounding sample size before accepting the conclusion built on top of it.
A modest, carefully gathered sample that has been checked for these things is worth substantially more than an enormous one collected however happened to be convenient. A study confident enough to describe its own sampling method in detail is usually the study most worth believing, since it has nothing to hide behind a large total count.
When a bigger sample is the right call
None of this argues against collecting a large sample. A large, well-gathered sample beats a small, well-gathered one, shrinking the scatter around an already trustworthy number and making it easier to detect a real, modest-sized effect against the noise. That extra sensitivity is the whole reason large studies are run at all.
The caution applies to the temptation to treat size as a substitute for checking the method, when it is really a separate benefit that only pays off once the method is sound. A very large sample gathered carelessly amounts to a large number of copies of the same underlying flaw, and recognising that is the reason to ask about a method before asking about a size.