Statistics is supposed to be the part of research that makes a finding trustworthy. In practice, it’s often the part that quietly undermines it. A striking share of published findings across psychology, medicine, and the social sciences have failed to replicate in follow-up studies — and a large body of methodological research now points to the same handful of statistical missteps showing up again and again, often made by well-intentioned researchers who simply weren’t taught to spot them.
None of this requires dishonesty. Most statistical mistakes in published research come from flexible, seemingly reasonable analytic choices — made one at a time, each looking harmless in isolation — that compound into results the data never actually supported. Understanding where these mistakes creep in is one of the highest-leverage skills a researcher can build, whether you’re finishing a dissertation, running a mixed methods study, or preparing a paper that needs to survive peer review.
This guide walks through the most common statistical errors in research, why they happen even to careful researchers, and what to do instead.
1. P-Hacking: Testing Until Something Sticks
P-hacking refers to running multiple analyses — different subgroups, variables, or model specifications — and reporting only the ones that produced a statistically significant result, without disclosing everything else that was tried.
It rarely looks like outright manipulation. It looks like: trying the analysis with and without one “odd” outlier, testing three outcome variables and writing up the one that mattered, or adding a covariate because it happened to push a p-value under 0.05. Each decision seems minor. But research modeling this behavior has shown that a study with just a few flexible choices — a handful of outcome variables, a couple of covariate options, an optional sample extension — can push the true false-positive rate well above the nominal 5% threshold, even though every individual test was run at a conventional significance level.
How to avoid it: Pre-register your primary outcome, your analysis plan, and your stopping rule before looking at the data. Report every test you ran, not only the ones that crossed the threshold, and be transparent when an analysis plan had to be adapted mid-study, along with the reasoning why.
2. HARKing: Hypothesizing After the Results Are Known
Closely related to p-hacking, HARKing happens when a researcher discovers an unexpected pattern in the data and then writes the paper as though that pattern was the original hypothesis all along. This erases the crucial distinction between confirmatory research (testing a pre-specified idea) and exploratory research (discovering a pattern worth testing later) — and it inflates the apparent strength of a finding that hasn’t actually been confirmed by anything.
How to avoid it: There’s nothing wrong with exploratory findings — they’re valuable. Just label them as exploratory, and treat them as a hypothesis for a future confirmatory study, not as confirmed evidence in the current one.
3. Overemphasizing P-Values Over Effect Size
A p-value tells you how surprising your data would be if the null hypothesis were true. It does not tell you how large, meaningful, or practically important the effect actually is. A study with a huge sample size can produce a tiny, practically meaningless effect that is nonetheless “statistically significant” — while a smaller, more meaningful effect in a modest sample might not cross the 0.05 threshold at all.
This is the classic confusion between statistical significance and practical significance: a result can be statistically real and still be too small to matter for policy, practice, or theory.
How to avoid it: Always report effect sizes and confidence intervals alongside p-values, and interpret the size of an effect in the context of what would actually be meaningful in your field — not just whether a threshold was crossed.
4. Misinterpreting a Non-Significant Result as “No Effect”
One of the most persistent misunderstandings in applied statistics is treating a non-significant result as proof that no effect exists. In reality, a non-significant p-value often just means the study didn’t have enough statistical power to detect the effect — it’s an absence of evidence, not evidence of absence.
How to avoid it: Run a power analysis before data collection to determine whether your sample size is even capable of detecting the effect size you care about. If a result comes back non-significant, discuss the study’s power honestly rather than concluding the effect doesn’t exist.
5. Multiple Comparisons Without Correction
Running many statistical tests on the same dataset — comparing several subgroups, several outcomes, or several time points — dramatically increases the chance that at least one test will appear significant purely by chance. Without adjusting for this, a researcher testing 20 independent hypotheses at the standard 0.05 threshold should expect roughly one false positive by chance alone, even if none of the underlying effects are real.
How to avoid it: Apply an appropriate correction — a Bonferroni adjustment for a small number of comparisons, or a Benjamini-Hochberg false discovery rate correction for larger sets of tests — and disclose the total number of comparisons made, not just the significant ones.
6. Outcome Switching
Outcome switching occurs when the outcome a study reports as its primary result differs from the outcome specified in the original study protocol or pre-registration — usually because the originally planned outcome didn’t produce a significant result. This is a particularly serious issue in clinical and health research, where it can meaningfully distort the evidence base that later informs practice guidelines.
How to avoid it: Register your primary and secondary outcomes before data collection begins, and if a deviation becomes necessary, disclose it explicitly along with the reason, rather than silently reporting a different outcome as though it were the plan all along.
7. Treating Correlation as Causation
It’s an old warning, but it remains one of the most common errors in applied research, particularly with observational data. Two variables moving together doesn’t establish that one causes the other — a third, unmeasured variable may be driving both, or the causal direction may run the opposite way from what’s assumed.
How to avoid it: Be precise with language — “associated with” rather than “leads to” or “causes” — unless your design (a randomized experiment, or a robust quasi-experimental method) actually supports a causal claim. Where causal inference is the goal but a true experiment isn’t feasible, methods like instrumental variables, regression discontinuity, or difference-in-differences designs exist specifically to strengthen causal claims from observational data — but they come with their own assumptions that need to be justified, not simply asserted.
8. Small Sample Sizes and Overreliance on Standard Errors
Small samples produce noisy, unstable estimates — a single unusual data point can shift results substantially. Compounding this, standard errors and confidence intervals are often misunderstood: a 95% confidence interval does not mean there’s a 95% probability the true value falls within that specific interval; it describes the long-run behavior of the method across repeated sampling.
How to avoid it: Conduct an a priori sample size calculation tied to a meaningful effect size, and interpret confidence intervals as describing plausible ranges given the method’s long-run properties — not as a direct probability statement about a single result.
9. Ignoring Assumptions Behind the Statistical Test
Every statistical test carries assumptions — normality, independence of observations, equal variances, and others — that determine whether its results are valid. Running a t-test or ANOVA without checking these assumptions is common, and can produce misleading results when the assumptions are meaningfully violated.
How to avoid it: Check assumptions before interpreting results, and know the appropriate alternative (a non-parametric test, a robust standard error correction, a transformed variable) for situations where the standard assumptions don’t hold.
10. Data Dredging in Large or Complex Datasets
With large datasets — particularly in fields now working with big data, survey panels, or administrative records — it’s tempting to explore dozens of variable combinations looking for “something interesting.” Without a pre-specified analysis plan, this kind of unstructured exploration will reliably surface spurious patterns that look meaningful but don’t replicate.
How to avoid it: Separate exploratory analysis clearly from confirmatory analysis. If you’re exploring a large dataset without a pre-registered plan, say so explicitly, and treat anything you find as a hypothesis for a follow-up confirmatory study — ideally on an independent sample.
Building Better Statistical Habits
The researchers who avoid these pitfalls consistently aren’t necessarily the most statistically sophisticated — they’re the ones who build a few habits into their workflow as standard practice:
- Plan the analysis before collecting data, including which outcomes matter most and what would count as a meaningful effect.
- Pre-register when possible — even an informal, time-stamped analysis plan shared with a supervisor or co-author creates useful accountability.
- Report everything you tried, not just what worked, including tests, models, and variables that didn’t make it into the final paper.
- Consult a statistician early, not just at the write-up stage — many of these errors are far easier to prevent at the design stage than to fix afterward.
- Read your results in terms of size and meaning, not just significance thresholds.
Final Thoughts
Most statistical mistakes in published research aren’t the result of bad intentions — they’re the result of flexible, individually reasonable-seeming choices made without a plan to constrain them. The fix isn’t becoming a statistics expert overnight; it’s building a small set of disciplined habits — pre-registration, transparent reporting, attention to effect size, and honest handling of non-significant results — into the research process from the start.
Statistical rigor isn’t a formality standing between you and publication. It’s what makes a finding worth publishing in the first place.

