Statistical errors rarely announce themselves. A p-value calculated with the wrong test, an effect size never reported, a handful of extra analyses run quietly until something reaches significance, none of these look dramatic in the moment. But they are exactly the kind of issue that can unravel a defense, draw a rejection from a journal, or, in more serious cases, contribute to a finding that simply does not replicate. Understanding the statistical errors that trip up graduate students specifically, and knowing how to avoid them before they become embedded in your thesis, is one of the highest-value things you can do before your analysis chapter is finalized.
This guide covers the specific quantitative data analysis mistakes thesis committees and journal reviewers flag most often, why they happen even among careful researchers, and the practical habits that prevent them.
Why these errors are so common, even among careful researchers
Most statistical errors in graduate research are not the result of carelessness or dishonesty. They happen because statistical training is often compressed into a single course years before a student's actual analysis, because deadline pressure creates incentives to find significant results, and because the line between exploring data and confirming a hypothesis becomes genuinely blurry in the middle of real analysis. Understanding this helps frame the goal correctly: not achieving statistical perfection, but building habits that catch these errors before they reach your final document.
Error 1: Running the wrong statistical test for your data
This is among the most fundamental, and most common, errors in student research. Using a test designed for one type of data or research question on data or questions it was not built for produces results that are, at best, difficult to interpret and, at worst, actively misleading. It shows up in familiar ways: running a t-test to compare three or more groups rather than an ANOVA, which inflates the risk of a false positive across multiple comparisons; using Pearson correlation on ordinal data where Spearman's correlation is more appropriate; or treating Likert-scale data as fully continuous in analyses that assume interval-level measurement, without justification or acknowledgment of the ongoing debate around this practice in your field.
The fix is to decide on your statistical tests based on your research questions and data types before you begin analysis, not after seeing preliminary results, and to confirm your choice against a statistics reference or your advisor before running your final analysis.
Error 2: Ignoring or failing to check statistical assumptions
Every parametric statistical test carries assumptions, normality, homogeneity of variance, independence of observations, and violating these assumptions without acknowledgment can invalidate your results even when your test selection was otherwise appropriate.
Build assumption checking into your analysis workflow as a mandatory step, not an optional one. Run normality tests, inspect residual plots, and check variance homogeneity before interpreting your main results, and have a documented plan for what you will do, transform your data, switch to a non-parametric alternative, or proceed with acknowledgment and justification, if assumptions are violated.
Error 3: P-hacking and data dredging
P-hacking and data dredging are treated as serious methodological concerns in academic writing, not minor technical missteps, and it is worth understanding exactly what these terms mean and why they matter. P-hacking refers to running many statistical tests, or the same test with slightly different specifications, until one produces a statistically significant result, then reporting only that result as though it were the single planned analysis. Data dredging, sometimes called fishing, refers to exploring a dataset extensively without a pre-specified hypothesis, then presenting whatever pattern emerges as though it had been predicted in advance. Both practices greatly increase the chance of finding false positives, since running enough tests on any dataset will eventually yield a statistically significant result just by chance, even if the data are completely random.
Guarding against this means pre-registering your hypotheses and planned analyses before collecting or analyzing data where your program or field supports the practice, and, if you do run exploratory analyses beyond your pre-specified plan, which is legitimate and often valuable, labelling them explicitly as exploratory in your write-up rather than presenting them as confirmatory findings. It also means applying appropriate corrections for multiple comparisons, such as Bonferroni or False Discovery Rate corrections, when running numerous tests on the same dataset, and being transparent about your full analytical process, including analyses that did not reach significance, rather than selectively reporting only significant findings.
"Running enough tests on any dataset will eventually yield a statistically significant result just by chance, even if the data are completely random."
Error 4: Reporting statistical results incompletely
Incomplete or inaccurate reporting can undermine even a well-run analysis. Reporting results correctly, in the way dissertation committees expect, means including a specific, consistent set of information every time you present a statistical finding: the test statistic, degrees of freedom, the exact p-value, effect size rather than significance alone, and confidence intervals where applicable to the conventions of your field.
The mistakes that undermine this tend to repeat themselves. Reporting p-values without effect sizes leaves readers unable to judge whether a significant result is practically meaningful. Rounding p-values in a way that obscures their actual value, reporting "p < .05" when the actual value was .049, right at the threshold, versus .001, deeply significant, communicates very different levels of evidence under an identical label. Misusing the term "trend toward significance" for results that did not reach conventional thresholds is widely considered poor practice in contemporary statistical reporting. And failing to report non-significant results at all distorts the overall picture of your findings, which can constitute a form of selective reporting in its own right.
Error 5: Confusing statistical significance with practical importance
A statistically significant result tells you that an observed effect is unlikely to be due to chance alone, given your sample size. It does not, by itself, tell you whether that effect is large, meaningful, or practically important. With a sufficiently large sample, even trivially small effects can reach statistical significance.
Always report and interpret effect size alongside significance, and discuss your findings' practical or clinical significance explicitly in your discussion, separate from their statistical significance.
Error 6: Conflating correlation and causation
This is a well-known error in principle, but one that recurs again and again in practice, often in subtle choices of language rather than explicit claims about causation. Phrases like "this led to," "resulted in," or "caused," applied to correlational or observational data, even unintentionally, overstate what your design can actually support.
Match your language precisely to your design. Correlational and observational studies support language like "was associated with" or "predicted," not causal language, unless your specific design, a true experiment with random assignment, genuinely supports causal inference.
Error 7: Inadequate handling of missing data
Simply deleting all cases with any missing data, listwise deletion, is common but can introduce bias, reduce statistical power unnecessarily, and is increasingly viewed skeptically by methodologically informed reviewers, especially when a substantial portion of a dataset is affected.
Report the amount of missing data and analyze whether it appears to be missing at random or systematically related to other variables. Depending on the pattern and amount of missingness, consider appropriate approaches such as multiple imputation, and always report your chosen approach and rationale transparently.
Error 8: Failing to acknowledge underpowered studies
Running a study with a sample too small to reliably detect the effect size you are looking for increases the risk of both false negatives, missing a real effect, and, counterintuitively, exaggerated effect size estimates among the significant results that do emerge by chance.
Where possible, run a power analysis before collecting data, so you can estimate a sample size adequate for your expected effect size and statistical power. If your final sample is smaller than ideal due to practical constraints, acknowledge this limitation explicitly and discuss its implications for interpreting your results, rather than presenting underpowered findings without qualification.
Building better statistical habits into your research process
The researchers who avoid these errors consistently share a few habits rather than a single trick. They decide their full analysis plan before collecting or analyzing data, including specific hypotheses, planned tests, and how they will handle assumption violations and missing data. They keep a detailed analysis log of every test they run, even exploratory analyses that never make it into the final write-up, so they can accurately and transparently describe their entire analytical process if asked. They have a second, statistically literate reader review the analysis before finalizing the results chapter, ideally someone not deeply invested in a particular outcome. And when in doubt, they consult a statistician or their program's statistical support resources early, not after they have already committed significant time to an analysis path that may need to change.
Rigor doesn't slow your work, it protects it
Good statistical practice may seem like it adds more friction to an already demanding process, but the alternative, finding a fundamental analytical error during your defense or from a journal reviewer, takes far more time and credibility than building these habits in from the start. Treating statistical rigor as integral to good research, rather than a bureaucratic hurdle layered on top of it, protects both your findings and your standing as a careful, trustworthy researcher.
Statistical errors are often invisible to the researcher who made them, which is exactly why a second, expert set of eyes matters.
Protect your research from costly analytical errors
Get expert data consulting and statistical review for your thesis or dissertation, catching errors in test selection, assumption checking, and results reporting before they reach your committee or a journal reviewer.
See research & data analysis services →