Hypothesis Testing in Biology: The 5-Step Method (2026)

Quick answer: Hypothesis testing in biology is a five-step loop: state H₀ (no effect) and H₁, set α (usually 0.05), pick a test matched to your data and design, compute the statistic and p-value, then decide — reject H₀ if p < α. The whole machinery exists because biology is noisy: you can never "prove" an effect, only bound how surprising your data would be if there were none. The general logic is in our hypothesis testing guide; here it’s applied to wet-lab and field data.

The five steps on a real bio example

Question: does a new fertilizer change plant height after 8 weeks? 20 plants per treatment.

  1. H₀: mean height fertilized = mean height control (the fertilizer does nothing). H₁: the means differ.
  2. α = 0.05 — the false-positive rate you accept up front, never after seeing the data.
  3. Test: two independent groups, one continuous outcome → independent-samples t-test (check roughly normal residuals; with n=20 per group it’s forgiving).
  4. Compute: t = difference in means ÷ standard error of that difference; software gives p directly.
  5. Decide: p = 0.003 < 0.05 → reject H₀; the data are unlikely under "no effect." Report the means, not just "significant."

Matching the test to the biological design

DesignOutcomeTest
Two independent groups (treated vs control)ContinuousIndependent t-test (Mann-Whitney if very non-normal)
Same subjects before/after, paired organsContinuousPaired t-test
Three+ groups (doses, species, sites)ContinuousOne-way ANOVA + post-hoc (Tukey)
Counts in categories (infected/not by sex)CountsChi-square / Fisher’s exact (small n)
Association between two measurementsContinuousCorrelation / regression

The three mistakes biology reviews catch

  • Pseudoreplication: treating 50 cells from one culture flask as n=50. The replicate is the flask, not the cell — inflate n and every p-value becomes meaningless.
  • p-hacking by iteration: adding animals or repeating the assay until p dips under 0.05. Decide n and stopping rules before starting; a p-value near 0.06 is “not demonstrated,” not “almost there.”
  • Confusing significance with size: a huge n makes a 0.4% difference “significant.” Always report effect size with a confidence interval next to the p — the meaning of p is covered step by step in our p-value explainer.

Context for the wider toolkit — data types, distributions, sampling — is in our introduction to biostatistics, and why biology leans on statistics at all in this overview. To run these tests yourself, point-and-click, our jamovi guide walks t-tests and ANOVA with zero code.

For one-to-one coaching on designing your thesis experiments and analyses, Ampersand Academy teaches biostatistics and R one-to-one.

Frequently asked questions

What does p < 0.05 actually mean in a biology experiment?

If there were truly no effect, results at least this extreme would occur less than five percent of the time. It is a statement about the data under the null hypothesis – not the probability the hypothesis is true.

Why can a study never prove the alternative hypothesis?

Hypothesis tests only measure how incompatible the data are with the null. Evidence accumulates across replication and design quality; a single significant result supports, but cannot prove, a biological claim.

What is pseudoreplication and why is it fatal?

It is counting subsamples from one biological unit as independent replicates – 50 cells from one flask as n=50. The true sample size is the flask, so the test dramatically overstates certainty.

When should I use ANOVA instead of multiple t-tests?

Whenever you compare three or more groups. Running all pairwise t-tests inflates the false-positive rate; ANOVA gives one overall test, then post-hoc comparisons like Tukey’s control for multiplicity.

Is p = 0.06 a failed experiment?

No. It means the effect was not demonstrated at your chosen threshold with this sample size. Report the effect size and confidence interval – the result may be biologically meaningful and detectable with more power.