Quick answer: Hypothesis testing in biology is a five-step loop: state H₀ (no effect) and H₁, set α (usually 0.05), pick a test matched to your data and design, compute the statistic and p-value, then decide — reject H₀ if p < α. The whole machinery exists because biology is noisy: you can never "prove" an effect, only bound how surprising your data would be if there were none. The general logic is in our hypothesis testing guide; here it’s applied to wet-lab and field data.
The five steps on a real bio example
Question: does a new fertilizer change plant height after 8 weeks? 20 plants per treatment.
- H₀: mean height fertilized = mean height control (the fertilizer does nothing). H₁: the means differ.
- α = 0.05 — the false-positive rate you accept up front, never after seeing the data.
- Test: two independent groups, one continuous outcome → independent-samples t-test (check roughly normal residuals; with n=20 per group it’s forgiving).
- Compute: t = difference in means ÷ standard error of that difference; software gives p directly.
- Decide: p = 0.003 < 0.05 → reject H₀; the data are unlikely under "no effect." Report the means, not just "significant."
Matching the test to the biological design
| Design | Outcome | Test |
|---|---|---|
| Two independent groups (treated vs control) | Continuous | Independent t-test (Mann-Whitney if very non-normal) |
| Same subjects before/after, paired organs | Continuous | Paired t-test |
| Three+ groups (doses, species, sites) | Continuous | One-way ANOVA + post-hoc (Tukey) |
| Counts in categories (infected/not by sex) | Counts | Chi-square / Fisher’s exact (small n) |
| Association between two measurements | Continuous | Correlation / regression |
The three mistakes biology reviews catch
- Pseudoreplication: treating 50 cells from one culture flask as n=50. The replicate is the flask, not the cell — inflate n and every p-value becomes meaningless.
- p-hacking by iteration: adding animals or repeating the assay until p dips under 0.05. Decide n and stopping rules before starting; a p-value near 0.06 is “not demonstrated,” not “almost there.”
- Confusing significance with size: a huge n makes a 0.4% difference “significant.” Always report effect size with a confidence interval next to the p — the meaning of p is covered step by step in our p-value explainer.
Context for the wider toolkit — data types, distributions, sampling — is in our introduction to biostatistics, and why biology leans on statistics at all in this overview. To run these tests yourself, point-and-click, our jamovi guide walks t-tests and ANOVA with zero code.
For one-to-one coaching on designing your thesis experiments and analyses, Ampersand Academy teaches biostatistics and R one-to-one.
Frequently asked questions
What does p < 0.05 actually mean in a biology experiment?
If there were truly no effect, results at least this extreme would occur less than five percent of the time. It is a statement about the data under the null hypothesis – not the probability the hypothesis is true.
Why can a study never prove the alternative hypothesis?
Hypothesis tests only measure how incompatible the data are with the null. Evidence accumulates across replication and design quality; a single significant result supports, but cannot prove, a biological claim.
What is pseudoreplication and why is it fatal?
It is counting subsamples from one biological unit as independent replicates – 50 cells from one flask as n=50. The true sample size is the flask, so the test dramatically overstates certainty.
When should I use ANOVA instead of multiple t-tests?
Whenever you compare three or more groups. Running all pairwise t-tests inflates the false-positive rate; ANOVA gives one overall test, then post-hoc comparisons like Tukey’s control for multiplicity.
Is p = 0.06 a failed experiment?
No. It means the effect was not demonstrated at your chosen threshold with this sample size. Report the effect size and confidence interval – the result may be biologically meaningful and detectable with more power.

