
Quick answer: Sample size is decided by a power analysis, not by lab convention: with the expected effect size, α (usually 0.05) and desired power (usually 0.8), the analysis returns the minimum n — the number of biological units, not subsamples. Underpowered studies chase noise; overpowered ones spend animals and budget to confirm trivia. This guide shows the calculation, the inputs, and the honest workarounds when n can’t be reached. Replication rules first: study design and sampling.
The four inputs (and the one people guess wrong)
| Input | Where it comes from | Typical value |
|---|---|---|
| Effect size | Pilot data or the smallest effect that would matter scientifically | d = 0.5–0.8, or from prior literature |
| Significance level α | Field convention | 0.05 |
| Power (1 − β) | Field convention | 0.80–0.90 |
| Test type | Design: t-test, ANOVA, proportions | Two-sided t-test for two groups |
Worked example: expecting a medium effect (d = 0.5) between two treatment groups at α = 0.05 and power = 0.80, a two-sided t-test needs about n = 64 per group. Halve the expected effect to d = 0.25 and the requirement jumps to ~250 per group — effect size drives everything, which is why pilots are worth their cost. Free tools: G*Power, the pwr package in R, or jamovi’s power module — menu-driven, no code (see the jamovi first-analysis guide).
When n can’t be reached: honest options
- Sharpen the measurement: paired designs and repeated measures slash variance — a paired t-test needs roughly half the n of a two-group design for the same power.
- Reduce noise, not ambition: randomize blocks (batches, cages, days) so known variance sources stop inflating the error term.
- Reframe the claim: if only n = 15 per group is feasible, the study demonstrates the large effects only — say so in the limitations, with the detectable effect size stated.
- Never iterative stopping: adding units until p < .05 inflates false positives; the testing logic that punishes this is in the hypothesis testing guide.
Report power thinking in the methods: expected effect, α, power, achieved n — the transparency reviewers increasingly demand. The measurement machinery behind effect sizes is in the effect size guide, and the p-value/power partnership in the p-value explainer. For one-to-one help planning a thesis study before the first sample, Ampersand Academy teaches biostatistics one-to-one.
Frequently asked questions
How do I calculate the sample size for my experiment?
Run a power analysis: combine the expected effect size, your alpha (usually 0.05), target power (usually 0.8) and the planned test. Tools like G*Power, R’s pwr package or jamovi modules return the minimum n per group.
Why does a smaller expected effect need a much larger sample?
Small effects are harder to see through random noise. Halving the effect size roughly quadruples the required n, which is why pilots and literature-based estimates matter before data collection.
Is 0.8 power a strict rule?
It is convention: an 80 percent chance of detecting a real effect of the assumed size. Studies that can afford 0.9 should use it; the honest failure mode is stating the achieved power in limitations.
Does a paired design really need fewer subjects?
Usually yes. Pairing removes between-subject variance from the comparison, so the same power is reached with roughly half the participants or animals of an independent-groups design.
Can I check power after the study instead of before?
Post hoc power calculated from your own p-value is circular and uninformative. What reviewers accept is the detectable effect size at your achieved n, stated in the limitations.
