
Quick answer: A valid study needs the right design (experimental vs observational), the right sampling (random — anything else invites bias), and the right n. The four decisions that make or break biological research: replicate at the correct unit, randomize treatments, control for confounders, and pre-specify your analysis. The measurement foundations — data types and scales — are in our data types guide.
Design families: pick before you collect
| Design | What it looks like | Strengths / caution |
|---|---|---|
| Completely randomized experiment | Treatments assigned by chance to homogeneous units | Gold standard for causation; needs enough units |
| Randomized block | Blocks (batches, days, cages) each containing all treatments | Removes batch noise; analyze the blocking factor |
| Repeated measures | Same subjects measured across time/conditions | Powerful but needs order effects controlled |
| Cross-sectional observational | One time point, measure exposure + outcome | Fast; shows association, never causation |
| Cohort / case-control | Follow exposed vs unexposed; or sample by outcome | Ethical way to study harmful exposures; confounding risk |
Sampling: the part reviewers check first
- Simple random: every unit equally likely — the default when a population list exists.
- Systematic: every k-th unit along a transect; fine if the starting point is random and the landscape has no periodicity.
- Stratified: split the population (by habitat, sex, age) and sample randomly within each — guarantees small subgroups are represented.
- Cluster: sample groups (ponds, plots, hospitals) then measure within — cheaper at scale, analyzed with the cluster as the sampling unit.
Whatever the method, the failure modes are the same: convenience sampling (the easy-to-reach sites), survivorship bias (measuring only the plants that lived), and pseudoreplication (see below). If you can’t randomize, you can still randomize where you place quadrats or which flask gets which position in the incubator.
Replication, randomization, and the n question
Three rules with teeth: the biological unit is the replicate — one patient, one flask, one field plot; cells within a flask are subsamples and get averaged first. Randomize every assignable factor — shelf position, handling order, plate position — so they can’t masquerade as treatment effects. And plan n before starting: a power analysis (expected effect size, α, desired power 0.8) tells you the minimum n; collecting “until it looks significant” is the definition of p-hacking, and the testing machinery that consumes this data is covered in our hypothesis testing in biology guide.
Foundations: what biostatistics covers and why in the introduction to biostatistics; interpreting the p-values your analysis produces in the p-value explainer. For one-to-one help designing your thesis study — power analysis included — Ampersand Academy teaches biostatistics and research methods one-to-one.
Frequently asked questions
What is the difference between experimental and observational studies?
In experiments the researcher assigns treatments (usually randomized), which supports causal claims. In observational studies exposure happens naturally, so associations may be confounded by factors that drove both exposure and outcome.
What is stratified sampling and when should I use it?
Divide the population into meaningful strata – habitat type, sex, age class – and randomly sample within each. Use it when subgroups matter to your question and simple random sampling might miss the small ones.
How many replicates do I need?
Run a power analysis before collecting data: with your expected effect size, chosen alpha (usually 0.05) and power (usually 0.8), it returns the minimum n. The biological unit is the replicate – not the subsamples within it.
What is pseudoreplication again, in one line?
Treating repeated measurements of the same biological unit as if they were independent samples, which inflates your n and makes p-values meaninglessly small.
Why is randomization so important in experiments?
Randomization makes known and unknown confounders – shelf position, handling order, incubator gradients – statistically harmless, so differences between groups are attributable to the treatment rather than placement.
