
Quick answer: A p-value is the probability of seeing data at least this extreme if the null hypothesis were true. Small p → your data would be very surprising under “no effect” → reject the null. A p-value is not the probability the hypothesis is true, not the size of an effect, and 0.05 is a convention, not a law. This guide makes it concrete with worked biology examples; the general logic is in our hypothesis testing guide.
Worked example 1: the coin-flip logic
Claim: a mutant strain has biased sex ratios. You count 1000 offspring: 530 male. H₀: ratio is 50/50. A fair coin gives 530+ heads in 1000 flips about 0.7% of the time (binomial test) — so p = 0.007. At α = 0.05 you reject H₀: the strain genuinely skews male. Notice what p is not saying: not that there’s a 99.3% chance the strain is biased — it’s the probability of the data under no bias.
Worked example 2: two treatment groups
Drug A lowers blood pressure by a mean of 12 mmHg, placebo by 4 mmHg; independent t-test gives p = 0.02. Read it as: “if the drug truly did nothing, a difference this large (or larger) would appear in only 2% of repeated experiments.” That’s the entire meaning — it says nothing about how well the drug works; the mean difference of 8 mmHg says that. Report both, always: the effect size with confidence interval is the science, the p-value is the noise bar.
What p-values are NOT — the five misreadings
| Misreading | Reality |
|---|---|
| “p = 0.03 → 97% chance H₁ is true” | p is P(data | H₀), not P(H₁ | data) — different directions |
| “p = 0.04 is meaningful, p = 0.06 is not” | They are nearly identical evidence; the 0.05 line is a publishing convention, not nature |
| “Small p = big effect” | With huge n, tiny effects get tiny p — check effect size |
| “Non-significant = no effect” | It means “not demonstrated at this sample size,” not proven absent |
| “p = 0.001 means the result will replicate” | Replication also depends on design, power and rigor — not just p |
Reading a results line like a reviewer
A good results sentence names test, statistic, p, and effect: “Treatment increased yield (t(38) = 2.71, p = 0.010; mean difference 1.8 g, 95% CI 0.5–3.1).” If a paper reports only “significant (p<0.05)” with no direction or size, be suspicious. And when the same hypothesis is tested repeatedly — ten assays, ten comparisons — raw p-values overstate evidence; corrections like Bonferroni or FDR exist for exactly that, and the distributions behind all of it are in our distributions guide.
To run these tests without touching code — jamovi computes p-values, CIs and effect sizes in three clicks — see our jamovi first-analysis guide. For one-to-one instruction on choosing and interpreting tests for your own thesis data, Ampersand Academy teaches biostatistics one-to-one.
Frequently asked questions
What is a p-value in one sentence?
The probability of observing data at least as extreme as yours if the null hypothesis were true. Small p means your data would be rare under no-effect, not that the null has a small chance of being true.
Is 0.05 a magic threshold?
No. It became a publishing convention through historical use. Evidence is continuous – p = 0.049 and p = 0.051 are practically the same result, and many fields now report exact p-values with effect sizes.
Can a p-value tell me the size of an effect?
No. Effect size comes from the estimate itself – a mean difference, odds ratio or correlation. A very large sample makes a trivial effect produce a very small p-value.
What does p = 0.5 mean?
Your data are entirely ordinary under the null hypothesis. It does not prove the null true – the study may simply lack the power to detect a real effect.
Why do multiple comparisons need corrections like Bonferroni?
Testing many hypotheses at once guarantees some small p-values by chance. Corrections raise the bar per test so the overall false-positive rate stays at your chosen alpha.
