Quick answer: A correlation coefficient, written r, measures how strongly two numeric variables move together on a straight-line scale. It runs from −1 to +1: zero means no linear relationship, +1 means a perfect positive line, and −1 a perfect negative one. The coefficient is unitless and symmetric, so the correlation between height and weight equals that between weight and height. Crucially, r says nothing about causation, and it only captures straight-line relationships.
Part of our statistics series. If you want to see how a correlation is turned into a prediction, read effect size and correlation and regression in jamovi next.
Reading the size of r
The sign tells you the direction; the distance from zero tells you the strength. These bands are conventions, not laws, and they shift by discipline — what counts as “strong” in psychology differs from physics.
| Value of r | Strength | Everyday reading |
|---|---|---|
| 0.00 to ±0.19 | Very weak | Essentially no usable linear pattern |
| ±0.20 to ±0.39 | Weak | A hint of a trend, much scatter |
| ±0.40 to ±0.59 | Moderate | A real relationship with plenty of noise |
| ±0.60 to ±0.79 | Strong | Tight enough to predict roughly |
| ±0.80 to ±1.00 | Very strong | Points hug the line closely |
A common error is to read r as a percentage. It is not: r = 0.5 is not “50% related”. Square it to get r², the coefficient of determination, which is a proportion — an r of 0.5 means the line explains about 25% of the variation in the data. That single squaring step is what turns a flattering number into an honest one.
Pearson, Spearman and when to switch
The default coefficient is Pearson’s r, which measures a straight-line relationship. When the data curve, or when the values are ranks rather than measurements, Spearman’s ρ (rho) or Kendall’s τ is the right tool.
| Coefficient | Measures | Use when |
|---|---|---|
| Pearson r | Linear association between two numeric variables | Both variables are continuous and roughly linear |
| Spearman ρ | Monotonic association using ranks | Data are ordinal, skewed or have outliers |
| Kendall τ | Concordance of pairs | Small samples, many tied ranks |
| Point-biserial | One numeric, one binary variable | Comparing a score against a yes/no group |
# R: Pearson and Spearman on the same pair
cor.test(df$height, df$weight, method = "pearson")
cor.test(df$height, df$weight, method = "spearman")
# R-squared from a fitted line
fit <- lm(weight ~ height, data = df)
summary(fit)$r.squared
Always plot before you trust a number. The classic warnings are Anscombe’s quartet and the Datasaurus: four datasets with identical r can look completely different, and a single outlier can drag a coefficient from 0.2 to 0.9. The visual inspection steps mirror those in the R Commander graphs guide.
Correlation is not causation — and not even linearity
Two variables can correlate strongly because a third drives both. Shoe size and reading ability correlate in children only because age drives both. Two more traps: a strong correlation may be curved rather than straight (Pearson r understates it), and restricting the range of your data — say, sampling only tall people — collapses a relationship that exists in the full population. Before reporting an r, ask what else could move both numbers and whether your sample covers enough of the range.
Reporting r properly
A complete report gives the direction, the size, the sample and a confidence interval: “height and weight were positively correlated, r(48) = 0.71, 95% CI [0.54, 0.82], p < .001″. The degrees of freedom are n − 2 for a Pearson correlation. Pair the coefficient with a scatter plot and, when the relationship is used for prediction, the explained variance. For how intervals are read, see confidence intervals explained, and for the test behind the p-value, hypothesis testing in statistics.
Common mistakes
- Calling r a percentage. Square it for the proportion of variance explained.
- Ignoring curvature. A strong curved relationship can give a low Pearson r; switch to Spearman or fit a curve.
- Letting one outlier decide. Always plot, and consider a rank coefficient.
- Restricting the range. Sampling a narrow slice of the data shrinks r artificially.
- Jumping to cause. Correlation alone never establishes that one variable makes the other change.
Learning to read your own results critically? Ampersand Academy teaches statistics one-to-one, using your own dataset as the curriculum.
Frequently asked questions
What does a correlation coefficient of 0.7 mean?
It means a strong positive linear relationship. Squared, it explains about 49 percent of the variation in the outcome. It does not mean the two variables are causally related.
Is a correlation coefficient of 0.5 good?
It is moderate. In many fields 0.5 is a useful relationship, but it explains only about 25 percent of the variance, so plenty of scatter remains. Judge it against typical values in your field.
What is the difference between Pearson and Spearman correlation?
Pearson measures a straight-line relationship between raw values. Spearman measures a monotonic relationship using ranks, so it copes better with skew and outliers.
Does correlation imply causation?
No. A correlation can arise from a shared third cause, from coincidence, or from reversed direction. Only a controlled design or an experiment can support a causal claim.
How do I calculate r in Excel or R?
In Excel use the CORREL function on the two ranges. In R use cor for the coefficient and cor.test for the coefficient with a p-value and confidence interval.

