Pearson Correlation Coefficient Calculator
Paste paired X and Y values and get the Pearson correlation r, with the strength of the relationship explained.
Enter at least 3 pairs as x, y with a comma between them.
Data points (n)
r squared (variance explained)
Strength
r runs from −1 to +1. Correlation measures linear association only — it never proves one variable causes the other.
Two variables can move together without one causing the other — ice cream sales and drowning deaths both rise in summer, yet nobody blames the ice cream. The Pearson correlation coefficient is the standard tool for measuring how strongly two variables move together in a straight-line relationship.
The Pearson Correlation Coefficient Calculator above takes your paired X and Y data and returns r, the coefficient between −1 and +1, along with r-squared and a plain-language interpretation of the strength.
This guide explains what r measures, how to read its value, what r-squared adds, and the traps that make correlation one of the most misused numbers in statistics.
What Does the Pearson Correlation Coefficient Calculator Do?
The calculator computes Pearson's r from paired observations you paste in — one x, y pair per line. From those pairs it derives the means, the covariance structure, and finally r itself.
The headline result is r to four decimal places. Below it, three rows show the number of data points, r-squared (the share of Y's variation associated with X), and a verbal strength rating from "very weak" to "very strong" with the direction included.
It needs at least three pairs, and it rejects degenerate inputs — identical X values or identical Y values — where no meaningful correlation can exist.
How to Use the Pearson Correlation Coefficient Calculator
Paste your data into the text box with one pair per line, x and y separated by a comma — for example, "1, 2.1" on the first line, "2, 4.0" on the second. Blank lines are ignored.
Press Calculate. The r value, r-squared, and strength interpretation appear in the result panel within a second. If a line is malformed or you have fewer than three pairs, an error message tells you exactly which line to fix.
Press Reset to clear the box and analyze a new dataset. Keep the original ordering of pairs intact — each x must stay matched with its own y, or the result is meaningless.
What r Actually Measures
Pearson's r measures the strength and direction of a linear relationship between two variables. A value near +1 means larger X reliably goes with larger Y along a straight line; near −1 means larger X goes with smaller Y; near 0 means no linear pattern at all.
The key word is linear. A perfect U-shaped relationship — Y high at both extremes of X and low in the middle — can have r near zero despite an obvious pattern. The coefficient is blind to curves.
It is also standardized: r has no units and always falls between −1 and +1, which is what makes correlations comparable across completely different domains, from stock prices to test scores.
The Formula Behind r
The coefficient is the covariance of X and Y divided by the product of their standard deviations — essentially, the co-movement normalized by each variable's own spread.
The formula is:
r = Σ((x − x̄)(y − ŷ)) / √(Σ(x − x̄)² · Σ(y − ŷ)²)
Each term (x − x̄)(y − ŷ) is positive when both variables are on the same side of their means and negative when they oppose. Summing these and normalizing yields a number whose sign captures direction and whose magnitude captures consistency.
Reading the Strength Scale
As a rough guide: below 0.2 is very weak, 0.2 to 0.4 is weak, 0.4 to 0.6 is moderate, 0.6 to 0.8 is strong, and above 0.8 is very strong. The sign indicates direction — positive means the variables rise together.
These cutoffs are conventions, not laws. In physics, 0.9 might be disappointingly low; in psychology, 0.4 can be a celebrated finding. Judge strength against what your field normally produces.
The calculator applies these bands automatically and reports both the band and the direction, so "strong negative linear relationship" tells you everything the number implies at a glance.
What r-Squared Adds
Squaring r gives r², the proportion of Y's variance statistically associated with X. An r of 0.8 becomes an r² of 0.64 — about 64 percent of the variation in Y moves with X.
r² is often more intuitive than r because it reads as a percentage. But that percentage describes association, not causation — 64 percent "explained" does not mean X causes 64 percent of Y.
The calculator reports r² with this plain-language gloss so the number is useful even if you have never studied variance decomposition.
Worked Example: Near-Perfect Positive Correlation
First: take five pairs — (1, 2.1), (2, 4.0), (3, 6.2), (4, 7.9), (5, 10.1). The means are x̄ = 3 and ŷ ≈ 6.06.
Then: the numerator sums (x − 3)(y − 6.06) across all pairs, and the denominator normalizes by the spreads. The calculator returns r ≈ 0.9994.
Then: r² ≈ 0.9987, so about 99.9 percent of Y's variation is associated with X — a very strong positive linear relationship, as the near-straight-line data suggests.
Worked Example: A Moderate Relationship
First: consider study hours and exam scores for six students: (2, 55), (3, 62), (4, 58), (5, 75), (6, 70), (7, 85).
Then: the general trend rises, but individual points scatter — the (4, 58) student underperforms the trend while (5, 75) overperforms it.
Then: the calculator gives r ≈ 0.90, a strong positive relationship with r² ≈ 0.81. Hours clearly relate to scores, but the scatter shows other factors matter too.
Worked Example: No Linear Relationship
First: take pairs (1, 5), (2, 3), (3, 7), (4, 2), (5, 6) — Y bouncing with no trend as X grows.
Then: positive and negative (x − x̄)(y − ŷ) terms roughly cancel in the numerator.
Then: r comes out near 0.10 — very weak. The calculator correctly reports essentially no linear relationship, even though Y is clearly varying.
Correlation Is Not Causation
The most abused sentence in statistics exists because r is so easy to compute and so easy to misread. A strong correlation between two variables never, by itself, shows that one causes the other.
Three explanations always compete: X causes Y, Y causes X, or a third variable Z causes both — the ice cream and drowning case, where summer heat is Z. There is also pure coincidence, which thrives in small samples.
Establishing causation needs controlled experiments, time ordering, and a plausible mechanism — none of which a correlation coefficient provides. Treat r as a clue that directs investigation, never as a verdict.
Outliers and Why They Distort r
A single extreme point can drag r dramatically. One far-flung pair aligned with the trend can inflate a weak correlation into a strong one; one point against the trend can collapse a strong correlation.
Always plot your data before trusting r. A scatterplot reveals outliers, curves, and clusters that the single number hides — the number summarizes the plot, it does not replace it.
If an outlier reflects a data-entry error, fix it. If it is genuine, report r with and without it so readers see how much rides on one observation.
Common Correlation Mistakes
Applying r to curved relationships is the classic error — a perfect parabola scores near zero and the analyst concludes "no relationship." Check linearity visually first.
Restricted range is subtler: correlating test scores with performance only among hired employees (who all scored highly) understates the true relationship, because the X range was truncated by selection.
Finally, small samples produce jumpy r values. With only a handful of pairs, r can swing wildly — treat correlations from tiny datasets as suggestive, not solid.
How to Interpret Your Result Correctly
Read r in three parts: sign (direction), magnitude (strength band), and r² (share of associated variation). "r = −0.72" means a strong negative linear relationship with about 52 percent of Y's variation tied to X.
Then ask the causation question explicitly and answer it honestly: is there a plausible mechanism, or could a third variable explain this? The calculator cannot answer that — only domain knowledge and study design can.
Check the sample size before getting excited. Strong r from three points is barely evidence; the same r from three hundred points is a finding.
Where Correlation Calculations Are Useful
Science uses r for exploratory analysis — screening which variables move together before designing experiments. Finance uses it for diversification: assets with low or negative correlation smooth a portfolio's ride.
Medicine correlates biomarkers with outcomes to find diagnostic candidates, and marketing correlates ad spend with sales to allocate budgets. In every case, r is the starting filter, not the final proof.
Quality control correlates process settings with defect rates to find which knobs actually matter. A strong r between oven temperature and defect rate is worth investigating — with an experiment, not a press release.
Frequently Asked Questions
1. What is the Pearson Correlation Coefficient Calculator?
It computes Pearson's r from paired X and Y data you paste in. It returns r, r-squared, and a plain-language strength rating like "strong positive linear relationship."
2. What does the r value mean?
r ranges from −1 to +1. Values near +1 indicate a strong positive linear relationship, near −1 a strong negative one, and near 0 no linear relationship. The sign gives direction; the magnitude gives strength.
3. What is the formula for Pearson's r?
r = Σ((x − x̄)(y − ŷ)) / √(Σ(x − x̄)² · Σ(y − ŷ)²). It is the covariance of X and Y divided by the product of their standard deviations.
4. What is r-squared?
r² is the proportion of Y's variance associated with X. An r of 0.8 gives r² = 0.64, meaning about 64 percent of Y's variation moves with X. It describes association, not causation.
5. How do I format my data?
One pair per line with a comma between x and y, like "1, 2.1". Blank lines are ignored, and you need at least three pairs.
6. Does correlation prove causation?
No. A strong r never shows that one variable causes the other. Third variables, reverse causation, and coincidence can all produce strong correlations. Causation requires experiments and mechanism, not just r.
7. What counts as a strong correlation?
By convention, 0.6 to 0.8 is strong and above 0.8 is very strong — but standards vary by field. In physics 0.9 may disappoint; in social science 0.4 can be notable.
8. Why is my r near zero when the plot shows a pattern?
Probably the pattern is curved. Pearson's r only detects straight-line relationships — a U-shape or parabola can score near zero despite an obvious pattern. Plot first, then compute.
9. How do outliers affect r?
Enormously. One extreme point can inflate a weak correlation into a strong one or collapse a strong one toward zero. Always inspect a scatterplot, and consider reporting r with and without suspicious points.
10. How many data points do I need?
At least three for the math to run, but many more for a trustworthy result. Small samples make r jumpy — treat tiny-sample correlations as suggestive only.
11. What if all my X values are identical?
The calculator rejects the input. With no variation in X, there is nothing for Y to co-vary with, so correlation is undefined.
12. Can r be exactly 1 or −1?
Yes, when the points fall exactly on a straight line (with a nonzero slope). Real data rarely does this — exact ±1 usually signals a constructed example or a duplicated variable.
13. What is the difference between Pearson and Spearman correlation?
Pearson measures linear relationships using actual values; Spearman measures monotonic relationships using ranks. Spearman handles curves that consistently rise or fall, while Pearson needs straight lines.
14. Should I standardize my variables first?
Not necessary — r is already standardized and unit-free. Standardizing X and Y beforehand gives the same r, which is part of what makes the coefficient so convenient.
15. Can I use this for categorical data?
No. Pearson's r needs numeric variables where means and differences are meaningful. For categories, use chi-square tests or appropriate association measures instead.