Linear Regression Calculator
Find the best-fit line through your data. Paste your (x, y) points, and get the slope, intercept, correlation, and a prediction for any x value, with the sums and formulas laid out.
The least-squares line minimizes the total squared vertical distance to your points. A correlation near 1 or −1 means a tight straight-line pattern; near 0 means x tells you almost nothing about y.
Data rarely falls on a perfect line, but it often huddles near one. Whether you are tracking how study hours relate to exam scores, how advertising spend relates to sales, or how temperature relates to ice cream revenue, the question is the same: what straight line best describes this cloud of points, and how much should I trust it?
The Linear Regression Calculator answers with the least-squares best-fit line. Paste your (x, y) points, and it returns the slope, the intercept, the correlation coefficient, and a prediction for any x you name, with the underlying sums and formulas shown in full.
This guide explains what the calculator computes, how to format your data, what each statistic means, and walks through four worked examples from a textbook-perfect line to a cautionary tale about extrapolation. Fifteen FAQs close out the tour.
What Does the Linear Regression Calculator Do?
The Linear Regression Calculator fits the straight line y = mx + b that minimizes the sum of squared vertical distances to your data points. You enter the points as x,y pairs, one per line, in the text box, plus an optional x value for prediction.
It returns the slope m, the intercept b, the correlation r, r-squared, a strength label for the relationship, and the predicted y at your chosen x. The steps panel lists the four fundamental sums and shows exactly how they combine into each statistic.
How to Use the Linear Regression Calculator
Type or paste your points into the data box, one pair per line, with a comma between x and y. The default five points give you a working example to experiment with before entering your own data. Add your prediction x in the second panel if you want a forecast.
Press Calculate and read the headline equation first: that is your line. Then check r to see how tightly the points hug it, and read the prediction row if you asked for one. Reset restores the sample data.
Formatting Your Data Points
Each line needs exactly two numbers separated by a comma, like 3.5,12. Extra spaces are harmless, and blank lines are skipped. You need at least two points, and the x values cannot all be identical, since a vertical stack of points has no definable slope.
Keep your units consistent across every line. If x is advertising dollars in thousands on line one, it must be thousands on every line. Mixed units are the quietest way to produce a confident-looking wrong answer.
What the Slope Really Tells You
The slope m is the expected change in y for each one-unit increase in x. A slope of 0.6, as in the default data, means each extra unit of x adds about 0.6 units of y on average. It is a rate, and it inherits the units of y-per-x.
Sign matters more than size. A positive slope means the quantities rise together; a negative slope means one falls as the other rises; a slope near zero means x barely moves y at all. Never read causation into it: slope describes association, not mechanism.
What the Intercept Really Tells You
The intercept b is the line’s predicted y when x equals zero. Sometimes that is meaningful, such as a base price before any usage, and sometimes it is pure extrapolation, such as predicting weight at zero height.
Treat the intercept as trustworthy only when x = 0 sits inside or near your data’s range. Far outside it, the intercept is a mathematical artifact of extending the line, not a measurement of anything real.
Worked Example: The Default Five Points
Use the points (1,2), (2,4), (3,5), (4,4), (5,5) with prediction x = 6.
First: accumulate the sums. Σx = 15, Σy = 20, Σxy = 66, Σx² = 55, with n = 5.
Then: slope m = (5×66 − 15×20) / (5×55 − 225) = 30/50 = 0.6.
Next: intercept b = (20 − 0.6×15) / 5 = 11/5 = 2.2.
Then: prediction at x = 6 gives y = 0.6×6 + 2.2 = 5.8.
Answer: y = 0.6x + 2.2, with a predicted value of 5.8 at x = 6.
Worked Example: A Near-Perfect Line
Take the points (0,1), (1,3), (2,5), (3,7), which lie exactly on y = 2x + 1.
First: accumulate carefully. Σx = 6, Σy = 16, Σxy = 0 + 3 + 10 + 21 = 34, Σx² = 14, with n = 4.
Then: slope m = (4×34 − 6×16) / (4×14 − 36) = (136 − 96) / (56 − 36) = 40/20 = 2.
Next: intercept b = (16 − 2×6) / 4 = 4/4 = 1.
Answer: y = 2x + 1 with r = 1 exactly, which confirms the calculator recovers a perfect line flawlessly when the data truly lies on one.
Worked Example: A Negative Relationship
Take price-discount points (0,100), (10,88), (20,75), (30,63), where x is percent discount and y is units’ profit index.
First: the slope comes out negative, roughly −1.24, meaning each extra discount point costs about 1.24 index points of profit.
Then: the correlation is strongly negative, near −0.99, so the line describes the data well.
Answer: a downward-sloping line with strong negative correlation, the classic trade-off pattern.
Worked Example: When r Is Near Zero
Take the scattered points (1,5), (2,3), (3,6), (4,4), (5,5).
First: the calculator fits a line anyway, because least squares always returns some line.
Then: r comes out near 0.14, labeled weak, and r² near 0.02.
Answer: the line exists but explains almost nothing; x tells you barely anything about y here, and predictions from it would be little better than guessing the mean.
Reading the Correlation Coefficient
The correlation r runs from −1 to 1. Values near 1 or −1 mean the points hug a straight line; values near 0 mean no linear pattern. The calculator labels the strength for you: very strong, strong, moderate, or weak.
Square it to get r², the fraction of y’s variation explained by x. An r of 0.77 gives r² of about 0.60, meaning 60% of the ups and downs in y are accounted for by the line, with 40% left to other factors.
The Extrapolation Trap
The prediction box accepts any x, including values far beyond your data. The mathematics will happily extend the line, but reality often refuses to follow: trends saturate, break, or reverse outside the observed range.
Treat predictions inside your data’s x-range as interpolation, which is relatively safe, and anything outside as extrapolation, which deserves explicit skepticism. If a forecast matters, say how far outside the data it reaches.
A practical rule: predictions up to about ten percent beyond the data range are usually defensible, while anything doubling the range is speculation dressed as arithmetic. When in doubt, collect more data instead of stretching the line.
Common Regression Mistakes
The biggest mistake is reading causation into correlation: the line says x and y move together, not that x causes y. The second is extrapolating without warning. The third is ignoring outliers, since a single wild point can drag the whole line off course.
Also check that a straight line is the right shape at all. If the scatterplot curves, a line will systematically miss at the ends and the middle, and no r value rescues a wrong model choice.
Where Linear Regression Is Useful
Businesses forecast sales from ad spend, scientists calibrate instruments from known standards, and analysts estimate trends in everything from housing prices to crop yields. Anywhere two quantitative variables move together, regression quantifies the relationship.
It is also the gateway to fancier modeling. Multiple regression, logistic regression, and most machine learning start from the least-squares idea, so fluency here pays dividends far beyond straight lines.
Even outside formal analysis, the instinct is useful. Comparing two phone plans, estimating delivery times from distance, or judging whether practice is improving your scores all involve fitting a mental line through noisy points, which is regression done intuitively.
How to Interpret Your Result Correctly
Read the equation first, then r, then r². A strong r with a sensible slope is a green light; a weak r means the line is decorative. Check whether your prediction x sits inside the data range before quoting the forecast.
Finally, plot the points with the line if you can. The human eye catches curvature, outliers, and clusters that numbers alone can hide, and a thirty-second scatterplot is the best regression diagnostic ever invented.
Report all three numbers together when you share results: the equation, r, and the sample size. A slope without its correlation is a claim without evidence, and a correlation without n hides how much data backs it.
Frequently Asked Questions
1. What is linear regression?
It is the method of fitting the straight line that best describes the relationship between two quantitative variables, where “best” means minimizing the sum of squared vertical distances from the points to the line.
2. How do I enter my data?
One x,y pair per line in the text box, with a comma between the two numbers. Blank lines are ignored, and you need at least two points with different x values.
3. What does the slope mean?
The expected change in y for a one-unit increase in x. A slope of 2 means y rises about 2 units per unit of x; a negative slope means y falls as x rises.
4. What does the intercept mean?
The line’s predicted y at x = 0. It is meaningful only when zero sits near your data; otherwise it is just where the extended line happens to cross the axis.
5. What is the correlation coefficient r?
A number from −1 to 1 measuring how tightly the points follow a straight line. Near ±1 is tight, near 0 is no linear relationship, and the sign gives the direction.
6. What is r-squared?
The square of r, interpreted as the fraction of variation in y explained by x. An r² of 0.60 means the line accounts for 60% of y’s ups and downs.
7. Why is my r so low?
Either the true relationship is weak, the relationship is nonlinear, or outliers are distorting the fit. Plot the data to see which story fits before trusting the line.
8. Can I predict outside my data range?
The calculator will compute it, but treat it as extrapolation and say so. Trends often change beyond the observed range, so distant forecasts carry real risk.
9. Does regression prove causation?
No. It measures association only. Two variables can track each other because of a shared third cause, coincidence, or reverse causation; the line cannot tell these apart.
10. How do outliers affect the line?
Strongly, because squaring punishes large misses. One bad point can tilt the slope and inflate or deflate r. Inspect your data and investigate outliers before deleting any.
11. What if all my x values are identical?
The calculator refuses, correctly: with no horizontal spread there is no slope to estimate. You need variation in x to learn how y responds to it.
12. How many points do I need?
Two is the mathematical minimum, but two points always fit perfectly and teach nothing. Aim for at least ten to fifteen points before taking r seriously.
13. What are the sums in the steps panel?
Σx, Σy, Σxy, and Σx² over your n points. Every least-squares formula is built from these four totals, so showing them lets you verify each statistic by hand.
14. When is a straight line the wrong model?
When the scatterplot curves, fans out, or clusters. A low r on curved data does not mean no relationship; it means no straight-line relationship. Consider a curve or a transformation instead.
15. How do I check my regression by hand?
Recompute the four sums from your points, plug them into the slope and intercept formulas shown in the steps, and confirm the equation matches. Then verify r with the correlation formula as a final check.