Statistics

Linear Regression & the Least Squares Method

Learn how to fit a line of best fit to data using the least squares method. Understand correlation coefficients, residuals, and how to interpret regression output.

V
Vectora Team
STEM Education
11 min read
2026-04-10

Try the interactive simulation for this topic

Explore this concept with our interactive 3D simulation.
Launch Simulator

What Is Linear Regression?

Linear regression is a statistical technique that finds the straight line which best fits a set of data points. The "best" line minimises the total squared distance from each data point to the line — this is the least squares criterion.

It answers the question: Given a set of (x,y)(x, y) data, what linear relationship y=a+bxy = a + bx best describes the trend?

Regression Playground

Add, drag and remove data points to see the least squares line update in real time. Watch how outliers affect the slope, intercept and R² value.
Launch Playground

Learning Goals: By the end of this guide, you should be able to:

  1. Calculate the equation of the least squares regression line from data.
  2. Interpret the slope and intercept in context.
  3. Understand and calculate the correlation coefficient rr and R2R^2.
  4. Use residuals to assess how well the model fits.

The Least Squares Regression Line

For a dataset of nn points (x1,y1),(x2,y2),…,(xn,yn)(x_1, y_1), (x_2, y_2), \ldots, (x_n, y_n), the line of best fit is:

y^=a+bx\hat{y} = a + bx

where the slope bb and intercept aa are:

b=n∑xiyi−∑xi∑yin∑xi2−(∑xi)2b = \frac{n\sum x_i y_i - \sum x_i \sum y_i}{n\sum x_i^2 - \left(\sum x_i\right)^2} a=yˉ−bxˉa = \bar{y} - b\bar{x}

Here xˉ\bar{x} and yˉ\bar{y} are the means of xx and yy respectively.

Key property: The regression line always passes through the point (xˉ,yˉ)(\bar{x}, \bar{y}).


Correlation Coefficient

The Pearson correlation coefficient rr measures the strength and direction of the linear relationship:

r=n∑xiyi−∑xi∑yi(n∑xi2−(∑xi)2)(n∑yi2−(∑yi)2)r = \frac{n\sum x_i y_i - \sum x_i \sum y_i}{\sqrt{\left(n\sum x_i^2 - (\sum x_i)^2\right)\left(n\sum y_i^2 - (\sum y_i)^2\right)}}
Value of rrInterpretation
r=1r = 1Perfect positive linear relationship
0.7<r<10.7 < r < 1Strong positive correlation
0<r<0.70 < r < 0.7Weak to moderate positive correlation
r=0r = 0No linear correlation
r<0r < 0Negative correlation (analogous ranges)

The coefficient of determination R2=r2R^2 = r^2 tells you the proportion of variation in yy explained by the model.


Worked Example

Data: (1,2), (2,4), (3,5), (4,4), (5,5)(1, 2),\ (2, 4),\ (3, 5),\ (4, 4),\ (5, 5)

Step 1: Compute the sums: ∑x=15\sum x = 15, ∑y=20\sum y = 20, ∑xy=67\sum xy = 67, ∑x2=55\sum x^2 = 55, n=5n = 5.

Step 2: Slope: b=5(67)−15(20)5(55)−152=335−300275−225=3550=0.7b = \frac{5(67) - 15(20)}{5(55) - 15^2} = \frac{335 - 300}{275 - 225} = \frac{35}{50} = 0.7

Step 3: Intercept: xˉ=3\bar{x} = 3, yˉ=4\bar{y} = 4, so a=4−0.7×3=1.9a = 4 - 0.7 \times 3 = 1.9

Answer: y^=1.9+0.7x\hat{y} = 1.9 + 0.7x


Residuals

A residual is the difference between an observed value and the predicted value:

ei=yi−y^ie_i = y_i - \hat{y}_i
  • Positive residual: the point is above the line
  • Negative residual: the point is below the line
  • If the model fits well, residuals should be randomly scattered around zero with no visible pattern.

Common Mistakes

  1. Extrapolating beyond the data range — The regression line is only reliable within the range of your data. Predicting far outside this range is unreliable.
  2. Confusing correlation with causation — A strong rr value means the variables are associated linearly, NOT that one causes the other.
  3. Using a linear model for non-linear data — Always plot the data first. If the scatter plot shows curvature, a linear model is inappropriate.

Exam Tips (A-Level / AP / IB)

  • For a quick sanity check, the slope bb should have the same sign as rr.
  • In context questions, always interpret the slope: "For each additional unit of xx, yy increases/decreases by bb units on average."
  • When asked about the reliability of a prediction, mention whether the predicted xx value is within the data range (interpolation) or outside it (extrapolation).

Frequently Asked Questions

When should I use yy on xx vs xx on yy regression?

Use yy on xx (y^=a+bx\hat{y} = a + bx) when you want to predict yy from a known xx. Use xx on yy when predicting xx from known yy. The two lines are generally different unless r=±1r = \pm 1.

What if my data has outliers?

Outliers can dramatically affect the regression line. Consider whether the outlier is a genuine data point or an error. Report results both with and without the outlier for transparency.


References & Further Reading

This article was created by the Vectora Editorial Team and is reviewed for alignment with AP, IB, and A-Level curricula. Content is based on standard academic sources in chemistry, physics, biology, and mathematics.

Published: 2026-04-10

For corrections or suggestions, contact support@vectora.one.