The best line, and how uncertain it is

The method of least squares picks the line that minimises the sum of the squared vertical distances between the points and the line. The result has a closed form: the slope is m = (n·Σxy − Σx·Σy)/(n·Σx² − (Σx)²) and the intercept follows. In the lab, though, the number alone is not enough, because the slope is nearly always the physical quantity being sought: a spring constant from force and extension, a resistance from voltage and current, g from a plot of 2h against t².

That is why the uncertainties matter. If each measurement of y has its own σ, the fit is weighted by 1/σ² and the uncertainties of slope and intercept come straight from the error bars. χ² then measures how far the points stray from the line in units of their bars: divided by the degrees of freedom it should be around 1 if line and uncertainties agree; much larger means the bars are underestimated or the relation is not linear; much smaller, that they are overestimated.

With no σ on the data, the only estimate there is comes from the scatter of the points about the line, with n − 2 degrees of freedom because two parameters were taken from the same data. The correlation coefficient r says how well the points line up, but it is not an uncertainty: an r of 0.999 can sit alongside a slope uncertain by 5 % if there are few points.

Common mistakes

  • Using r or R² as the measure of how good the result is: they say how aligned the data are, not how precise the slope is. That is what σ_m and σ_q are for.
  • Forcing the line through the origin when the law does not call for it, or not forcing it when it does: an intercept consistent with zero within its uncertainty is the check, not a detail to remove.
  • Swapping x and y: regression minimises vertical distances, so the quantity measured with less uncertainty goes on the x axis.
  • Linearising and forgetting to transform the uncertainties: if t² goes on the x axis, its uncertainty changes too, to 2t·σ_t.

Frequently asked questions

What is the difference between a weighted and an unweighted fit?

In a weighted fit each point counts in proportion to 1/σ²: a precise measurement pulls the line harder than an imprecise one, and the parameter uncertainties come from the stated σ. In an unweighted fit every point counts the same and the uncertainties are estimated from the scatter. If all the σ are equal, the two lines coincide.

How do I read χ²?

Look at χ² divided by the degrees of freedom, that is the number of points less the parameters fitted. A value near 1 says the points stray from the line as much as their uncertainties lead you to expect. Well above 1: uncertainties underestimated or a non-linear law. Well below 1: uncertainties overestimated.

When should I use the line through the origin?

When theory says y is proportional to x, like voltage and current in a resistor or force and extension of a spring. Only one parameter is estimated and the slope comes out more precise. If in doubt, fit the full line and check that the intercept is consistent with zero.

How do I get a physical quantity from the slope?

Write the law in linear form and identify the slope. For a spring, extension = F/k: with force on the x axis the slope is 1/k, so k = 1/m and its uncertainty is σ_m/m². For more involved transformations, feed the slope and its σ into the error propagation calculator.

Can I paste data from a spreadsheet?

Yes. Copying two or three columns from Excel or LibreOffice brings them in separated by tabs, which are read directly, decimal commas included. A CSV file with comma-separated values and decimal points works as well.

How this calculation works

With weights wᵢ = 1/σᵢ² (all equal to 1 when there are no σ) and the sums S = Σw, Sx = Σwx, Sy = Σwy, Sxx = Σwx², Sxy = Σwxy, Δ = S·Sxx − Sx²: slope m = (S·Sxy − Sx·Sy)/Δ, intercept q = (Sxx·Sy − Sx·Sxy)/Δ, σ_m² = S/Δ, σ_q² = Sxx/Δ. In an unweighted fit the variances are multiplied by s² = Σ(yᵢ − m·xᵢ − q)²/(n − 2). Line through the origin: m = Sxy/Sxx and σ_m² = 1/Sxx, times s² with n − 1 degrees of freedom when unweighted. χ² = Σwᵢ(yᵢ − m·xᵢ − q)². r is Pearson's correlation coefficient of the data and R² = r².