An ordinary least-squares slope summarizes a conditional linear association, with an interpretation tied to its units and design.
Linear regression slopes: define the comparison a coefficient makes
Start with the target
A courier team records route distance and delivery minutes. Before fitting a line, define whether the target is average minutes for the sampled routes, a prediction for a future route, or the change that would result from altering distance. The first two can be studied from observed pairs; the last needs a causal design. State the route population, measurement window and units. The estimand lesson prevents a tidy coefficient from answering the wrong operational question.
Read the fitted slope narrowly
In a one-predictor line, the slope is the change in fitted mean minutes per additional distance unit within the observed range. The intercept is the fitted value at distance zero, which can be outside the data and physically meaningless. With multiple predictors, a coefficient compares observations at fixed values of the included predictors; that condition may be implausible if the fields move together. Do not convert association into an intervention claim. A confounder graph gives a better starting point for that question.
Inspect support before inference
Plot or tabulate distances, route types and observed times. A single distant route can set the apparent slope. A narrow distance range can produce unstable extrapolation. Fit on a declared training period; assess residuals and out-of-period error before trusting a release. Residual checks separate a good-looking line from a usable error model. A holdout baseline addresses prediction, while a coefficient interval addresses sampling uncertainty under model assumptions.
Keep the arithmetic reproducible
For a single predictor with nonzero variation, divide the centered cross-product by the centered distance sum of squares; subtract slope times mean distance from mean time to get the intercept. Save the exact rows and filtering rule. The tiny calculation below is for inspection, not a substitute for a design-aware standard error. If several routes belong to the same depot, route-level independence is suspect; clustered uncertainty uses the depot as the sampling unit.
Implementation
def fit_route_line(distances, minutes):
if len(distances) != len(minutes) or len(distances) < 2:
raise ValueError("paired route observations required")
distance_mean = sum(distances) / len(distances)
minute_mean = sum(minutes) / len(minutes)
denominator = sum((distance - distance_mean) ** 2 for distance in distances)
if denominator == 0:
raise ValueError("distance has no variation")
slope = sum((distance - distance_mean) * (minute - minute_mean)
for distance, minute in zip(distances, minutes)) / denominator
return minute_mean - slope * distance_mean, slope
intercept, slope = fit_route_line([2, 4, 6, 8, 10], [5, 9, 10, 15, 17])
assert round(intercept, 6) == 2.2
assert round(slope, 6) == 1.5
Performance and operating cost
The one-predictor calculation is O(n) time and O(1) extra space for n routes. General regression solvers use matrix methods whose cost depends on row count and feature count. Computing a coefficient is cheap; obtaining credible uncertainty requires a sampling design, dependence check and enough support in the distances of interest.
Common Mistakes
- Calling the observational slope the effect of sending a driver farther.
- Interpreting an intercept far outside observed route distances.
- Reporting a precise slope after filtering influential routes without a declared rule.
- Treating repeated routes from one depot as independent samples.
Read next
- Regression diagnostics: residual pattern, scale and influential routes
- Mean-response and prediction intervals: identify whose uncertainty is covered
- Population, estimand and sampling frame: name the quantity before calculating
- Standard error and cluster bootstrap: resample the independent unit
- Regression baselines and honest holdout metrics
