Distance-based clustering depends on feature units; fit a scaling rule on the analysis population before interpreting which records are near each other.
Distance scaling before clustering
Expose the unit problem
A warehouse profile includes average backlog of 1,200 tickets and a handoff delay of 8 minutes. In raw Euclidean distance, a backlog change of 100 overwhelms a delay change of 2 even if operations consider both important. Scaling changes the geometry and therefore the clusters. Choose features for the question first, then choose weights or a scale that reflects meaningful variation. K-means optimizes distance under that geometry.
Fit the transform on the intended cohort
The teaching code uses training means and population standard deviations to standardize each numeric feature. A zero-variance feature carries no distance information and is rejected. If the clusters will score later warehouses, save the fitted means, scales and feature order; do not refit per incoming batch. If the goal is a one-time descriptive map of a fixed cohort, that cohort can define the transform, but future claims need a frozen state.
Treat outliers and categories deliberately
One extreme warehouse can inflate standard deviation and compress ordinary sites. A median-based scale or capped feature may be better, but caps must be fitted without looking at future evaluation rows. Encoding a category as arbitrary integers creates fake distances; use a suitable representation or a distance designed for mixed data. The code accepts only numeric feature columns and makes no claim that standard scaling is always right.
Test sensitivity to design choices
Re-run clustering with a defensible alternative feature set or scaling method. If membership changes sharply, the segmentation is design-dependent. A lower within-cluster distance alone cannot tell whether segments help staffing or routing decisions. Compare cluster profiles with operational outcomes that were not used to construct the distance, while avoiding post-hoc claims of causal effect. The project captures that audit.
Keep prediction and interpretation distinct
Clusters describe similarity under a chosen feature space; they do not discover natural types guaranteed to exist outside it. Scaling fitted on combined training and future data can also contaminate a predictive downstream evaluation. The pipeline lesson applies when clusters become model features.
Implementation
from math import sqrt
def fit_numeric_scale(warehouse_profiles, features):
if not warehouse_profiles or not features:
raise ValueError("profiles and features required")
state = {}
for feature in features:
values = [profile[feature] for profile in warehouse_profiles]
mean = sum(values) / len(values)
scale = sqrt(sum((value - mean) ** 2 for value in values) / len(values))
if scale == 0:
raise ValueError("constant feature: " + feature)
state[feature] = (mean, scale)
return state
def scaled_profile(profile, state):
return {feature: (profile[feature] - mean) / scale
for feature, (mean, scale) in state.items()}
warehouses = [{"backlog": 800, "handoff_minutes": 4},
{"backlog": 1200, "handoff_minutes": 8},
{"backlog": 1600, "handoff_minutes": 12}]
scale = fit_numeric_scale(warehouses, ("backlog", "handoff_minutes"))
middle = scaled_profile(warehouses[1], scale)
assert abs(middle["backlog"]) < 1e-12
assert abs(middle["handoff_minutes"]) < 1e-12Performance and operating cost
For N profiles and P features, fitting means and scales costs O(N × P) time and O(P) fitted-state space; the displayed value lists use O(N) temporary space per feature. Distance calculations add their own cluster-specific cost.
Common Mistakes
- Do not cluster mixed-unit columns without a justified distance rule.
- Do not encode categories as arbitrary numeric distances.
- Do not refit the scaler independently for each production batch.
Read next
- K-means objective and restart stability
- Principal components and variance retention
- Leakage-safe preprocessing: fit every learned transform inside the training fold
- Project: select a model and audit warehouse segments
Continue the workflow: DBSCAN core, border and noise points.
Continue the workflow: Nearest-neighbor distance and local support.
