Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Distance scaling before clustering

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Distance-based clustering depends on feature units; fit a scaling rule on the analysis population before interpreting which records are near each other.

Expose the unit problem

A warehouse profile includes average backlog of 1,200 tickets and a handoff delay of 8 minutes. In raw Euclidean distance, a backlog change of 100 overwhelms a delay change of 2 even if operations consider both important. Scaling changes the geometry and therefore the clusters. Choose features for the question first, then choose weights or a scale that reflects meaningful variation. K-means optimizes distance under that geometry.

Fit the transform on the intended cohort

The teaching code uses training means and population standard deviations to standardize each numeric feature. A zero-variance feature carries no distance information and is rejected. If the clusters will score later warehouses, save the fitted means, scales and feature order; do not refit per incoming batch. If the goal is a one-time descriptive map of a fixed cohort, that cohort can define the transform, but future claims need a frozen state.

Treat outliers and categories deliberately

One extreme warehouse can inflate standard deviation and compress ordinary sites. A median-based scale or capped feature may be better, but caps must be fitted without looking at future evaluation rows. Encoding a category as arbitrary integers creates fake distances; use a suitable representation or a distance designed for mixed data. The code accepts only numeric feature columns and makes no claim that standard scaling is always right.

Test sensitivity to design choices

Re-run clustering with a defensible alternative feature set or scaling method. If membership changes sharply, the segmentation is design-dependent. A lower within-cluster distance alone cannot tell whether segments help staffing or routing decisions. Compare cluster profiles with operational outcomes that were not used to construct the distance, while avoiding post-hoc claims of causal effect. The project captures that audit.

Keep prediction and interpretation distinct

Clusters describe similarity under a chosen feature space; they do not discover natural types guaranteed to exist outside it. Scaling fitted on combined training and future data can also contaminate a predictive downstream evaluation. The pipeline lesson applies when clusters become model features.

Implementation

python
from math import sqrt

def fit_numeric_scale(warehouse_profiles, features):
    if not warehouse_profiles or not features:
        raise ValueError("profiles and features required")
    state = {}
    for feature in features:
        values = [profile[feature] for profile in warehouse_profiles]
        mean = sum(values) / len(values)
        scale = sqrt(sum((value - mean) ** 2 for value in values) / len(values))
        if scale == 0:
            raise ValueError("constant feature: " + feature)
        state[feature] = (mean, scale)
    return state

def scaled_profile(profile, state):
    return {feature: (profile[feature] - mean) / scale
            for feature, (mean, scale) in state.items()}

warehouses = [{"backlog": 800, "handoff_minutes": 4},
              {"backlog": 1200, "handoff_minutes": 8},
              {"backlog": 1600, "handoff_minutes": 12}]
scale = fit_numeric_scale(warehouses, ("backlog", "handoff_minutes"))
middle = scaled_profile(warehouses[1], scale)
assert abs(middle["backlog"]) < 1e-12
assert abs(middle["handoff_minutes"]) < 1e-12

Performance and operating cost

For N profiles and P features, fitting means and scales costs O(N × P) time and O(P) fitted-state space; the displayed value lists use O(N) temporary space per feature. Distance calculations add their own cluster-specific cost.

Common Mistakes

  • Do not cluster mixed-unit columns without a justified distance rule.
  • Do not encode categories as arbitrary numeric distances.
  • Do not refit the scaler independently for each production batch.

Read next

Continue the workflow: DBSCAN core, border and noise points.

Continue the workflow: Nearest-neighbor distance and local support.

ai-data
machine-learning
Storage details