Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Principal components and variance retention

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Principal component analysis rotates centered numeric features toward directions of greatest variance, then projects records onto selected directions without using outcome labels.

Center and scale before rotating

A warehouse dataset has backlog and handoff duration. PCA on raw columns can point almost entirely toward whichever column has larger units. Define the feature set and a training-only scaling policy first. Fit PCA on the same permitted training cohort, then apply its fixed center, scale and axes to later records. The first axis maximizes variance under that transformed geometry, not business value. Distance scaling explains the unit choice.

Read a component as a weighted direction

The code computes the leading direction of a two-feature covariance matrix through repeated multiplication and normalization. Its projection is a score along that direction; reversing the sign of the direction reverses all scores without changing the represented subspace. Do not label a component "operational quality" from the sign alone. Inspect feature loadings and the source records before assigning an interpretation.

Choose dimensions by downstream purpose

Retaining 92% of training variance means the selected components reconstruct much of the scaled input variation. It does not guarantee preservation of a rare predictive signal or small-group structure. A low-variance feature can be essential to detect costly incidents. Compare downstream holdout performance or reconstruction error under the intended use, with PCA fitted inside each development fold. Model selection handles the final test boundary.

Watch missing values and drift

PCA requires a declared missing-value policy. Filling future records from statistics fitted on future batches changes the projection over time. If warehouse processes shift, old components may no longer summarize the same pattern; monitor transformed feature distributions and reconstruction error. A visually pleasing two-dimensional scatter is not proof of natural clusters. K-means asks a separate partition question.

Use a numerically sound solver in larger work

The short power-iteration example makes the leading-axis idea inspectable for two features. It does not implement a full PCA decomposition or handle nearly equal eigenvalues with careful numerical diagnostics. Use a tested linear algebra implementation for high-dimensional or sparse data, and store feature order and fitted transform with any downstream model. The project audits what the projection is used for.

Implementation

python
from math import sqrt

def leading_axis_two_features(training_profiles, iterations=40):
    if len(training_profiles) < 3 or any(len(profile) != 2 for profile in training_profiles):
        raise ValueError("at least three two-feature profiles required")
    means = tuple(sum(profile[column] for profile in training_profiles) /
                  len(training_profiles) for column in range(2))
    centered = [(left - means[0], right - means[1])
                for left, right in training_profiles]
    covariance = [
        [sum(left * left for left, _ in centered),
         sum(left * right for left, right in centered)],
        [sum(left * right for left, right in centered),
         sum(right * right for _, right in centered)],
    ]
    axis = (1.0, 1.0)
    for _ in range(iterations):
        rotated = (covariance[0][0] * axis[0] + covariance[0][1] * axis[1],
                   covariance[1][0] * axis[0] + covariance[1][1] * axis[1])
        length = sqrt(rotated[0] ** 2 + rotated[1] ** 2)
        if length == 0:
            raise ValueError("no variance in the training profiles")
        axis = (rotated[0] / length, rotated[1] / length)
    return means, axis

profiles = [(1.0, 2.0), (2.0, 3.0), (3.0, 5.0), (4.0, 6.0)]
means, axis = leading_axis_two_features(profiles)
scores = [sum((value - mean) * direction
              for value, mean, direction in zip(profile, means, axis))
          for profile in profiles]
assert abs(sum(scores)) < 1e-10
assert abs(sum(direction ** 2 for direction in axis) - 1) < 1e-10

Performance and operating cost

For N two-feature profiles and I power iterations, this teaching fit is O(N + I) time and O(N) space. General PCA cost depends on N, feature count and solver; projection of M new records onto D retained components costs O(M × P × D).

Common Mistakes

  • Do not infer predictive value from explained variance alone.
  • Do not refit PCA on later evaluation records.
  • Do not interpret the arbitrary sign of a component as a fixed business direction.

Read next

ai-data
machine-learning
Storage details