Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Compression Pareto review and shadow check

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A compression candidate should be compared on target error and measured serving resources, with a shadow period before it changes live decisions.

Keep all dimensions visible

A student, pruned head and quantized artifact may trade target error, p95 latency and peak memory differently. A candidate dominates another only if it is no worse on every declared measure and better on at least one. The code finds nondominated candidates from an invented device profile; it does not decide which tradeoff the business can accept. The baseline profile fixes hardware and measurement rules.

Apply hard acceptance bounds first

A model that exceeds the memory budget or breaches a safety-critical miss limit is not rescued by a low average latency. Keep error by site and label, artifact compatibility and security checks as gates. Then compare the surviving candidates’ operating costs. Per-label denominators prevent a small rare class from disappearing in a pooled score.

Shadow the selected artifact

Run the candidate beside the live model on the same incoming cases without changing actions. Log both outputs, missing scores, latency and later outcomes. A shadow comparison tests serving and predictive behavior under the current policy; if the candidate would alter review decisions, observed outcomes may reflect current interventions. The shadow guide states that boundary.

Measure cold and loaded conditions

A smaller file may cold-start faster yet run no faster after warm-up. Benchmark under expected concurrency and image sizes, with p95 and deadline misses. Watch peak resident memory and crash rate during the shadow period. Serving contracts include these failures.

Name promotion and rollback

Record chosen candidate, device profile, model version, target cohort, acceptance limits, pilot owner and old artifact retained for rollback. Do not choose a candidate on the final test and then call that same test independent evidence. The release project assembles the decision packet.

Implementation

python
# Same target cohort and serving device for all candidates.
candidates = [
    {"name": "baseline", "error": 0.07, "p95_ms": 84, "peak_mb": 620},
    {"name": "student", "error": 0.08, "p95_ms": 51, "peak_mb": 290},
    {"name": "pruned", "error": 0.09, "p95_ms": 62, "peak_mb": 340},
    {"name": "int8", "error": 0.075, "p95_ms": 56, "peak_mb": 230},
]
metrics = ("error", "p95_ms", "peak_mb")

def dominates(left, right):
    return (all(left[key] <= right[key] for key in metrics) and
            any(left[key] < right[key] for key in metrics))

frontier = [candidate["name"] for candidate in candidates
            if not any(dominates(other, candidate) for other in candidates
                       if other is not candidate)]
assert set(frontier) == {"baseline", "student", "int8"}

Performance and operating cost

A direct Pareto screen over C candidates and M metrics takes O(C squared times M) time and O(C) output storage; C is usually small. Measuring each model on the same target set and hardware costs far more. Shadow inference temporarily doubles or increases serving work and must fit available capacity.

Common Mistakes

  • Do not compare latency from different devices or loads.
  • Do not choose a candidate on the final test and reuse that test as an independent release check.
  • Do not call a shadow gain proven policy value when the live action would change.

Read next

ai-data
machine-learning
Storage details