Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Inference contracts: preserve preprocessing and measure tail latency

Last updated: 6 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Deployment binds weights to input transforms and class mapping, then measures latency under realistic request load.

Package the decision

Publish model weights, preprocessing version, dimensions, channel order, class names and threshold together. A server that resizes differently or reverses channels may return plausible scores with the wrong meaning. Validate file type, decoded dimensions and maximum bytes before allocating large tensors. Tensor contracts] are API constraints, not training-only checks.

Separate score and action

Return a score only with a declared class order. Convert logits to probabilities when the endpoint promises them, then apply a threshold chosen under a review-capacity policy. Include an abstain path for unsupported or poor-quality inputs. Threshold analysis] connects model output to the operational action.

Measure loaded behavior

Warm the model, then measure decode, preprocessing, queue, forward pass and serialization separately. Report median and high-percentile latency under concurrent traffic, rather than one idle request. Batch inference may raise throughput while making one request wait longer. State maximum batch and timeout, and reject overload predictably.

Monitor failure safely

Count decode failures, invalid shapes, abstentions and score shifts by relevant slice. Avoid raw personal images in routine logs. Keep a rollback artifact and known-good requests; compare outputs and class names with the manifest before shifting traffic.

Implementation

python
def predict_receipt(image_bytes, receipt_model, preprocess, class_names):
    if len(image_bytes) > 4_000_000:
        raise ValueError("image exceeds request limit")
    image_tensor = preprocess(image_bytes)
    if image_tensor.shape != (3, 224, 224):
        raise ValueError("unexpected dimensions")
    receipt_model.eval()
    with torch.inference_mode():
        logits = receipt_model(image_tensor.unsqueeze(0))
        probabilities = logits.softmax(dim=1)[0]
    return dict(zip(class_names, probabilities.tolist()))

Performance and operating cost

Prediction costs preprocessing and one model forward pass. Concurrent requests add queue and memory pressure; batching exchanges lower compute per image for additional wait time. Measure p95 under intended concurrency.

Common Mistakes

  • Do not deploy weights without class mapping and preprocessing.
  • Do not treat an idle request as a latency guarantee.
  • Do not log sensitive images during routine errors.

Read next

Continue the workflow: Shadow and canary rollout: compare a candidate without losing a rollback.

Continue the workflow: Model inference budget and device profile.

Continue the workflow: Causal target shifts and padding-aware loss.

Continue the workflow: Project: adapt a receipt model to a new capture device.

Continue the workflow: Recurrent hidden-state resets and truncated gradients.

Continue the workflow: Project: audit receipt confidence and review handoff.

Continue the workflow: Post-training quantization and calibration contracts.

Continue the workflow: Autoregressive KV cache position and mask parity.

Continue the workflow: Project: segment clipped edges and folds on receipts.

Continue the workflow: Project: attribute service incidents with a time-safe graph.

Continue the workflow: Project: forecast warehouse queues with a causal temporal model.

Continue the workflow: Project: detect total, date and merchant fields on receipts.

Continue the workflow: Project: classify machine alarms from short audio windows.

Continue the workflow: Project: detect conveyor jams from timestamped video clips.

Continue the workflow: Project: audit a neural inventory dispatch policy before release.

Continue the workflow: Project: inspect pallet damage from 3D point scans.

Continue the workflow: Project: test sparse experts for service-event classification.

Continue the workflow: Project: retrieve learning articles with two neural towers.

Continue the workflow: Project: audit a warehouse-shift PPO dispatcher.

Continue the workflow: Project: adapt a service-event classifier with low-rank weights.

Continue the workflow: Project: route uncertain receipt defects to human review.

ai-data
deep-learning
Storage details