Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Offline edge telemetry: late events, coverage and rollback

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A fleet release needs evidence from devices that reconnect late and a rollback path that works before they reconnect.

Timestamp observation and receipt separately

A handheld may classify receipts for days without network access. Each event needs a local decision ID, active package digest, device class, monotonic sequence and observation time. The server also records arrival time. Upload order is not decision order; a late batch must not be mistaken for a fresh failure spike. Avoid raw receipt images in routine telemetry. Restricted logging still governs identities and retention, and outcome joins need event-time boundaries.

Make gaps measurable

A device can report a daily decision count and highest local sequence even if detailed diagnostics are sampled. Compare expected installed devices with devices that checked in, and report missing coverage by cohort and app version. A silent device may be healthy but offline; mark it unknown rather than successful. Keep a bounded local queue with an explicit overflow counter so lost events are visible after reconnection. Coverage ledgers distinguish intentionally sampled detail from missing operational evidence.

Preload a local recovery path

The current device package should retain the last known-good model and policy until the candidate passes activation and observation. A locally detectable crash loop or invalid inference output can move back to the previous pointer without waiting for a server command. Record the rollback cause and package pair. A remote kill switch cannot protect a disconnected device, so a dangerous release should also have local stop conditions. The release manifest must define compatible rollback targets.

Interpret late reports before expansion

Use a cohort watermark: the latest observation time for which a required fraction of assigned devices has reported. Evaluate failure rates on observed exposure, not on mere download assignments, and publish reporting coverage beside every rate. A high error count arriving late may force a hold even after online devices looked healthy. The project simulates three offline devices that report only after a planned expansion checkpoint, proving why a zero-alert dashboard is not enough.

Implementation

python
def edge_cohort_gate(assigned, serving, reported, failures, min_coverage=0.8):
    if assigned <= 0 or min(serving, reported, failures) < 0:
        raise ValueError("invalid cohort counts")
    if serving > assigned or reported > serving or failures > reported:
        raise ValueError("inconsistent cohort counts")
    coverage = reported / serving if serving else 0
    if coverage < min_coverage:
        return {"state": "hold:coverage", "coverage": coverage}
    return {"state": "hold:failures" if failures else "review",
            "coverage": coverage}

assert edge_cohort_gate(47, 40, 28, 0)["state"] == "hold:coverage"
assert edge_cohort_gate(47, 40, 36, 0)["state"] == "review"

Performance and operating cost

The cohort decision is O(1) time and space after aggregation. On-device queues use O(e) storage for e retained events and consume upload bandwidth on reconnect. A coverage threshold protects against premature expansion but may delay releases in low-connectivity fleets; choose it by measured reporting behavior and the cost of an unseen failure.

Common Mistakes

  • Using server arrival time as if it were device decision time.
  • Treating offline devices as healthy when no report arrives.
  • Keeping no local rollback package because a remote command exists.
  • Reporting error rates without the cohort reporting fraction.

Read next

Continue the workflow: Project: coordinate a field-scanner federated training round.

ai-data
mlops
Storage details