Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: trace a missing revenue slice to one release

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Investigate a missing regional revenue slice using committed generations, input versions and bounded telemetry, then publish a corrected answer.

Create a controlled discrepancy

Build an account-day revenue mart from order generation 48 and rate revision 8. Publish generation 49, then introduce a failed generation-50 attempt that omitted one region because a filter used the wrong field. The serving index must keep pointing to generation 49 while the failed candidate remains inspectable. The incident begins with a consumer query returning 7,300 cents where the approved control expects 9,100.

Follow the answer backward

Record the query generation, index pointer, table snapshot, successful run ID and its exact inputs. Compare the failed candidate separately; do not mistake the most recent attempt for the published answer. Run-scoped lineage should show whether the missing slice originated in source input, transform logic or index publication.

Use signals without exploding labels

Alert on consumer-visible freshness and a regional reconciliation failure using bounded labels such as pipeline, stage and failure class. Put run-50 in the trace and restricted log, not a metric dimension. The cardinality budget should remain unchanged as the number of orders rises. Verify no raw customer identifier enters diagnostic attributes.

Repair from pinned inputs

Correct the transform, pin the accepted order and rate generations, and rerun the affected interval. Compare totals, unique account-day keys and the region distribution before publishing generation 51. Switch the table and index pointers only after both are validated. A partial index load must not become visible. Keep the failed attempt and correction reason in the manifest.

Hand over an evidence chain

Deliver the original query, observed mismatch, alert timestamp, trace link key, published lineage, failed attempt, corrected checksums and first good query. Reproduce the investigation from the artifacts alone. Record detection and repair durations separately; a green worker dashboard after rerun is weaker evidence than the corrected consumer answer.

Implementation

python
attempts = [
    {"run": "run-49", "status": "committed", "generation": 49, "regional_cents": 7300},
    {"run": "run-50", "status": "failed", "generation": 50, "regional_cents": 0},
    {"run": "run-51", "status": "committed", "generation": 51, "regional_cents": 9100},
]

def visible_total(records, published_generation):
    releases = [item for item in records if item["generation"] == published_generation
                and item["status"] == "committed"]
    if len(releases) != 1:
        raise ValueError("invalid published generation")
    return releases[0]["regional_cents"]

assert visible_total(attempts, 49) == 7300
assert visible_total(attempts, 51) == 9100

Performance and operating cost

The reference scan is O(R) time for R attempts and O(R) retained evidence; a generation index makes lookups O(1) expected. The exercise also pays for pinned input snapshots and one extra candidate output. That is deliberate: discarding failure evidence to save metadata makes the exact cause harder to prove.

Common Mistakes

  • Do not display an uncommitted candidate as the current answer.
  • Do not repair from moving inputs without recording the new versions.
  • Do not equate a successful rerun with a validated consumer query.

Read next

ai-data
data-engineering
Storage details