Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Inference cost ledger: allocate shared capacity without false precision

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Attribute model-serving spend to traffic and capacity while keeping idle time, shared hosts and sampling limits visible.

Choose the billable boundary

A receipt-risk service pays for instances, accelerators, memory, storage, network and supporting queues. A request counter alone does not identify the cost of idle replicas or shared hosts. Collect provider billing by resource and time window, then attach immutable deployment, model and route identities to usage measurements. Keep a separate unallocated bucket for idle or unknown capacity. Shared serving capacity explains why one tenant’s usage can affect another’s latency and bill.

Allocate with a stated rule

For an exclusive endpoint, attribute its resource cost to its model for the window. For a shared pool, allocate a measured share by GPU time, CPU time or another defensible usage signal, and report the residual as shared overhead. If only request counts exist, label the result an estimate and keep the method visible. Never force allocations to sum to a false level of precision by hiding idle capacity. Telemetry coverage tells readers how much usage was actually observed.

Normalize carefully

Cost per completed prediction can be useful, but count retries, rejected inputs, asynchronous completions and manual fallback separately. A cheap route that drops hard cases can look efficient while shifting cost to reviewers. Compare cost by model digest, hardware class, traffic cohort and quality outcome, then pair cost with p99 latency and error rate. Async result ledgers prevent submitted jobs from being mistaken for completed predictions.

Turn the ledger into a decision

Set a per-route budget with a lookback window and an owner. If cost rises, split the change into traffic volume, unit price, utilization and request work. A new model that raises per-request compute may still lower total business cost if it reduces review workload; state both measures. Budget gates choose an action, and the project traces a retry loop that makes the unit-cost graph misleading.

Implementation

python
def allocate_shared_cost(attributable_cost, idle_cost, measured_usage):
    if min(attributable_cost, idle_cost) < 0 or any(
            value < 0 for value in measured_usage.values()):
        raise ValueError("negative cost or usage")
    used = sum(measured_usage.values())
    if used <= 0:
        return {"unallocated": attributable_cost + idle_cost}
    allocated = {model: attributable_cost * value / used
                 for model, value in measured_usage.items()}
    allocated["unallocated"] = idle_cost
    return allocated

share = allocate_shared_cost(470.0, 47.0,
                             {"receipt-risk": 3.0, "invoice-risk": 2.0})
assert share["receipt-risk"] == 282.0
assert share["invoice-risk"] == 188.0
assert share["unallocated"] == 47.0
assert allocate_shared_cost(47.0, 5.0, {}) == {"unallocated": 52.0}

Performance and operating cost

Allocation is O(m) time and O(m) output space for m models. The sample divides a known attributable charge by measured usage and preserves idle cost as unallocated; production ledgers must first separate those charges, or the rule will misattribute idle spend. High-cardinality model and route labels increase telemetry storage, so aggregate at the finest level needed for a decision.

Common Mistakes

  • Dividing the entire instance bill by successful requests without accounting for idle time.
  • Calling a request-count split exact when request compute varies.
  • Counting submissions as completions in an asynchronous route.
  • Optimizing serving spend while ignoring manual-review cost and quality.

Read next

ai-data
mlops
Storage details