Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: protect analytics reads during a tenant burst

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Design and test an admission policy that protects customer dashboards while allowing bounded batch progress during a shared-cluster burst.

Record the workload contract

Two tenants share an analytics cluster. North runs long exploratory scans; South has a dashboard with a 14-second end-to-end target. A maintenance extract must complete within 36 minutes. Set per-class concurrency, queue size, deadline and measured usage budget. Admission rules apply before query execution, while row policies protect data access.

Simulate the burst

Start two North scans, enqueue a third, then submit two South dashboard queries and one maintenance extract. Reject the North request that exceeds its waiting bound. Expire a queued dashboard refresh whose deadline passes, then admit its newer replacement. Record enqueue, start and finish time for each accepted request. A dashboard that runs quickly after a long wait still misses its target.

Enforce tenant boundaries

Derive tenant identity from authenticated context, execute row-policy tests under both tenant identities and confirm North cannot read South’s account rows. Apply cost limits to measured work units as well as estimated units. Tenant fairness must not accidentally allow a query to borrow authorization along with idle compute.

Publish only a consistent result

Run dashboard queries against a named serving-index generation and disclose its source cutoff. A backlog can make a fast query stale; compare the answer’s freshness with the product contract. If the index rebuild and interactive queries share workers, reserve enough capacity for both or provide a dated last-good response. Freshness evidence belongs in the result.

Review the failure transcript

Deliver admitted, queued, expired and rejected request lists; queue and runtime percentiles; per-tenant work units; policy and row-access test results; and maintenance completion time. Inject one estimate that is too low and verify a runtime cap still stops it. The policy passes only when South’s latency target and the maintenance deadline both hold under the burst.

Implementation

python
requests = [
    {"id": "north-71", "tenant": "north", "units": 31},
    {"id": "north-72", "tenant": "north", "units": 29},
    {"id": "south-73", "tenant": "south", "units": 5},
]
running_cap = {"north": 2, "south": 1}

def schedule_by_tenant(items, caps):
    active, accepted = {}, []
    for request in items:
        tenant = request["tenant"]
        if active.get(tenant, 0) < caps[tenant]:
            accepted.append(request["id"])
            active[tenant] = active.get(tenant, 0) + 1
    return accepted

assert schedule_by_tenant(requests, running_cap) == ["north-71", "north-72", "south-73"]

Performance and operating cost

The reference pass is O(Q) time for Q requests and O(T) counter state for T tenants. A full scheduler also maintains bounded waiting queues and per-window usage ledgers. Capacity reserved for dashboards and maintenance adds cost, but an uncontrolled burst can exceed the user-facing deadline even when the average cluster utilization appears reasonable.

Common Mistakes

  • Do not let a long scan fill every shared worker.
  • Do not report latency without queue time and freshness.
  • Do not treat tenant resource limits as a replacement for data-access policy.

Read next

ai-data
data-engineering
Storage details