Design and test an admission policy that protects customer dashboards while allowing bounded batch progress during a shared-cluster burst.
Project: protect analytics reads during a tenant burst
Record the workload contract
Two tenants share an analytics cluster. North runs long exploratory scans; South has a dashboard with a 14-second end-to-end target. A maintenance extract must complete within 36 minutes. Set per-class concurrency, queue size, deadline and measured usage budget. Admission rules apply before query execution, while row policies protect data access.
Simulate the burst
Start two North scans, enqueue a third, then submit two South dashboard queries and one maintenance extract. Reject the North request that exceeds its waiting bound. Expire a queued dashboard refresh whose deadline passes, then admit its newer replacement. Record enqueue, start and finish time for each accepted request. A dashboard that runs quickly after a long wait still misses its target.
Enforce tenant boundaries
Derive tenant identity from authenticated context, execute row-policy tests under both tenant identities and confirm North cannot read South’s account rows. Apply cost limits to measured work units as well as estimated units. Tenant fairness must not accidentally allow a query to borrow authorization along with idle compute.
Publish only a consistent result
Run dashboard queries against a named serving-index generation and disclose its source cutoff. A backlog can make a fast query stale; compare the answer’s freshness with the product contract. If the index rebuild and interactive queries share workers, reserve enough capacity for both or provide a dated last-good response. Freshness evidence belongs in the result.
Review the failure transcript
Deliver admitted, queued, expired and rejected request lists; queue and runtime percentiles; per-tenant work units; policy and row-access test results; and maintenance completion time. Inject one estimate that is too low and verify a runtime cap still stops it. The policy passes only when South’s latency target and the maintenance deadline both hold under the burst.
Implementation
requests = [
{"id": "north-71", "tenant": "north", "units": 31},
{"id": "north-72", "tenant": "north", "units": 29},
{"id": "south-73", "tenant": "south", "units": 5},
]
running_cap = {"north": 2, "south": 1}
def schedule_by_tenant(items, caps):
active, accepted = {}, []
for request in items:
tenant = request["tenant"]
if active.get(tenant, 0) < caps[tenant]:
accepted.append(request["id"])
active[tenant] = active.get(tenant, 0) + 1
return accepted
assert schedule_by_tenant(requests, running_cap) == ["north-71", "north-72", "south-73"]Performance and operating cost
The reference pass is O(Q) time for Q requests and O(T) counter state for T tenants. A full scheduler also maintains bounded waiting queues and per-window usage ledgers. Capacity reserved for dashboards and maintenance adds cost, but an uncontrolled burst can exceed the user-facing deadline even when the average cluster utilization appears reasonable.
Common Mistakes
- Do not let a long scan fill every shared worker.
- Do not report latency without queue time and freshness.
- Do not treat tenant resource limits as a replacement for data-access policy.
