Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Tenant query budgets and fairness

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A tenant query budget limits one tenant’s concurrent work and cumulative resource use so a shared analytics service can preserve latency for others.

Identify the tenant at the trusted edge

The tenant key must come from authenticated request context, not a user-editable SQL comment. Enforce row access separately from compute limits: a fair scheduler does not stop a query from reading another tenant’s rows. Row policy tests must still run under every service identity and job type.

Cap active and queued work

Give each tenant a maximum running count, waiting count and query deadline. A single large tenant should not fill the global queue while smaller tenants wait. Use per-tenant queues or weighted scheduling with an explicit minimum share. A tenant at its cap can wait within its own bound, but should not consume another tenant’s reserved slot.

Track usage over a window

Concurrency alone does not stop a tenant from submitting one expensive scan after another. Record bytes scanned, CPU time or charged work units against a rolling budget. Decide whether unused allowance carries over; an unlimited rollover creates a future burst. Include retries and cancelled work according to a written billing and abuse policy. Admission should check the budget before work begins.

Avoid false precision in estimates

A query planner can underestimate skew, cache misses and join expansion. Treat estimates as admission hints and enforce runtime limits. When the actual cost exceeds the estimate, charge the measured usage and stop or throttle according to contract. Keep the reason visible to the user; an opaque timeout makes it impossible to distinguish fairness from a broken query.

Validate isolation with contention

Start two expensive scans for tenant North, one dashboard read for tenant South and a maintenance job. The South read should meet its queue target even when North reaches its cap. Verify global memory stays bounded and that maintenance eventually runs. Observe both per-tenant and total utilization; a reservation that stays idle during a heavy burst may waste capacity unless safe borrowing is defined.

Implementation

python
limits = {"north": 2, "south": 1}
running = {"north": 2, "south": 0}
remaining_units = {"north": 16, "south": 47}

def may_start(tenant, estimated_units):
    return (tenant in limits and running[tenant] < limits[tenant]
            and estimated_units <= remaining_units[tenant])

assert not may_start("north", 8)
assert may_start("south", 8)
assert not may_start("south", 52)

Performance and operating cost

The reference checks use O(1) expected map lookups and O(T) stored policy state for T tenants. A scheduler needs bounded queues and accounting for active requests; its runtime cost is usually small beside query execution. Reserved capacity raises predictable cost, while controlled borrowing improves utilization. Keep authorization enforcement separate from this resource policy.

Common Mistakes

  • Do not accept a tenant identifier supplied only in query text.
  • Do not mistake compute fairness for row-level access control.
  • Do not rely solely on estimated cost without runtime limits.

Read next

ai-data
data-engineering
Storage details