Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Query admission queues and deadlines

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Query admission decides whether a request runs, waits or fails before it consumes scarce memory and worker capacity.

Put a bound on waiting

A BI dashboard may issue dozens of interactive queries while a batch analyst starts a broad scan. If all run immediately, memory pressure can slow or fail everyone. Give each workload a concurrency limit and a queue limit. When the queue is full, reject with a clear retry response rather than accumulating unbounded requests. A serving index helps only the queries it was built for; admission still protects shared compute.

Treat deadline as an end-to-end budget

A 12-second query target includes queue wait, execution and result delivery. A query that waited 11 seconds should not start a 30-second scan merely because a worker became free. Record enqueue time and absolute deadline; expire stale work before starting. This prevents old dashboard refreshes from crowding out the user’s latest request.

Reserve capacity for critical reads

Separate interactive serving, scheduled extracts and exploratory scans into workload groups. A reserved slice protects critical reads during a batch burst, while controlled borrowing can use idle capacity when policy allows. Avoid a single high-priority queue that permanently starves maintenance work. Define minimum progress or maximum wait for every accepted class.

Account for cost before admission

Estimate scanned bytes, memory and partition breadth using the query plan or a coarse request class. An estimate is imperfect, so combine it with runtime caps and cancellation. Reject a query that exceeds a tenant or class budget even when the cluster is idle; otherwise one expensive request can consume the next interval’s allowance. Cost attribution closes the loop after execution.

Test a burst and recovery

Queue five interactive requests behind two running slots, add one batch scan and then cancel the oldest interactive request. Verify that no more than two run, the expired request never starts, and an admitted batch request eventually progresses. Report queue time and execution time separately. A low execution p95 can hide users waiting in admission.

Implementation

python
from collections import deque

queue = deque([{"id": "q-47", "deadline": 52},
               {"id": "q-48", "deadline": 61},
               {"id": "q-49", "deadline": 58}])

def admit_ready(waiting, now, free_slots):
    started = []
    while waiting and len(started) < free_slots:
        request = waiting.popleft()
        if request["deadline"] > now:
            started.append(request["id"])
    return started

assert admit_ready(queue, 55, 2) == ["q-48", "q-49"]

Performance and operating cost

Admission is O(Q) time in the worst case when Q expired entries must be removed and O(Q) memory for a bounded queue. Executing a query can be much more expensive than this control path. A strict queue limit protects coordinator memory; deadlines prevent stale work from occupying future capacity. Measure rejection, wait and runtime by workload class.

Common Mistakes

  • Do not measure only execution time while ignoring queue delay.
  • Do not start a request whose end-to-end deadline has already passed.
  • Do not allow an unbounded queue to masquerade as capacity.

Read next

ai-data
data-engineering
Storage details