Compile only validated semantic fields into a parameterized read-only query with a fixed tenant and row boundary.
Read-only query compilation: identifiers, parameters and scope
Compile from approved identifiers
Database drivers parameterize values, not table and column identifiers. Map each semantic field to an approved physical column; never paste a model-supplied identifier into SQL. Build the query shape from a small allowlisted grammar and bind service, time and tenant values as parameters. The compiled statement should be read-only and carry a hard result limit. Typed plans provide the input contract that makes this compiler manageable.
Bind access beyond the prompt
A user asking for all incidents does not gain access to another tenant because the model phrased the request confidently. Derive tenant and role from authenticated application state, not from the natural-language question. Filter rows at the database boundary and consider column-level policy for fields such as personal notes. A retrieved runbook cannot authorize a broader query. Trusted action gates apply the same separation to effectful tools.
Constrain cost and cardinality
A harmless SELECT can still scan a huge table or return too much private data. Require a bounded time window, row cap and indexed filter where the product needs them. Estimate or reject expensive plans before execution. A count question and a row-list question have different result types and should not share an improvised compiler path. Keep query timeout, logging and pagination policy outside generated prose. The example below shows a single approved row-list shape rather than claiming to handle arbitrary SQL.
Test with adversarial wording
Try quoted SQL fragments, nonexistent fields, alternate service aliases, null timestamps and limit requests above policy. Verify the resulting statement and parameters, plus returned rows under two tenant identities. Compare semantic answers, not only whether the query executed. The incident query project includes an ambiguous question that must stop for clarification before any statement is compiled.
Implementation
ALLOWED_COLUMNS = {"incident_id": "incident_id",
"opened_at": "opened_at", "status": "status"}
def compile_incident_rows(plan, tenant_id):
if not tenant_id or plan["dataset"] != "incidents":
raise ValueError("authorized tenant and dataset are required")
fields = plan["projection"]
if not fields or any(field not in ALLOWED_COLUMNS for field in fields):
raise ValueError("unknown or disallowed projection")
if not 1 <= plan["limit"] <= 47:
raise ValueError("limit outside policy")
columns = ", ".join(ALLOWED_COLUMNS[field] for field in fields)
statement = (f"SELECT {columns} FROM incidents "
"WHERE tenant_id = ? AND service_id = ? "
"AND opened_at >= ? ORDER BY opened_at DESC LIMIT ?")
parameters = (tenant_id, plan["service_id"],
plan["after_utc"], plan["limit"])
return statement, parameters
plan = {"dataset": "incidents", "projection": ["incident_id"],
"service_id": "gateway-west", "after_utc": "2026-09-27T00:00:00Z",
"limit": 47}
sql, values = compile_incident_rows(plan, "tenant-82")
assert "tenant_id = ?" in sql and values[0] == "tenant-82"
assert "gateway-west" not in sql
Performance and operating cost
Compiling p projected columns takes O(p) time and space for the statement. Database work depends on indexes, selectivity and returned rows; a tenant-service-time index can avoid a full scan for this shape. This code returns a statement and parameters but does not establish an authenticated connection or enforce database-level row policy by itself.
Common Mistakes
- Interpolating model-supplied column names into SQL.
- Trusting a tenant ID supplied by the question.
- Assuming a SELECT cannot overload a database.
- Using parameter placeholders for values while leaving identifiers unchecked.
Read next
- Natural-language queries: intent, schema grounding and ambiguity
- Project: answer an incident question through a governed query plan
- Trusted action gates for retrieval-assisted answers
- Quantities in text: value, unit, range and source span
- Text classification evaluation: inspect slices and allow abstention
