Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Natural-language queries: intent, schema grounding and ambiguity

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Convert a user question into a typed, reviewable query plan before any database syntax or execution is considered.

Separate wording from data shape

“Show gateway-west incidents after the rollout” expresses a dataset, service filter and time boundary. It does not name a table or prove which rollout is meant. Map each phrase to an approved semantic field, operator and typed value. Store the original question and the schema version that supplied the mapping. If “rollout” could mean two release events, ask for clarification rather than inventing a date. Entity linking with a NIL decision offers the same discipline for an ambiguous service name.

Ground fields in a governed catalog

A model may propose a plausible column that is absent, sensitive or semantically wrong. The application should accept only published field IDs from a catalog that states type, access level and allowed operations. A free-text “customer” field may mean account owner, ticket submitter or affected tenant; choosing one silently changes the answer. Keep the mapping from question span to field ID for review. Quantity and unit contracts matter when a question asks for duration or a numeric threshold.

Treat uncertainty as a plan state

Validate required slots before compiling anything: dataset, allowed projection, filters, time range and row limit. Mark unsupported aggregation or missing entity identity as unresolved. Do not let an untrusted document define new schema names or permissions. A typed plan allows the execution layer to reject unsafe operations without trying to understand arbitrary generated SQL. Retrieved-text trust keeps runbook prose from becoming query authority.

Evaluate semantic correctness

Exact string comparison between SQL queries can punish equivalent plans, while one matching result on a small fixture can reward an accidental filter. Review the intended field bindings and test execution on cases that distinguish close interpretations. Include aliases, date boundaries, null values and users with different access scopes. The compiler lesson checks the plan against an allowlist; the project checks answer behavior end to end.

Implementation

python
def validate_incident_plan(plan, allowed_fields):
    if plan.get("dataset") != "incidents":
        return {"state": "review", "reason": "unknown-dataset"}
    if not plan.get("service_id") or not plan.get("after_utc"):
        return {"state": "review", "reason": "missing-constraint"}
    projection = plan.get("projection", [])
    if not projection or any(field not in allowed_fields for field in projection):
        return {"state": "review", "reason": "field-scope"}
    if plan.get("limit", 0) not in range(1, 48):
        return {"state": "review", "reason": "row-limit"}
    return {"state": "validated", "plan": dict(plan)}

candidate = {"dataset": "incidents", "service_id": "gateway-west",
             "after_utc": "2026-09-27T00:00:00Z",
             "projection": ["incident_id", "opened_at"], "limit": 47}
assert validate_incident_plan(candidate,
                              {"incident_id", "opened_at"})["state"] == "validated"
assert validate_incident_plan({**candidate, "projection": ["secret_note"]},
                              {"incident_id", "opened_at"})["state"] == "review"

Performance and operating cost

Checking p projected fields takes expected O(p) time and O(p) space for the returned plan copy. This example checks shape and field names, not timestamp validity or user authorization. A production boundary must parse typed values, bind the authenticated principal and reject unsupported operations before database access.

Common Mistakes

  • Treating a plausible generated column name as a real schema field.
  • Guessing an ambiguous rollout date.
  • Calling a plan safe because its SQL parses.
  • Letting retrieved text or model output expand the user’s field access.

Read next

ai-data
natural-language-processing
Storage details