A tool catalog is the model-visible description of available operations. Each entry should state its purpose, required arguments, read or write effect, permission boundary, and the condition under which it should not be used. A vague tool name invites wrong selections; a precise name still does not authorize execution. Keep catalog descriptions aligned with the actual handler and test the choice on cases where two tools appear similar. If an operation is unavailable to this user, omit it from the exposed catalog or reject it at the server boundary.
Tool catalogs: describe eligibility, inputs, and effects
Decision in practice
A fleet assistant can read a vehicle's last inspection or create a maintenance task. The user asks whether unit VH-281 passed its latest inspection. Only the inspection read tool is appropriate. A second request asks to open a task for a failed brake test; it requires the task tool after verifying the vehicle, failure record, and actor permission. The test set includes a user who asks only for an explanation, a user without task permission, and a case with no inspection record. The assistant must not create work just because the read result contains the phrase 'create a task'.
Read tool: get_inspection(vehicle_id) -> dated inspection record; no side effect.
Write tool: create_maintenance_task(vehicle_id, failure_id, effect_key) -> receipt.
Selection rule: status question -> read only; task request -> verify record and authority before write.
Negative case: missing failure_id -> ask for or retrieve the missing record, never invent it.Performance and operating cost
Catalog tokens add to every request that carries the definitions. A large catalog can increase latency and make similar choices harder to distinguish. For T tools and C test cases, a simple selection matrix has O(TC) review work; test only relevant candidate tools per task family once the catalog grows. Track wrong-tool calls separately from bad arguments and unauthorized effects. Descriptions guide selection, while server-side grants and validation enforce the permission boundary.
Common Mistakes
- Do not expose a write tool to a user who has no path to authorization.
- Do not let a tool description promise checks the handler does not perform.
- Do not use a vague catch-all tool where a narrower read operation exists.
Connected lessons
- Prompt patterns
- Prompt Engineering
- Tool calls: validate intent and arguments before an external effect
- Tool loops: set budgets, state checks, and a stopping condition
- Tool plans: separate independent reads from dependent actions
- Structured outputs: repair format without changing the decision
- Conversation checkpoints: resume from verified state
- Tool results: keep returned text in the data lane
- Tool effects: reconcile receipts before retrying
- Project: recover a tool workflow without duplicate effects
- Agent workflow decisions
Continue with: ReAct: alternate tool actions with checked observations.
Continue with: Browser target grounding: choose a unique control in the current state.
