FIELD NOTE / SYSTEMS REFERENCE
LLM Tool Calling: The Complete Field Reference
A decision-first field reference for writing reliable tool definitions, controlling calls, handling failures, preserving attribution, and shipping provider-neutral agent loops.
The implementation is not the hard part. The hard part is preserving intent, identity, authority, and useful failure information while a probabilistic caller crosses a deterministic boundary.
Most tool-calling documentation begins with a conceptual loop. Production incidents begin with a message that had the right shape but the wrong meaning. Start by making the protocol mechanically legible, then use the entries as decision records.
How do I write a tool definition that the model actually uses reliably?
WHEN THIS BITES YOU · The tool passes unit tests, then production traffic selects a neighboring tool or fills a field with a plausible value that no human meant.
THE PATTERN
Treat the definition as an inference-time routing contract. Give the tool one job, a concrete positive trigger, a concrete negative trigger, typed parameters, bounded enums, and defaults for values the user did not specify. Put the operational rule in description, not only in a parameter name. The model sees the tool name, top-level description, parameter descriptions, required list, and enum values as a combined decision surface.
What you wrote vs what the model sees: You wrote a schema for a validator. The model reads it as a compressed policy document. A field named `date` tells it almost nothing; ‘the requested appointment date in YYYY-MM-DD; never infer a date from today unless the user says today’ does.
FAILURE MODES
Right tool, wrong boundary
Symptom: The model calls `search_customers` for a request that needs `get_customer`, then invents a filter. The descriptions overlap.
Fix: Give each tool a mutually exclusive purpose and say when not to use it.
Type-valid nonsense
Symptom: An enum accepts `standard`, but production sends `standard_plus`; the model silently chooses the nearest value.
Fix: Model the real domain or reject the request explicitly; do not hide unsupported values.
Developer-language drift
Symptom: A description says ‘hydrate the principal object’ and argument accuracy collapses outside the team that wrote it.
Fix: Write descriptions as user-intent rules with examples and exclusions.
A billing lookup tool worked perfectly in staging because every test used the word ‘invoice’. Production users said ‘what did we charge them?’ The model selected a payment-history tool because the description was written for developers, not for user language.
Parallelism looks like a performance feature until a tool has state. Then it becomes an ordering question. The provider can tell you which calls exist; it cannot know which writes your business process must serialize.
When should I use parallel tool calls versus sequential tool calls?
WHEN THIS BITES YOU · The model emits two calls in one response and your executor assumes the array is ordered, or two stateful writes race each other.
THE PATTERN
Parallel calls are safe when operations are independent, read-only or idempotent, and their results can be joined without order. Sequential calls are required when tool B consumes tool A's output, either call mutates shared state, or the second call must observe a committed first result. Represent every call by its provider-issued id; never use array position as identity.
What you wrote vs what the model sees: The model can infer some dependencies from descriptions, but it cannot guarantee execution order. A sentence saying ‘use the customer id from the previous call’ is not a transaction boundary.
FAILURE MODES
Array-order coupling
Symptom: The executor maps result[0] to the first call it expected, but the provider or middleware reordered calls.
Fix: Join results by tool-call id and preserve the id through every adapter.
Hidden dependency
Symptom: `reserve_inventory` and `create_order` run together; the order is created before the reservation commits.
Fix: Expose the dependency as a state transition and execute the second call only after the first result validates.
Parallel writes
Symptom: Two profile updates race and the last write wins even though the model produced both calls intentionally.
Fix: Parallelize reads; serialize writes behind an idempotency key and version check.
A payment reconciliation worker parallelized ‘mark settled’ and ‘send receipt’. Under retry load, receipts were sent for transactions that later failed settlement. Sequential calls added 80 ms and removed the race.
How do I handle tool call failures without breaking the agent loop?
WHEN THIS BITES YOU · The loop either stops on the first timeout or retries a side effect until the downstream system sees duplicates.
THE PATTERN
Return a typed result envelope even on failure: `{ status: 'failed' | 'pending' | 'succeeded', retryable, code, message, data? }`. Give the model enough information to choose a next action, but never expose secrets or raw stack traces. Retry only when the operation is retryable, the call is idempotent, the budget remains, and the failure class is known.
What you wrote vs what the model sees: A tool error is not the same as a bad plan. A valid result can still lead to a wrong next action, so execution success must not be your only loop health signal.
FAILURE MODES
Silence after failure
Symptom: The model receives no tool result and answers as if the tool never ran.
Fix: Inject a structured failure result with status, safe reason, and allowed next actions.
Retrying a side effect
Symptom: A timeout after a charge request looks retryable, so the loop charges twice.
Fix: Use idempotency keys and distinguish unknown outcome from confirmed failure.
Infinite recovery
Symptom: The model sees ‘try again’ and repeats the same call until token or cost limits stop it.
Fix: Track attempts by operation and failure class; after the cap, escalate or answer with uncertainty.
A payment API returned ‘pending’ with HTTP 200. The agent schema only had `success: boolean`, so it retried three times. The fix was an explicit outcome enum plus idempotency keys.
Argument hallucination is structurally different from a hallucinated sentence. A sentence can be reviewed for plausibility. An argument can be plausible enough to pass a type checker and still point at the wrong person, account, date, or resource.
How do I control whether the model uses a tool, forces a tool, or avoids tools?
WHEN THIS BITES YOU · A router forces a tool for every request and the model invents arguments instead of admitting that the user asked for explanation, not execution.
THE PATTERN
Use three modes deliberately: `auto` when the model should choose, `required` when at least one tool action is part of the contract, and a named tool only when the application has already established the exact operation. Use `none` for turns that must remain explanatory, confirmation-only, or privacy-safe. Treat tool choice as a policy decision, not a confidence slider.
What you wrote vs what the model sees: The application knows the workflow state. The model knows linguistic intent. Forcing one without the other creates a false certainty that appears as a malformed or fabricated argument.
FAILURE MODES
Forced wrong tool
Symptom: The model fills a required `account_id` with a guess because the router forced an account lookup.
Fix: Let the model ask a clarification question or route to `none`.
Required without confirmation
Symptom: A required tool is selected before the user confirms a destructive action.
Fix: Separate planning from execution and require an explicit confirmation state.
Streaming surprise
Symptom: The UI renders partial text while the model later switches into a tool call, leaving two competing answers.
Fix: Treat a tool-call delta as a mode switch and buffer user-visible text until the turn resolves.
A support workflow forced `lookup_subscription` for every billing question. The model fabricated subscription ids for users who only wanted an explanation of a policy. Allowing `none` reduced false lookups more than adding another prompt paragraph.
How do I pass tool results back correctly for different providers?
WHEN THIS BITES YOU · The tool ran successfully, but the next model turn cannot associate the result with the call that produced it—or the result is so large that reasoning degrades.
THE PATTERN
Preserve the provider call id, tool name where required, result status, and a bounded payload. Store the full result outside the prompt and return a compact model-facing projection with identifiers, totals, exceptions, and links to trace data. When multiple calls complete, return one result for each call id in the provider's required message structure; never concatenate them into one unlabeled blob.
What you wrote vs what the model sees: Your executor has an object. The provider expects a message or content block with a particular role and id relationship. A semantically correct result in the wrong envelope is still invisible to the model.
FAILURE MODES
Unattributed result
Symptom: The model says it has no result even though the backend returned 200.
Fix: Echo the exact tool-call id and provider-specific result role.
Full payload echo
Symptom: A 14 KB result pushes earlier instructions and evidence out of the useful context window.
Fix: Summarize for the model; retain the full payload in a trace store.
Mixed result types
Symptom: One call returns an object and another returns a string; the adapter serializes both as text and loses fields.
Fix: Use a stable result envelope and explicit content type.
A 14 KB customer profile result caused the model to ignore the first half of the conversation on the next turn. Returning a 420-token projection preserved the decision fields and kept the trace complete.
How do I prevent the model from hallucinating tool arguments?
WHEN THIS BITES YOU · The arguments pass JSON Schema, look reasonable in logs, and still identify a real person, account, or date that the user never supplied.
THE PATTERN
Constrain the space, then validate the meaning. Use enums for finite choices, optional fields with explicit defaults when omission is safe, and descriptions that state provenance rules. Validate syntax, schema, authorization, existence, and cross-field consistency before execution. Dates, identifiers, and nested objects deserve semantic validation because type correctness does not prove that the value exists or belongs to the current user.
What you wrote vs what the model sees: The model can produce a valid string that is not a valid identifier in your system. The execution layer must own that distinction.
FAILURE MODES
Plausible identifier
Symptom: A patient id matches the required pattern but belongs to nobody in the authorized tenant.
Fix: Resolve identifiers against an authorized index before execution; reject unknown values.
Date invention
Symptom: ‘Next Friday’ becomes a date in the server timezone rather than the user's timezone.
Fix: Normalize timezone and ask for clarification when a relative date is ambiguous.
Nested object drift
Symptom: The model places `currency` inside `amount` because both descriptions mention money.
Fix: Use shallow schemas where possible and validate cross-field shape before calling.
A medical scheduling assistant generated patient ids that matched the format and passed type validation. Existence and tenant checks caught the problem before the scheduling API saw the request.
A context window is not a memory model. It is a bounded serialization budget shared by instructions, history, calls, results, wrappers, and the next decision. Measure the conversation you actually send.
How do I chain tool calls across multiple turns without context window pollution?
WHEN THIS BITES YOU · The agent reaches turn eight, still has every raw tool response, and starts repeating earlier reasoning because the context budget is full.
THE PATTERN
Treat each turn as an auditable checkpoint. Keep the current objective, verified facts, pending actions, tool ids, failure state, and next allowed transitions. Summarize completed work into a state record; retain raw tool payloads in external trace storage. Prune only after result attribution is captured. In practice, measure the chain rather than choosing a universal depth: quality often degrades before the advertised context limit when large results and repeated instructions compete for attention.
What you wrote vs what the model sees: The context window counts messages, tool arguments, tool results, hidden wrappers, and repeated system instructions. ‘The model supports 128K’ is not a chain budget.
FAILURE MODES
Raw-result accumulation
Symptom: Every response is appended verbatim and later turns spend tokens re-reading old payloads.
Fix: Checkpoint verified facts and externalize raw results.
Broken attribution after pruning
Symptom: A result remains but its tool-use id was removed, so the provider rejects the next message.
Fix: Prune complete exchanges, never one half of a tool-use/tool-result pair.
Reasoning displacement
Symptom: The model's earlier plan is truncated before the tool results, causing a repeated or contradictory plan.
Fix: Persist the plan state separately and re-inject only the current checkpoint.
An agentic pipeline hit context limits on turn eight because nobody measured tool-result size. The repair was a checkpoint record plus a trace pointer, not a larger model.
What is different between OpenAI, Anthropic, and Gemini tool calling in practice?
WHEN THIS BITES YOU · The same adapter works for one provider, then silently drops a result or misreads a malformed call after a provider switch.
THE PATTERN
Keep a provider adapter at the boundary. Normalize calls into `{ id, name, arguments, status }`, normalize results into `{ id, status, content }`, and keep provider-specific serialization out of the agent loop. Test successful calls and malformed calls as separate contracts.
What you wrote vs what the model sees: The conceptual loop is shared. Message roles, content-part names, and malformed-call behavior are not.
FAILURE MODES
Role translation omission
Symptom: The executor returns a valid object under the wrong role and the provider treats it as user text.
Fix: Use contract tests for every role and content-part mapping.
No-result call
Symptom: A tool emits no body on success and the adapter omits the result message entirely.
Fix: Send an explicit empty success result with the call id.
Malformed JSON
Symptom: The model emits partial arguments and the executor tries to repair them by string concatenation.
Fix: Return a structured parse failure and let the loop decide whether to retry or ask.
The cross-provider bug that lasted longest was not a missing tool. It was an empty successful result that one adapter dropped and another encoded explicitly.
PROVIDER COMPARISON · THE THREE BREAKING DIFFERENCES
| Difference | OpenAI | Anthropic | Gemini |
|---|---|---|---|
| Normal successful call | tool call + tool message with tool_call_id | tool_use block + tool_result block with same id | functionCall part + functionResponse part with same name |
| Missing required argument | validate before execution; ask or retry | input schema failure; return a tool_result error | function call may be incomplete; validate before response |
| Malformed JSON | adapter returns parse failure | tool input must be parsed before tool_result | function args are a structured part; reject malformed content |
Share a thought
Comments appear immediately. Email is optional and never shown.
No comments yet. Be the first to share a thought.