/ ENGINEERING SYSTEMS
Back to Field Notes

FIELD NOTE / SYSTEMS REFERENCE

LLM Tool Calling: The Complete Field Reference

15 min readUpdated 2026-09-30

A decision-first field reference for writing reliable tool definitions, controlling calls, handling failures, preserving attribution, and shipping provider-neutral agent loops.

Public-safe engineering note · no client or confidential implementation details

The implementation is not the hard part. The hard part is preserving intent, identity, authority, and useful failure information while a probabilistic caller crosses a deterministic boundary.

DIAGRAM 1 / TOOL CALLING PROTOCOL AT MESSAGE LEVEL
USERMODELTOOL EXECUTIONTURN 1 · SINGLErole: usertool_use · id: call_01name + arguments · no final textexecute(call_01)TURN 2 · RESULTtool_result · id: call_01status + bounded contentsame idfinal text responsePARALLELcall_02 + call_03two results, matched by idThe model's tool-use turn is not a user-visible answer. The result must preserve the id before the model can continue.

Most tool-calling documentation begins with a conceptual loop. Production incidents begin with a message that had the right shape but the wrong meaning. Start by making the protocol mechanically legible, then use the entries as decision records.

01

How do I write a tool definition that the model actually uses reliably?

WHEN THIS BITES YOU · The tool passes unit tests, then production traffic selects a neighboring tool or fills a field with a plausible value that no human meant.

THE PATTERN

Treat the definition as an inference-time routing contract. Give the tool one job, a concrete positive trigger, a concrete negative trigger, typed parameters, bounded enums, and defaults for values the user did not specify. Put the operational rule in description, not only in a parameter name. The model sees the tool name, top-level description, parameter descriptions, required list, and enum values as a combined decision surface.

What you wrote vs what the model sees: You wrote a schema for a validator. The model reads it as a compressed policy document. A field named `date` tells it almost nothing; ‘the requested appointment date in YYYY-MM-DD; never infer a date from today unless the user says today’ does.

FAILURE MODES

Right tool, wrong boundary

Symptom: The model calls `search_customers` for a request that needs `get_customer`, then invents a filter. The descriptions overlap.

Fix: Give each tool a mutually exclusive purpose and say when not to use it.

Type-valid nonsense

Symptom: An enum accepts `standard`, but production sends `standard_plus`; the model silently chooses the nearest value.

Fix: Model the real domain or reject the request explicitly; do not hide unsupported values.

Developer-language drift

Symptom: A description says ‘hydrate the principal object’ and argument accuracy collapses outside the team that wrote it.

Fix: Write descriptions as user-intent rules with examples and exclusions.

PRODUCTION NOTE

A billing lookup tool worked perfectly in staging because every test used the word ‘invoice’. Production users said ‘what did we charge them?’ The model selected a payment-history tool because the description was written for developers, not for user language.

DIAGRAM 2 / TOOL DEFINITION ANATOMY
tool: schedule_appointmentdescriptionUse when the user asks to book or movean appointment. Do not use for policy questions.parameters.propertiesdate · string · YYYY-MM-DDtimezone · enum · [user, UTC]duration · integer · default 30required: [date]only values the user must supplyPrimary routing signalTell the model when to use the tool and when not to.Argument mapping surfaceDescriptions map user language to values. Write for inference.The hallucination boundaryRequired fields without user evidence invite invention.

INTERACTIVE / TOOL DEFINITION BUILDER

LIVE CRITIQUE · UNRELIABLE IN PRODUCTION

OPEN · Description must state when to use and not use the tool.

OPEN · Enum must include values present in production data.

OPEN · Only user-supplied values belong in required; duration and timezone have defaults.

Parallelism looks like a performance feature until a tool has state. Then it becomes an ordering question. The provider can tell you which calls exist; it cannot know which writes your business process must serialize.

02

When should I use parallel tool calls versus sequential tool calls?

WHEN THIS BITES YOU · The model emits two calls in one response and your executor assumes the array is ordered, or two stateful writes race each other.

THE PATTERN

Parallel calls are safe when operations are independent, read-only or idempotent, and their results can be joined without order. Sequential calls are required when tool B consumes tool A's output, either call mutates shared state, or the second call must observe a committed first result. Represent every call by its provider-issued id; never use array position as identity.

What you wrote vs what the model sees: The model can infer some dependencies from descriptions, but it cannot guarantee execution order. A sentence saying ‘use the customer id from the previous call’ is not a transaction boundary.

FAILURE MODES

Array-order coupling

Symptom: The executor maps result[0] to the first call it expected, but the provider or middleware reordered calls.

Fix: Join results by tool-call id and preserve the id through every adapter.

Hidden dependency

Symptom: `reserve_inventory` and `create_order` run together; the order is created before the reservation commits.

Fix: Expose the dependency as a state transition and execute the second call only after the first result validates.

Parallel writes

Symptom: Two profile updates race and the last write wins even though the model produced both calls intentionally.

Fix: Parallelize reads; serialize writes behind an idempotency key and version check.

PRODUCTION NOTE

A payment reconciliation worker parallelized ‘mark settled’ and ‘send receipt’. Under retry load, receipts were sent for transactions that later failed settlement. Sequential calls added 80 ms and removed the race.

03

How do I handle tool call failures without breaking the agent loop?

WHEN THIS BITES YOU · The loop either stops on the first timeout or retries a side effect until the downstream system sees duplicates.

THE PATTERN

Return a typed result envelope even on failure: `{ status: 'failed' | 'pending' | 'succeeded', retryable, code, message, data? }`. Give the model enough information to choose a next action, but never expose secrets or raw stack traces. Retry only when the operation is retryable, the call is idempotent, the budget remains, and the failure class is known.

What you wrote vs what the model sees: A tool error is not the same as a bad plan. A valid result can still lead to a wrong next action, so execution success must not be your only loop health signal.

FAILURE MODES

Silence after failure

Symptom: The model receives no tool result and answers as if the tool never ran.

Fix: Inject a structured failure result with status, safe reason, and allowed next actions.

Retrying a side effect

Symptom: A timeout after a charge request looks retryable, so the loop charges twice.

Fix: Use idempotency keys and distinguish unknown outcome from confirmed failure.

Infinite recovery

Symptom: The model sees ‘try again’ and repeats the same call until token or cost limits stop it.

Fix: Track attempts by operation and failure class; after the cap, escalate or answer with uncertainty.

PRODUCTION NOTE

A payment API returned ‘pending’ with HTTP 200. The agent schema only had `success: boolean`, so it retried three times. The fix was an explicit outcome enum plus idempotency keys.

INTERACTIVE / FAILURE MODE DIAGNOSTIC

ROOT CAUSERetry state is not visible to the model or executor.

CONFIRMING LOGRepeated identical name + arguments with no new evidence.

FIXAdd attempt budget and idempotency at schema/execution; return a terminal status in the prompt.

FREQUENCY · OFTEN MISSED

Argument hallucination is structurally different from a hallucinated sentence. A sentence can be reviewed for plausibility. An argument can be plausible enough to pass a type checker and still point at the wrong person, account, date, or resource.

04

How do I control whether the model uses a tool, forces a tool, or avoids tools?

WHEN THIS BITES YOU · A router forces a tool for every request and the model invents arguments instead of admitting that the user asked for explanation, not execution.

THE PATTERN

Use three modes deliberately: `auto` when the model should choose, `required` when at least one tool action is part of the contract, and a named tool only when the application has already established the exact operation. Use `none` for turns that must remain explanatory, confirmation-only, or privacy-safe. Treat tool choice as a policy decision, not a confidence slider.

What you wrote vs what the model sees: The application knows the workflow state. The model knows linguistic intent. Forcing one without the other creates a false certainty that appears as a malformed or fabricated argument.

FAILURE MODES

Forced wrong tool

Symptom: The model fills a required `account_id` with a guess because the router forced an account lookup.

Fix: Let the model ask a clarification question or route to `none`.

Required without confirmation

Symptom: A required tool is selected before the user confirms a destructive action.

Fix: Separate planning from execution and require an explicit confirmation state.

Streaming surprise

Symptom: The UI renders partial text while the model later switches into a tool call, leaving two competing answers.

Fix: Treat a tool-call delta as a mode switch and buffer user-visible text until the turn resolves.

PRODUCTION NOTE

A support workflow forced `lookup_subscription` for every billing question. The model fabricated subscription ids for users who only wanted an explanation of a policy. Allowing `none` reduced false lookups more than adding another prompt paragraph.

05

How do I pass tool results back correctly for different providers?

WHEN THIS BITES YOU · The tool ran successfully, but the next model turn cannot associate the result with the call that produced it—or the result is so large that reasoning degrades.

THE PATTERN

Preserve the provider call id, tool name where required, result status, and a bounded payload. Store the full result outside the prompt and return a compact model-facing projection with identifiers, totals, exceptions, and links to trace data. When multiple calls complete, return one result for each call id in the provider's required message structure; never concatenate them into one unlabeled blob.

What you wrote vs what the model sees: Your executor has an object. The provider expects a message or content block with a particular role and id relationship. A semantically correct result in the wrong envelope is still invisible to the model.

FAILURE MODES

Unattributed result

Symptom: The model says it has no result even though the backend returned 200.

Fix: Echo the exact tool-call id and provider-specific result role.

Full payload echo

Symptom: A 14 KB result pushes earlier instructions and evidence out of the useful context window.

Fix: Summarize for the model; retain the full payload in a trace store.

Mixed result types

Symptom: One call returns an object and another returns a string; the adapter serializes both as text and loses fields.

Fix: Use a stable result envelope and explicit content type.

PRODUCTION NOTE

A 14 KB customer profile result caused the model to ignore the first half of the conversation on the next turn. Returning a 420-token projection preserved the decision fields and kept the trace complete.

06

How do I prevent the model from hallucinating tool arguments?

WHEN THIS BITES YOU · The arguments pass JSON Schema, look reasonable in logs, and still identify a real person, account, or date that the user never supplied.

THE PATTERN

Constrain the space, then validate the meaning. Use enums for finite choices, optional fields with explicit defaults when omission is safe, and descriptions that state provenance rules. Validate syntax, schema, authorization, existence, and cross-field consistency before execution. Dates, identifiers, and nested objects deserve semantic validation because type correctness does not prove that the value exists or belongs to the current user.

What you wrote vs what the model sees: The model can produce a valid string that is not a valid identifier in your system. The execution layer must own that distinction.

FAILURE MODES

Plausible identifier

Symptom: A patient id matches the required pattern but belongs to nobody in the authorized tenant.

Fix: Resolve identifiers against an authorized index before execution; reject unknown values.

Date invention

Symptom: ‘Next Friday’ becomes a date in the server timezone rather than the user's timezone.

Fix: Normalize timezone and ask for clarification when a relative date is ambiguous.

Nested object drift

Symptom: The model places `currency` inside `amount` because both descriptions mention money.

Fix: Use shallow schemas where possible and validate cross-field shape before calling.

PRODUCTION NOTE

A medical scheduling assistant generated patient ids that matched the format and passed type validation. Existence and tenant checks caught the problem before the scheduling API saw the request.

A context window is not a memory model. It is a bounded serialization budget shared by instructions, history, calls, results, wrappers, and the next decision. Measure the conversation you actually send.

DIAGRAM 3 / TOOL CHAIN CONTEXT GROWTH
typical 8K limit128K limit still reachable12345678TOKENSTURNsystem · history · tool calls · tool resultsprune here: preserve checkpoint, externalize raw results
07

How do I chain tool calls across multiple turns without context window pollution?

WHEN THIS BITES YOU · The agent reaches turn eight, still has every raw tool response, and starts repeating earlier reasoning because the context budget is full.

THE PATTERN

Treat each turn as an auditable checkpoint. Keep the current objective, verified facts, pending actions, tool ids, failure state, and next allowed transitions. Summarize completed work into a state record; retain raw tool payloads in external trace storage. Prune only after result attribution is captured. In practice, measure the chain rather than choosing a universal depth: quality often degrades before the advertised context limit when large results and repeated instructions compete for attention.

What you wrote vs what the model sees: The context window counts messages, tool arguments, tool results, hidden wrappers, and repeated system instructions. ‘The model supports 128K’ is not a chain budget.

FAILURE MODES

Raw-result accumulation

Symptom: Every response is appended verbatim and later turns spend tokens re-reading old payloads.

Fix: Checkpoint verified facts and externalize raw results.

Broken attribution after pruning

Symptom: A result remains but its tool-use id was removed, so the provider rejects the next message.

Fix: Prune complete exchanges, never one half of a tool-use/tool-result pair.

Reasoning displacement

Symptom: The model's earlier plan is truncated before the tool results, causing a repeated or contradictory plan.

Fix: Persist the plan state separately and re-inject only the current checkpoint.

PRODUCTION NOTE

An agentic pipeline hit context limits on turn eight because nobody measured tool-result size. The repair was a checkpoint record plus a trace pointer, not a larger model.

08

What is different between OpenAI, Anthropic, and Gemini tool calling in practice?

WHEN THIS BITES YOU · The same adapter works for one provider, then silently drops a result or misreads a malformed call after a provider switch.

THE PATTERN

Keep a provider adapter at the boundary. Normalize calls into `{ id, name, arguments, status }`, normalize results into `{ id, status, content }`, and keep provider-specific serialization out of the agent loop. Test successful calls and malformed calls as separate contracts.

What you wrote vs what the model sees: The conceptual loop is shared. Message roles, content-part names, and malformed-call behavior are not.

FAILURE MODES

Role translation omission

Symptom: The executor returns a valid object under the wrong role and the provider treats it as user text.

Fix: Use contract tests for every role and content-part mapping.

No-result call

Symptom: A tool emits no body on success and the adapter omits the result message entirely.

Fix: Send an explicit empty success result with the call id.

Malformed JSON

Symptom: The model emits partial arguments and the executor tries to repair them by string concatenation.

Fix: Return a structured parse failure and let the loop decide whether to retry or ask.

PRODUCTION NOTE

The cross-provider bug that lasted longest was not a missing tool. It was an empty successful result that one adapter dropped and another encoded explicitly.

PROVIDER COMPARISON · THE THREE BREAKING DIFFERENCES

DifferenceOpenAIAnthropicGemini
Normal successful calltool call + tool message with tool_call_idtool_use block + tool_result block with same idfunctionCall part + functionResponse part with same name
Missing required argumentvalidate before execution; ask or retryinput schema failure; return a tool_result errorfunction call may be incomplete; validate before response
Malformed JSONadapter returns parse failuretool input must be parsed before tool_resultfunction args are a structured part; reject malformed content

INTERACTIVE / PROVIDER BEHAVIOR TESTER

OpenAI

tool call + tool message with tool_call_id

HANDLING · Adapter-specific serialization; application-level retry and authorization remain shared.

Anthropic

tool_use block + tool_result block with same id

HANDLING · Adapter-specific serialization; application-level retry and authorization remain shared.

Gemini

functionCall part + functionResponse part with same name

HANDLING · Adapter-specific serialization; application-level retry and authorization remain shared.

Share a thought

Comments appear immediately. Email is optional and never shown.

Markdown & code fences supported
AUTH VIA: PUBLIC FORM

No comments yet. Be the first to share a thought.

01 / SYSTEM VIEWInspect the architecture Follow a request through retrieval, validation, and audit.02 / EVIDENCEReview the case studies See the engineering decisions and measured outcomes.03 / OPEN CHANNELStart a conversation Talk through a production GenAI problem.