← Blog

Typed tool use with Claude and Zod: the agent loop, and the usage you forget to bill

One Zod schema per tool gives the API schema, runtime checks, and types. Plus the agent loop detail billing needs: usage summed across every iteration.

Tool use turns one user request into several model calls: the model asks for a tool, you run it, you send the result back, and it may ask for another before it answers. Two things go wrong in that loop more often than anything else. The model sends a tool input that does not match what your code expects, and your billing only counts the last call — so an agent that took four round trips gets metered as one. Both are fixable with a small amount of structure.

One Zod schema, used three ways

Declare each tool's input once, with Zod, and let that single schema do three jobs: generate the JSON Schema the API needs, validate what the model actually sends at runtime, and type the argument your execute function receives.

export interface ToolDefinition<S extends z.ZodType = z.ZodType> {
  name: string;
  /** Be prescriptive about WHEN to call it, not just what it does. */
  description: string;
  schema: S;
  execute(input: z.infer<S>): string | Promise<string>;
}

function toApiTool(tool: ToolDefinition): Anthropic.Tool {
  const schema = z.toJSONSchema(tool.schema) as Record<string, unknown>;
  delete schema.$schema; // metadata key, not part of the tool's input schema
  return { name: tool.name, description: tool.description, input_schema: schema };
}

Without this, the JSON Schema you send, the validation you (maybe) do, and your TypeScript types are three hand-maintained copies that drift apart. With it, changing a tool's input is one edit. The description comment matters too: tell the model when to call a tool, not just what it does, or it will call it at the wrong moments.

Errors the model can read and recover from

The model's tool input is untrusted data that is usually right. When it is wrong, do not throw — return the problem as a tool result flagged is_error, in words the model can act on:

const parsed = tool.schema.safeParse(block.input);
if (!parsed.success) {
  // Tell the model what was wrong so it can correct itself on the next call.
  return {
    type: "tool_result",
    tool_use_id: block.id,
    content: `Invalid input: ${z.prettifyError(parsed.error)}`,
    is_error: true,
  };
}

The same pattern covers an unknown tool name and a tool that throws while running. In each case the loop keeps going and the model usually fixes its call on the next turn, instead of the whole request dying on one malformed argument.

The loop: sum usage on every iteration

The kit runs the loop by hand rather than through the SDK's beta tool runner, specifically to keep three control points: validation with model-readable errors, a hard iteration cap, and per-iteration usage accounting. That last one is the one that hits your margin:

for (let iteration = 1; iteration <= maxIterations; iteration++) {
  const response = await anthropic().messages.create({ model, max_tokens, tools: apiTools, messages });
  addUsage(usage, usageFromMessage(response.usage)); // every call, not just the last

  if (response.stop_reason !== "tool_use") {
    return { message: response, text: messageText(response), usage, iterations: iteration };
  }
  messages.push({ role: "assistant", content: response.content });
  // ...execute each tool_use block, push the results as the next user turn
}
throw new Error(`model was still requesting tools after ${maxIterations} iterations`);

Every iteration resends the growing conversation — the original request plus every tool call and result so far — so input tokens climb with each round trip. A four-iteration agent can easily cost several times what its final message alone suggests. Meter only that final response and you undercount the expensive requests specifically, which is exactly the error that makes a pricing tier look profitable when it is not. Sum usage across the loop and record the total per request, the same way the per-user metering does for plain chat.

Things that bite

  • Cap the iterations. A model stuck re-calling a tool will loop until something stops it. The cap here defaults to 10 and throws past it — a bounded failure instead of an unbounded bill.
  • Handle pause_turn. When a server-side tool hits its own limit, the response stops with pause_turn; resend the assistant content to let it resume rather than treating it as a final answer.
  • Return strings the model can use. Tool results go back into the conversation as input tokens. A tool that returns a huge JSON blob makes every later iteration more expensive; return what the model needs, not everything you have.
  • Structured output and tools are different jobs. If you just need a typed object back, use structured output with Zod — no loop, one call.

Shipwright — the Next.js and Claude starter kit this blog documents — ships this loop as runWithTools with defineTool, so typed tools, recoverable errors, the iteration cap, and loop-wide usage accounting come built in. Try a metered request in the live demo.

Shipwright is a Next.js 16 + Claude starter kit that ships these patterns already done.

Try the live demo →