Validating tool calls: declaring a schema is not enough

Validating tool calls: declaring a schema is not enough

The advertised shape and the received data

A tool is a function the model can request. Its input schema describes the arguments: which fields it needs and their types. If a function records a body measurement, declaring that weight is a number does not automatically make that function check every incoming value. Your program needs to perform that check.

Provider constraints can help generate correctly shaped data. The harness still needs to enforce its contract before acting, especially when it supports different providers. A declaration describes what is accepted; a validator compares that declaration with a specific call.

Three questions before execution

These are three different checks. First, did the provider finish the response? Next, can the arguments be read as a JSON object? Finally, does that object match the schema? A complete response can contain invalid arguments; readable JSON can too.

  1. Provider events reach the parser, which reconstructs the pending call.
  2. After the parser result, the harness checks completeness and reads the arguments.
  3. The validator checks fields and types against that tool’s schema.
  4. Valid arguments enter the handler; invalid arguments return an error without running it.
  5. The result goes back to the model, which can correct the call or answer.

One place to check the arguments

The dispatcher, the function that selects executable code using the tool name, is a good place for this check. Reads and writes pass through the same point. If each handler invents its own checks for types and required fields, adding a new tool can leave a gap nobody notices.

Define each tool once and use that definition both to advertise it to the provider and to check incoming arguments. This pattern does not need a server: Gymnasia performs the check locally. Notice that the handler only runs after the validator accepts the call:

when a complete tool call arrives:
    find the tool definition
    parse the arguments as a JSON object
    errors = validate(arguments, definition.input_schema)
    if errors exist:
        return a tool result with the failing fields and reasons
    return run_handler(arguments)

append the tool result to the conversation
ask the model for the next turn, within the round limit

OpenAI, Anthropic and Google Interactions have different formats for requesting tools and receiving results. Once the call is reconstructed, Coach tools reach the same executor. The OpenAI-compatible provider uses that executor too. There is no need for four versions of the input rules.

An error that lets the turn continue

Gymnasia records weight using a tool called write_measurement. This example fails: data.weight_kg contains text, even though it looks like a number. The date and the rest of the object can be correct:

{"date":"2024-04-11","data":{"weight_kg":"75,5"}}

The result says that the tool did not run and identifies data.weight_kg as a field that must be a number. The harness sends that result back to the model with the matching call identity. In the next round, the model can send:

{"date":"2024-04-11","data":{"weight_kg":75.5}}

The harness does not convert the text or execute again on its own. The model proposes a correction and the new call is checked again. Valid arguments run; another invalid call returns another error. This happens within the existing round limit. A known argument error allows continuation; a write with an uncertain outcome needs different handling and must not be repeated as though nothing happened.

Reproducible cases in Gymnasia

These cases come from executor tests and reproducible fixtures, provider responses prepared for tests. They are not transcripts of private conversations or a measurement of how frequently a model makes mistakes.

Problematic inputPrevious behaviorWith central validation
A measurement sent as JSON inside a string, with a number written as text.The measurement contract accepted the legacy format and normalized the value.The model must send an object with numeric values, as the schema declares. The existing measurement remains intact while the call is corrected.
A meal written as “ desayuno ”.The handler normalized whitespace and capitalization.Coach receives the allowed enum values and can resend “Desayuno”.
Unreadable JSON in OpenAI for a tool with no required fields.The parser replaced it with an empty object, which could reach the handler.An argument error is returned without running the tool. A legitimate empty object still passes if the schema allows it.
An optional numeric field sent as null.The validator skipped the value.The value is rejected when its type does not allow null, including nested objects or arrays.

Shape alone cannot decide

A schema checks shape, not the entire intent or product rules. A date can be a string and describe an impossible day. An identifier can have the correct type and point to a nonexistent exercise. Dates, references and exercise sets still require domain validation. Passing the schema does not grant permission for an action that needs confirmation.

Gymnasia’s own validator covers the types and constraints declared in its tool catalogue, not the entire JSON Schema specification. It adds no new dependency to the mobile bundle and does not transform arguments. Extra fields are allowed unless the schema prohibits them with additionalProperties set to false. An optional field may be omitted; that does not turn null into a number. The JSON Schema references for objects and null document these rules.

If your schemas start needing unions, references or new constraints, extend the validator and its contract or adopt a library that supports them. Do not advertise rules the program never checks. Tools that declare a string containing JSON also need to parse and validate its contents.

The contract can also change during transport. In a real test with OpenAI Responses, asking for weight alone made the provider require every measurement and the model fill the others with 0.01. In a temporary trial with strict:false, the schema kept those fields optional and the model sent only weight. Local validation remained active. The OpenAI documentation explains the strict default.

Testing the contract and its effects

Unit tests should cover valid arguments without modification, missing required fields, wrong types, enums, limits and nested objects. Gymnasia also compares its validator with Ajv in tests for the schema subset it uses. Generating arbitrary arguments, or fuzzing, looks for inputs that cause exceptions or disagreements.

The execution property is observable: an invalid call does not load data, resolve catalogues, create IDs or produce a handler write. Integration tests replay an invalid measurement, the error result, a corrected call and a single write. The E2E runs that sequence from the chat with simulated OpenAI, Anthropic and Google providers, checks storage and reloads the app. This demonstrates local behavior for those responses, not that a real model always corrects itself.

See the implementation and tests in the Gymnasia PR. Central validation has been merged and published on the web.

QA with a real OpenAI model also found that transport changed the optional-field contract. The additional fix is now published on the web. Final QA with the published code, without temporary network intervention, recorded only 75.9 kg and preserved the other measurements after reloading.

A series about building agents

This article belongs to a series about building an agent and its harness through the practical example of a gym application. Each article develops a decision you can apply to your own agent.

Read the series index

Continue reading

Last posts -->

Have you seen these projects?

Gymnasia

Gymnasia Gymnasia
Expo
React Native
TypeScript
OpenAI
Anthropic

Fitness app with two agents that run entirely on the device, with no backend, so the user's data never leaves the phone. A BYOK conversational coach with local tools and a remote system prompt with offline fallback, plus a food estimator that pulls macronutrients out of a photo of the plate, with barcode scanning against OpenFoodFacts.

LangGraph Deep Researcher

LangGraph Deep Researcher LangGraph Deep Researcher
Python
LangGraph
FastAPI
React
TypeScript
Docker

Multi-agent research system built with LangGraph. A supervisor breaks your question down into topics and launches search sub-agents in parallel; each one compresses its findings before handing them to a writer agent that produces the final sourced markdown report. Live streaming over WebSockets, a configurable model per role and bring-your-own API keys that are never persisted server-side.

Tau

Tau Tau
Python
LangChain

Multi-agent tutoring system for secondary school students, with one agent per subject and course material written and validated by a team of teachers. It was used with real students at a private school in Spain and at a secondary school in Colombia.

View all projects -->
>_ Available for projects

Do you have an AI project?

Let's talk.

maximofn@gmail.com

Machine Learning and AI specialist. I develop solutions with generative AI, intelligent agents and custom models.

Do you want to watch any talk?

Last talks -->

Do you want to improve with these tips?

Last tips -->

Use this locally

Hugging Face spaces allow us to run models with very simple demos, but what if the demo breaks? Or if the user deletes it? That's why I've created docker containers with some interesting spaces, to be able to use them locally, whatever happens. In fact, if you click on any project view button, it may take you to a space that doesn't work.

Flow edit

Flow edit Flow edit

FLUX.1-RealismLora

FLUX.1-RealismLora FLUX.1-RealismLora
View all containers -->
>_ Available for projects

Do you have an AI project?

Let's talk.

maximofn@gmail.com

Machine Learning and AI specialist. I develop solutions with generative AI, intelligent agents and custom models.

Do you want to train your model with these datasets?

short-jokes-dataset

HuggingFace

Dataset with jokes in English

Use: Fine-tuning text generation models for humor

231K rows 2 columns 45 MB
View on HuggingFace →

opus100

HuggingFace

Dataset with translations from English to Spanish

Use: Training English-Spanish translation models

1M rows 2 columns 210 MB
View on HuggingFace →

netflix_titles

HuggingFace

Dataset with Netflix movies and series

Use: Netflix catalog analysis and recommendation systems

8.8K rows 12 columns 3.5 MB
View on HuggingFace →
View more datasets -->