Tool errors: how an agent recovers and when it must stop

Tool errors: how an agent recovers and when it must stop

An error is a result, but it is not the requested data

A weight lookup can return 76 kg, report a missing record or fail because storage cannot be read. All three cases return something from the function, but only the first contains the requested data. When everything travels as text without an explicit status, the model has to infer what happened from the wording.

An expected error belongs to the domain: a date with no record, an unknown field or arguments that fail the schema, the accepted shape of a request. An unexpected exception is different: a read or write interrupted inside the application. Separating them lets the agent correct a request without treating every failure as a provider outage.

The harness must distinguish these cases before calling the model again. A sentence starting with “Error” is not a contract. Nor should user-supplied text be classified as a failure because it contains is_error: data remains data.

A contract that explains how to recover

Execution returns two pieces: the content the model will see and an explicit error flag controlled by the program. The following example is an illustrative internal payload, not a mandatory field in every API. The message stays in Spanish, as in the case study. Besides the stable not_found code, it describes the scope of the failure and suggests a recovery path.

{
  "kind": "tool_error",
  "version": 1,
  "is_error": true,
  "error": "not_found",
  "message": "No hay registro para esa fecha. No implica que no haya registros en otras fechas.",
  "recovery": "choose_alternative",
  "recoverable": true
}

correct_arguments allows correcting the arguments; choose_alternative allows another query when enough information is available; stop_turn requires stopping. These are labels in this application's contract. The model can propose an alternative, but the harness still validates every call and enforces the round limit.

Execution result
  1. Invalid arguments → correct
  2. Known absence → choose an alternative
  3. Uncertain write → stop

“No record on this date” does not mean “no records exist.” That precision prevents a local absence from becoming a global claim. The error flag distinguishes the status; the message helps the model reason about its next step.

The same failure across three providers

Adapters, the pieces that translate the local contract into each API, preserve the identifier connecting a call with its result. The wrapper changes; the decision about whether execution failed does not.

  • Anthropic Messages: the tool_result block carries is_error: true and safe failure content.
  • Google Interactions: function_result also accepts is_error. This comparison specifically concerns Interactions.
  • OpenAI Responses: function_call_output.output contains the serialized error JSON. We do not add an unsupported top-level is_error field.

The flag survives replaying an already recorded result. It is not inferred again from its text. A result recovered from history therefore keeps the meaning it had when it was executed.

Where the try/catch belongs

Validating before execution prevents effects from invalid arguments. Then the exception boundary must wrap the actual handler execution, the function that performs the operation. Catching only the provider call misses local failures.

validate the tool name and arguments
if validation fails:
    return a marked error with recovery = correct_arguments

try:
    result = execute the handler
    return result with its explicit success or error status
catch an unexpected exception:
    if the tool only reads:
        return a safe error that allows an alternative
    else:
        stop the turn locally
        say that completion cannot be confirmed
        do not repeat the action automatically

This pseudocode summarizes the decision, not the entire loop. Expected failures follow the normal contract; exceptions become safe messages. The original exception may help local diagnosis, but it must not appear as a success confirmation or expose internal details to the model. A fatal local failure is not labeled “Provider error” either.

A real conversation with Coach

Exploratory test on October 4, 2026, in Gymnasia's web app, version 1.51.0, using OpenAI gpt-6-luna. Dates and weights were fictitious. Reading measurements retrieves a saved weight for one day without changing records.

User, translated: look up the weight for 2020-01-02. If it is missing, look up 2020-01-01. Do not create or change data.

Coach, translated: The record for 2020-01-02 is missing. 2020-01-01 does have a saved weight: 76 kg.

The user's request is shortened here. Actual requests showed the whole sequence: read_measurement for day 2, a marked error with choose_alternative, a read for day 1 and the final answer. We checked the content sent to the provider, not just the visible text.

What we did not reproduce: in a separate isolated test, the same model received the old unmarked text format and understood that a field was missing. We did not observe it confusing that error with valid data.

For a missing date, the old format produced “You have no measurements recorded for 2020-01-02.” The new contract produced “There is no measurement record for 2020-01-02. This does not tell us whether you have records on other dates.” These are translations of the Spanish responses. The new message was also more precise: this comparison does not isolate the flag's effect or prove that every model improves. The contract's value is making the program's decision explicit, even when a model already interprets plain text correctly.

An uncertain write changes the decision

A read can be repeated without creating a record. An interrupted write may have completed before its confirmation failed. “It failed” does not always mean “it did nothing.” Retrying by default can duplicate an action.

In the same session, we induced a controlled persistence failure while saving a fictitious weight. The model was real; the storage failure was injected for QA. Coach displayed this message, translated from Spanish:

Gymnasia cannot confirm whether the action completed. To avoid duplicating it, it has not repeated it. Check your data before requesting it again.

There was one model request, no additional round and no automatic retry. The 2020-01-03 record was absent after reloading. We removed the injection, then a write for 2020-01-04 with 77.4 kg succeeded and could be read after another reload. This verifies storage recovery in that scenario; it does not prove that every failed write leaves data unchanged.

The conservative message remains necessary when the effect is uncertain. If an action must support safe retries, it needs an additional strategy, such as a stable identity and a durable result record. Changing the error message cannot provide that guarantee.

Tests that protect this contract

Deterministic tests check the contract without depending on how a model responds that day. Real QA adds behavioral evidence, but does not replace those boundaries.

  • Contract: invalid arguments return a marked error without executing the handler; user text resembling an error remains a valid result.
  • Fake-provider integration: OpenAI, Anthropic and Google receive the correct format; a corrected call persists one measurement.
  • Write regression: an exception containing network timeout does not trigger transport retries or confirm success.
  • Persistence: replay preserves the result status; reloading distinguishes the confirmed record from the one that was not saved.

The delivery passed 933 Vitest tests, 11 development-store tests and 3 mobile-boundary tests, plus type-checking and fake-provider E2E tests. The exploratory test described here ran on the web with one model. It is neither a native-device test nor a statistical evaluation of all three providers.

Build an agent, step by step

This article belongs to the series on building an agent and its harness through a gym application. Previous installments explain how to declare tools, validate calls and close the loop; this one adds a decision that must live in the program: when a failure allows continuing and when it requires stopping.

Visit the series index and the Gymnasia implementation used as the case study.

Continue reading

Last posts -->

Have you seen these projects?

Gymnasia

Gymnasia Gymnasia
Expo
React Native
TypeScript
OpenAI
Anthropic

Fitness app with two agents that run entirely on the device, with no backend, so the user's data never leaves the phone. A BYOK conversational coach with local tools and a remote system prompt with offline fallback, plus a food estimator that pulls macronutrients out of a photo of the plate, with barcode scanning against OpenFoodFacts.

LangGraph Deep Researcher

LangGraph Deep Researcher LangGraph Deep Researcher
Python
LangGraph
FastAPI
React
TypeScript
Docker

Multi-agent research system built with LangGraph. A supervisor breaks your question down into topics and launches search sub-agents in parallel; each one compresses its findings before handing them to a writer agent that produces the final sourced markdown report. Live streaming over WebSockets, a configurable model per role and bring-your-own API keys that are never persisted server-side.

Tau

Tau Tau
Python
LangChain

Multi-agent tutoring system for secondary school students, with one agent per subject and course material written and validated by a team of teachers. It was used with real students at a private school in Spain and at a secondary school in Colombia.

View all projects -->
>_ Available for projects

Do you have an AI project?

Let's talk.

maximofn@gmail.com

Machine Learning and AI specialist. I develop solutions with generative AI, intelligent agents and custom models.

Do you want to watch any talk?

Last talks -->

Do you want to improve with these tips?

Last tips -->

Use this locally

Hugging Face spaces allow us to run models with very simple demos, but what if the demo breaks? Or if the user deletes it? That's why I've created docker containers with some interesting spaces, to be able to use them locally, whatever happens. In fact, if you click on any project view button, it may take you to a space that doesn't work.

Flow edit

Flow edit Flow edit

FLUX.1-RealismLora

FLUX.1-RealismLora FLUX.1-RealismLora
View all containers -->
>_ Available for projects

Do you have an AI project?

Let's talk.

maximofn@gmail.com

Machine Learning and AI specialist. I develop solutions with generative AI, intelligent agents and custom models.

Do you want to train your model with these datasets?

short-jokes-dataset

HuggingFace

Dataset with jokes in English

Use: Fine-tuning text generation models for humor

231K rows 2 columns 45 MB
View on HuggingFace →

opus100

HuggingFace

Dataset with translations from English to Spanish

Use: Training English-Spanish translation models

1M rows 2 columns 210 MB
View on HuggingFace →

netflix_titles

HuggingFace

Dataset with Netflix movies and series

Use: Netflix catalog analysis and recommendation systems

8.8K rows 12 columns 3.5 MB
View on HuggingFace →
View more datasets -->