An error is a result, but it is not the requested data
A weight lookup can return 76 kg, report a missing record or fail because storage cannot be read. All three cases return something from the function, but only the first contains the requested data. When everything travels as text without an explicit status, the model has to infer what happened from the wording.
An expected error belongs to the domain: a date with no record, an unknown field or arguments that fail the schema, the accepted shape of a request. An unexpected exception is different: a read or write interrupted inside the application. Separating them lets the agent correct a request without treating every failure as a provider outage.
The harness must distinguish these cases before calling the model again. A sentence starting with “Error” is not a contract. Nor should user-supplied text be classified as a failure because it contains is_error: data remains data.
A contract that explains how to recover
Execution returns two pieces: the content the model will see and an explicit error flag controlled by the program. The following example is an illustrative internal payload, not a mandatory field in every API. The message stays in Spanish, as in the case study. Besides the stable not_found code, it describes the scope of the failure and suggests a recovery path.
{
"kind": "tool_error",
"version": 1,
"is_error": true,
"error": "not_found",
"message": "No hay registro para esa fecha. No implica que no haya registros en otras fechas.",
"recovery": "choose_alternative",
"recoverable": true
}correct_arguments allows correcting the arguments; choose_alternative allows another query when enough information is available; stop_turn requires stopping. These are labels in this application's contract. The model can propose an alternative, but the harness still validates every call and enforces the round limit.
- Invalid arguments → correct
- Known absence → choose an alternative
- Uncertain write → stop
“No record on this date” does not mean “no records exist.” That precision prevents a local absence from becoming a global claim. The error flag distinguishes the status; the message helps the model reason about its next step.
The same failure across three providers
Adapters, the pieces that translate the local contract into each API, preserve the identifier connecting a call with its result. The wrapper changes; the decision about whether execution failed does not.
- Anthropic Messages: the
tool_resultblock carriesis_error: trueand safe failure content. - Google Interactions:
function_resultalso acceptsis_error. This comparison specifically concerns Interactions. - OpenAI Responses:
function_call_output.outputcontains the serialized error JSON. We do not add an unsupported top-levelis_errorfield.
The flag survives replaying an already recorded result. It is not inferred again from its text. A result recovered from history therefore keeps the meaning it had when it was executed.
Where the try/catch belongs
Validating before execution prevents effects from invalid arguments. Then the exception boundary must wrap the actual handler execution, the function that performs the operation. Catching only the provider call misses local failures.
validate the tool name and arguments
if validation fails:
return a marked error with recovery = correct_arguments
try:
result = execute the handler
return result with its explicit success or error status
catch an unexpected exception:
if the tool only reads:
return a safe error that allows an alternative
else:
stop the turn locally
say that completion cannot be confirmed
do not repeat the action automaticallyThis pseudocode summarizes the decision, not the entire loop. Expected failures follow the normal contract; exceptions become safe messages. The original exception may help local diagnosis, but it must not appear as a success confirmation or expose internal details to the model. A fatal local failure is not labeled “Provider error” either.
A real conversation with Coach
Exploratory test on October 4, 2026, in Gymnasia's web app, version 1.51.0, using OpenAI gpt-6-luna. Dates and weights were fictitious. Reading measurements retrieves a saved weight for one day without changing records.
User, translated: look up the weight for 2020-01-02. If it is missing, look up 2020-01-01. Do not create or change data.
Coach, translated: The record for 2020-01-02 is missing. 2020-01-01 does have a saved weight: 76 kg.
The user's request is shortened here. Actual requests showed the whole sequence: read_measurement for day 2, a marked error with choose_alternative, a read for day 1 and the final answer. We checked the content sent to the provider, not just the visible text.
What we did not reproduce: in a separate isolated test, the same model received the old unmarked text format and understood that a field was missing. We did not observe it confusing that error with valid data.
For a missing date, the old format produced “You have no measurements recorded for 2020-01-02.” The new contract produced “There is no measurement record for 2020-01-02. This does not tell us whether you have records on other dates.” These are translations of the Spanish responses. The new message was also more precise: this comparison does not isolate the flag's effect or prove that every model improves. The contract's value is making the program's decision explicit, even when a model already interprets plain text correctly.
An uncertain write changes the decision
A read can be repeated without creating a record. An interrupted write may have completed before its confirmation failed. “It failed” does not always mean “it did nothing.” Retrying by default can duplicate an action.
In the same session, we induced a controlled persistence failure while saving a fictitious weight. The model was real; the storage failure was injected for QA. Coach displayed this message, translated from Spanish:
Gymnasia cannot confirm whether the action completed. To avoid duplicating it, it has not repeated it. Check your data before requesting it again.
There was one model request, no additional round and no automatic retry. The 2020-01-03 record was absent after reloading. We removed the injection, then a write for 2020-01-04 with 77.4 kg succeeded and could be read after another reload. This verifies storage recovery in that scenario; it does not prove that every failed write leaves data unchanged.
The conservative message remains necessary when the effect is uncertain. If an action must support safe retries, it needs an additional strategy, such as a stable identity and a durable result record. Changing the error message cannot provide that guarantee.
Tests that protect this contract
Deterministic tests check the contract without depending on how a model responds that day. Real QA adds behavioral evidence, but does not replace those boundaries.
- Contract: invalid arguments return a marked error without executing the handler; user text resembling an error remains a valid result.
- Fake-provider integration: OpenAI, Anthropic and Google receive the correct format; a corrected call persists one measurement.
- Write regression: an exception containing
network timeoutdoes not trigger transport retries or confirm success. - Persistence: replay preserves the result status; reloading distinguishes the confirmed record from the one that was not saved.
The delivery passed 933 Vitest tests, 11 development-store tests and 3 mobile-boundary tests, plus type-checking and fake-provider E2E tests. The exploratory test described here ran on the web with one model. It is neither a native-device test nor a statistical evaluation of all three providers.
Build an agent, step by step
This article belongs to the series on building an agent and its harness through a gym application. Previous installments explain how to declare tools, validate calls and close the loop; this one adds a decision that must live in the program: when a failure allows continuing and when it requires stopping.
Visit the series index and the Gymnasia implementation used as the case study.