Parallel tools: the provider proposes, your app decides what can overlap

Parallel tools: the provider proposes, your app decides what can overlap

The model can already return a batch

OpenAI, Anthropic and Google can return several tool calls in one turn. These are requests: the provider does not execute your application code. If you await each call inside a loop, every call waits for the previous one. Three reads that need the network can pay three successive waits even when their arguments are already complete.

A model batch does not prove independence. The app needs its own rules about which tools read, which modify data and which require confirmation for an external effect. That decision belongs to the code responsible for the data.

Writes are barriers

Gymnasia already declares every tool effect in a shared catalog. Eight tools read: discover personal fields, read their descriptions or values, query measurements, meals and routines, and search foods or exercises. Consecutive reads can overlap. Saving personal data, recording a measurement, adding food and creating a routine are local writes; proposing an issue is an external write.

The scheduler allows up to four active reads. Every write waits for earlier calls to finish and runs alone. An unknown tool is also a barrier. The sequence read, read, write, read drains the first two, waits for the write and then starts the final read. It never moves a read across a write.

Completion order is not result order

Call identities and occurrences are assigned in the original order before execution starts. The result of call zero returns to position zero even if it finishes last. Adapters preserve provider IDs, history, thought signatures and the journal that prevents repeated effects during a retry.

A recoverable error marks only its own result and sibling reads complete. A fatal error stops new reads, waits for those already running and ends the turn before a later write. If several reads fail fatally, the first in original position is reported. A write with an uncertain outcome is still never retried.

OpenAI Responses, Anthropic, Google Interactions and the OpenAI-compatible adapter share this scheduler. We keep the provider’s ability to return several calls: disabling it would require more model rounds to obtain the reads and would not replace client barriers.

A phone can overlap waits

This change creates no threads. JavaScript executes synchronous work on its thread and can start several asynchronous operations while waiting for their results. Searching foods in memory does not distribute computation across cores; waiting for exercise catalog shards creates an opportunity to overlap work. The React Native performance documentation explains JavaScript thread constraints.

Four is a conservative active-tool limit awaiting native profiling. It is not a global HTTP request limit: one tool can download several pages. It does not establish a particular memory or battery cost. The benefit depends on caches, the phone and its connection.

The comparison runs through Coach

We exported the same Expo web app before and after the change. Coach uses its actual executor and catalog. A fake provider returns three searches: press, curl and sentadilla. Each search shard download waits for a controlled 200 ms. Timing starts when the tool-call turn is delivered and ends when the app requests continuation with all three results. It excludes real model generation time.

After an unmeasured warmup pair, we measured 30 cold/warm pairs in Chromium 145.0.7632.6, viewport 390 × 844. Cold clears only search shards: the manifest and exercise pages stay cached. Warm keeps everything. The median averages the middle two values; p95 uses the nearest upper rank.

VariantCold: median / p95Warm: median / p95Concurrent cold downloads
Sequential696.00 / 859.86 ms57.94 / 160.83 ms1
Parallel263.59 / 278.73 ms46.72 / 62.91 ms3

The cold median drops by about 62% in this scenario. The test also checks nonempty exercise results, ordered IDs, a visible answer and no page exceptions. Warm differences include local work and execution noise, so they do not support a general performance promise. These are controlled web measurements, not APK timings or the complete response time of a model.

Measure before promising a speedup

Deterministic tests check the limit, barriers and order with promises finishing out of order. A fake clock turns waits of 100, 200 and 300 ms into 300 ms concurrently versus 600 ms sequentially. All four dialects cover recoverable and fatal errors, plus two actual writes to the same date followed by a read. One hundred generated sequences search for limit or barrier violations.

The previous reference is compiled as Staging 1.50.4. The implementation was merged and web 1.51.0 passed QA with real OpenAI calls: three searches in one turn kept their order; a read error did not cancel the other reads, and the last saved test value survived a reload. The Production 1.51.0 APK is compiled, verified and published. The native comparison remains pending. On a phone, alternate both APKs on the same device with equivalent data and queries. Separate cold and warm caches, record median, p95 and errors, and use Android traces for CPU, memory and frames. Battery needs repeated long sessions with equivalent brightness and connectivity.

The reproducible method and samples are documented alongside the code. The app receives no new performance telemetry, and no personal session data is exported to obtain these numbers.

Reviewable implementation and tests in the Gymnasia PR. Reproducible method, samples and comparison limits.

Continue with the series

Index of the series about building an agent. This article completes the loop discussion: receive a batch, check its effects and return every result with its identity.

Continue reading

Last posts -->

Have you seen these projects?

Gymnasia

Gymnasia Gymnasia
Expo
React Native
TypeScript
OpenAI
Anthropic

Fitness app with two agents that run entirely on the device, with no backend, so the user's data never leaves the phone. A BYOK conversational coach with local tools and a remote system prompt with offline fallback, plus a food estimator that pulls macronutrients out of a photo of the plate, with barcode scanning against OpenFoodFacts.

LangGraph Deep Researcher

LangGraph Deep Researcher LangGraph Deep Researcher
Python
LangGraph
FastAPI
React
TypeScript
Docker

Multi-agent research system built with LangGraph. A supervisor breaks your question down into topics and launches search sub-agents in parallel; each one compresses its findings before handing them to a writer agent that produces the final sourced markdown report. Live streaming over WebSockets, a configurable model per role and bring-your-own API keys that are never persisted server-side.

Tau

Tau Tau
Python
LangChain

Multi-agent tutoring system for secondary school students, with one agent per subject and course material written and validated by a team of teachers. It was used with real students at a private school in Spain and at a secondary school in Colombia.

View all projects -->
>_ Available for projects

Do you have an AI project?

Let's talk.

maximofn@gmail.com

Machine Learning and AI specialist. I develop solutions with generative AI, intelligent agents and custom models.

Do you want to watch any talk?

Last talks -->

Do you want to improve with these tips?

Last tips -->

Use this locally

Hugging Face spaces allow us to run models with very simple demos, but what if the demo breaks? Or if the user deletes it? That's why I've created docker containers with some interesting spaces, to be able to use them locally, whatever happens. In fact, if you click on any project view button, it may take you to a space that doesn't work.

Flow edit

Flow edit Flow edit

FLUX.1-RealismLora

FLUX.1-RealismLora FLUX.1-RealismLora
View all containers -->
>_ Available for projects

Do you have an AI project?

Let's talk.

maximofn@gmail.com

Machine Learning and AI specialist. I develop solutions with generative AI, intelligent agents and custom models.

Do you want to train your model with these datasets?

short-jokes-dataset

HuggingFace

Dataset with jokes in English

Use: Fine-tuning text generation models for humor

231K rows 2 columns 45 MB
View on HuggingFace →

opus100

HuggingFace

Dataset with translations from English to Spanish

Use: Training English-Spanish translation models

1M rows 2 columns 210 MB
View on HuggingFace →

netflix_titles

HuggingFace

Dataset with Netflix movies and series

Use: Netflix catalog analysis and recommendation systems

8.8K rows 12 columns 3.5 MB
View on HuggingFace →
View more datasets -->