The model can already return a batch
OpenAI, Anthropic and Google can return several tool calls in one turn. These are requests: the provider does not execute your application code. If you await each call inside a loop, every call waits for the previous one. Three reads that need the network can pay three successive waits even when their arguments are already complete.
A model batch does not prove independence. The app needs its own rules about which tools read, which modify data and which require confirmation for an external effect. That decision belongs to the code responsible for the data.
Writes are barriers
Gymnasia already declares every tool effect in a shared catalog. Eight tools read: discover personal fields, read their descriptions or values, query measurements, meals and routines, and search foods or exercises. Consecutive reads can overlap. Saving personal data, recording a measurement, adding food and creating a routine are local writes; proposing an issue is an external write.
The scheduler allows up to four active reads. Every write waits for earlier calls to finish and runs alone. An unknown tool is also a barrier. The sequence read, read, write, read drains the first two, waits for the write and then starts the final read. It never moves a read across a write.
Completion order is not result order
Call identities and occurrences are assigned in the original order before execution starts. The result of call zero returns to position zero even if it finishes last. Adapters preserve provider IDs, history, thought signatures and the journal that prevents repeated effects during a retry.
A recoverable error marks only its own result and sibling reads complete. A fatal error stops new reads, waits for those already running and ends the turn before a later write. If several reads fail fatally, the first in original position is reported. A write with an uncertain outcome is still never retried.
OpenAI Responses, Anthropic, Google Interactions and the OpenAI-compatible adapter share this scheduler. We keep the provider’s ability to return several calls: disabling it would require more model rounds to obtain the reads and would not replace client barriers.
A phone can overlap waits
This change creates no threads. JavaScript executes synchronous work on its thread and can start several asynchronous operations while waiting for their results. Searching foods in memory does not distribute computation across cores; waiting for exercise catalog shards creates an opportunity to overlap work. The React Native performance documentation explains JavaScript thread constraints.
Four is a conservative active-tool limit awaiting native profiling. It is not a global HTTP request limit: one tool can download several pages. It does not establish a particular memory or battery cost. The benefit depends on caches, the phone and its connection.
The comparison runs through Coach
We exported the same Expo web app before and after the change. Coach uses its actual executor and catalog. A fake provider returns three searches: press, curl and sentadilla. Each search shard download waits for a controlled 200 ms. Timing starts when the tool-call turn is delivered and ends when the app requests continuation with all three results. It excludes real model generation time.
After an unmeasured warmup pair, we measured 30 cold/warm pairs in Chromium 145.0.7632.6, viewport 390 × 844. Cold clears only search shards: the manifest and exercise pages stay cached. Warm keeps everything. The median averages the middle two values; p95 uses the nearest upper rank.
| Variant | Cold: median / p95 | Warm: median / p95 | Concurrent cold downloads |
|---|---|---|---|
| Sequential | 696.00 / 859.86 ms | 57.94 / 160.83 ms | 1 |
| Parallel | 263.59 / 278.73 ms | 46.72 / 62.91 ms | 3 |
The cold median drops by about 62% in this scenario. The test also checks nonempty exercise results, ordered IDs, a visible answer and no page exceptions. Warm differences include local work and execution noise, so they do not support a general performance promise. These are controlled web measurements, not APK timings or the complete response time of a model.
Measure before promising a speedup
Deterministic tests check the limit, barriers and order with promises finishing out of order. A fake clock turns waits of 100, 200 and 300 ms into 300 ms concurrently versus 600 ms sequentially. All four dialects cover recoverable and fatal errors, plus two actual writes to the same date followed by a read. One hundred generated sequences search for limit or barrier violations.
The previous reference is compiled as Staging 1.50.4. The implementation was merged and web 1.51.0 passed QA with real OpenAI calls: three searches in one turn kept their order; a read error did not cancel the other reads, and the last saved test value survived a reload. The Production 1.51.0 APK is compiled, verified and published. The native comparison remains pending. On a phone, alternate both APKs on the same device with equivalent data and queries. Separate cold and warm caches, record median, p95 and errors, and use Android traces for CPU, memory and frames. Battery needs repeated long sessions with equivalent brightness and connectivity.
The reproducible method and samples are documented alongside the code. The app receives no new performance telemetry, and no personal session data is exported to obtain these numbers.
Reviewable implementation and tests in the Gymnasia PR. Reproducible method, samples and comparison limits.
Continue with the series
Index of the series about building an agent. This article completes the loop discussion: receive a batch, check its effects and return every result with its identity.