iv — Reliability
How runs survive failure.
The orchestration layer is engineered for reliability: agents keep working when models return bad output, upstream systems slow down, or a machine fails mid-run.
- i
Every task carries a policy - timeout, retry, budget, idempotency - and runs in a container drawn from a warm pool, so there are no cold-start delays.
- ii
There are no checkpoints to fall out of sync. Executor state is a fold of the run's event log, so a killed executor rebuilds every run and finishes it.
- iii
Structured outputs are schema-validated at every boundary; malformed responses are repaired or fail typed before they touch your data.
- iv
Large data never travels inline. Files and tables move as handles, so a run costs the same whether it processes a hundred rows or fifty thousand.
- v
No agent framework is baked into the executor. The harness lives inside the runner image, and swapping it touches zero engine code.
Failures are surfaced immediately, never silent.
Model output is untrusted input.
Nothing that comes back from a model is trusted by default. Every output is checked at the boundary before it is used.
- Task bodies speak JSON-RPC over stdio from a warm container pool, and reach resources only through a Unix socket minted for that run.
- Events are appended before execution; recovery folds the log rather than re-running the model.
- When an agent's output fails its schema, a truncated error diff goes back to the model - two attempts, then a typed failure into the run log.
- Every value carries an envelope: producer, causing event, taint, and budget spent. Taint is recorded today, not yet enforced.
failure handling
What happens when a run fails.
invalid model output
It is caught and rejected.
Every result has a declared shape, and one that does not match is not passed along. The error goes back to the model to correct, twice at most, and then the run fails with a clear record of which step failed and why.
report.spec@1 · repair 2 of 2 exhausted · typed failure written
model output is untrusted input→the machine fails mid-run
The run resumes where it stopped.
No state lives only in the engine's memory. Every event is appended to one log and the state of a run is read back from it, so an interrupted run continues from where it stopped instead of starting over.
executor restarted · run resumed from seq 41 · completed
the log it recovers from→approval required
The run waits for the decision.
A step can be gated on a human approval. The run suspends until somebody approves or denies it — surviving restarts in the meantime — and either decision is written to the same record as everything else.
awaiting approval · 2 approvers · run suspended durably
the record it keeps→