2026-07-18 · Updated 2026-08-27 · 6 min read
Build auditable agent workflows with handoff
One run directory holds the manifest, per-step responses, session handoff, doctor diagnostics, and fail-closed recovery that survives interruption.
By Juno AI INC · workflow-runner · session-handoff
An auditable workflow is one whose evidence survives every process boundary: the end of a step, the end of the run, a crash in the middle, and the handoff to whoever reviews or continues the work. Workflow Runner writes that evidence into one run directory, so review, diagnosis, recovery, and continuation never depend on a terminal that has already scrolled away. Start from reviewed input, then let the artifacts carry the record:
Lint flags response and log template anti-patterns using the same parser the run will use, and the dry run writes the rendered artifacts without executing any step.
Treat the run directory as the record
Every run owns one directory — .juno_task/specs/workflows/<workflow_id>/<run_id> by default, or --out-dir when you place it deliberately — and the reviewer's map of it is short:
NNN_<step>.stdout.txt,NNN_<step>.stderr.txt, andNNN_<step>.response.txt— each step's streams in step order, plus the final answer for agent steps.manifest.jsonandmanifest.yaml— every step's status, exit code, command preview, and session id, the run status, and the failed-step list.run_contract.json— the run's single checkpoint and attempt index; each attempt's manifest is archived underattempts/<attempt_id>/.summary.md,summary.stdout.txt, andsummary.stderr.txt— the rendered summary and its own streams, next tosummary.command.sh.
For agent steps the response is the answer: stdout is the canonical response, successful stderr stays in artifacts and is printed only when the step fails, and a detected agent command that exits zero with an empty response is marked failed — silence cannot masquerade as completion. Template {{ steps.<id>.response }} downstream instead of reaching for stderr, and let lint catch that mistake before launch. Steps that declare typed receipts expose their paths as {{ receipts.<id>.path }} or the JUNO_WORKFLOW_RECEIPT_<ID> variable, never as duplicated literals that can drift from the contract.
With --tmux the observer is exactly that — a detached view that never detaches the producer. Step streams also land in workflow.live.log, and manifest.json records the observer's session, live log, and attach command, so even the view is part of the record.
Stop automation at boundaries you chose in advance
By default a failed step is recorded in the manifest while the process still exits zero and later steps keep running. That is deliberate — evidence first — and it means the exit code alone must never be your review. Put fail_workflow: true on the steps where continuing past failure would be worse than stopping: validation gates, destructive commands, and any step whose consumers would poison themselves on a bad response.
Choose those boundaries in the YAML before the run, while the consequences are still cheap. After the fact, audit the manifest: a green exit with a non-empty failed-step list is a finding, not a success, and the diagnostics in the next section usually explain it.
Diagnose from artifacts, not the console
After any run — green, red, or merely suspicious — inspect the run directory with doctor, short alias dr:
Doctor checks that the manifest and artifact paths exist, that no successful agent step carries an empty response, and that agent commands were not accidentally quieted; it labels successful stderr as log or audit noise rather than summary input. It resolves the newest hash-bound attempt manifest through the same evidence resolver that review uses, keeping the root manifest.json only as a legacy fallback.
Most symptoms have a direct artifact move. A downstream value that arrives empty points first at the producing step's response.txt — a detected agent command that exits zero with an empty response is failed by contract, so an empty template value is a step worth reading, not a silent success. An exit code that surprises you means checking whether the failed step lacked fail_workflow: true. A noisy console is a presentation choice: --no-print-step-stdout --print-output summary keeps the invoking terminal quiet while every artifact still lands on disk.
Recover an interrupted run without inventing success
When the producer dies before writing terminal metadata — a crash, a kill, a lost connection — the hash-bound checkpoints of completed steps remain in the run contract. Verify what is recoverable before changing anything:
Recovery refuses active, partial, non-contiguous, cross-run, or drifted evidence; when it succeeds it appends an interrupted attempt manifest recording the verified prefix and the first invalid step, and it never infers semantic completion. Resume exactly at the reported first invalid step: --from-step accepts a step id or name, a zero-based index, or -1 for the final step, and before dispatch it re-verifies the unchanged workflow, variables, rendered commands, frozen inputs, producer digests, and receipt hashes — only then are predecessors marked reused_verified.
Never edit a historical run to make its evidence reusable. A harness-only correction — you fixed the workflow, not the work — uses a fresh output directory, amendment_mode: harness_only_validation, and --amends-run PRIOR_RUN; add --from-step to revalidate and import the prior successful prefix instead of replaying it, and read the printed plan and the manifest's amendment_plan to see which steps were revalidated versus executed.
Hand off the session, not the terminal
Continuation is part of the record. Steps that invoke yylo, yy, or ypl capture session metadata automatically — opt out per step with capture_session: false — so later steps and templates can reference {{ steps.<id>.session_id }} and even resume a prior agent mid-run. At the end the runner prints every detected session id and persists the final successful agent session to the same continue-scope env file and registry yy cc reads:
Run yy cc from the shell scope that produced the run. When the handoff should point at a different step, set top-level continue_from_step: <step-id> in the workflow; the selection is strict and fails loudly when that step produced no session id.
Independent fan-out has a different evidence owner: Parallel Runner writes per-item status and aggregation artifacts rather than an ordered run directory, and choosing between the two shapes is its own decision. Within an ordered run, keep the run directory with the review evidence — the manifest and response artifacts are the record, and every console, including the tmux observer, is only a view.