2026-07-18 · Updated 2026-08-28 · 7 min read
Use YYLO Ledger as the source of truth for agent work
Give agents durable task memory instead of hand-moved board cards with dependency-aware readiness, required responses, commit evidence, and hash-chained history in Git.
By Juno AI INC · yylo-ledger · task-truth
An agent with no readable task state improvises one: it re-reads a chat transcript, guesses what finished, and re-does or skips work accordingly. The usual fix is a board, but a board whose state lives in someone else's database only moves the improvisation. A card dragged to Done is a claim. Nothing in the repository confirms it, no script can check it without an integration and credentials, and the audit trail is an activity log you cannot diff. YYLO Ledger keeps task truth the way source control keeps code truth: as plain files in your repository, read and written through one shell-friendly CLI that humans and agents share.
Four semantics turn those files into task memory. Dependencies are declared data, not prose. Readiness is computed from that data instead of asserted by a person. Every state transition requires a recorded response, and completion carries commit evidence. Underneath all four, an append-only, hash-chained history keeps every change auditable. Each semantic has a concrete command and a concrete refusal, which is what makes the state trustworthy for unattended work.
Keep task state where the work happens
Canonical current state is one safe Markdown file per task under .juno_task/tasks/: YAML frontmatter for the structured fields (id, status, blocked_by, commit_hash, timestamps), with hidden marker comments bounding the body and response so free-form text can never corrupt the structure. The files are versioned by the same Git history as the code they track, and two properties keep that mergeable: status updates never rename a task file, and different tasks never share one — so concurrent agents working separate tasks in separate worktrees produce independent diffs, not conflicts.
Querying stays script-friendly. list, search, ready, and get emit NDJSON by default, with json, xml, and table behind -f, so pipelines and agents parse one stable shape:
The speed layer is honest about being disposable: search runs on a SQLite cache under .juno_task/cache/, and deleting or corrupting it only triggers a rebuild from the canonical Markdown. State never comes from the cache, and it never comes from the history ledger either — the current file is the truth, everything else is derived.
Declare dependencies as data
Ordering that lives in a paragraph ("do this after that") is invisible to every machine and most humans. The ledger makes it a field. blocked_by is a list of task IDs that must reach a terminal status first; related_tasks is the non-blocking twin, linking context an agent should read without ever gating execution. Declare edges at creation, inline in the body, or later:
deps DEPLOY_TASK answers the question a blocked agent actually has: which blockers are met, which are unmet, who depends on this task, and what its priority score is. Two graph rules keep the data meaningful. A blocker that does not exist, or never terminates, is simply never resolved — the dependent never becomes ready, and completing it anyway is refused (below). And cycles fail loudly: order reports Dependency cycle detected with the task IDs involved rather than inventing a schedule through the loop.
For sequencing, order returns a topological sort of open tasks, and order --scores adds priority scores that rank a task by how much downstream work it unblocks — useful when several tasks are simultaneously ready and you must choose.
Make readiness a computed answer
ready lists tasks whose status is actionable (backlog, todo, in_progress) and whose every blocker exists and is terminal (done or archive). It is a query over declared state, not an opinion, and it is deterministic: ready --sort asc orders by last_modified with the task ID as a tie-breaker, and explicit status filters preserve their given group order — the same contract list and search honor, so scripts can rely on it.
Readiness is also enforced, in both directions, before any byte is written:
- Marking a task
done(or archiving it) while a declared blocker is missing or non-terminal is refused outright —cannot complete task X; unmet blockers: ...— so a completion cannot strand a graph that lies. - Reopening a task that other completed work depends on is refused too, because that would retroactively block a finished dependent.
A task missing from ready is therefore a diagnosis, not a mystery: deps TASK_ID names the blocker still holding it. When readiness becomes the admission rule for running many tasks at once — quotas, isolation, per-run evidence — that is parallel execution territory.
Require a response for every transition
Every mark carries the agent's own account of the work, and the CLI treats that account as mandatory, not optional: mark STATUS TASK_ID without --response or --response-file is a usage error, and nothing changes. The response is stored as durable task truth, retrieved by get alongside the body — not a chat message that scrolls away.
Responses hold real Markdown, so use the file forms for anything with fences, backticks, or variables — the Kanban wrapper reference documents --body-file and --response-file, both accepting - for stdin. Durability also has a safety edge: recorded responses persist and re-enter later prompts and sessions as memory-surface input, so they belong to the same review discipline as any text an agent will trust later — keep secrets and unbounded shell output out of them.
A strong response states what changed, the exact commands and checks that passed, and the remaining risk. It does not paste uncurated logs, and it never claims success from an exit code nobody observed.
Bind completion to a commit
A response is a claim; the commit is the evidence. mark done --commit abc123 stores the hash on the task's commit_hash field, tying the recorded account to a diff reviewable in Git. Omitting --commit succeeds only with a reminder that the evidence link is missing — the flag is recommended precisely because get TASK_ID returns body, response, and commit together, and a done task with no commit hash is an audit gap you can see immediately.
Read them as a pair. If the response claims a passing gate but the committed diff shows unrelated edits, the completion failed your review even with a zero exit code — the same triad — response, commit, diff — that disciplines each cycle of a bounded agent loop. Auditing the whole chain — from the body's stated intent through the recorded response to the validation it cites to the hash it carries — is the task-evidence guide's dedicated walkthrough.
Keep history hash-chained and repairable
Current state is one file per task; history is a separate append-only ledger per task under .juno_task/ledger/, written as segmented NDJSON events that are chained by hash. A discontinuity or a hash mismatch is an integrity error, not silently missing history. history TASK_ID replays it:
Every mutation follows the same discipline: it locks one task, optionally compares --expected-revision and fails closed on a stale revision, atomically replaces current state first, appends the ledger event, verifies persistence, and refreshes the disposable cache. create, update, mark, and archive can also emit a task-scoped receipt with --receipt-file — before and after hashes, changed paths, and the ledger event ID — when you need the mutation itself as an artifact.
Because task files are plain text, direct edits are possible, and the ledger notices: reconcile --check flags any task file whose content no longer matches the ledger tip, reconcile records the edit as a real event, and doctor verifies markers, paths, hashes, and ledger integrity across the board:
If a mutation is interrupted mid-flight, canonical current state wins and the next mutation or reconcile converges the ledger — recovery is a repair of derived records, never a guess about the task itself.
That is the whole contract: dependency-aware readiness you can query, transitions that must say what happened, completions bound to commits, and history that verifies itself. Agents that can trust those four things stop improvising memory and start composing work — in parallel once readiness is the admission rule, and in bounded loops once the response-commit-diff triad is the habit.