Lesson 6: The backlog is a queue, not a list
What you'll learn: how to turn the list of work you mean to do into a file an agent can work one item at a time: a status, an acceptance line, a depends list, and a model tier per item, plus one pick rule, so the next item is never your decision.
What you need: the specs/ directory from Lesson 1, one home per stack from Lesson 4 or a single repository, the standing authority file from Lesson 5, and five to-dos you have been meaning to get to.
Time: twenty minutes once, five per item after. Cost: zero new dollars.
The idea#
You keep a list. It has forty lines, three of them half done, and the agent you hand it to starts at the top, asks what the second line means, and by the fourth has done a little of each. The list was written for you, by you, in the order they occurred to you, and every gap in it is filled by a person who remembers the stand-up. The agent was not at the stand-up. A ticket board already solved this: a ticket is ready only when it carries acceptance criteria, one engineer pulls one ticket, and work in progress is capped. Do that in a file. Every item gets a status, an acceptance line you could check, the items it depends on, and the tier of model it deserves, and the file carries one rule for which item comes next. A backlog is what agile calls that pile of work; the version an agent can drain is a queue.
Why start here#
Lessons 1 through 5 gave the agent a spec, a decision record, a check, one home, and a file that says what it may do. None of them says which task, and that question, asked at the end of every session, is the one you still answer by hand. Since the end of May the queue file in my own setup has taken over twelve hundred edits, among the most of any file I own. The queue answers what comes next; Lesson 7 adds the gates before you look away.
The template#
Save this as BACKLOG.md beside the standing file; the field names match the public schema in Further reading:
# Queue
## Pick rule
Read this whole file. Take the lowest-numbered item whose status is ready
and whose depends_on items are all done; if none qualifies, say so and
stop. Before taking an item whose pr is empty, search the merge history
for its id; if the work is already there, mark it done and pick again.
One item per session: set it doing, finish it, set it review, and stop.
Never edit an item you are not holding. Work you discover is a new item
at the bottom with status draft, never a wider current one.
Status: draft -> ready -> doing -> review -> done. I move draft to ready
and review to done; you move ready to doing to review. For code, done
means merged. For a queue-only investigation, done means I have reviewed
and accepted its draft children. blocked waits for me.
Tier: judge (the acceptance needs a decision), build (a checkable
acceptance line), mechanical (the steps are already listed in the item).
## B-001 health() returns ok
- repo: api
- status: ready
- depends_on: []
- tier: build
- spec: specs/api/002-health.md
- acceptance: python -m pytest api/tests -q passes with a test asserting health() returns {"ok": true}
- pr: -
- notes: -
## B-002 Work out what the status page needs
- repo: specs
- status: ready
- depends_on: [B-001]
- tier: judge
- acceptance: this item is done when it has appended B-00N children as draft, each with an acceptance line I can run, and it has asked me nothing it could have found in the repository
- pr: -
- notes: an investigate item; its output is items, not code. Context: nothing exists yet. The page lives in web/ and shows when a nightly job, to be created in infra/, last ran; the job writes its last-run time to a file api/ can read
Add this to the file your agent reads at session start, under the lines from Lessons 1 through 5:
Before starting work, read BACKLOG.md and take the item its pick rule names. It is the one file outside your home you may edit, and only the item you hold. When it is finished, set its status and stop.
Each part, and the mistake it prevents#
The pick rule is the whole difference between a list and a queue. A list is read by someone who chooses; a queue names its next item itself: lowest number with its dependencies done, arithmetic instead of you.
One item per session prevents the three-half-done afternoon. It is a ticket board's work-in-progress limit, applied to one engineer.
Status prevents the file from lying. The agent moves the item it holds to doing and to review; you move it to done when you merge the code. For a queue-only investigation, you review and accept its draft children before marking the investigation done. That acceptance leaves the children as draft until you decide which are ready. Review belongs to you until Lesson 7 gives it to a pipeline.
The acceptance line is the ticket's definition of done, the one field that cannot be vague. "Improve the status page" is a list entry; "the status page returns 200 and shows the last run time" is an item.
Depends lets items be written out of order and worked in order; the pick rule reads it, so a wrong depends list is a wrong queue.
Tier is the one sentence of routing this series spends. Deciding what the work is and doing it want different models: the expensive one judges, the cheap one types. Written on the item, the choice is made once.
The investigate item is how the queue grows its own items. Work you cannot write an acceptance line for still gets an item; its acceptance is that it appends children that have one. The agent researches and writes items, not code, and you strike or edit the children before any is marked ready.
A worked example#
The homes from Lesson 4, the standing file, Lesson 5's authority file, and the queue above with three build items behind the two shown. A fresh session, given only the standing file and "work the next item":
B-001 is done and set to review. `python -m pytest api/tests -q` exited 0 in
this session with 1 test passing. The work is a local commit only: `api/` has
no remote, so nothing is pushed.
**What shipped** (Home: api, spec `specs/api/002-health.md`)
- `api/src/health.py` has `health()`, which returns `{"ok": True}`.
- `api/tests/test_health.py` asserts that return value.
- Commit `30c7333` is on branch `b-001-health`, cut from `main`. It holds only
those two files, added by path.
- In `BACKLOG.md` I changed only B-001's status and notes.
- Nothing under the spec's Out of scope changed (no HTTP server, `web/`, or
`infra/`).
**Choices I made**
- The test adds `api/src` to `sys.path` itself. That way it imports `health`
whether pytest runs from the project root, as the verify command does, or
from inside `api/`.
- I left `pr` as `-` and put the commit and branch in notes, since there is no
PR to link.
**What the spec could be read to demand that I left out**
- Nothing. It asks only for the function and a test.
**Follow-ups that deserve their own spec**
- `api/` has no remote, so the branch can't be pushed or reviewed as a PR
until one exists.
- There is no `.gitignore`, so `__pycache__/` shows as untracked. I left it
alone.
[one bullet about a commit hook on my machine omitted]
I stopped after this one item, as the pick rule says. B-002 is ready now that
B-001 is in review, but its dependency isn't `done` (merged) yet, so it isn't
eligible until you merge.
I merged that branch and set B-001 to done. A fresh session, the same prompt, and the pick rule landed on the investigate item:
I took B-002 and set it to `review`. Its output is items, not code, so I added
six draft children (B-006 to B-011) to `BACKLOG.md`, which is the only file I
edited. The file on disk matches what I wrote, and no other repo changed.
Home: specs. The work is in the queue file, not in any repo, so nothing was
committed or pushed.
**What shipped.** Each child is confined to one home, and each has an
acceptance line you can run:
- **B-006 (specs):** a spec for the last-run file contract.
- **B-007 (api):** a `last_run()` reader function, with pytest.
- **B-008 (infra):** a nightly job script that writes the file, checked with a
temp path.
- **B-009 (web):** a `status.html` page, checked with grep.
- **B-010 (api):** an HTTP route, only if you choose the route option below.
- **B-011 (infra):** scheduling the job, which waits on your approval of a
cloud resource.
**Choices I made, and why.**
- I set the file format to `{"last_run": "<ISO 8601 UTC>"}`, with the path
taken from an environment variable. A missing file means the job has never
run. Without a fixed contract the three homes can't build in parallel.
- I made B-006 the root, because the layout skill says the spec comes first.
B-008 and B-009 also cite specs that don't exist yet
(`specs/infra/001-nightly-job.md`, `specs/web/002-status-page.md`), so their
notes say each needs its own spec first.
- I gave B-008's acceptance a temp-file check so it runs locally with no
cloud.
**What the spec could be read to demand that I left out.**
- I did not write the specs themselves, because a spec is a separate change in
the specs home.
- B-011 doesn't pretend to have an executable acceptance line. Its line is
your written approval.
**Two questions the repo can't answer.**
1. **Where does the last-run file physically live?** `infra/` and `api/` are
separate repos with no shared storage. I'd pick an object in a bucket that
`infra/` creates. That is a new cloud resource, so it needs your approval.
2. **How does `web/` read it?** The options are a static `/last-run.json`
published beside the site, or an `api/` HTTP route. I'd pick static, since
it adds no server or web framework. If you agree, drop B-010.
**Follow-ups worth their own spec.**
- A staleness rule: when should the page flag the job as overdue?
- Job failure reporting, since a last-run time says nothing about whether the
run succeeded.
I stopped at one item per session, as the pick rule says. B-003 is the next
ready item.
Two of the six, as the session wrote them into the file:
## B-007 api: last_run() reads the last-run file
- repo: api
- status: draft
- depends_on: [B-006]
- tier: build
- spec: specs/api/003-last-run.md
- acceptance: python -m pytest api/tests -q passes with tests asserting
last_run(path) returns {"last_run": "<ts>"} for a valid file and
{"last_run": None} for a missing file
- pr: -
- notes: new function in api/src/last_run.py; health() untouched. No HTTP
layer here (see B-010).
## B-009 web: status page shows the last run
- repo: web
- status: draft
- depends_on: [B-006]
- tier: build
- spec: specs/web/002-status-page.md
- acceptance: web/public/status.html exists and contains the text "Last run"
and a fetch of "/last-run.json"
- pr: -
- notes: needs its own spec first. Default: static page that fetches
/last-run.json and renders "Last run: <ts>" or "never". OWNER QUESTION: is
/last-run.json served as a static file (job publishes it beside the site) or
through an api HTTP route (B-010)? I would pick static, since it adds no
server to api/.
No code in the second run: six draft children, each in one home with an acceptance line, and two questions the repository could not answer. The first version of that item left out whether the job existed, and the session stopped to ask exactly that; the fix was one clause on the item.
How it scales#
The file grows by items, not by rules. Mine holds several hundred items across fourteen repositories and the pick rule is still the paragraph above; what grew was a planning pass that reads the blocked and draft items whenever fewer than five are ready and marks the answered ones ready. Who does the work: you write acceptance lines and set tiers: your definition of finished and your money. The agent proposes children from investigate items, drafts items from what it discovered mid-task, and moves status on the item it holds.
From the field#
On July 28 the loop in my setup picked three items in a row whose work had already merged. The queue said open; the repositories said done. Each turn went to confirming the shipped work was the item. The loop had done exactly what the file told it; the tool that recorded status had been reverting it on its own, so the file was lying, and a loop that reads a lying file does the wrong work faithfully. The guard was two minutes at the top of every pick: search the merge history for the item's id, and close the item if the work is already there.
Exercise#
- Save
BACKLOG.mdbeside the standing file and add the line. - Replace B-001 and B-002 with five items of your own, numbered from B-001. Each gets an acceptance line you could check and a depends list. The one you cannot write a line for becomes an investigate item, with the acceptance from B-002.
- Set a tier on each. If you cannot say whether an item needs judgment, it is an investigate item.
- Open a fresh session with the standing file only and say "work the next item". Watch it: it should name the item, do it, set the status, and stop.
- Review that one result; merge its code and mark the item done. If a later watched run reaches the investigate item, review and accept its draft children before marking the investigation done. Edit the children before marking any ready.
You're done when#
An item closed without a single mid-flight clarification from you, while you watched. Two results fall short: the session asked a question the item should have answered, which means the acceptance line was vague, so fix the item and not the agent; or it kept going into a second item, which means the pick rule was not in front of it. The eyes-off version waits for Lesson 7, which teaches the isolation that makes it safe.
Common mistakes#
- An item with no acceptance line. "Improve the admin page" is a wish, not an item.
- The agent editing items it is not holding.
- A status left stale after a merge. The pick rule believes the file, so the file has to be true.
- Batching two small items in one session because they were close together.
- A judgment item marked mechanical. The cheap tier may finish it, confidently and wrong.
Further reading#
The queue above is the small end of loop-harness, the public schema and generator skill I run; it adds the project profile block a loop needs to run unattended. The Kanban Guide covers work-in-progress limits and explicit workflow policies. Bill Wake's INVEST is the older checklist for a ticket worth taking; the acceptance line is its T.
Next lesson#
The pipeline reviews so you don't: a disposable copy of the repository, a critic pass against the spec, a live check that a merged change reached the live site, and the deny rule from Lesson 5's must-ask lines, so the first time you look away, something else is looking.
Questions this post answers
- What is the difference between a to-do list and a queue an AI agent can work?
- A list is ordered by when each line occurred to you and read by a person who fills in the gaps. A queue is a file where every item carries a status, an acceptance line you could check, the items it depends on, and the tier of model it deserves, plus one rule that names the next item. The agent reads the file, takes one item, closes it, and stops. Nothing about what to do next lives in your head.
- What goes in a queue item?
- An id, a status, a depends list, a tier, an acceptance line, and a pointer to the spec if one exists. The acceptance line is the contract: if you cannot write one you could check, the item is not ready, and it becomes an investigate item whose only output is children that do have one.
- How does the queue handle work I have not figured out yet?
- It gets an item too, with a different acceptance line: this item is done when it has produced child items, each with an acceptance line I can check. The agent researches and writes items instead of code. You strike or edit the children before any of them is marked ready.
- Why does each item name a model tier?
- Because deciding what the work is and doing the work want different models. An investigate item needs judgment and gets the expensive tier; an item whose steps are already enumerated gets the cheap one. Writing the tier on the item once beats choosing it in the moment, and it is the field that ages fastest, so it lives in the file and not in a lesson.
Previous lesson: Lesson 5: What the agent may do without asking