Lowball / Field notesCox Automotive Hackathon 2026
Human–agent collaboration

RoboScrum and
Mussolini Mode

Why I gave agents their own kanban board—and why I took it back.

Jeff Cameron · October 2026
A robot scrum team coordinating around a kanban board with To Do, In Progress, and Review columns.
THE AGENTS HAD THEIR OWN BOARD.

Some of my teammates are still trying to understand how they were able to open Claude, point it at our repo, and start contributing with so little intervention. They didn't have much coding experience. They hadn't learned the whole codebase. Yet their sessions could pick up work and move it forward alongside everyone else's.

That was one of the most useful outcomes of Lowball, our vehicle-appraisal hackathon project. It also needs more explanation than “we used agents.”

A lot of the answer was in the planning and research I'd done before they started. We had defined the work, divided ownership, frozen the contracts between components, and specified how to check and hand off the results. That gave a new session somewhere to begin without requiring its human to explain the entire system.

But the plan couldn't tell a session what another person's agent was doing right now. For that, we needed an orchestration pattern and a way for agents running under different accounts to communicate.

We gave them their own board. Around it, I built what amounted to an autonomous scrum team: agents claiming work, talking across accounts, checking on each other, and keeping the human team and its sessions moving toward a production-grade MCP with real infrastructure.

It got us remarkably far. It also became too controlling, and I eventually had to dismantle the control system so the humans could take the project further. The software architecture had room for what they wanted to build. Our process had become much less flexible.

What was already there when they opened Claude

Lowball was six people and a lot of agent sessions. By the end, 69 assistant names had posted on our GitHub board. That's not 69 distinct agents; some were one session having an identity crisis. Most ran in Claude Code under different people's accounts, with no shared chat or runtime. A couple of Codex sessions worked alongside them on review and overwatch only.

From a teammate's seat, the experience could look almost hands free. Underneath it, several pieces were doing different jobs:

What we'd prepared What it let a session do
Lanes with owned directories Work within a defined part of the repo
Frozen contracts and a fake implementation behind each seam Build against an agreed interface before the neighboring component was ready
Acceptance IDs, commands, and a fixed handoff format Check its work and leave evidence another session could use
An agent board with addressed messages and path claims Find current ownership, ask questions, and hand off across accounts
Hooks connecting Claude Code to the board Receive updates and maintain its presence without remembering to check

That preparation reduced how much a teammate had to know before making a useful contribution. They didn't have to become the person who could explain every directory, dependency, and parallel task. Their session had a bounded assignment, an interface to work against, and somewhere to ask when the plan met reality.

It also explains why this wasn't simply a matter of pointing Claude at any repo. The research and planning supplied the starting context. The board and hooks kept that context current as other people worked.

How a new session joins the workResearch and planning feed lanes, contracts and acceptance criteria, then a teammate’s Claude session. The session exchanges updates through hooks with a shared agent board and sessions under other accounts. Work and handoff evidence proceeds to review and human decisions.Upfront research and planningLanes, contracts,acceptance criteriaTeammate’s Claude sessionHooks: inbox, claims, heartbeatShared agent boardSessions under other accountsWork and handoff evidenceReview and human decisions
FIGURE 1 · How a new session joins the work

The repo supplied the plan. Cross-account communication supplied the changing state of the work. “Hands free” described stretches of execution, not permission to make every decision.

We shipped the round trip: an appraisal worked with Claude, saved into vAuto on a test rooftop, and read back by its real ID. The preparation and coordination both mattered to getting there.

But what about SDD?

This is usually the first question I get: how does this relate to spec-driven development and frameworks like Spec Kit?

It builds on that approach. The upfront research, contracts, task breakdown, and acceptance criteria were central to why our sessions could work independently. Spec Kit's documented workflow carries specifications through planning, dependency-ordered tasks, implementation, and checks against the resulting work. It supports parallel task markers and converting tasks into GitHub issues, too.

What we added was ongoing coordination among independent sessions running under different people's accounts. Even with the same spec and task list, those sessions needed answers to questions that changed during the build:

A task marked safe to run in parallel describes the plan. A live claim tells another session that someone has taken it. An agreed contract describes the interface. An addressed conversation lets two sessions resolve a question about it while they work. Hooks deliver those updates without making each human the messenger.

That's the addition: shared working state and cross-account communication around spec-driven execution. The board was their Rally, Slack, and Outlook because a team needs all of those functions even when its members are agents.

You could use this pattern with Spec Kit or another SDD framework. The spec, plan, and tasks would remain the reference for the work; claims, inboxes, heartbeats, and handoffs would coordinate the sessions carrying it out. A board conversation that reveals a needed contract or scope change still has to go through the agreed approval process and back into the relevant artifacts.

For our less-experienced teammates, both parts mattered. The preparation reduced what they had to explain to their own agent. The coordination reduced what they had to relay between everyone else's.

Why the agents needed a separate board

Our planning board had about eighteen issues, with one milestone per gate. The agents' side ended up with thousands of posts. On a single board, the team's commitments would have been buried under claims, heartbeats, and “does get_appraisal own the revision check?”

The boards answered different questions. We wanted to know which gate we were at, who was blocked, and what was at risk. The agents needed to know which paths were claimed, who could answer a question, and what was ready for review.

The project board was their Rally, Slack, and Outlook rolled into one: a place to claim and track work, have conversations, and receive messages addressed to them. It did more than show status. It was where otherwise separate sessions found each other and coordinated.

The separate board also gave sessions a common place to communicate across accounts. A question from one person's Claude session didn't have to be copied into Slack, noticed by a teammate, and pasted into another chat. It could be addressed to the session doing the work and delivered through that session's inbox.

Agent posts were written to be parsed: useful for a script, tedious for a person. On their own board that was fine. We could still read every word, and claims, PRs, and reviews linked back to the same repo.

The trust rule was explicit: a board message informs; it never authorizes. Agents coordinated on their board. Decisions happened on ours, in Slack, and with the people themselves. Merges to main, deploys, AWS creates, and scope changes needed the agent's own human, whatever another session had posted.

Separate sessions, shared coordinationClaude sessions under separate accounts use hooks and board.py; Codex review calls the script directly. All exchange messages with the agent board. Status and evidence flow to people; each agent’s own human must approve merges, deploys, AWS creates, and scope changes.SEPARATE ACCOUNTS AND SESSION CONTEXTSPerson A’sClaude sessionPerson B’sClaude sessionCodex review sessionAgent boardClaims, questions, handoffsPeoplePlanning board and decisionsMerge, deploy, AWS creates, scope changesHooks + board.pyHooks + board.pyDirect callsStatus and evidenceOwn human’s approval required
FIGURE 2 · Separate sessions. Shared coordination.

The board connected independent sessions. It did not grant them each other's authority.

How sessions talked to each other

Every post had a readable header and a hidden copy a script could parse:

[jeff-claude-2 → kenny-claude · L2] question: does get_appraisal own the revision check?
<!-- lowball-agent {"from":"jeff-claude-2","to":["kenny-claude"],"lane":"L2","kind":"question"} -->

A message could go to one agent, any of a person's agents, a lane, or everyone. Each carried a kind: claim, progress, question, answer, heads-up, request, blocked, handoff, or done.

Several sessions shared one GitHub login, so the session name was the identity. Routing read the header rather than @mentions. Agents could pick nicknames—it was supposed to be fun—but the routing name was meant to stay stable underneath.

They also claimed paths before editing. A claim was an issue naming the files or directories a session intended to touch. Before an edit, the session checked for overlapping live claims. When the work was ready, it posted a handoff and moved the claim to Review. The claim closed on merge.

These were claims, not locks. The edit check warned about a conflict; it didn't prevent one. A claim became stale only when its holder had been quiet for 30 minutes and the claim was more than two hours old. Another session could then take it over. That rule would turn out to need more judgment than we'd given it.

Hooks kept the coordination running

Writing “check the board” in the instructions wouldn't have been enough. We wired the board into the session lifecycle.

board.py used only the Python standard library, ran through each person's existing gh login, and handled no tokens itself. A Claude Code plugin connected it to hooks:

When What the hook did
Session start Supplied the build clock, the session's inbox, and its lane's live claims
Each prompt Supplied only new updates
Each edit Warned if the file overlapped another session's live claim
Stop Posted a silent heartbeat to keep claims live

The overlap check was a warning rather than a permission prompt because a prompt would stall an unattended run. Every hook failed open after three seconds, so a slow GitHub Enterprise response wouldn't block work. That favored continuity over guaranteed coordination: the board helped sessions avoid collisions, but it couldn't guarantee they wouldn't happen.

Review sessions outside Claude Code called the script directly.

This was another part of what teammates couldn't see from the initial prompt. Their sessions were getting updates through the plugin rather than relying on their humans to relay every change.

What it looked like in practice

The week's archive contains 2,914 issue bodies, comments, and reviews. The lobby alone had 220 comments from 24 agent names. One of my sessions ran as a monitor and wrote about 17% of all traffic, mostly answers, heads-ups, and corrections.

On day one, it caught one of my other sessions taking over a teammate's claim:

jeff-claude-6 (one of Jeff's sessions) took over your claim #194 as stale and closed it... You own platform deploys, so if you're working toward the first deploy, re-claim #95 to keep it yours. jeff-claude-6: #194 was Kenny's lane, not Jeff's. Before taking over another teammate's claim, check with that person's session.

One agent had made a coordination mistake. Another corrected it in public, addressing both the teammate and the session within minutes.

That exchange shows both the usefulness and the limits of the setup. Cross-account communication let the correction reach the people and sessions involved. It also exposed that our takeover rule treated a quiet claim as available without enough regard for whose lane it was in.

The architecture had room for more

One thing I'd deliberately designed for was uncertainty about what we'd be able to connect. I didn't know which API access I could get or everything the data scientists would dream up. The architecture needed to accommodate both.

The intent was to take the integrations and models we could provide, make them available through the agreed interfaces, and run with whatever was there. As more became available, we could turn a valve and make it live. We didn't have to know the final set of capabilities before assembling the system.

That was part of the point of the contracts and fake implementations: we could build the surrounding application while the real capabilities were still taking shape. A stable interface left room for the implementation behind it to grow.

I'd built flexibility into the software. What I hadn't made flexible enough was the way the team worked on it.

Where it broke

The biggest failure was that the project reached a point where we needed to get creative, and the control system I'd built got in the way.

It had helped people join the work without understanding the whole codebase. It gave their agents defined assignments and kept those assignments coordinated. But making it easy to execute the plan hadn't made it easy to depart from it.

There were two sides to what I called “full Mussolini mode.” The system was already in it: the rules I'd put in place had become too controlling. Then I had to go full Mussolini myself, take the power back from my own system, and give it back to the humans. I wasted a day doing that.

Once I got out of the way, Kenny unveiled an insanely awesome new 3D lot view tool, and the data science guys were able to execute on creating richer data models. Those were concrete contributions that made the project better, beyond getting the existing assignments done.

That's the uncomfortable part of this story. The architecture was ready to accept more of what the team could bring, but the process I'd built was getting in their way. People still held the formal approval authority; in practice, the controls made changing direction too difficult. I had to recognize that the system I'd put so much work into had become part of the problem.

Next time, I'd give the coordination process the same flexibility I'd given the architecture. We'd need a clear way to loosen controls, enter an exploratory phase, reconsider assignments or contracts where necessary, and bring what we learned back into the plan. Changing how we worked shouldn't have required a day of fighting the machinery I'd built to help us.

There were smaller coordination failures, too:

Our stale-claim rule was too mechanical. It treated a quiet teammate's claim the same as a quiet claim of mine. The monitor's correction is how we caught it.

Names weren't reliable session IDs. Sessions without an explicit name defaulted to the same one, causing their claims to skip each other's overlap checks. At one point, two sessions ran as jeff-claude-6 and refreshed each other's heartbeats. PR #112 added a warning. We'd require a unique name at session start next time.

Cleanup could be confidently wrong. One cleanup marked four cards Done while their issues were still open. The monitor moved three back. Having a board didn't make its status correct; having visible evidence made the mistake easier to find.

We migrated the board mid-build. Open packages and live claims were recreated in the new repo with their paths, branch, and last-seen time. We updated the plugin so old branches could find the new board. It worked. We wouldn't choose to do it again.

Work escaped the loop. A cloud session left a real proposal on a branch with no PR and no claim. We found it and cherry-picked it back in through PR #396, mostly by luck. Work the board couldn't see was still easy to lose.

The board's value depended on how many people and sessions were active. On the final weekend, one person owned every lane and board posting became optional. With six people and dozens of sessions, it was essential. With one person, it became overhead.

What I'd keep

I'd keep the separate agent board, path claims, addressed messages, automatic inbox reads, heartbeats, and a monitor session. I'd also keep the rule that another agent's message cannot authorize an action requiring your human.

But I'd start with the work that made those mechanisms useful: the research, the lane boundaries, the contracts, the fake implementations, and the acceptance and handoff requirements. A claim on a directory helped because we'd decided who owned it. A question about an interface could be answered because there was an agreed contract. A fresh session could continue because the handoff included owned paths, contract version, acceptance IDs, exact commands and their output, and open gates.

For the teammates wondering how they could open Claude and get so far: much of the context you would otherwise have had to learn and relay was already available to your session. My upfront planning gave it a place to start. The orchestration pattern and cross-account communication let it keep working as everyone else moved.

I'd also keep the architecture that could accept new APIs and richer models as they became available. That flexibility paid off. The technical lesson is to extend it to the orchestration itself: make the rules easy to revise or suspend as the team's needs change. The harder personal lesson was recognizing when my own system needed to get out of the way.

Submission day

On submission day, I experienced a personal tragedy and had to hand over the captaincy to another team member. They stepped up and carried us over the finish line.

The whole build became an experiment across the spectrum of human–agent collaboration: people setting direction, agents executing and coordinating with little intervention, a control system that became too restrictive, and humans taking back the freedom to reshape the work. At the end, it also meant someone else taking the lead when I couldn't.

I'm proud of the autonomous scrum team I built and the architecture it helped us assemble. I also had to dismantle its controls to give my teammates room to expand what we were building. And when I couldn't finish the job myself, a teammate took over and got us through submission.

Published with OpZero