bargello

Journal

What my agents did at 4:30am

For the past few months I've been spending my evenings building software with AI agents, and the thing that surprised me most is how little of that time goes on building. Most of it goes on reading what an agent did and deciding whether to let it carry on.

The coding agents I use all work in roughly the same way. They take a step, then stop and ask - can I run this, does this look right? Each question takes ten seconds to answer, and I answer it, and the next one arrives a minute later. With a few agents running at once I've spent whole evenings doing little else, and finished with a lot of half-built things and nothing a customer could use. I've started calling it the approval tax.

Feeling fast and being fast

There's some evidence that the feeling of speed is misleading. In early 2025 the research group METR ran a randomised trial with 16 experienced open-source developers working on 246 real tasks in projects they knew well. With AI tools they took 19% longer. Before the study they expected to be 24% faster, and afterwards they still believed they'd been about 20% faster.

To be fair to the tools, that was a year and a half ago, and METR itself now thinks the result is out of date. When they tried to repeat the study in early 2026, some developers wouldn't do the tasks without AI at all, which says something on its own, and their rough new estimate points to a speed-up of around 18% (with wide error bars). I believe that - the models I use now are far better than the ones I started with. But the part I find most interesting hasn't changed. People couldn't tell how they'd spent their time, and checking work feels like progress whether or not it is.

It isn't only engineers. A review of how people describe vibe coding in blogs and write-ups, presented at ICSE 2026, found that most of them saw the result as "fast but flawed", and that testing was often skipped altogether.

The other way to get it wrong

The obvious fix is to stop asking, and I tried a version of that too. I set my agents to pick work back up on their own overnight whenever my usage allowance reset. At 4:30 one morning they did, and my Mac woke up with the voice assistant switched on. Nothing was damaged, but I realised I didn't properly know what they had permission to touch, and I added guardrails the next day.

So I don't think zero approvals is the answer either. Some decisions should stop and wait for a person, like spending money or anything you can't undo. The problem is that today's tools ask about everything with the same urgency, so the decisions that matter get lost among the ones that don't. An engineer who wants a hand on every line might be fine with that. Someone who wants a booking site working by Saturday isn't, and shouldn't have to be.

That's the bar I'm holding Bargello to. It should ask you about money and about anything that goes out under your name, and handle the rest. It should get a business open and taking payments rather than produce another nice-looking demo, and tell you plainly when something isn't working. If it makes your evenings longer, it has failed, however clever the model behind it is.

Bargello isn't open yet, and I won't say much about how it works until it is. Until then I'll write here every week about building with AI and what it actually takes to get from an idea to a first sale.

Where does the approval tax hit you hardest? I'd like to know - reply, or join the waitlist at bargello.ai.

Sources

  1. METR, early-2025 developer trial (10 July 2025). 16 developers, 246 tasks, 19% slower with AI tools; they expected to be 24% faster and believed afterwards they had been 20% faster.
  2. METR, uplift update (24 February 2026). Developers are likely sped up now; some would not do tasks without AI; a rough new estimate of about 18% faster, with a confidence interval from 38% faster to 9% slower.
  3. Fawzy, Tahir and Blincoe, "Vibe Coding in Practice", a grey literature review of 101 practitioner sources, ICSE-SEIP 2026.

Be first in when Bargello goes live.