I’ve been running a lot of agentic coding sessions lately, and the sessions that go well have one thing in common: I gave the agent a number to hit, not a vibe to chase.
The bad version looks like “improve the audio pipeline” or “clean this up.” An agent handed that kind of instruction will do something — it has to, agents are compulsive doers — but it has no way to know when it’s done, so it stops when it runs out of ideas rather than when the problem is solved. The good version looks like “get audio round-tripping fidelity to 99%, and keep iterating against a test that measures it.” Same task, wildly different outcome, because now the agent has a number to close a gap on instead of a mood to satisfy.
This is training an ML model with extra steps, and once I noticed that I couldn’t unsee it. Training an ML model without a loss function doesn’t work — there’s no signal telling the optimizer which direction is “better,” so it just wanders. Prompting an agent without one doesn’t either, for the same reason: without a measurable target, “iterate until it’s good” quietly collapses into “stop when tired.” Give the agent a metric — a percentage, a test suite, a check it can run on its own output — and iteration turns into something closer to gradient descent: try, measure, adjust, repeat, until the number clears the bar.
I’d actually run into a version of this problem years before agentic coding was a category. The multi-agent translation simulation I built back in 2023 had a forward-translation bot, a back-translation bot, and an evaluation bot looping until “the translation draft stabilizes (if it does!)” — and in the demo, it never did, because none of the bots had a concrete stop condition, just each other to wait on. “Stabilized” was a vibe, not a number, so nothing ever stabilized. Different domain, same missing piece: no loss function, no convergence, no way to know the job was done.
That earlier miss is also where the swarm framing turns out to generalize better than I expected. Once each subagent in a swarm has its own measurable target, you can split work across a lot of them without losing the plot. The pattern I keep coming back to for coding specifically: spin up several cheaper, faster subagents to handle the parallelizable pieces of a task, and have each one test its own output before it reports back — a local loss function, one per subagent, scoped to its slice of the problem. Then an orchestrator agent merges everything and runs one more integration pass over the merged result, because a pile of locally-passing pieces isn’t automatically a working whole. Parallelize everything you can. Just make sure someone re-tests the whole thing once it’s back together.
That last clause is the part I skipped the first few times, and paid for. Individually-tested subagent output can still merge into something broken — a shared file gets touched twice, an interface drifts between two branches of work that never talked to each other — and the only way to catch that is a final pass that treats the merged result as a fresh unknown, not a sum of already-verified parts.
Where this actually earns its keep is in how much you can hand off once the goal is scoped this precisely. A few weeks ago I gave an agent a real-time voice chat feature to build — proximity-based audio, so voices fade as avatars drift apart — framed as “research the approach, then implement it, and don’t stop until it works end to end.” I walked away for dinner. Not a check-in, not a glance at the terminal — hours. When I came back, there was working progress: a real approach researched, implemented, and mostly functioning, rather than a pile of half-finished scaffolding waiting on me to unblock it.
What changed wasn’t the model. It was where I spent my attention — upfront, on making the goal checkable, instead of mid-flight, supervising every step. Once the target is well-defined, closing the gap becomes the agent’s problem, the same way a training run doesn’t need someone watching every epoch tick by. My job shifted from babysitting the process to scoping the goal well enough that babysitting stopped being the bottleneck. Set the goal before dinner. Check the diff after.
I don’t think this always works — some problems genuinely resist a clean metric, and I’ve had plenty of sessions since where I couldn’t articulate a loss function crisply enough and paid for it in aimless iteration, agents cheerfully producing motion without progress. But when I can name the number, the rest mostly takes care of itself, which is a strange thing to have learned from software, and a stranger thing still to have first learned from watching translation bots wait on each other in a room that never quite got going.