#agentic-workflow

4 notes

Point an Agent at Your Primitives

2026-07-09 · 6 min read

I’ve been redesigning how agents interact with Codex, the Bible translation platform I work on, and I keep arriving at the same conclusion from different directions: the fewer tools you hand an agent, the more useful it becomes.

That sounds backwards. The instinct when you’re wiring an LLM up to your app is to give it a tool for everything — draft-cell, add-comment, flag-verse, assign-reviewer, one narrow function per feature, mirroring your UI one button at a time. I did this. It’s tempting because each tool is easy to write and easy to reason about in isolation. But a big tool list doesn’t compose. The agent has to guess which of forty near-synonymous actions you meant, the tools drift out of sync with the schema the moment someone ships a migration, and every new feature means another tool to write, document, and keep working. You end up maintaining two apps: the one your users touch and the one your agent touches, and they’re never quite the same shape.

What’s worked much better is collapsing almost everything down to a couple of general read and write tools, scoped to whatever permissions the calling user already has. The agent reads the state of a project the same way a human collaborator would, and writes to it the same way — no bespoke verb for every noun in the schema. If it needs to add a comment, it writes a comment record. If it needs to update a draft, it writes a draft record. The permission boundary does the safety work a narrow tool surface used to do, more reliably, because it’s the same boundary the rest of the app already respects. My rule of thumb these days: the fewer skills the better — that’s my life motto.

The same instinct showed up somewhere I didn’t expect: file import. Every translation partner who comes to Codex arrives with a different mess — some odd export from a decades-old tool, a spreadsheet with the columns in a different order than the last one, a zip of audio files with timestamps embedded in the filenames a different way each time. For a while I did what you’d expect: write a custom importer per partner, hand-map fields, ship it, wait for the next partner to show up with a format that didn’t fit any of my existing code. Individually reasonable, collectively absurd. Why would I build a custom importer for everybody — that’s so dumb. Nobody has time to maintain forty importers for forty file dialects that only exist because forty people exported data forty different ways.

So I stopped. Now the files just get dropped in, and an agent looks at them and writes the import logic on the fly, against the same general read/write primitives the rest of the system uses. It doesn’t need a parse-partner-x-format tool — just the raw files and a sense of what a well-formed record looks like on the other end. It can figure out the rest, which is exactly what I was doing by hand, just slower and with worse pattern recognition than a model that’s seen a thousand CSVs.

The moment this really clicked, though, came from a partner, not from my own code. He mentioned, almost in passing, that he’d basically stopped using the platform’s UI. Instead he tells his own agent: “you have the API key, solve this problem.” Whatever shape his data was in, whatever oddball task he needed done, he wasn’t waiting for a feature request to land on my roadmap — he was pointing an agent at the API and letting it write the one-off integration itself.

That reframed the importer problem for me. I’d been treating “build a custom importer per partner” as the cost of doing business, when the fix was to stop being the one who has to write the importer at all. If the app exposes its primitives — read, write, the actual shape of the data — over an API, then each partner’s own agent can write its own integration code, on demand, for exactly the file format it happens to have that day. Funnily enough, I’d already done the equivalent of this the same day, for a genuinely strange batch of audio, timestamps, and transcription files that didn’t map onto any importer I’d built — I just pointed an agent at the primitives and let it sort out the mapping. As long as the primitives are there in the app, if you can just point your agent at it and have it do stuff, that would be super cool. It already is.

There’s a broader shift underneath this that isn’t specific to translation tooling at all. A well-understood, genuinely low-complexity integration — say, auto-syncing translation progress to a partner’s own project-tracking board — used to sit in the backlog forever. Not because it was hard, exactly, but because “hard” was never the real constraint; the constraint was that it was tedious, one-off, and not obviously worth a sprint against everything else competing for the same week. A year ago I’d have said no to this integration. Today it’s an afternoon, because I’m not the one writing the glue code line by line — I’m reviewing what an agent wrote against primitives that already existed.

That’s the pattern I keep running into, in multi-agent translation and swarm-translation and everywhere else I’ve experimented with multi-agent setups: the leverage isn’t in building smarter individual tools, it’s in building fewer, more general ones and trusting an agent — the app’s, or increasingly, the partner’s own — to compose them into whatever the moment requires. The bespoke integration business was never a good business to be in. It just used to be the only option.

Permalink →

Steering, Not Batch-and-Correct

2026-07-08 · 6 min read

I’ve started thinking about AI-assisted translation the way I think about flying a plane. Not takeoff and landing — the boring middle part, where the whole job is thousands of tiny course corrections. A pilot doesn’t fly a thousand miles off heading and then swing back. They nudge, constantly, and each nudge changes the trajectory for everything that follows. That’s steering. The alternative — fly wherever the plane happens to go, then fix the path after landing — isn’t really an alternative, it’s a category error. But it’s more or less how most people still do AI-assisted translation: generate a draft, then post-edit it.

The generate-then-fix pattern is everywhere in machine translation tooling, and for good reason — it’s simple, it maps onto existing editorial workflows, and for a lot of content it’s fine. But I’ve become convinced it’s the wrong default for anything that depends on in-context learning, which is most of what I work on. Here’s the distinction I keep coming back to: when you correct a translation in context, before the model has moved on, that correction becomes part of the context for every subsequent prediction. When you batch-generate first and correct afterward, the correction lands in a document, not in the model’s working context for the next verse. It’s informative to a human reviewer. It’s close to invisible to the model.

Concretely: imagine you finish Genesis 1:1, and you corrected three things — a term choice, a syntactic pattern, an idiom. Am I gonna have to correct those same three things in every other verse where they show up? That would be crazy. If the correction is made in-context and the model sees it before generating Genesis 1:2, it doesn’t have to be. If it’s made afterward, in a separate editing pass, it does — you’ll be fixing the same idiom for the two-hundredth time in Deuteronomy, wondering why the model “isn’t learning.” It isn’t that the model can’t learn. It’s that batch editing never gave it the chance.

Whether this matters depends a lot on the stakes. For low-stakes content — app UI copy, a button label, something a user skims and moves past — generate-and-fix is genuinely fine. Nobody’s faith or family history hinges on a “Submit” button. But for high-trust, educational, or scriptural content, the calculus flips, because a fluent wrong answer is a different kind of failure than an obviously bad one. A garbled, clunky translation gets caught — it looks wrong, so someone checks it. A confidently fluent wrong translation doesn’t announce itself; it reads naturally and gets trusted. I’ve come to think a confident wrong answer is worse than no answer at all, and steering is partly a response to that asymmetry: it’s a way of keeping a human correction in the loop before fluency gets a chance to disguise an error, rather than trying to audit fluency after the fact. This is the same instinct behind treating translation quality as something that has to be actively converged on rather than assumed — I wrote about the social, consensus-driven version of that idea in Consensus Translation.

I have a real case that convinced me this isn’t just theory. On a full Bible translation project in Portuguese, we saw both patterns play out on the same underlying system. One workflow involved continuous correction — a translator reviewing and correcting essentially every prediction as it was generated, feeding straight back into the next one. Another workflow, used by some translators on the same team, was to batch-draft a chunk of verses first and edit them afterward, document-style. The batch-and-edit group came away convinced the AI wasn’t learning from their corrections — the same errors kept resurfacing chapter after chapter. But the model wasn’t failing to learn; it was never shown the correction at a moment when it could use it. It’s like flying a thousand miles and then trying to correct it, instead of correcting as you go — the destination doesn’t move just because you finally noticed you were off heading.

That case study is what pushed me toward a rougher framework I’ve been using since, mostly as a sequencing heuristic rather than anything rigorous: golden path and silver path. Golden path is for when someone’s actually watching — a human reviewing and correcting every single prediction in real time. Because that attention is expensive, you spend it on the material most worth learning from: the most distinctive, highest-coverage constructions, the ones whose corrections will generalize furthest downstream, translated first. Silver path is for when nobody’s watching as closely — the model runs ahead on material that’s comparatively easy or low-risk, and anything that looks uncertain gets flagged for a human to review later rather than corrected live. Golden path when someone’s watching, silver path when they’re not.

I don’t think this framework is settled — it’s closer to a rule of thumb I’ve been testing against real projects than a methodology I’d defend in the abstract, and I’d guess the boundary between “distinctive enough to deserve golden-path attention” and “safe enough for silver path” needs to get a lot more precise before this generalizes past the projects I’ve watched it work on. But it’s already changed how I sequence work, in the same spirit as the simplification argument in Occam’s Razor and Bible Translation: before adding more review infrastructure, ask whether the correction you’re about to make is actually reaching the next prediction, or just the next document.

Permalink →

Give the Agent a Loss Function

2026-07-04 · 5 min read

I’ve been running a lot of agentic coding sessions lately, and the sessions that go well have one thing in common: I gave the agent a number to hit, not a vibe to chase.

The bad version looks like “improve the audio pipeline” or “clean this up.” An agent handed that kind of instruction will do something — it has to, agents are compulsive doers — but it has no way to know when it’s done, so it stops when it runs out of ideas rather than when the problem is solved. The good version looks like “get audio round-tripping fidelity to 99%, and keep iterating against a test that measures it.” Same task, wildly different outcome, because now the agent has a number to close a gap on instead of a mood to satisfy.

This is training an ML model with extra steps, and once I noticed that I couldn’t unsee it. Training an ML model without a loss function doesn’t work — there’s no signal telling the optimizer which direction is “better,” so it just wanders. Prompting an agent without one doesn’t either, for the same reason: without a measurable target, “iterate until it’s good” quietly collapses into “stop when tired.” Give the agent a metric — a percentage, a test suite, a check it can run on its own output — and iteration turns into something closer to gradient descent: try, measure, adjust, repeat, until the number clears the bar.

I’d actually run into a version of this problem years before agentic coding was a category. The multi-agent translation simulation I built back in 2023 had a forward-translation bot, a back-translation bot, and an evaluation bot looping until “the translation draft stabilizes (if it does!)” — and in the demo, it never did, because none of the bots had a concrete stop condition, just each other to wait on. “Stabilized” was a vibe, not a number, so nothing ever stabilized. Different domain, same missing piece: no loss function, no convergence, no way to know the job was done.

That earlier miss is also where the swarm framing turns out to generalize better than I expected. Once each subagent in a swarm has its own measurable target, you can split work across a lot of them without losing the plot. The pattern I keep coming back to for coding specifically: spin up several cheaper, faster subagents to handle the parallelizable pieces of a task, and have each one test its own output before it reports back — a local loss function, one per subagent, scoped to its slice of the problem. Then an orchestrator agent merges everything and runs one more integration pass over the merged result, because a pile of locally-passing pieces isn’t automatically a working whole. Parallelize everything you can. Just make sure someone re-tests the whole thing once it’s back together.

That last clause is the part I skipped the first few times, and paid for. Individually-tested subagent output can still merge into something broken — a shared file gets touched twice, an interface drifts between two branches of work that never talked to each other — and the only way to catch that is a final pass that treats the merged result as a fresh unknown, not a sum of already-verified parts.

Where this actually earns its keep is in how much you can hand off once the goal is scoped this precisely. A few weeks ago I gave an agent a real-time voice chat feature to build — proximity-based audio, so voices fade as avatars drift apart — framed as “research the approach, then implement it, and don’t stop until it works end to end.” I walked away for dinner. Not a check-in, not a glance at the terminal — hours. When I came back, there was working progress: a real approach researched, implemented, and mostly functioning, rather than a pile of half-finished scaffolding waiting on me to unblock it.

What changed wasn’t the model. It was where I spent my attention — upfront, on making the goal checkable, instead of mid-flight, supervising every step. Once the target is well-defined, closing the gap becomes the agent’s problem, the same way a training run doesn’t need someone watching every epoch tick by. My job shifted from babysitting the process to scoping the goal well enough that babysitting stopped being the bottleneck. Set the goal before dinner. Check the diff after.

I don’t think this always works — some problems genuinely resist a clean metric, and I’ve had plenty of sessions since where I couldn’t articulate a loss function crisply enough and paid for it in aimless iteration, agents cheerfully producing motion without progress. But when I can name the number, the rest mostly takes care of itself, which is a strange thing to have learned from software, and a stranger thing still to have first learned from watching translation bots wait on each other in a room that never quite got going.

Permalink →

I Don't Write Code Anymore

2026-07-01 · 5 min read

I still open editors. I still read diffs, more carefully than I ever read my own code when I was the one typing it. But somewhere in the last year the actual verb changed underneath me. I don’t write code anymore. I make sure it works.

The shift didn’t happen all at once. It started with Copilot doing the boring parts — closing brackets, guessing the next line of a for-loop I’d already half-written in my head. That felt like a faster typist next to me, not a different kind of work. Then it was whole functions, whole files, whole features described in a paragraph and returned as a PR. At some point I noticed I was spending almost all of my time reading rather than writing, and that the reading was the actual job now, not a chore attached to it.

Here’s the part that took longer to admit: delegation only works if I understand the requirements and constraints well enough to notice when the output is wrong. Easy to skip in practice, especially when the code compiles, the tests pass, and the diff looks plausible. Plausible is not correct. If I’m fuzzy on what a function is actually supposed to guarantee — the edge cases, the invariants, the thing a teammate would ask about in review — I’m not verifying anything. I’m rubber-stamping a guess.

And the uncomfortable symmetry is: to the extent you’re confused or missing context, you’ll be like an AI model hallucinating. Not metaphorically. The same failure mode — filling a gap with something fluent and wrong because stopping to say “I don’t know” felt worse than producing an answer — applies to me approving a PR I only half understood as much as it applies to the model that wrote it. The model hallucinates when its context window is missing the constraint it needed. I hallucinate approval when my mental model is missing it too. Verification is only as good as the understanding behind it, and understanding doesn’t arrive for free just because you stopped typing.

So the job, as I’d have described it a week ago, became: stay close enough to the requirements to be a reliable check, not a rubber stamp. Read the code like I’m going to be paged about it at 2am, because I am. That felt like a tidy place to land — the kind of clean insight that makes a good closing paragraph.

It did not stay tidy.

Update, one week later: I want to complicate what I wrote above, because the version of “verification” I described a week ago already looks quaint.

I don’t write code directly anymore, full stop. What I actually do most days: watch a user demo something that’s broken, or describe it to me, and transcribe that into something an agent can act on. The agent extracts the tickets, works through them, opens PRs. I check whether the fix actually fixes the thing the user showed me. That’s most of my week now — not “review code an AI drafted” so much as “run a small verification desk for a coworker I never see write anything.”

The part I didn’t see coming: I no longer know how to onboard a human contributor into this. Explaining a task clearly enough for a person to pick it up — writing the ticket, answering follow-up questions, reviewing their PR style, waiting on their PR — is often just slower than having the agent do it end to end. Not slightly slower. Slower in a way that makes me hesitate before I even open the “assign to teammate” dropdown. I’m like, can I say, “Claude, don’t work on all the tickets, leave some for the people?” That doesn’t make any sense.

I can gesture at a reframe: maybe human contribution has to move up a level, toward defining problems well rather than solving them — deciding which demo is worth turning into a ticket, deciding what “fixed” even means for something ambiguous, holding the judgment calls the agent can’t make on its own. Probably true. It’s also easy to say and hard to build a team around, and I haven’t built it. I don’t have a good answer for what a junior engineer’s first month looks like when the fastest path from “user is annoyed” to “verified fix” doesn’t reliably pass through a human writing code anywhere in the middle.

I keep circling the same tension from the top of this post, one layer down. Verification only works if you understand the requirements — but if defining the problem well is now the whole job, I’m not sure verification was ever a separate skill from understanding, and I’m not sure I’ve figured out how to teach the second one to someone else. I don’t have a tidy close for this. It’s genuinely unresolved, and I’d rather say that than dress the reframe above up as a solution instead of a name for the gap.

This rhymes with the crowd-sourcing-plus-AI shape I sketched in The Path Forward — human judgment and AI throughput each doing the part they’re actually good at. And Translators’ Notes carries the same question in a different outfit: at what point does “reviewing AI output” quietly become the whole job, for a translator or an engineer, and what happens to everyone whose job used to be the part before that.

Permalink →