#product

3 notes

Point an Agent at Your Primitives

2026-07-09 · 6 min read

I’ve been redesigning how agents interact with Codex, the Bible translation platform I work on, and I keep arriving at the same conclusion from different directions: the fewer tools you hand an agent, the more useful it becomes.

That sounds backwards. The instinct when you’re wiring an LLM up to your app is to give it a tool for everything — draft-cell, add-comment, flag-verse, assign-reviewer, one narrow function per feature, mirroring your UI one button at a time. I did this. It’s tempting because each tool is easy to write and easy to reason about in isolation. But a big tool list doesn’t compose. The agent has to guess which of forty near-synonymous actions you meant, the tools drift out of sync with the schema the moment someone ships a migration, and every new feature means another tool to write, document, and keep working. You end up maintaining two apps: the one your users touch and the one your agent touches, and they’re never quite the same shape.

What’s worked much better is collapsing almost everything down to a couple of general read and write tools, scoped to whatever permissions the calling user already has. The agent reads the state of a project the same way a human collaborator would, and writes to it the same way — no bespoke verb for every noun in the schema. If it needs to add a comment, it writes a comment record. If it needs to update a draft, it writes a draft record. The permission boundary does the safety work a narrow tool surface used to do, more reliably, because it’s the same boundary the rest of the app already respects. My rule of thumb these days: the fewer skills the better — that’s my life motto.

The same instinct showed up somewhere I didn’t expect: file import. Every translation partner who comes to Codex arrives with a different mess — some odd export from a decades-old tool, a spreadsheet with the columns in a different order than the last one, a zip of audio files with timestamps embedded in the filenames a different way each time. For a while I did what you’d expect: write a custom importer per partner, hand-map fields, ship it, wait for the next partner to show up with a format that didn’t fit any of my existing code. Individually reasonable, collectively absurd. Why would I build a custom importer for everybody — that’s so dumb. Nobody has time to maintain forty importers for forty file dialects that only exist because forty people exported data forty different ways.

So I stopped. Now the files just get dropped in, and an agent looks at them and writes the import logic on the fly, against the same general read/write primitives the rest of the system uses. It doesn’t need a parse-partner-x-format tool — just the raw files and a sense of what a well-formed record looks like on the other end. It can figure out the rest, which is exactly what I was doing by hand, just slower and with worse pattern recognition than a model that’s seen a thousand CSVs.

The moment this really clicked, though, came from a partner, not from my own code. He mentioned, almost in passing, that he’d basically stopped using the platform’s UI. Instead he tells his own agent: “you have the API key, solve this problem.” Whatever shape his data was in, whatever oddball task he needed done, he wasn’t waiting for a feature request to land on my roadmap — he was pointing an agent at the API and letting it write the one-off integration itself.

That reframed the importer problem for me. I’d been treating “build a custom importer per partner” as the cost of doing business, when the fix was to stop being the one who has to write the importer at all. If the app exposes its primitives — read, write, the actual shape of the data — over an API, then each partner’s own agent can write its own integration code, on demand, for exactly the file format it happens to have that day. Funnily enough, I’d already done the equivalent of this the same day, for a genuinely strange batch of audio, timestamps, and transcription files that didn’t map onto any importer I’d built — I just pointed an agent at the primitives and let it sort out the mapping. As long as the primitives are there in the app, if you can just point your agent at it and have it do stuff, that would be super cool. It already is.

There’s a broader shift underneath this that isn’t specific to translation tooling at all. A well-understood, genuinely low-complexity integration — say, auto-syncing translation progress to a partner’s own project-tracking board — used to sit in the backlog forever. Not because it was hard, exactly, but because “hard” was never the real constraint; the constraint was that it was tedious, one-off, and not obviously worth a sprint against everything else competing for the same week. A year ago I’d have said no to this integration. Today it’s an afternoon, because I’m not the one writing the glue code line by line — I’m reviewing what an agent wrote against primitives that already existed.

That’s the pattern I keep running into, in multi-agent translation and swarm-translation and everywhere else I’ve experimented with multi-agent setups: the leverage isn’t in building smarter individual tools, it’s in building fewer, more general ones and trusting an agent — the app’s, or increasingly, the partner’s own — to compose them into whatever the moment requires. The bespoke integration business was never a good business to be in. It just used to be the only option.

Permalink →

Every demo I give starts with the same reflex from the audience: they want to see the model translate a hard verse well. That’s the “sparkle button” moment, and I used to think it was the whole pitch. Increasingly I don’t. A lot of the friction in a real translation project has nothing to do with translation quality at all — it’s project management. Who’s done Genesis. Who’s approved Exodus. Which reviewer has the current draft of Leviticus, and has anyone told them Numbers is ready. If we can point AI at that, it’s arguably more valuable than a better draft of any single verse, because it’s friction every project pays regardless of how good the translation itself is.

I don’t have a rigorous number for how large that share is, so I want to be careful here rather than confident. The figure I keep coming back to — heard from Reinier, who built Paratext, the software most of the field’s translation project management already runs on — puts project-management overhead at somewhere around 36% of total translation friction. I’m citing that because it matches everything I’ve seen firsthand, not because I’ve independently audited it, and I’d rather flag that than dress up a remembered conversation as a verified statistic.

Whatever the precise number, the shape of the claim matches what I’d expect from watching real projects. Translators lose time to status-chasing, handoff ambiguity, and duplicate work far more than they lose time to a model producing a mediocre first draft — a bad draft gets fixed in minutes; a lost handoff can stall a book for weeks. If that’s roughly right, it reframes where the leverage is. The sparkle button is a nice demo. The unglamorous version — an agent that knows who has what, what’s blocking what, and what needs a nudge — is the one that actually compounds across a whole Bible.

I have one real data point that makes this concrete rather than theoretical, and I want to present it more conservatively than the raw comparison suggests, because the comparison is easy to oversell. A team of two translators completed a full Bible translation in about a year. Sagamore Institute’s September 2022 report, A Study of Cost and Use of Funds in Bible Translation — commissioned by the Maclellan Foundation for illuminNations Resource Partners, based on self-reported data validated against audited financials — put the field average at roughly 15.8 years and $937,446 per complete Bible.

Even being conservative about it, that’s something like two, three, or four times faster than the field average, not fifteen times faster, and I think the smaller number is the honest one. The comparison isn’t apples to apples in several ways that matter: this was a high-resource language with existing reference translations to draw on, the two translators were already professionals, and “done” here means digital-ready, not print-ready — print typically adds another one and a half to two and a half years of typesetting, formatting, and final checks on top. A fifteen-year average almost certainly includes projects in far lower-resource languages, with far less existing scaffolding, and I don’t want to imply this team solved a problem that’s actually much harder in most of the places Bible translation happens.

What I do think the comparison shows is that the acceleration is real and worth taking seriously, even after you strip out every favorable condition you can find. And notably, neither of the two speed levers here — the model’s draft quality, or the reduced project-management overhead — did all the work alone. The team wasn’t just generating faster; per the pattern I described in Steering, Not Batch-and-Correct, they were correcting in context, which is itself a project-management behavior as much as a translation one — it’s about how work flows from one verse to the next, not just how good any single prediction is. I’d guess if I could actually decompose their year into “time saved on drafting” versus “time saved on not losing track of where things stood,” the second number would surprise people who assume this is purely a model-quality story.

None of this is a claim that Bible translation project management is a solved problem, or that 36% (or any other number) is the right way to carve it up. It’s more that I keep being drawn back to the boring part of the workflow as the place with the most unclaimed leverage, in a field where everyone — myself included, most days — wants to talk about the model instead.

Permalink →

#design

Here’s a quick reference list of questions to ask when designing a product, when you want to get to the core of your product, trim the fat, sort out the signal from the noise, etc.

VISCERAL IMPACT & DESTINY

  1. “What would make someone’s eyes light up when they first see this?”
  2. “How does this change people’s relationship with [the domain]?”
  3. “What limitation are people accepting as ‘just how it is’ that we could eliminate?”
  4. “If this succeeds wildly, what industry becomes obsolete?”

EXPERIENCE PERFECTIONISM

  1. “What moment in using this feels even slightly awkward or inelegant?”
  2. “Where are we making the user do something the computer should know?”
  3. “What detail would seem obsessive to include but users would love once they discover it?”
  4. “Which part of this experience isn’t bringing joy yet?”

PRODUCT TRUTH

  1. “What’s the ‘one thing’ this product is about that everything else serves?”
  2. “Which feature are we keeping because we’re afraid to remove it?”
  3. “What would this look like if we had to remove half the features?”
  4. “Where are we compromising on excellence to accommodate edge cases?”

REVOLUTIONARY POTENTIAL

  1. “How could this be 10x better, not just 2x better?”
  2. “What crazy idea for this product would be dismissed as impossible but would be magical if it worked?”
  3. “What would make the mainstream mock this at first but later seem obvious?”
  4. “How does this contribute to making technology more human?”
Permalink →