Point an Agent at Your Primitives
2026-07-09 · 6 min readI’ve been redesigning how agents interact with Codex, the Bible translation platform I work on, and I keep arriving at the same conclusion from different directions: the fewer tools you hand an agent, the more useful it becomes.
That sounds backwards. The instinct when you’re wiring an LLM up to your app is to give it a tool for everything — draft-cell, add-comment, flag-verse, assign-reviewer, one narrow function per feature, mirroring your UI one button at a time. I did this. It’s tempting because each tool is easy to write and easy to reason about in isolation. But a big tool list doesn’t compose. The agent has to guess which of forty near-synonymous actions you meant, the tools drift out of sync with the schema the moment someone ships a migration, and every new feature means another tool to write, document, and keep working. You end up maintaining two apps: the one your users touch and the one your agent touches, and they’re never quite the same shape.
What’s worked much better is collapsing almost everything down to a couple of general read and write tools, scoped to whatever permissions the calling user already has. The agent reads the state of a project the same way a human collaborator would, and writes to it the same way — no bespoke verb for every noun in the schema. If it needs to add a comment, it writes a comment record. If it needs to update a draft, it writes a draft record. The permission boundary does the safety work a narrow tool surface used to do, more reliably, because it’s the same boundary the rest of the app already respects. My rule of thumb these days: the fewer skills the better — that’s my life motto.
The same instinct showed up somewhere I didn’t expect: file import. Every translation partner who comes to Codex arrives with a different mess — some odd export from a decades-old tool, a spreadsheet with the columns in a different order than the last one, a zip of audio files with timestamps embedded in the filenames a different way each time. For a while I did what you’d expect: write a custom importer per partner, hand-map fields, ship it, wait for the next partner to show up with a format that didn’t fit any of my existing code. Individually reasonable, collectively absurd. Why would I build a custom importer for everybody — that’s so dumb. Nobody has time to maintain forty importers for forty file dialects that only exist because forty people exported data forty different ways.
So I stopped. Now the files just get dropped in, and an agent looks at them and writes the import logic on the fly, against the same general read/write primitives the rest of the system uses. It doesn’t need a parse-partner-x-format tool — just the raw files and a sense of what a well-formed record looks like on the other end. It can figure out the rest, which is exactly what I was doing by hand, just slower and with worse pattern recognition than a model that’s seen a thousand CSVs.
The moment this really clicked, though, came from a partner, not from my own code. He mentioned, almost in passing, that he’d basically stopped using the platform’s UI. Instead he tells his own agent: “you have the API key, solve this problem.” Whatever shape his data was in, whatever oddball task he needed done, he wasn’t waiting for a feature request to land on my roadmap — he was pointing an agent at the API and letting it write the one-off integration itself.
That reframed the importer problem for me. I’d been treating “build a custom importer per partner” as the cost of doing business, when the fix was to stop being the one who has to write the importer at all. If the app exposes its primitives — read, write, the actual shape of the data — over an API, then each partner’s own agent can write its own integration code, on demand, for exactly the file format it happens to have that day. Funnily enough, I’d already done the equivalent of this the same day, for a genuinely strange batch of audio, timestamps, and transcription files that didn’t map onto any importer I’d built — I just pointed an agent at the primitives and let it sort out the mapping. As long as the primitives are there in the app, if you can just point your agent at it and have it do stuff, that would be super cool. It already is.
There’s a broader shift underneath this that isn’t specific to translation tooling at all. A well-understood, genuinely low-complexity integration — say, auto-syncing translation progress to a partner’s own project-tracking board — used to sit in the backlog forever. Not because it was hard, exactly, but because “hard” was never the real constraint; the constraint was that it was tedious, one-off, and not obviously worth a sprint against everything else competing for the same week. A year ago I’d have said no to this integration. Today it’s an afternoon, because I’m not the one writing the glue code line by line — I’m reviewing what an agent wrote against primitives that already existed.
That’s the pattern I keep running into, in multi-agent translation and swarm-translation and everywhere else I’ve experimented with multi-agent setups: the leverage isn’t in building smarter individual tools, it’s in building fewer, more general ones and trusting an agent — the app’s, or increasingly, the partner’s own — to compose them into whatever the moment requires. The bespoke integration business was never a good business to be in. It just used to be the only option.