#engineering

6 notes

Point an Agent at Your Primitives

2026-07-09 · 6 min read

I’ve been redesigning how agents interact with Codex, the Bible translation platform I work on, and I keep arriving at the same conclusion from different directions: the fewer tools you hand an agent, the more useful it becomes.

That sounds backwards. The instinct when you’re wiring an LLM up to your app is to give it a tool for everything — draft-cell, add-comment, flag-verse, assign-reviewer, one narrow function per feature, mirroring your UI one button at a time. I did this. It’s tempting because each tool is easy to write and easy to reason about in isolation. But a big tool list doesn’t compose. The agent has to guess which of forty near-synonymous actions you meant, the tools drift out of sync with the schema the moment someone ships a migration, and every new feature means another tool to write, document, and keep working. You end up maintaining two apps: the one your users touch and the one your agent touches, and they’re never quite the same shape.

What’s worked much better is collapsing almost everything down to a couple of general read and write tools, scoped to whatever permissions the calling user already has. The agent reads the state of a project the same way a human collaborator would, and writes to it the same way — no bespoke verb for every noun in the schema. If it needs to add a comment, it writes a comment record. If it needs to update a draft, it writes a draft record. The permission boundary does the safety work a narrow tool surface used to do, more reliably, because it’s the same boundary the rest of the app already respects. My rule of thumb these days: the fewer skills the better — that’s my life motto.

The same instinct showed up somewhere I didn’t expect: file import. Every translation partner who comes to Codex arrives with a different mess — some odd export from a decades-old tool, a spreadsheet with the columns in a different order than the last one, a zip of audio files with timestamps embedded in the filenames a different way each time. For a while I did what you’d expect: write a custom importer per partner, hand-map fields, ship it, wait for the next partner to show up with a format that didn’t fit any of my existing code. Individually reasonable, collectively absurd. Why would I build a custom importer for everybody — that’s so dumb. Nobody has time to maintain forty importers for forty file dialects that only exist because forty people exported data forty different ways.

So I stopped. Now the files just get dropped in, and an agent looks at them and writes the import logic on the fly, against the same general read/write primitives the rest of the system uses. It doesn’t need a parse-partner-x-format tool — just the raw files and a sense of what a well-formed record looks like on the other end. It can figure out the rest, which is exactly what I was doing by hand, just slower and with worse pattern recognition than a model that’s seen a thousand CSVs.

The moment this really clicked, though, came from a partner, not from my own code. He mentioned, almost in passing, that he’d basically stopped using the platform’s UI. Instead he tells his own agent: “you have the API key, solve this problem.” Whatever shape his data was in, whatever oddball task he needed done, he wasn’t waiting for a feature request to land on my roadmap — he was pointing an agent at the API and letting it write the one-off integration itself.

That reframed the importer problem for me. I’d been treating “build a custom importer per partner” as the cost of doing business, when the fix was to stop being the one who has to write the importer at all. If the app exposes its primitives — read, write, the actual shape of the data — over an API, then each partner’s own agent can write its own integration code, on demand, for exactly the file format it happens to have that day. Funnily enough, I’d already done the equivalent of this the same day, for a genuinely strange batch of audio, timestamps, and transcription files that didn’t map onto any importer I’d built — I just pointed an agent at the primitives and let it sort out the mapping. As long as the primitives are there in the app, if you can just point your agent at it and have it do stuff, that would be super cool. It already is.

There’s a broader shift underneath this that isn’t specific to translation tooling at all. A well-understood, genuinely low-complexity integration — say, auto-syncing translation progress to a partner’s own project-tracking board — used to sit in the backlog forever. Not because it was hard, exactly, but because “hard” was never the real constraint; the constraint was that it was tedious, one-off, and not obviously worth a sprint against everything else competing for the same week. A year ago I’d have said no to this integration. Today it’s an afternoon, because I’m not the one writing the glue code line by line — I’m reviewing what an agent wrote against primitives that already existed.

That’s the pattern I keep running into, in multi-agent translation and swarm-translation and everywhere else I’ve experimented with multi-agent setups: the leverage isn’t in building smarter individual tools, it’s in building fewer, more general ones and trusting an agent — the app’s, or increasingly, the partner’s own — to compose them into whatever the moment requires. The bespoke integration business was never a good business to be in. It just used to be the only option.

Permalink →

The Pipeline That Wrote This Post

2026-07-08 · 4 min read

I want to make a slightly awkward claim up front: I didn’t sit down and decide to write this post. A pipeline on my laptop decided it was worth writing, and I mostly agreed.

Here’s the setup. I have an on-device transcriber that arms itself passively whenever the microphone is in use for anything else — a call, a voice memo, a meeting, doesn’t matter what app. It listens, transcribes locally, and then does something I care about more than the transcription itself: it discards the raw audio. No recording survives. What gets saved is the text, dropped into Notes, and that’s it. If you want to know what was said, you get a transcript. If you want to hear my voice saying it, you’re out of luck, because that copy never existed past the few seconds it took to turn sound into words.

That’s step one, and it runs quietly, all day, without me thinking about it.

Step two is where it gets more interesting, and slightly stranger. A scheduled job — a Claude automation that fires once a day — reads that day’s transcripts and goes looking for anything shareable: an anecdote, a lesson, a half-formed idea I said out loud in a meeting and immediately forgot. It doesn’t try to write anything polished. It produces a set of short, structured candidate notes, one per idea, each tagged with a letter and led with a hook — the kind of thing you’d pitch yourself in one sentence before deciding whether it’s worth the next thirty minutes. Some days there are none. Some days there are six, and four of them are bad.

Step three is me. I read the day’s candidates the way I’d skim a slush pile — fast, unsentimental, looking for the one or two that still sound interesting once the meeting adrenaline has worn off. Most get discarded. A few get flagged.

Step four is an AI agent taking a flagged candidate and turning it into an actual draft — structure, examples, the whole thing. This post is a step-four output. This post came out of it.

I’ll admit there’s something faintly unsettling about a system that’s quietly always listening for your own good ideas. Not unsettling in a surveillance sense — I built the privacy boundary specifically so it wouldn’t be that — but in the sense that it changes how I relate to my own stray thoughts. I used to lose most of them. Now there’s a nonzero chance that something I mumbled in a standup gets mined, structured, and handed back to me a day later as “here’s a hook, do you want this.” It’s a strange kind of externalized memory, closer to what I was reaching for in Memorize Anything with AI Music — offloading recall onto a system so the interesting part of my brain doesn’t have to hold onto it — except this one runs on everything I say, not just what I deliberately feed it.

The audio-discard choice is the part I’d defend hardest if someone pushed back on this. It would be trivial to keep the recordings — better transcription accuracy, re-processing later, all the usual arguments. I didn’t want any of that badly enough to keep a standing archive of my own and other people’s voices sitting on disk. Text is legible, editable, and it forgets the timbre, the room, the identifiable stuff. That’s not a limitation I haven’t gotten around to fixing. It’s the actual design, the same instinct that shows up in how I’d want any multi-agent system — like the ones in multi-agent translation — to be legible about what it’s doing with your input, not just capable.

I don’t fully know yet whether mining my own days for content is a good habit or a slightly cursed one. I’m going to keep running it and see what it thinks is worth saying next.

Permalink →

Give the Agent a Loss Function

2026-07-04 · 5 min read

I’ve been running a lot of agentic coding sessions lately, and the sessions that go well have one thing in common: I gave the agent a number to hit, not a vibe to chase.

The bad version looks like “improve the audio pipeline” or “clean this up.” An agent handed that kind of instruction will do something — it has to, agents are compulsive doers — but it has no way to know when it’s done, so it stops when it runs out of ideas rather than when the problem is solved. The good version looks like “get audio round-tripping fidelity to 99%, and keep iterating against a test that measures it.” Same task, wildly different outcome, because now the agent has a number to close a gap on instead of a mood to satisfy.

This is training an ML model with extra steps, and once I noticed that I couldn’t unsee it. Training an ML model without a loss function doesn’t work — there’s no signal telling the optimizer which direction is “better,” so it just wanders. Prompting an agent without one doesn’t either, for the same reason: without a measurable target, “iterate until it’s good” quietly collapses into “stop when tired.” Give the agent a metric — a percentage, a test suite, a check it can run on its own output — and iteration turns into something closer to gradient descent: try, measure, adjust, repeat, until the number clears the bar.

I’d actually run into a version of this problem years before agentic coding was a category. The multi-agent translation simulation I built back in 2023 had a forward-translation bot, a back-translation bot, and an evaluation bot looping until “the translation draft stabilizes (if it does!)” — and in the demo, it never did, because none of the bots had a concrete stop condition, just each other to wait on. “Stabilized” was a vibe, not a number, so nothing ever stabilized. Different domain, same missing piece: no loss function, no convergence, no way to know the job was done.

That earlier miss is also where the swarm framing turns out to generalize better than I expected. Once each subagent in a swarm has its own measurable target, you can split work across a lot of them without losing the plot. The pattern I keep coming back to for coding specifically: spin up several cheaper, faster subagents to handle the parallelizable pieces of a task, and have each one test its own output before it reports back — a local loss function, one per subagent, scoped to its slice of the problem. Then an orchestrator agent merges everything and runs one more integration pass over the merged result, because a pile of locally-passing pieces isn’t automatically a working whole. Parallelize everything you can. Just make sure someone re-tests the whole thing once it’s back together.

That last clause is the part I skipped the first few times, and paid for. Individually-tested subagent output can still merge into something broken — a shared file gets touched twice, an interface drifts between two branches of work that never talked to each other — and the only way to catch that is a final pass that treats the merged result as a fresh unknown, not a sum of already-verified parts.

Where this actually earns its keep is in how much you can hand off once the goal is scoped this precisely. A few weeks ago I gave an agent a real-time voice chat feature to build — proximity-based audio, so voices fade as avatars drift apart — framed as “research the approach, then implement it, and don’t stop until it works end to end.” I walked away for dinner. Not a check-in, not a glance at the terminal — hours. When I came back, there was working progress: a real approach researched, implemented, and mostly functioning, rather than a pile of half-finished scaffolding waiting on me to unblock it.

What changed wasn’t the model. It was where I spent my attention — upfront, on making the goal checkable, instead of mid-flight, supervising every step. Once the target is well-defined, closing the gap becomes the agent’s problem, the same way a training run doesn’t need someone watching every epoch tick by. My job shifted from babysitting the process to scoping the goal well enough that babysitting stopped being the bottleneck. Set the goal before dinner. Check the diff after.

I don’t think this always works — some problems genuinely resist a clean metric, and I’ve had plenty of sessions since where I couldn’t articulate a loss function crisply enough and paid for it in aimless iteration, agents cheerfully producing motion without progress. But when I can name the number, the rest mostly takes care of itself, which is a strange thing to have learned from software, and a stranger thing still to have first learned from watching translation bots wait on each other in a room that never quite got going.

Permalink →

I Don't Write Code Anymore

2026-07-01 · 5 min read

I still open editors. I still read diffs, more carefully than I ever read my own code when I was the one typing it. But somewhere in the last year the actual verb changed underneath me. I don’t write code anymore. I make sure it works.

The shift didn’t happen all at once. It started with Copilot doing the boring parts — closing brackets, guessing the next line of a for-loop I’d already half-written in my head. That felt like a faster typist next to me, not a different kind of work. Then it was whole functions, whole files, whole features described in a paragraph and returned as a PR. At some point I noticed I was spending almost all of my time reading rather than writing, and that the reading was the actual job now, not a chore attached to it.

Here’s the part that took longer to admit: delegation only works if I understand the requirements and constraints well enough to notice when the output is wrong. Easy to skip in practice, especially when the code compiles, the tests pass, and the diff looks plausible. Plausible is not correct. If I’m fuzzy on what a function is actually supposed to guarantee — the edge cases, the invariants, the thing a teammate would ask about in review — I’m not verifying anything. I’m rubber-stamping a guess.

And the uncomfortable symmetry is: to the extent you’re confused or missing context, you’ll be like an AI model hallucinating. Not metaphorically. The same failure mode — filling a gap with something fluent and wrong because stopping to say “I don’t know” felt worse than producing an answer — applies to me approving a PR I only half understood as much as it applies to the model that wrote it. The model hallucinates when its context window is missing the constraint it needed. I hallucinate approval when my mental model is missing it too. Verification is only as good as the understanding behind it, and understanding doesn’t arrive for free just because you stopped typing.

So the job, as I’d have described it a week ago, became: stay close enough to the requirements to be a reliable check, not a rubber stamp. Read the code like I’m going to be paged about it at 2am, because I am. That felt like a tidy place to land — the kind of clean insight that makes a good closing paragraph.

It did not stay tidy.

Update, one week later: I want to complicate what I wrote above, because the version of “verification” I described a week ago already looks quaint.

I don’t write code directly anymore, full stop. What I actually do most days: watch a user demo something that’s broken, or describe it to me, and transcribe that into something an agent can act on. The agent extracts the tickets, works through them, opens PRs. I check whether the fix actually fixes the thing the user showed me. That’s most of my week now — not “review code an AI drafted” so much as “run a small verification desk for a coworker I never see write anything.”

The part I didn’t see coming: I no longer know how to onboard a human contributor into this. Explaining a task clearly enough for a person to pick it up — writing the ticket, answering follow-up questions, reviewing their PR style, waiting on their PR — is often just slower than having the agent do it end to end. Not slightly slower. Slower in a way that makes me hesitate before I even open the “assign to teammate” dropdown. I’m like, can I say, “Claude, don’t work on all the tickets, leave some for the people?” That doesn’t make any sense.

I can gesture at a reframe: maybe human contribution has to move up a level, toward defining problems well rather than solving them — deciding which demo is worth turning into a ticket, deciding what “fixed” even means for something ambiguous, holding the judgment calls the agent can’t make on its own. Probably true. It’s also easy to say and hard to build a team around, and I haven’t built it. I don’t have a good answer for what a junior engineer’s first month looks like when the fastest path from “user is annoyed” to “verified fix” doesn’t reliably pass through a human writing code anywhere in the middle.

I keep circling the same tension from the top of this post, one layer down. Verification only works if you understand the requirements — but if defining the problem well is now the whole job, I’m not sure verification was ever a separate skill from understanding, and I’m not sure I’ve figured out how to teach the second one to someone else. I don’t have a tidy close for this. It’s genuinely unresolved, and I’d rather say that than dress the reframe above up as a solution instead of a name for the gap.

This rhymes with the crowd-sourcing-plus-AI shape I sketched in The Path Forward — human judgment and AI throughput each doing the part they’re actually good at. And Translators’ Notes carries the same question in a different outfit: at what point does “reviewing AI output” quietly become the whole job, for a translator or an engineer, and what happens to everyone whose job used to be the part before that.

Permalink →

I’ve spent most of my career assuming open source is a default good — you build something, you put the code out, the community makes it better, everyone wins. I still believe that most of the time. But two things have been nagging at me lately, and I don’t think I can resolve them cleanly, so I’m writing them down instead.

Argument one: you’re publishing a map

The classic security argument for open source is “many eyes make bugs shallow” — enough people read the code, someone spots the flaw before an attacker does. That argument was always a little optimistic (plenty of famous vulnerabilities sat in plain sight in popular open-source libraries for years before anyone noticed), but it’s gotten shakier now that “reading the code” is something an AI agent can do at scale, continuously, for free, on every public repo it can find.

An attacker no longer has to be a person who understands your codebase. They can point an agent at your repo and ask it to reason about your auth flow, your input validation, your dependency graph, and your deploy config, and get back a prioritized list of soft spots — the same kind of architectural reasoning a senior engineer would do on a code review, except it scales to every repo on GitHub simultaneously and never gets bored. Open-sourcing your code doesn’t just risk a vulnerability being found; it hands over the reconnaissance step for free. As I’ve put it to myself while deciding what to open up: it just hands an AI a map of where you’re weak.

That’s part of why the auth and user-data module of one of my own projects has stayed closed-source, even from early on, well before “AI agent reads your repo” was a mainstream worry. The rest of the codebase — the interesting, differentiated parts — I’ve been comfortable sharing. But the module that decides who gets in and what happens to people’s data felt like a bad place to hand out a floor plan, and I still think that instinct was right, even if the reasoning behind it (attackers get smarter tools) has changed since I first had it.

Argument two: code without a business model dies anyway

The second argument is less about attackers and more about attrition. Plenty of genuinely good, well-maintained open-source projects still fold — not because the code was bad or the maintainers stopped caring, but because nothing was paying for the maintainers’ time. Goodwill is real, but it’s not a renewable resource at the rate a serious project needs it. Sooner or later the project dies anyway without a business model behind it, openness or no openness.

If that’s true, the security argument above is almost moot for a lot of projects — the vulnerability never gets exploited because the maintainers already walked away and the repo is quietly rotting in an unmaintained state, which is arguably worse than either alternative.

The tension I can’t fully resolve

Here’s what makes this uncomfortable for me specifically: I’ve argued at length that CC0 — zero restrictions, no strings attached — is the right license for openly licensed Bible texts, precisely because openness removes friction, invites local ownership, and lets communities build sustainably on top of a shared resource. I still believe that. I’ve also written about wanting Bible translation to move toward permissionless, decentralized tooling built and extended by the communities using it, not gatekept by a central authority.

So why would I turn around and be cagier about code than I am about content? Part of the answer might just be that content and code fail differently. A Bible text doesn’t have an attack surface. Nobody can use a translated verse to compromise a server or exfiltrate user data. The “many eyes” argument for content is basically just quality and reach — more translators, more dialects, more reviewers — with none of the security downside, because there’s no exploit sitting in a psalm. Code is different: openness and attack surface aren’t separable the way openness and translation quality are. Maybe that’s the whole answer and I’m overthinking it. Or maybe it’s a rationalization I’ve backed into because I already had a closed-source module and needed a reason.

I don’t have a tidy rule here, and I’m suspicious of anyone who does. If you’ve found a decision framework that actually holds up — something more principled than “vibes about which module feels sensitive” — I’d like to hear it, because right now I’m mostly running on instinct and a lingering unease about what an agent could do with a full read of my repo.

Permalink →

Somewhere around 2010, deep into a master’s degree in Ancient Greek linguistics, I was trying to build word2vec models over a Greek corpus using Gensim, which was — and largely still is — the workhorse library for this kind of thing. Gensim’s whole pitch was memory efficiency: instead of loading a corpus into RAM and training a model on the full thing in one pass, you streamed documents through it one at a time. The model updated incrementally as the data arrived, rather than requiring the entire corpus to exist in memory simultaneously before training could start.

At the time this felt like a plumbing detail, not an idea. But it planted a question I didn’t know I was carrying around for the next decade and a half: why train a model? Why don’t we figure out a streaming approach where we just give it the data it needs on the fly? Ancient Greek is a low-resource language by any modern standard — there wasn’t enough data to justify pretraining a dedicated model in the first place, and streaming was mostly a practical concession to a laptop that couldn’t hold the corpus in memory. But the concession turned out to be more interesting than the workaround it was standing in for.

That question is, more or less, how the translation tool I’ve been building handles languages it has never seen. It doesn’t train a per-language model. There’s no fine-tuning step, no “we support 40 languages, request the 41st.” Instead, at prediction time, it streams relevant translated examples to the model as context and asks it to produce a translation using those examples as its only local evidence for how the language behaves. The model isn’t learning the language in any durable sense; it’s being handed exactly the data it needs, on the fly, for this one prediction, and then that context evaporates. This is the same move as Gensim’s streaming corpus reader, just moved from “how do I fit training data in memory” to “how do I give a frozen model enough signal to be useful in a language it’s never encountered.”

The practical difference is enormous. Training or fine-tuning a model per language — even a small one — is the kind of thing that will bog your computer down for hours and then crash, especially if you’re iterating on ultra-low-resource languages where you don’t have the corpus size to justify the cost in the first place. Streaming relevant examples at inference time is orders of magnitude faster, because you’ve swapped a training job for a retrieval-and-context-assembly step. It’s the same logic I was chasing in Minimal Translation Memory for AI Agents — find the smallest, most generalizable set of examples that covers the most ground, and hand the model only that, rather than everything you have.

I want to be honest about where this breaks down, though, because it’s a real limitation and not a rounding error. The streaming approach works today because a human is validating each prediction live — approving, correcting, or rejecting the model’s output as it’s generated. That human-in-the-loop step is quietly doing a lot of work: it’s the error-correction mechanism that a trained model would otherwise have baked in through gradient descent over many examples. Remove the human, and you no longer have a guarantee that streamed context is sufficient on its own to produce a reliable translation. Whether streaming-plus-retrieval can substitute for training without a validator in the loop is, as far as I can tell, an open research problem — not a solved one I’m just being modest about. I’d put it in the same bucket as the clustering evaluation problem I never resolved in Contextual Vectors for Lexical Meaning: a piece I know is missing, not a piece I’ve quietly decided doesn’t matter.

What convinces me this is a real pattern, rather than a translation-specific hack I’ve talked myself into, is that I’ve watched the identical lesson surface somewhere completely unrelated. Years ago, I was on the other side of a job interview where a candidate was asked to process seventy million XML records. The naive approach — load everything, validate everything — didn’t just run slowly, it fell over. The fix was to stream: validate records as they arrived, discard what you didn’t need to keep, never hold the whole seventy million in memory at once. It’s the same shape as few-shot prompting, just wearing different clothes: you don’t dump your entire knowledge base into a context window and hope the model sorts it out, you stream the handful of examples that are actually relevant to this prediction and let the rest stay where it lives.

That’s one of maybe four lessons a developer has to learn, and I seem to have learned it twice, fifteen years apart, in two fields that have nothing to do with each other. I don’t think that’s a coincidence so much as a sign that “stream the relevant slice instead of loading the whole thing” is one of those ideas that’s basically fractal — it shows up at the level of a corpus reader, a translation model, an XML validator, and a prompt, and it’s probably going to keep showing up wherever the next constraint is memory, time, or data scarcity.

Permalink →