I’ve started thinking about AI-assisted translation the way I think about flying a plane. Not takeoff and landing — the boring middle part, where the whole job is thousands of tiny course corrections. A pilot doesn’t fly a thousand miles off heading and then swing back. They nudge, constantly, and each nudge changes the trajectory for everything that follows. That’s steering. The alternative — fly wherever the plane happens to go, then fix the path after landing — isn’t really an alternative, it’s a category error. But it’s more or less how most people still do AI-assisted translation: generate a draft, then post-edit it.
The generate-then-fix pattern is everywhere in machine translation tooling, and for good reason — it’s simple, it maps onto existing editorial workflows, and for a lot of content it’s fine. But I’ve become convinced it’s the wrong default for anything that depends on in-context learning, which is most of what I work on. Here’s the distinction I keep coming back to: when you correct a translation in context, before the model has moved on, that correction becomes part of the context for every subsequent prediction. When you batch-generate first and correct afterward, the correction lands in a document, not in the model’s working context for the next verse. It’s informative to a human reviewer. It’s close to invisible to the model.
Concretely: imagine you finish Genesis 1:1, and you corrected three things — a term choice, a syntactic pattern, an idiom. Am I gonna have to correct those same three things in every other verse where they show up? That would be crazy. If the correction is made in-context and the model sees it before generating Genesis 1:2, it doesn’t have to be. If it’s made afterward, in a separate editing pass, it does — you’ll be fixing the same idiom for the two-hundredth time in Deuteronomy, wondering why the model “isn’t learning.” It isn’t that the model can’t learn. It’s that batch editing never gave it the chance.
Whether this matters depends a lot on the stakes. For low-stakes content — app UI copy, a button label, something a user skims and moves past — generate-and-fix is genuinely fine. Nobody’s faith or family history hinges on a “Submit” button. But for high-trust, educational, or scriptural content, the calculus flips, because a fluent wrong answer is a different kind of failure than an obviously bad one. A garbled, clunky translation gets caught — it looks wrong, so someone checks it. A confidently fluent wrong translation doesn’t announce itself; it reads naturally and gets trusted. I’ve come to think a confident wrong answer is worse than no answer at all, and steering is partly a response to that asymmetry: it’s a way of keeping a human correction in the loop before fluency gets a chance to disguise an error, rather than trying to audit fluency after the fact. This is the same instinct behind treating translation quality as something that has to be actively converged on rather than assumed — I wrote about the social, consensus-driven version of that idea in Consensus Translation.
I have a real case that convinced me this isn’t just theory. On a full Bible translation project in Portuguese, we saw both patterns play out on the same underlying system. One workflow involved continuous correction — a translator reviewing and correcting essentially every prediction as it was generated, feeding straight back into the next one. Another workflow, used by some translators on the same team, was to batch-draft a chunk of verses first and edit them afterward, document-style. The batch-and-edit group came away convinced the AI wasn’t learning from their corrections — the same errors kept resurfacing chapter after chapter. But the model wasn’t failing to learn; it was never shown the correction at a moment when it could use it. It’s like flying a thousand miles and then trying to correct it, instead of correcting as you go — the destination doesn’t move just because you finally noticed you were off heading.
That case study is what pushed me toward a rougher framework I’ve been using since, mostly as a sequencing heuristic rather than anything rigorous: golden path and silver path. Golden path is for when someone’s actually watching — a human reviewing and correcting every single prediction in real time. Because that attention is expensive, you spend it on the material most worth learning from: the most distinctive, highest-coverage constructions, the ones whose corrections will generalize furthest downstream, translated first. Silver path is for when nobody’s watching as closely — the model runs ahead on material that’s comparatively easy or low-risk, and anything that looks uncertain gets flagged for a human to review later rather than corrected live. Golden path when someone’s watching, silver path when they’re not.
I don’t think this framework is settled — it’s closer to a rule of thumb I’ve been testing against real projects than a methodology I’d defend in the abstract, and I’d guess the boundary between “distinctive enough to deserve golden-path attention” and “safe enough for silver path” needs to get a lot more precise before this generalizes past the projects I’ve watched it work on. But it’s already changed how I sequence work, in the same spirit as the simplification argument in Occam’s Razor and Bible Translation: before adding more review infrastructure, ask whether the correction you’re about to make is actually reaching the next prediction, or just the next document.