#methodology

12 notes

In a room of around forty consultants, one I respect raised five objections to “AI in Bible translation” in the space of about ten minutes: it can’t use the Greek and Hebrew, it has no dictionary, it doesn’t know genre, you can’t give it a translator’s prompt, and you can’t feed it small, controllable batches of text. Every one of those is a fair complaint. None of them is about the tool we’re actually using.

They’re objections to Google Translate.

Seq2seq machine translation—the architecture behind Google Translate and its era-mates—really did work that way. It mapped strings to strings with no access to source-language grammar as an object you could reason about, no lexical resources beyond whatever got baked into training, no sense of genre or register beyond statistical correlation, and no mechanism for a human to say “translate this the way we discussed last week” and have it stick. If that’s the AI you have in mind, the objections aren’t just fair, they’re the whole story. I’ve spent enough time watching seq2seq output wander off-register mid-sentence to know the scar tissue is earned.

But a modern LLM can read the Hebrew, hold a conversation about a difficult word’s semantic range, look up a lexicon entry you hand it in context, take a translator’s brief as a literal system prompt, and work in exactly the batch size a team wants—a verse, a paragraph, a pericope. It can be shown the genre and asked to translate accordingly. None of this is hypothetical; it’s the ordinary operating mode of the tools I use most weeks. So when I hear the same five objections raised against “AI,” I want to redirect: every objection I hear to “AI in Bible translation” is actually an objection to Google Translate. Which means the objection is sound, but the target has moved.

I don’t think this makes the skepticism illegitimate. Consultants built their caution around real failure modes, over years of watching tools get chosen from the top down and pushed into low-resource languages even after the field had already seen them produce fluent-sounding nonsense. That caution is a feature of the field, not a bug to route around. The useful move isn’t to say “you’re wrong, AI is great now”—it’s to ask which specific objection still holds against a system that can read source text, hold a controllable-length conversation, and be handed a prompt like any other collaborator. Some of them, updated for the new architecture, still land. Most of the five I listed above don’t survive contact with what an LLM can actually do, though—which is a separate question from whether the output is good, a question I’ve tried to be more precise about before.

That precision question is where I get more cautious, not less. There’s a pull toward settling all of this with a rigorous benchmark—controlled comparison, blinded reviewers, speed and quality scored against a fixed rubric. I understand the appeal; it’s the responsible-sounding thing to do. But I’ve come to distrust it for a specific reason: by the time you finish the rigorous benchmark, the thing you benchmarked is obsolete. The workflow, the model, the prompting approach—all of it will have shifted meaningfully in the months a careful study takes to design, run, and write up. What you publish is a precise measurement of a tool nobody is using anymore, dressed up in the authority of a number. I’d rather have a translator’s rough, unrigorous sense of “this draft used to take me a day and a half, now it takes an afternoon” than a beautifully controlled study of last year’s pipeline. It’s less defensible on paper and more true in practice—false precision is still false.

That same instinct toward “faster at what, exactly” applies to the bigger debate hovering over all of this: is AI-assisted translation faster than the traditional process? I’ve started to think that’s the wrong question, or at least an underspecified one. Traditional translation and AI-assisted translation quite often aren’t racing toward the same finish line. A team producing a digital-only translation for a community that needs Scripture in an app this year has a different goal than a team producing a fully checked, print-ready Bible meant to anchor a language’s translation for a generation. Comparing their speed head-to-head is a bit like timing a sprinter against a marathoner and declaring one of them slow. “Is AI faster than traditional translation?” is the wrong question. Faster at what—and toward whose goal? This is the same move I made in thinking through Occam’s razor and translation: the interesting question usually isn’t which process is objectively better, it’s what inputs and outputs a given goal actually requires, and whether the process in front of you is shaped for that goal or just inherited from a different one.

I don’t think any of this settles the underlying argument about whether AI belongs in Bible translation workflows at all—that’s a longer conversation, and a more contested one than I’m trying to resolve here. But I’d like the argument to be about the tool that’s actually in the room. Google Translate earned its objections. The thing translators are using now is a different animal, and it deserves to be evaluated as one—carefully, skeptically, and on its own terms, rather than refuted by proxy.

Permalink →

Every demo I give starts with the same reflex from the audience: they want to see the model translate a hard verse well. That’s the “sparkle button” moment, and I used to think it was the whole pitch. Increasingly I don’t. A lot of the friction in a real translation project has nothing to do with translation quality at all — it’s project management. Who’s done Genesis. Who’s approved Exodus. Which reviewer has the current draft of Leviticus, and has anyone told them Numbers is ready. If we can point AI at that, it’s arguably more valuable than a better draft of any single verse, because it’s friction every project pays regardless of how good the translation itself is.

I don’t have a rigorous number for how large that share is, so I want to be careful here rather than confident. The figure I keep coming back to — heard from Reinier, who built Paratext, the software most of the field’s translation project management already runs on — puts project-management overhead at somewhere around 36% of total translation friction. I’m citing that because it matches everything I’ve seen firsthand, not because I’ve independently audited it, and I’d rather flag that than dress up a remembered conversation as a verified statistic.

Whatever the precise number, the shape of the claim matches what I’d expect from watching real projects. Translators lose time to status-chasing, handoff ambiguity, and duplicate work far more than they lose time to a model producing a mediocre first draft — a bad draft gets fixed in minutes; a lost handoff can stall a book for weeks. If that’s roughly right, it reframes where the leverage is. The sparkle button is a nice demo. The unglamorous version — an agent that knows who has what, what’s blocking what, and what needs a nudge — is the one that actually compounds across a whole Bible.

I have one real data point that makes this concrete rather than theoretical, and I want to present it more conservatively than the raw comparison suggests, because the comparison is easy to oversell. A team of two translators completed a full Bible translation in about a year. Sagamore Institute’s September 2022 report, A Study of Cost and Use of Funds in Bible Translation — commissioned by the Maclellan Foundation for illuminNations Resource Partners, based on self-reported data validated against audited financials — put the field average at roughly 15.8 years and $937,446 per complete Bible.

Even being conservative about it, that’s something like two, three, or four times faster than the field average, not fifteen times faster, and I think the smaller number is the honest one. The comparison isn’t apples to apples in several ways that matter: this was a high-resource language with existing reference translations to draw on, the two translators were already professionals, and “done” here means digital-ready, not print-ready — print typically adds another one and a half to two and a half years of typesetting, formatting, and final checks on top. A fifteen-year average almost certainly includes projects in far lower-resource languages, with far less existing scaffolding, and I don’t want to imply this team solved a problem that’s actually much harder in most of the places Bible translation happens.

What I do think the comparison shows is that the acceleration is real and worth taking seriously, even after you strip out every favorable condition you can find. And notably, neither of the two speed levers here — the model’s draft quality, or the reduced project-management overhead — did all the work alone. The team wasn’t just generating faster; per the pattern I described in Steering, Not Batch-and-Correct, they were correcting in context, which is itself a project-management behavior as much as a translation one — it’s about how work flows from one verse to the next, not just how good any single prediction is. I’d guess if I could actually decompose their year into “time saved on drafting” versus “time saved on not losing track of where things stood,” the second number would surprise people who assume this is purely a model-quality story.

None of this is a claim that Bible translation project management is a solved problem, or that 36% (or any other number) is the right way to carve it up. It’s more that I keep being drawn back to the boring part of the workflow as the place with the most unclaimed leverage, in a field where everyone — myself included, most days — wants to talk about the model instead.

Permalink →

Steering, Not Batch-and-Correct

2026-07-08 · 6 min read

I’ve started thinking about AI-assisted translation the way I think about flying a plane. Not takeoff and landing — the boring middle part, where the whole job is thousands of tiny course corrections. A pilot doesn’t fly a thousand miles off heading and then swing back. They nudge, constantly, and each nudge changes the trajectory for everything that follows. That’s steering. The alternative — fly wherever the plane happens to go, then fix the path after landing — isn’t really an alternative, it’s a category error. But it’s more or less how most people still do AI-assisted translation: generate a draft, then post-edit it.

The generate-then-fix pattern is everywhere in machine translation tooling, and for good reason — it’s simple, it maps onto existing editorial workflows, and for a lot of content it’s fine. But I’ve become convinced it’s the wrong default for anything that depends on in-context learning, which is most of what I work on. Here’s the distinction I keep coming back to: when you correct a translation in context, before the model has moved on, that correction becomes part of the context for every subsequent prediction. When you batch-generate first and correct afterward, the correction lands in a document, not in the model’s working context for the next verse. It’s informative to a human reviewer. It’s close to invisible to the model.

Concretely: imagine you finish Genesis 1:1, and you corrected three things — a term choice, a syntactic pattern, an idiom. Am I gonna have to correct those same three things in every other verse where they show up? That would be crazy. If the correction is made in-context and the model sees it before generating Genesis 1:2, it doesn’t have to be. If it’s made afterward, in a separate editing pass, it does — you’ll be fixing the same idiom for the two-hundredth time in Deuteronomy, wondering why the model “isn’t learning.” It isn’t that the model can’t learn. It’s that batch editing never gave it the chance.

Whether this matters depends a lot on the stakes. For low-stakes content — app UI copy, a button label, something a user skims and moves past — generate-and-fix is genuinely fine. Nobody’s faith or family history hinges on a “Submit” button. But for high-trust, educational, or scriptural content, the calculus flips, because a fluent wrong answer is a different kind of failure than an obviously bad one. A garbled, clunky translation gets caught — it looks wrong, so someone checks it. A confidently fluent wrong translation doesn’t announce itself; it reads naturally and gets trusted. I’ve come to think a confident wrong answer is worse than no answer at all, and steering is partly a response to that asymmetry: it’s a way of keeping a human correction in the loop before fluency gets a chance to disguise an error, rather than trying to audit fluency after the fact. This is the same instinct behind treating translation quality as something that has to be actively converged on rather than assumed — I wrote about the social, consensus-driven version of that idea in Consensus Translation.

I have a real case that convinced me this isn’t just theory. On a full Bible translation project in Portuguese, we saw both patterns play out on the same underlying system. One workflow involved continuous correction — a translator reviewing and correcting essentially every prediction as it was generated, feeding straight back into the next one. Another workflow, used by some translators on the same team, was to batch-draft a chunk of verses first and edit them afterward, document-style. The batch-and-edit group came away convinced the AI wasn’t learning from their corrections — the same errors kept resurfacing chapter after chapter. But the model wasn’t failing to learn; it was never shown the correction at a moment when it could use it. It’s like flying a thousand miles and then trying to correct it, instead of correcting as you go — the destination doesn’t move just because you finally noticed you were off heading.

That case study is what pushed me toward a rougher framework I’ve been using since, mostly as a sequencing heuristic rather than anything rigorous: golden path and silver path. Golden path is for when someone’s actually watching — a human reviewing and correcting every single prediction in real time. Because that attention is expensive, you spend it on the material most worth learning from: the most distinctive, highest-coverage constructions, the ones whose corrections will generalize furthest downstream, translated first. Silver path is for when nobody’s watching as closely — the model runs ahead on material that’s comparatively easy or low-risk, and anything that looks uncertain gets flagged for a human to review later rather than corrected live. Golden path when someone’s watching, silver path when they’re not.

I don’t think this framework is settled — it’s closer to a rule of thumb I’ve been testing against real projects than a methodology I’d defend in the abstract, and I’d guess the boundary between “distinctive enough to deserve golden-path attention” and “safe enough for silver path” needs to get a lot more precise before this generalizes past the projects I’ve watched it work on. But it’s already changed how I sequence work, in the same spirit as the simplification argument in Occam’s Razor and Bible Translation: before adding more review infrastructure, ask whether the correction you’re about to make is actually reaching the next prediction, or just the next document.

Permalink →

I ran into something last week that unsettled me more than expected. I was testing whether a model could translate a passage using only its own reasoning, so I hid the source text behind a simple substitution cipher before handing it over — the idea being that if the model could still produce a faithful translation, it would have to be doing something closer to genuine cross-lingual reasoning than pattern completion. It decrypted the cipher and recited the verse anyway. Correctly. Immediately. The Bible is so memorized by frontier models that hiding the source text doesn’t work — the model just decrypts and recites it anyway.

That’s stranger than it first sounds. This wasn’t a model translating an encrypted string. It was a model recognizing that a scrambled string of characters was shaped like a passage it already had memorized verbatim, discarding the cipher as a distraction, and reciting from memory. That’s a different capability than translation, and it means a huge chunk of what looks like “the model translated this well” for biblical text is really “the model already knew this text in every language I asked for.” You can’t use Scripture to test whether a model can translate, because for Scripture specifically, the model usually doesn’t need to.

This is a genuine methodological problem, not just a fun fact. If I want to evaluate — or train — a system’s ability to reason from meaning to meaning rather than reciting a memorized target string, I need source material the model hasn’t swallowed whole. Obfuscating the surface form doesn’t work, as I just found out the hard way. What might actually work is changing what the “source” is in the first place.

Acting the scene instead of laundering the words

There’s a practice-level version of this same insight that I’ve been leaning on for a while, mostly with human translators and increasingly as an instruction to models: stop translating the words. Act the scene. Then say what just happened in your own language.

The reasoning is the same as in prioritizing semantics over structures: a word is not a translation unit, it’s a residue left behind by a meaning that already happened. Translating the residue is really just copying with extra steps — the failure mode the cipher experiment exposed at scale. A model (or a person) staring at a string of tokens finds the shortest path to an output resembling those tokens, and if it’s already memorized the “right answer,” the shortest path is recall, not comprehension. Acting the scene forces a detour around that shortcut. You have to know who is speaking, to whom, what they want, what’s at stake, before you’re allowed to say anything at all. Only then do you say what happened, in whatever language you’re working in.

The more rigorous version: don’t feed it the words at all

The practice-level fix is good discipline, but discipline is not architecture, and I’d rather not depend on a model or a translator “resisting the urge” to recall a memorized string. The more rigorous fix, and the one I actually want to build, is to stop handing the model literal source text as the thing to be translated at all.

Instead, represent the passage as a language-agnostic semantic description: pericopes broken down into nested units of who is speaking, to whom, and what is being communicated — speech acts nested inside speech acts, the way narrative and dialogue actually nest inside each other. This is the same move I described in the case for dynamic equivalence alignments — form → meaning → form instead of form → form — pushed one step further upstream. Instead of asking a model to align a translation back to source words after the fact, you never give it the source words as the object to reproduce in the first place. Its job becomes representing the meaning of a discourse unit faithfully, not pattern-matching against a corpus it already has memorized down to the verse numbers. If there’s no recognizable surface string to recite, “recite the memorized answer” stops being a viable shortcut, and the model actually has to do the thing you’re asking it to do. This is really an extension of the point I made in beyond word alignments: the word was never the right unit, and now I have a concrete reason — memorization contamination — to stop treating it as one even operationally.

Where this is honestly still rough

I don’t want to oversell this, because I’ve already tried a related methodology and it only sort of worked. I built a multimodal process I’ve been calling Familiarize → Internalize → Articulate — acting out a passage, drawing the scene, studying the characters and key terms — before ever producing a rendering, as a more embodied version of the “act the scene” instinct above. In practice it can shunt between layers of context in ways that are only sometimes helpful; moving between acting, drawing, and terminology study doesn’t reliably converge on a clearer semantic representation, it sometimes just adds noise. And the supporting resources for it — the character studies, the term glossaries, the scaffolding that’s supposed to make internalization tractable rather than vibes-based — are still unfinished.

So this isn’t a solved problem, it’s a diagnosis with a promising direction attached. The cipher experiment told me the failure mode is real and larger than I assumed. The semantic-description architecture is my best current guess at a fix that doesn’t rely on willpower. Whether nested speaker/addressee/content units are actually the right granularity, and whether a model can be made to represent meaning at that level reliably, is still very much open.

Permalink →

Somewhere around 2010, deep into a master’s degree in Ancient Greek linguistics, I was trying to build word2vec models over a Greek corpus using Gensim, which was — and largely still is — the workhorse library for this kind of thing. Gensim’s whole pitch was memory efficiency: instead of loading a corpus into RAM and training a model on the full thing in one pass, you streamed documents through it one at a time. The model updated incrementally as the data arrived, rather than requiring the entire corpus to exist in memory simultaneously before training could start.

At the time this felt like a plumbing detail, not an idea. But it planted a question I didn’t know I was carrying around for the next decade and a half: why train a model? Why don’t we figure out a streaming approach where we just give it the data it needs on the fly? Ancient Greek is a low-resource language by any modern standard — there wasn’t enough data to justify pretraining a dedicated model in the first place, and streaming was mostly a practical concession to a laptop that couldn’t hold the corpus in memory. But the concession turned out to be more interesting than the workaround it was standing in for.

That question is, more or less, how the translation tool I’ve been building handles languages it has never seen. It doesn’t train a per-language model. There’s no fine-tuning step, no “we support 40 languages, request the 41st.” Instead, at prediction time, it streams relevant translated examples to the model as context and asks it to produce a translation using those examples as its only local evidence for how the language behaves. The model isn’t learning the language in any durable sense; it’s being handed exactly the data it needs, on the fly, for this one prediction, and then that context evaporates. This is the same move as Gensim’s streaming corpus reader, just moved from “how do I fit training data in memory” to “how do I give a frozen model enough signal to be useful in a language it’s never encountered.”

The practical difference is enormous. Training or fine-tuning a model per language — even a small one — is the kind of thing that will bog your computer down for hours and then crash, especially if you’re iterating on ultra-low-resource languages where you don’t have the corpus size to justify the cost in the first place. Streaming relevant examples at inference time is orders of magnitude faster, because you’ve swapped a training job for a retrieval-and-context-assembly step. It’s the same logic I was chasing in Minimal Translation Memory for AI Agents — find the smallest, most generalizable set of examples that covers the most ground, and hand the model only that, rather than everything you have.

I want to be honest about where this breaks down, though, because it’s a real limitation and not a rounding error. The streaming approach works today because a human is validating each prediction live — approving, correcting, or rejecting the model’s output as it’s generated. That human-in-the-loop step is quietly doing a lot of work: it’s the error-correction mechanism that a trained model would otherwise have baked in through gradient descent over many examples. Remove the human, and you no longer have a guarantee that streamed context is sufficient on its own to produce a reliable translation. Whether streaming-plus-retrieval can substitute for training without a validator in the loop is, as far as I can tell, an open research problem — not a solved one I’m just being modest about. I’d put it in the same bucket as the clustering evaluation problem I never resolved in Contextual Vectors for Lexical Meaning: a piece I know is missing, not a piece I’ve quietly decided doesn’t matter.

What convinces me this is a real pattern, rather than a translation-specific hack I’ve talked myself into, is that I’ve watched the identical lesson surface somewhere completely unrelated. Years ago, I was on the other side of a job interview where a candidate was asked to process seventy million XML records. The naive approach — load everything, validate everything — didn’t just run slowly, it fell over. The fix was to stream: validate records as they arrived, discard what you didn’t need to keep, never hold the whole seventy million in memory at once. It’s the same shape as few-shot prompting, just wearing different clothes: you don’t dump your entire knowledge base into a context window and hope the model sorts it out, you stream the handful of examples that are actually relevant to this prediction and let the rest stay where it lives.

That’s one of maybe four lessons a developer has to learn, and I seem to have learned it twice, fifteen years apart, in two fields that have nothing to do with each other. I don’t think that’s a coincidence so much as a sign that “stream the relevant slice instead of loading the whole thing” is one of those ideas that’s basically fractal — it shows up at the level of a corpus reader, a translation model, an XML validator, and a prompt, and it’s probably going to keep showing up wherever the next constraint is memory, time, or data scarcity.

Permalink →

#translation

Three months ago, my team was discussing ways to gather translation data for new, undocumented languages that would be helpful for Bible translation, but which sidestepped the intimidation factor that accompanies the prospect of translating the Bible itself.

One idea that came to mind was to create a short novel that would cover the major genres of the Bible in a minimal form.

Here’s an example outline for a short novel that would cover the major genres of the Bible in a minimal form:

Title: “The Sojourner’s Tale”

  1. Prologue: A brief narrative introducing the main character, a sojourner named Elias, and his quest for wisdom and understanding.

  2. Chapter 1: Elias’ journey begins. He encounters a wise old man who shares three proverbs about life, faith, and perseverance.

  3. Chapter 2: Elias arrives in a city where he witnesses a heated argument between two religious leaders. He listens to their debate and reflects on the nature of truth and belief.

  4. Chapter 3: Elias attends a worship service where a liturgical ceremony is performed. The chapter includes a description of the ritual and its significance.

  5. Chapter 4: Elias discovers a letter from his father, which contains a genealogical record of his family and a personal message of love and guidance.

  6. Chapter 5: Elias meets a group of travelers who share a meal with him. During the meal, one of the travelers tells a parable about the importance of compassion and forgiveness.

  7. Chapter 6: Elias has an apocalyptic dream filled with symbolic imagery and prophecies about the future. He struggles to understand its meaning and seeks the advice of a local interpreter.

  8. Chapter 7: Elias encounters a prophet who delivers a message of repentance and hope. The prophet’s words are a mix of poetry and prose, and they deeply impact Elias.

  9. Chapter 8: Elias joins a group of worshippers in singing a hymn of praise. The lyrics of the song are included in the chapter.

  10. Chapter 9: Elias meets a teacher who shares a sermon on the nature of wisdom and the importance of living a righteous life. The sermon includes a short creed or statement of belief.

  11. Chapter 10: Elias faces a moral dilemma and seeks guidance from a trusted mentor. The mentor shares a piece of wisdom literature, similar in style to the book of Ecclesiastes, to help Elias navigate his decision.

  12. Epilogue: Elias reflects on his journey and the lessons he has learned. He pens a letter to his future descendants, sharing his experiences and encouraging them to seek wisdom and live according to the truths he has discovered.

This outline incorporates various genres found in the Bible, including narrative, proverbs, argument, liturgy, genealogy, parable, apocalyptic literature, prophecy, poetry, hymn, sermon, wisdom literature, and epistolary. While condensed, this structure provides a way to showcase the diverse literary forms found in the Bible within a single, cohesive story.

Linguistic levels

While this outline addresses genre coverage, there are also other linguistic levels that we need to consider. These include (among others):

  • Lexical coverage (Biblical key terms, proper nouns, etc.)
  • Semantic coverage (typical domains-specific meanings in the Bible)
  • Interpersonal coverage (typical interpersonal/transactional/political/intertextual/etc. meanings in the Bible)
  • Information structure, logical structure, etc.

One of the ways to make this data particularly useful would be to ensure that contrastive pairs are included in the data at each level. If one wording presents a pair of objects, then another should present only one object but not the other. If one wording realizes a command activity, then another one should realize a request activity using the same lexical items, etc.

In other words, attempting to control the data at each level ensures that the data addresses specific items of interest in translation. However, this is a bit of a double-edged sword, and the wrong way to do this would be to control for formal, grammatical patterns in the source language that are intentionally left behind in translation (cf. Prioritizing Semantics Over Structures).

Conclusion

The outline above lays out an example of how we might approach the task of Bible translation in a new, undocumented language without both the intimidation that normally accompanies the prospect of translating the Bible itself, and also without the security challenges that are implicated in translating the Bible in hostile contexts.

Permalink →

Occam's Razor and Bible Translation

2024-05-08 · 3 min read

#translation

Question

Is there a simpler approach to Bible translation?

Tesla’s Application of Occam’s Razor

Tesla faced a massive, patchwork C++ codebase for identifying labelled items in video input for full self-driving. Progress was slow, and scalability for the last 1% of edge cases seemed unattainable.

Then they applied Occam’s razor and shifted to training a model with extensive 360-degree video input and controls output (steering, brake, gas). This approach proved significantly better, improving 5x-10x each month as more video data was uploaded from drivers.

See discussion at 3:42 in this video. Thanks to Birch Champeon for sharing.

Applying Occam’s Razor to Bible Translation

I propose a similar shift in Bible translation.

Currently, we use large legacy apps and newer apps with numerous features and checks. With LLMs and agent-based translation, we can simplify the task.

Starting from first principles: what are the inputs and outputs?

If we move from Greek/Hebrew source texts to a semantic representation, semantics become inputs. A model could then apply these semantics to the collocation and colligation patterns in the target language, possibly by reranking outputs with context awareness.

Like Tesla, we need to crowdsource data to map semantics to target language expressions. One method could involve eliciting translations of semantic units from Greek/Hebrew, displayed in aligned gateway languages familiar to target language speakers.

Synthesizing Crowdsourced Data

However the data is elicited, we need to synthesize crowdsourced data into a coherent whole, i.e., complete, versioned, public domain Bible translations.

Adapting swarm intelligence to Bible translation could be the solution. This would involve numerous agents, each translating small, semantically coherent text units. These translations would be combined and refined by a central agent to ensure consistency and coherence across the translation.

This method would allow a more flexible and scalable translation process, distributing work among numerous agents, each responsible for a manageable portion of the text. This would enable faster, more efficient translation, and greater adaptability to local community needs and preferences.

Every approach has risks and trade-offs, but the prospects of applying Occam’s razor to Bible translation are exciting and promising.

Permalink →

#translation

This idea was co-ideated with Ben Scholtens

Double-Entry Translation, or Ledger Translation (LT), is a framework that adapts the principles of double-entry bookkeeping to ensure accuracy, consistency, and a clear mapping between the source and target texts in the Bible translation process.

Double-Entry or Ledger Translation Framework

The framework consists of the following components:

  1. Translation Accounts:

    • Source Text (ST): The original text, treated as an asset.
    • Translated Text (TT): The new translation, treated as a liability.
    • Translation Mapping (TM): The mapping between ST and TT, representing the transactions or “translation-actions” for translation-units at multiple levels of granularity (e.g., meaning-based alignments, sliding-window ngram alignments, etc.).
  2. Translation Journal: A chronological record of all translation entries, including the source text, multiple translated texts, and their corresponding translation mappings.

  3. Translation Ledger: A collection of “posted” translations and their reconciled mappings, perhaps organized by book, chapter, and verse, representing the accepted translations and alignments.

  4. Translation Process:

    • One or more translators create journal entries for each translation unit, including the source text (ST), multiple translated texts (TT1, TT2, etc.), and their respective translation mappings (TM1, TM2, etc.).
    • Each translation mapping “debits” the source text and “credits” its corresponding translated text, ensuring a balanced and accurate representation of the translation process. Anything credited to the translation can be debited from the source text, and vice versa for full accountability.
  5. Reconciliation and Review:

    • Editors review the translation journal entries to ensure accuracy and consistency between the source text, multiple translated texts, and their mappings.
    • Discrepancies or issues are discussed and resolved, with the goal of creating a reconciled translation—there can be more than one accepted version for any unit, though a published version might be expected to arrive at a single preferred version for a given project—and mapping for each translation unit.
    • The reconciled translation and mapping are “posted” to the translation ledger, linking to the relevant journal entries.
    • The ledger entry represents a balanced translation that accounts for the credits, debits by way of the linking or mapping between the source and target texts.
  6. Reporting and Iterative Publishing:

    • Regular reports summarize translation progress, including the number of translation units processed, reviewed, reconciled, and posted to the ledger.
    • These reports can be used as a mechanism for iterative publishing, allowing the translation team to share their work in progress and gather feedback without the pressure of producing a finished product.
    • Iterative publishing helps maintain transparency, engage stakeholders, encourage believers, and refine the translation over time.
  7. Breaking Down Translation Units and Machine Translation:

    • The LT Framework enables the translation process to be broken down into smaller, manageable units, such as phrases or sentences, rather than focusing on larger segments like paragraphs or chapters.
    • As the translation progresses, the growing dataset of source texts, translated texts, and their mappings can be used to train and refine statistical and neural machine translation models.
    • These models can provide suggestions and assistance to translators, improving efficiency and consistency as the project grows.

By allowing for multiple translations and mappings to be submitted via the journal, and then reconciling them into balanced ledger entries, the LT Framework promotes collaboration and ensures that the final translation is a result of careful consideration and consensus.

The linking of ledger entries to their corresponding journal entries maintains a clear “audit trail” and allows for a comprehensive understanding of the translation process. This approach enhances transparency, accountability, and the overall quality of the translated text. More intriguingly, the LT Framework is a decentralized, collaborative, and iterative framework that does not rely on external expertise, and that can adapt to various translation projects, languages, and contexts, making it a valuable tool for emerging Bible translation projects and teams around the world.

Addendum: Simplified Implementation of Double-Entry Principles

While the full Ledger Translation Framework promises a robust approach to managing translations, a simplified implementation can still offer many of the benefits of double-entry principles.

In this simplified approach, translations are stored as both “journals” and “ledgers”:

  1. Translation Journal: The journal represents the raw, unadulterated translations, capturing the direct output from the translators. This includes the source text, the translated text, and any relevant metadata or notes.

  2. Translation Ledger: The ledger represents the reviewed, reconciled, and approved translations. It contains the final, polished versions of the translated text, along with the corresponding source text and metadata.

The process flow for this simplified implementation would be as follows:

  1. Translators create entries in the Translation Journal, providing their raw translations for each unit of text.

  2. Editors and reviewers assess the journal entries, making corrections, and ensuring consistency and accuracy.

  3. Once a translation is approved, it is transferred from the Translation Journal to the Translation Ledger, becoming an official, accepted version of the translation. Ledgers are versioned, journals are not - though they may be versioned via source control management software.

  4. The Translation Ledger serves as the authoritative source for the translated text, while the Translation Journal maintains a record of the original translations and the process of refinement. Because the ledger is versioned, translators are able to make updates to the journal entries, which can then be reviewed and re-posted to the ledger under a new version number. Theoretically, each verse may be versioned, or each chapter, or each book, or the entire translation.

By maintaining both a journal and a ledger, this simplified approach still provides transparency, accountability, and a clear audit trail for the translation process. It allows for the tracking of changes and the preservation of the original translator’s work, while also presenting a clean, approved version of the text for use and distribution.

While not as comprehensive as the full Ledger or Double-Entry Translation Framework, this simplified implementation still captures the core principles of double-entry bookkeeping and can be a valuable tool for managing translations in a more streamlined manner while maintaining transparency regarding the source texts used, the translation process, and the published and versioned translated text.

Permalink →

Few-shot Translation Evaluation

2023-12-22 · 2 min read

#translation #ai

In the context of a translation project, where we aim to utilize only the resources at hand, let’s consider a simplified approach to the evaluation stage. This approach would involve the following steps:

  1. Start by taking the draft translation and identify the translation pairs that are most similar to the paired prediction.
  2. Evaluate the draft translation for any potential misuse of words.
    1. One possible method could be to use a rules-based matcher that identifies any tokens in the source+prediction pair that only occur in EITHER the source OR the target in any of the similar examples.
    2. It could also be beneficial to examine whether any words seem to be mistranslated based on the gold standard examples.
    3. To aid this process, aligning these examples by phonetics, orthography, and apparent semantic similarities could be helpful.
  3. Once the evaluation is complete, pass the task back to the translation bot, providing special notes about any words that appear to be mistranslated.

In essence, the process can be summarized as: [ most similar examples ] —> [ prediction ] —> [ most similar examples ]

The goal is to identify the most similar examples to the prediction, and then use those examples to evaluate the prediction. This process can be repeated iteratively, with the bot learning from the evaluation and then generating a new prediction and a new set of examples that will help the initial bot improve its output.

Permalink →

#translation

Linguistic structures are the formal patterns that realize linguistic meanings. In the process of translation, it is the meanings that are translated; the structures are left behind by design. Semantics are the priority in translation.

Note: between two languages, the semantic systems are not isomorphic. This mismatch means that we cannot assume that the precise semantic categories of one language may be transferred into another. Comparing the lexicogrammatical systems of gender in two languages should suffice to illustrate this point. However, the general semantic categories seem to be fairly universal (one would have to compare every language to actually establish this claim).

This fundamental observation must be factored into the translation process as we consider how AI tools might be leveraged for translation drafting.

For instance, rather than trying to account for every word in the source by translating a word in the target, we should attempt to account for the things the words are being used to realise, namely the entities, processes, traits, collocations, logical sets, speech acts, and more.

It is definitively better to traffic in semantic units such as these, since accounting for such meanings is what makes a translation good or bad.

By contrast, imagine if someone claimed that you needed to account for every character in the source text (all the alphas, omegas, deltas, and iotas). If you did somehow manage to do this, you would not be translating, but copying.

Permalink →

#translation #ai

In the realm of machine translation, achieving the highest level of accuracy is paramount. One approach to improve the precision of translations is through a method called “Token Matching Evaluation”. This method centres around the use of valid tokens, which are words or phrases that have appeared in sample translation pairs.

The Process

  1. Populate a list of valid tokens for back-translation: The first step is to create a list of valid tokens. These tokens are the words or phrases that the model is allowed to use for back-translations.

  2. Back-translate using valid tokens: The model is instructed to back-translate any given word only from the list of valid tokens. For example, using the {{select options=valid_tokens n=6}} syntax, the model selects a specific value from a chunk of 6 valid tokens.

    • If a word hasn’t even been translated once by a human, then any attempt would effectively be indistinguishable from a hallucination. Therefore, it is important that the valid tokens correspond to reality. For instance, if you need to translate the English word ‘Jerusalem’, you would find 5 sentences with ‘Jerusalem’ in them and those sentences would have Abanyom renderings, like ‘Yerusalem’. The only valid renderings must be contained in the sample sentences.

    • Any tokens not in the set of valid tokens are simply retained in their translated form in [square brackets] to indicate they have not been back-translated because they are unknown. This could be facilitated by having an agent send out messages to the translation team via WhatsApp, or using a simple gamified app.

  3. Gather samples based on token coverage: It might be beneficial to gather samples based on whether or not they help cover the missing tokens. Start with an initial set of 5 semantically similar examples, followed by a second set of 5 more sentences that include the English glosses not represented in the first five examples.

By following these steps, the output translation will more likely align with the sample translations provided in the prompt, thereby enhancing the accuracy of the translation.

Permalink →

#translation

(A Dynamic Equivalence Alignment is something like “Essentially Aligned Bible Translations”)

Token to token mapping is problematic, despite its ubiquity. It is theoretically the same as character to character mapping (a token is, with some exceptions, a set of characters delimited by whitespace). There is a fundamental distinction between structural units like characters, tokens, phrases, clusters, wordings, sets, etc., where the defining characteristic is composition (“what is this thing built out of?”) and semantic or functional units like entities, processes, predications, speech acts, definitions, scenarios, etc., where the defining characteristic is role or function (“what does this thing do in its immediate context?”). 

With regard to alignment, the question driving any alignment is “what does span A have in common with span B?

The answer depends on what type of spans you are trying to align. 

Span typeSpan A is…Span B is…Both spans have in common…What type of similarity is this?Does ER++ gold standard alignment data get it right?Does the SOTA algorithm get it right?
CharacterκaNothing obviousSuperficialn/an/a
ωoSome kind of non-meaningful phonological similarityn/an/a
TokenκαὶasNothing obvious.Depends on the type of token.nono
theGrammatical (sort of) and semanticnoyes
ἔπεσενfellGrammatical (partial) and analogous lexical semanticsyesyes
Formal Syntactic Structureκαὶ ἐγένετο… (clause)as he was… (adverbial phrase)Partial lexical, partial grammaticalDepends on how similar the two languages’ grammatical systems are.nono
ἐν  τῷ  σπείρειν (prep. phrase)scattering (part of compound verb ‘was scattering’)noyes
Semantic unitἐγένετο (event)he was (event)Semantic categoryA meaningful semantic unitnono
ἐν ταῖς πορείαις αὐτοῦ (circumstance)in the midst of his pursuits (circumstance)Semantic categoryn/an/a

The case against token mapping as a basic strategy

Let’s assume that tokens are whitespace delimited, though this is not a necessary condition for the argument I want to make here. My argument is this: we should attempt to map semantic or meaning-based similarity instead of formal or structural similarity, because semantic overlap is implied by a translation, while formal overlap is explicitly repudiated (otherwise no translation would be necessary, or the task would be simple deciphering, as in an interlinear).

The issue with trying to map tokens between texts of different languages and/or translation styles (even within the same language) is based on the fact that a “token” is a formal category, and there is no principled linguistic or theoretical reason to assume that a token in language 1 will correspond isomorphically to a token in language 2. This is related to the use of whitespace to delimit tokens, but any tokenization strategy could fall prey to the same problem. Conversely, using whitespace delimited tokens is in fact a possible path forward, but the crucial hurdle to overcome is that we must not think of our task as one of aligning formal structures. 

Consider the very nature of translation as a human activity. When one translates from text A to text B, the task is not to create a more-or-less graphically isomorphic re-presentation of text A in text B. The success of the task does not require there to be any structural or formal isomorphism between the two texts. Consider the following “translation” of John 3:16 into English (The Message), Chinese (CSB), and Arabic (KSS):

Source: For God so loved the world that he gave his only begotten son, that whoever believes in him should not perish but have everlasting life.

Target 1: This is how much God loved the world: He gave his Son, his one and only Son. And this is why: so that no one need be destroyed; by believing in him, anyone can have a whole and lasting life.

Target 2: 因为上帝爱世人,甚至将祂独一的儿子赐给他们,叫一切信祂的人不致灭亡,反得永生。

Target 3:لەبەر ئەوەی خودا ئەوەندە جیهانی خۆشویست، تەنانەت کوڕە تاقانەکەی بەختکرد، تاکو هەرکەسێک باوەڕی پێ بهێنێت لەناو نەچێت، بەڵکو ژیانی هەتاهەتایی هەبێت،

What is the one thing that none of these translations have in common? Structural isomorphism. There are degrees of formal similarity between them (somewhat similar overall length? Perhaps there are some similar phonological patterns (though I can’t tell, not knowing all the scripts). There are more formal similarities between the English versions, but this is, I would argue, only superficial. The reason these are translations (or perhaps we prefer the term paraphrase for same-language translation) is because they have semantic overlap. 

You cannot find any of the Greek tokens in either the English or Spanish texts!

Once we move to separate scripts, I hope it becomes clear that it is not a useful question to ask “Where does the word ‘God’ in the source appear in each of these target texts?” The word “God” does not appear in the other texts—that’s the whole point of translation. What we should be asking is something like “Where does the entity construed via the word ‘God’ get construed in the other texts?” This is a subtle difference, but it’s the difference between being able to answer the question and being stuck trying to accomplish a task that we are not well equipped to do in a scalable and reliable manner, and for which we have no formal evaluation apparatus.

When it comes to aligning whitespace-delimited tokens, just because it’s the state of the art, does not mean it is reliable or scalable, and it doesn’t mean we can’t do better with tools we already have.

Our gold standard data is not as reliable as we would like it to be. In addition, Clear Engine’s alignment relies on token-to-token alignment, and it is far from reliable out of the box. Even when improved by structure-based segmentation of units, the structures that each language uses to realize its semantics are always divergent to some degree.

Proposed path forward

The most basic implementation of semantics-based alignment for our purposes could take three forms (and these could also be combined into a single approach).

I think we should really pursue Option 3, since it best accords with the broader theoretical conception of translation, and it is likely to be even more useful than a word-level alignment in terms of the kinds of resources that can be retrieved.

Option 1: drop function words (pretty easy)

First, we can pursue the essentially aligned Bible translations approach. Initially this might mean we simply drop the so-called function words from the source language, and attempt to identify matching content words in the target languages. The downside of this coarse approach (i.e., simply dropping words we suspect will not be important), is that there are many likely exceptions. For example, a single Greek verb ἦλθον would align best with two English words “was going.” 

Option 2: replace source text words with glosses or lemmas (easiest?)

Second, we could use language-specific glosses for aligning to another language. We have English and Mandarin glosses for our Greek and/or Hebrew texts, and I suspect we can much more easily align and automate evaluation for the alignments between an ‘interlinearized’ English string. For example,

  • Greek: καὶ εἶπεν αὐτοῖς Ποῖα;οἱ δὲ εἶπαν αὐτῷ Τὰ περὶ Ἰησοῦ τοῦ Ναζαρηνοῦ,

  • English glosses: And He said to them What things; - And they said to Him The things concerning Jesus of Nazareth,

The English glosses in the above example are probably more straightforward to align to English translations. These glosses can then be decoded back into the Greek text from which they were generated.

Option 3: treat alignments as annotations (best)

Third, and more ideally, we would think about alignments as what they really are, just another form of annotation, where text A is “annotated” with text B. 

As the figure above illustrates, an annotation like “a semantic unit representing an entity” can be annotated across languages. It does not assume that the tokens in any language correspond isomorphically to tokens in another language, yet token- and character-level alignment is implicated by semantic alignment! I believe this is the best approach to alignment of the options listed above. 

This approach can be technically implemented by indexing the spans in each text that should be annotated with each semantic annotation. This means we do not need to even use IDs for alignment—alignment can be a character-level annotation — because the things being aligned (e.g., the blue rectangular box in the figure above) are not tokens. 

Summary

Let us make it our explicit goal to align meanings, not forms. The meanings are realized in forms, but the nature of the question changes because we are mapping form->meaning->form (which can be more or less objectively evaluated) instead of form->form (which implies the meaning without making it explicit). 

Our goal should not be to align: FORM → FORM

Our goal should be to align:

FORM → MEANING → FORM

ApproachAlignment taskImplementation
Formal or structuralFORM → FORM
e.g.,
character → character
token → token
whitespace break → whitespace break
Language-specific tokenization.
Every text must be parsed into individual forms that are as similar as possible to our Greek tokens to try to compare apples to apples
- Greek is whitespace tokenized (+/- punctuation)
- Hebrew is broken down to subword levels to try to match the Greek
- Chinese? Arabic? Malayalam?
Semantic or functionalFORM → MEANING → FORM
e.g.,
index → entity → index
index → process → index
index → circumstance → index
index → speech act (e.g., statement | command | question | request) → index
Meaning is linked to string index ranges in any text we want to align
- We identify meanings in the source text (or fall back to phrase structures if we do not have meanings, e.g., for Hebrew)
- Meanings are indexed to arbitrary character spans in the source text file (e.g., Mike and Ben’s suggested implementation)
- Meanings can be subsequently indexed to arbitrary strings in any text. This means that form-to-form mapping is implied by form-to-meaning-to-form, without the problematic aspects of trying to create a dataset that shows where Greek tokens appear in English texts, etc.
Permalink →