Streaming the Model Instead of Pretraining It

ryder · 2026-07-01 · 5 min read

Somewhere around 2010, deep into a master’s degree in Ancient Greek linguistics, I was trying to build word2vec models over a Greek corpus using Gensim, which was — and largely still is — the workhorse library for this kind of thing. Gensim’s whole pitch was memory efficiency: instead of loading a corpus into RAM and training a model on the full thing in one pass, you streamed documents through it one at a time. The model updated incrementally as the data arrived, rather than requiring the entire corpus to exist in memory simultaneously before training could start.

At the time this felt like a plumbing detail, not an idea. But it planted a question I didn’t know I was carrying around for the next decade and a half: why train a model? Why don’t we figure out a streaming approach where we just give it the data it needs on the fly? Ancient Greek is a low-resource language by any modern standard — there wasn’t enough data to justify pretraining a dedicated model in the first place, and streaming was mostly a practical concession to a laptop that couldn’t hold the corpus in memory. But the concession turned out to be more interesting than the workaround it was standing in for.

That question is, more or less, how the translation tool I’ve been building handles languages it has never seen. It doesn’t train a per-language model. There’s no fine-tuning step, no “we support 40 languages, request the 41st.” Instead, at prediction time, it streams relevant translated examples to the model as context and asks it to produce a translation using those examples as its only local evidence for how the language behaves. The model isn’t learning the language in any durable sense; it’s being handed exactly the data it needs, on the fly, for this one prediction, and then that context evaporates. This is the same move as Gensim’s streaming corpus reader, just moved from “how do I fit training data in memory” to “how do I give a frozen model enough signal to be useful in a language it’s never encountered.”

The practical difference is enormous. Training or fine-tuning a model per language — even a small one — is the kind of thing that will bog your computer down for hours and then crash, especially if you’re iterating on ultra-low-resource languages where you don’t have the corpus size to justify the cost in the first place. Streaming relevant examples at inference time is orders of magnitude faster, because you’ve swapped a training job for a retrieval-and-context-assembly step. It’s the same logic I was chasing in Minimal Translation Memory for AI Agents — find the smallest, most generalizable set of examples that covers the most ground, and hand the model only that, rather than everything you have.

I want to be honest about where this breaks down, though, because it’s a real limitation and not a rounding error. The streaming approach works today because a human is validating each prediction live — approving, correcting, or rejecting the model’s output as it’s generated. That human-in-the-loop step is quietly doing a lot of work: it’s the error-correction mechanism that a trained model would otherwise have baked in through gradient descent over many examples. Remove the human, and you no longer have a guarantee that streamed context is sufficient on its own to produce a reliable translation. Whether streaming-plus-retrieval can substitute for training without a validator in the loop is, as far as I can tell, an open research problem — not a solved one I’m just being modest about. I’d put it in the same bucket as the clustering evaluation problem I never resolved in Contextual Vectors for Lexical Meaning: a piece I know is missing, not a piece I’ve quietly decided doesn’t matter.

What convinces me this is a real pattern, rather than a translation-specific hack I’ve talked myself into, is that I’ve watched the identical lesson surface somewhere completely unrelated. Years ago, I was on the other side of a job interview where a candidate was asked to process seventy million XML records. The naive approach — load everything, validate everything — didn’t just run slowly, it fell over. The fix was to stream: validate records as they arrived, discard what you didn’t need to keep, never hold the whole seventy million in memory at once. It’s the same shape as few-shot prompting, just wearing different clothes: you don’t dump your entire knowledge base into a context window and hope the model sorts it out, you stream the handful of examples that are actually relevant to this prediction and let the rest stay where it lives.

That’s one of maybe four lessons a developer has to learn, and I seem to have learned it twice, fifteen years apart, in two fields that have nothing to do with each other. I don’t think that’s a coincidence so much as a sign that “stream the relevant slice instead of loading the whole thing” is one of those ideas that’s basically fractal — it shows up at the level of a corpus reader, a translation model, an XML validator, and a prompt, and it’s probably going to keep showing up wherever the next constraint is memory, time, or data scarcity.