MameLoshnLM adapts Llama 3.1 with 915 million Yiddish words

Uri Katz and four co-authors built the 8B model on curated Yiddish text after finding that a common web corpus misclassified Hebrew and machine-translated pages.

By · Published

Primary source: X

Why it matters

MameLoshnLM shows how mislabeled and translated web text can distort a model's apparent fluency. The team's corpus, benchmark and openly inspectable results offer a testable approach to language-specific model development, though the weights are restricted to non-commercial use.

An archivist holds open a Yiddish book beside crowded shelves of translated editions in a dim archive.

Uri Katz and four co-authors built MameLoshnLM by continuing to train Meta's Llama 3.1 8B on more than 915 million words of Yiddish. The work addresses a problem that a language label in a training dataset can conceal: much of the text may not represent the language well at all. The team's paper and model were available before Katz described the work in an October 6th thread on X. The thread previewed a poster presentation scheduled for October 7th at COLM 2026, which ran October 6th through 9th.

Photo of the MameLoshnLM poster presentation attached to Uri Katz’s October 6th, 2026 X post.
Uri Katz’s original post previews the MameLoshnLM poster presentation at COLM 2026. Original image from @UrikaUri.

Katz, the paper's first author, is a computer science PhD candidate at Bar-Ilan University. He previously earned master's degrees at Tel Aviv University and the Weizmann Institute, according to his academic profile. The project brings together Katz with Omer Goldman, Tomasz Limisiewicz, Reut Tsarfaty and Noah A. Smith, researchers whose work spans multilingual language modeling and natural-language processing.

The team's central evidence comes from its audit of the Yiddish portion of the mC4 web corpus. Researchers classified 143,708 pages: 42.2% came from sources they validated as Yiddish, at least 29.8% appeared machine-translated, and 21.9% were Hebrew pages classified as Yiddish. Those are the authors' estimates from source checks and language identification, not a universal measurement of every multilingual corpus. The figures show how a model can encounter a sizeable quantity of nominally Yiddish material while absorbing translated phrasing, mislabeled text or both.

MameLoshnLM's training corpus, Oytser, combines contemporary online writing with literary material. That mix addresses a mismatch in Yiddish: its Germanic grammar, Hebrew script and vocabulary shaped by Hebrew and Aramaic traditions are difficult to capture through a thin web sample alone. The researchers say Oytser contains more than 915 million words. The Yiddish Book Center permitted use of its Digital Yiddish Library, according to the project's model card.

The team also released Kashes, a benchmark covering translation, language analysis, information extraction and language understanding. Its dataset page documents test sets for tasks including translation, part-of-speech tagging and parsing. The model card reports five-shot results: MameLoshnLM's listed average is 62.6, compared with 56.8 for the original Llama 3.1 8B and 54.7 for Qwen3 8B. The comparison is not a single universal measure of language ability: it combines different tasks and metrics, and the authors assembled the benchmark used to report the gains. The detailed task scores are more revealing than the average; MameLoshnLM leads on several linguistic and translation evaluations, while Qwen3 scores higher on some understanding tasks.

The model has limits. It is a base model, not an instruction-tuned assistant, and its weights are available under a non-commercial license. Its card also warns that OCR from historical books can leave recognition errors and spelling variation in the training data. Researchers can load the weights from Hugging Face, but the project is not presented as a commercial product.

The audit puts measurement alongside model size. A multilingual model can generate readable text without reliably learning the lexical and grammatical features that distinguish a language. For Yiddish, the researchers' approach was to curate language-specific data, continue training an existing model and evaluate it on tasks built for that language. Whether the gains transfer beyond the team's benchmark remains a separate question; the released paper and evaluation resources let other researchers test that claim.

Reader comments

Conversation for this story loads after sign-in.