Nilky documents a floppy-disk model, failed tokenizer and unfinished follow-up

In a September 13th Hugging Face Community Article, the hobbyist says the model performed worse than a 1,000-parameter model, without providing a benchmark method or score.

By · Published

Primary source: Hugging Face

Why it matters

Tiny-model experiments can expose practical limits that polished release posts omit. Nilky's essay records a poor result, a tokenizer failure and an abandoned follow-up, while its missing benchmark and training details show how little can be concluded from the comparison alone.

A vintage 3.5-inch floppy disk sits beside a Raspberry Pi computer board on a desk, with a blurred monitor in the background displaying abstract code.

Nilky, a Hugging Face community contributor and hobbyist, published a retrospective account of several tiny language-model experiments on September 13th, 2026. Nilky aimed to make a language model small enough to fit on one floppy disk. In the essay, they say the resulting Single Floppy model performed worse than a 1,000-parameter model.

The source is a first-person Community Article. It does not announce a company, funding event or Hugging Face product. Nilky presents the project as a personal experiment shaped by an interest in open-source software and old electronics.

Nilky traces that interest in open source to using Raspberry Pi OS. Later experiments involving ChatGPT and DeepSeek led them toward models whose implementations could be inspected directly. Old hardware supplied the storage constraint. "I really love old electronics," Nilky wrote, adding that the essay itself was composed on a 2016 laptop.

A storage target without a benchmark

Nilky calls the result the Single Floppy model and says it "scores worse than a 1k model." The essay supplies no benchmark method, test set, metric or numerical score, so the comparison remains the author's assessment rather than a reproducible performance result.

It also provides no training recipe or detailed accounting of how the storage constraint affected model quality. Readers cannot determine from the essay which design choice caused the poor result or how the model compares with other tiny language models under equivalent conditions.

That limited disclosure changes what the experiment can establish. It records the goal and the author's verdict, rather than documenting exactly how a floppy-disk storage limit reduces performance.

The follow-ups failed differently

The experiment is useful as a record of constraints rather than a practical model release. Nilky says a later tokenizer failure stopped the floppyx3 attempt, while floppyx4 was never finished.

The essay does not explain the tokenizer failure or provide results from either follow-up. Nilky ends by considering the purchase of a dedicated PC for training, another indication that the post is a personal retrospective rather than a maintained development roadmap.

Nilky's account is unusually direct about the outcome. The model performed poorly by the creator's own comparison, the next tokenizer failed and a fourth experiment remained incomplete. The useful artifact is the candid record of those constraints and failures, with the technical limits of that record left plainly visible.

Reader comments

Conversation for this story loads after sign-in.