CrowdGPT recruits volunteer GPUs to train a 1B-parameter model

The open-source project routes local training through a central coordinator; its creator is testing a consumer-hardware approach, not claiming a finished chatbot.

By · Published

Primary source: Reddit

Why it matters

CrowdGPT is testing whether consumer machines can contribute to LLM pretraining, while its public model card still describes an in-progress model whose output may not be coherent. The distinction between a functioning volunteer network and a useful trained model remains the key test.

https://www.reddit.com/r/LLMDevs/comments/1wwrot7/crowdgpt_the_first_datacenterless_llm/

CrowdGPT is asking people to contribute their computers to train a roughly 1-billion-parameter language model, putting volunteer hardware at the center of an open-source pretraining experiment. The project’s creator, Vxtzq, described the effort in an October 4th Reddit post. The post calls CrowdGPT the “first datacenterless LLM,” a claim its own mechanics and earlier distributed-training work make difficult to sustain.

Vxtzq’s public DEV profile describes the creator as working in natural-language processing and computer vision and lists France as their location. In an earlier post, Vxtzq framed the project as a way for people to contribute compute to a shared model. They said the initial network was still being built and asked readers to identify what would keep them from running a client. The Reddit post is a renewed pitch for that project, rather than evidence of a newly released, finished model.

The CrowdGPT project splits training across users’ machines. Its site says a client downloads public training data, trains locally, then sends a model update to a coordinator for validation and aggregation. The project says submitted updates are checked for size, numerical errors and norm bounds before being incorporated. Contributors can run the client on supported NVIDIA or AMD GPUs, Apple Silicon, Windows systems using DirectML, or a CPU, which the project warns will be very slow. The GitHub repository publishes setup instructions and code for the client.

Diagram of CrowdGPT’s reported process: a client downloads public training data, trains locally, sends a model update to a coordinator for validation, and the coordinator aggregates updates.
CrowdGPT describes local training on contributors’ machines followed by coordinator-side validation and aggregation — AI explanatory diagram, not documentary evidence. RuntimeWire · AI-generated diagram.

That design makes “datacenterless” a narrower description than the Reddit headline suggests. CrowdGPT still relies on a coordinating server and public hosting for its model files and data; the distinction is that it aims to avoid a centralized GPU cluster doing the training. CrowdGPT’s website identifies Xevex as the host of its primary coordinator API. Its earlier description also acknowledged a lightweight central server that receives and merges client updates. The project is distributed at the level of compute, not wholly peer-to-peer.

Diagram showing volunteer machines training locally and sending model updates one way to a coordinator for validation and aggregation, with Xevex identified as the coordinator API host; public hosting for model files and data is shown separately.
CrowdGPT describes volunteer machines sending updates to a central coordinator for validation and aggregation, alongside public hosting for model files and data — AI explanatory diagram, not documentary evidence. RuntimeWire · AI-generated diagram.

The public model card for Crowd-v1 describes a roughly 1-billion-parameter Transformer and lists training as in progress. It reports 1.6 billion tokens trained, while also describing the model weights as randomly initialized and warning that the model should not be expected to produce coherent text. That leaves the project at the stage of demonstrating a training process, not offering a competitive ChatGPT alternative. The project’s model card also labels its inference code as testing-only.

CrowdGPT is entering a field with established distributed-training precedents. The Hivemind project has described decentralized deep learning across volunteer computers since 2020. Prime Intellect went further in a different direction: its 2024 account of INTELLECT-1 described a 10-billion-parameter model trained across five countries and three continents, using as many as 112 H100 GPUs. That effort relied on distributed GPU resources rather than establishing CrowdGPT’s specific consumer-volunteer model, but it makes a broad claim to be the first distributed LLM training effort untenable.

The narrower test is whether many ordinary machines can contribute usefully to pretraining when hardware, connection speeds and availability vary. CrowdGPT’s design tries to handle that variability through local work and server-side aggregation, while publishing the model architecture, training data and code for inspection. Its website offers instructions for running a client, but the training total on a model card does not reveal how many distinct machines contributed, how much of the reported training used volunteer hardware, or whether the resulting weights improve against a held-out evaluation. Those are the measures that will show whether the community-compute approach works beyond the mechanics of accepting updates.

For now, CrowdGPT is a small open-source experiment with an accessible client and a concrete target: train a shared 1-billion-parameter model using contributed consumer compute. Its strongest claim is the attempt to open pretraining to volunteers. Calling it the first datacenterless LLM says more than the project’s evidence establishes.

Reader comments

Conversation for this story loads after sign-in.