This work adapts existing open-weight models across decentralized GPUs without architectural changes, training on masked activations and compressed synchronization while occasional unmasked passes correct the gradients. It matches uncompressed training at over 40× the throughput on ~200 Mbps links.
“To speak of justice requires questioning the global distribution of power that decides who in fact can train these models and who is merely subjected to them.”
— Pope Leo XIV, Magnifica Humanitas (2026)
About
Pluralis is a research lab focused on decentralized AI. We believe the best path is where the models are developed, trained and served across global networks of many participants and owned by the collective.
We are currently carrying out open, multi-participant training runs; you can find information about previous runs here; the current run here, and can apply to join in the planning and development of future runs here.
Research
We show that Muon-trained models develop the low-rank weights exploited in Subspace Networks, despite Muon’s full-rank updates. NuMuon constrains the nuclear norm of its updates, improving compression and post-compression quality at billion-parameter scale while retaining Muon’s convergence.
We introduce training that is asynchronous across both data and pipeline parallelism, using weight look-ahead and EMA-corrected sparse averaging to handle stale updates. It matches synchronous training on models up to 1B parameters while significantly reducing communication.
We relax DiLoCo’s exact outer synchronization to approximate synchronization via mixing and gossip, factorizing it into a non-blocking step that overlaps computation with no staleness and a blocking step that tightens worker agreement. On billion-parameter language models in low-bandwidth settings, the method substantially improves compute utilization while matching DiLoCo’s training progress and is more robust to failures.
We introduce a fast online curvature estimator that tracks preconditioned Hessian behavior during billion-parameter Transformer training. It reveals depth-driven curvature surges behind loss spikes and motivates architecture warm-up: progressively growing depth to stabilize training without slowing convergence.
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism
NeurIPS 2025This work demonstrates that model-parallel training over low-bandwidth networks is possible, training an 8B LLaMA model on par with centralized training while transformer blocks are split across four locations connected only by standard internet links.
Pipeline parallelism trains large models by splitting them into stages, but idle “bubbles” slow training, especially when network latency is high. Our Nesterov method corrects stale updates and outperforms existing async techniques and the synchronous baseline.
Unextractable Protocol Models: Collaborative Training and Inference without Weight Materialization
NeurIPS 2025UPMs enable collaborative training and inference without ever materializing the full model weights for any participant, making decentralized models unextractable in practice.
We introduce a compression method for communication-efficient context parallelism that achieves over 95 % compression with negligible overhead and no convergence loss. By exploiting low-rank activation structure through learned mixtures of subspaces, it scales billion-parameter decentralized models to 100 K+ context lengths on 300 Mbps networks while matching centralized wall-clock convergence.
Technical Reports
Agora: a decentralized training system that shards models across heterogeneous, internet-connected participants while preserving fault tolerance and collective ownership. The report describes Pluralis-8B, an 8.6B-parameter open pretraining run trained on 500B tokens across 330 contributor nodes.
