“To speak of justice requires questioning the global distribution of power that decides who in fact can train these models and who is merely subjected to them.”

— Pope Leo XIV, Magnifica Humanitas (2026)

About

Pluralis is a research lab focused on decentralized AI. We believe the best path is where the models are developed, trained and served across global networks of many participants and owned by the collective.

We are currently carrying out open, multi-participant training runs; you can find information about previous runs here; the current run here, and can apply to join in the planning and development of future runs here.

Research

S. Ramasinghe, S. Siriwardhana, T. Ajanthan, H. Dolatabadi, C. Koneputugodage, G. Avraham, V. Shevchenko, J. Snewin, K. Pajak, H. Xi, A. Long

This work adapts existing open-weight models across decentralized GPUs without architectural changes, training on masked activations and compressed synchronization while occasional unmasked passes correct the gradients. It matches uncompressed training at over 40× the throughput on ~200 Mbps links.

H. Dolatabadi, T. Ajanthan, S. Ramasinghe, C. Koneputugodage, S. Siriwardhana, V. Shevchenko, K. Pajak, J. Snewin, G. Avraham, A. Long

We show that Muon-trained models develop the low-rank weights exploited in Subspace Networks, despite Muon’s full-rank updates. NuMuon constrains the nuclear norm of its updates, improving compression and post-compression quality at billion-parameter scale while retaining Muon’s convergence.

T. Ajanthan, S. Ramasinghe, G. Avraham, H. Dolatabadi, C. Koneputugodage, V. Shevchenko, Y. Zuo, A. Long

We introduce training that is asynchronous across both data and pipeline parallelism, using weight look-ahead and EMA-corrected sparse averaging to handle stale updates. It matches synchronous training on models up to 1B parameters while significantly reducing communication.

C. Koneputugodage, T. Ajanthan, S. Ramasinghe, H. Dolatabadi, S. Siriwardhana, G. Avraham, V. Shevchenko, K. Pajak, J. Snewin, A. Long

We relax DiLoCo’s exact outer synchronization to approximate synchronization via mixing and gossip, factorizing it into a non-blocking step that overlaps computation with no staleness and a blocking step that tightens worker agreement. On billion-parameter language models in low-bandwidth settings, the method substantially improves compute utilization while matching DiLoCo’s training progress and is more robust to failures.

S. Ramasinghe, T. Ajanthan, H. Dolatabadi, C. Koneputugodage, G. Avraham, V. Shevchenko, Y. Zuo, K. Pajak, A. Long

We introduce a fast online curvature estimator that tracks preconditioned Hessian behavior during billion-parameter Transformer training. It reveals depth-driven curvature surges behind loss spikes and motivates architecture warm-up: progressively growing depth to stabilize training without slowing convergence.

T. Ajanthan, S. Ramasinghe, Y. Zuo, G. Avraham, A. Long

Pipeline parallelism trains large models by splitting them into stages, but idle “bubbles” slow training, especially when network latency is high. Our Nesterov method corrects stale updates and outperforms existing async techniques and the synchronous baseline.

S. Ramasinghe, T. Ajanthan, H. Dolatabadi, G. Avraham, V. Shevchenko, Y. Zuo, C. Koneputugodage, A. Long

We introduce a compression method for communication-efficient context parallelism that achieves over 95 % compression with negligible overhead and no convergence loss. By exploiting low-rank activation structure through learned mixtures of subspaces, it scales billion-parameter decentralized models to 100 K+ context lengths on 300 Mbps networks while matching centralized wall-clock convergence.

Technical Reports

G. Avraham, V. Shevchenko, H. Dolatabadi, K. Pajak, J. Snewin, H. Xi, R. O’Donnell, T. Ajanthan, S. Ramasinghe, C. Koneputugodage, S. Siriwardhana, A. Long

Agora: a decentralized training system that shards models across heterogeneous, internet-connected participants while preserving fault tolerance and collective ownership. The report describes Pluralis-8B, an 8.6B-parameter open pretraining run trained on 500B tokens across 330 contributor nodes.

Backed by