Konstantin MishchenkoResearch Scientist, Meta FAIRDistributed Optimization StoriesICML Workshop: Protocol Learning
July 10, 2026InterContinental Grand Hotel, Seoul Parnas
Pluralis Research brought together academic and industry leaders during ICML to convene on the emerging paradigm of Protocol Learning - decentralized, communication-efficient, model-parallel training of foundation models.
Workshop
Researchers working on related topics joined us for a series of talks, followed by a selection of posters tackling various parts of the decentralized training stack. The workshop was organized in collaboration with Professor Namhoon Lee and POSTECH.
Protocol Learning
Training frontier foundation models today demands massive, co-located clusters of high-end GPUs - accessible only to a handful of the most well-resourced organizations. Protocol Learning removes this co-location requirement, enabling multi-participant training of foundation models across open, permissionless networks of globally distributed compute, where no single participant has, or can ever obtain, a full copy of the model.
This requires solving hard open problems in low-bandwidth model parallelism, asynchronous distributed optimization, supporting heterogeneous hardware, fault-tolerant training systems, Byzantine robustness, and trustless verification. The workshop convened the researchers advancing these building blocks to define the challenges ahead and chart a research roadmap for training the next generation of community-owned frontier models with self-sustaining economics.
Talks
Lightning Talks
Photos















Poster Sessions
Sungbin Shin, Hyunji Jung
Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis RotationZhiwei Bai
Adaptive Preconditioners Trigger Loss Spikes in AdamJin Lee
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUsAndrej Jovanović
LoRDO: Distributed Low-Rank Optimization with Infrequent CommunicationEgor Shulgin
General Analysis of LMO-based Optimizers: Beyond Bounded VarianceXingyu Qu
Can Muon Fine-tune Adam-Pretrained Models?Benjamin Thérien
MuLoCo: Muon is a Practical Inner Optimizer for DiLoCoPaul Janson
Stabilizing Native Low-Rank LLM PretrainingZhuoli Ouyang
RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based OptimizationJeffrey T. H. Wong (Imperial College London)
A3: an Analytical Low-Rank Approximation Framework for AttentionPhilip Zmushko (ISTA), Egor Petrov (Yandex Research)
One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM PretrainingHan Shi (Huawei)
POET-X: Memory-efficient LLM Training by Scaling Orthogonal TransformationDongyeop Lee
FRESCO: A Novel Consistency Control for Asynchronous Pipeline Parallel Training











