Skip to content

Training Architecture#

Two layers compose an Agora training run. Compute is the pipeline itself: workers, each holding one stage's parameters, running forward and backward with head workers originating microbatches from their assigned data shards, and every worker routing its output directly to a worker in the next stage. Coordination is the Metadata Store, a central service where peers publish the registration, progress, and matchmaking records that keep the swarm in agreement. The full system runs on heterogeneous, untrusted hardware without any one contributor holding the complete model weights.

Compute Layer
Training Pipeline
Workers hold one pipeline stage's parameters, process fwd/bwd, run periodic SPARTA state averaging within their stage.
Stage 0
Head
Pipe 0
H100 80G
4 layers + embed
Pipe 1
H100 80G
4 layers + embed
SPARTA
State averaging
Stage 1
Body 1
Pipe 0
RTX 4090
4 layers
Pipe 1
RTX 4090
4 layers
SPARTA
State averaging
…
Bodies 2–10
Pipe 0
mixed
4 layers each
Pipe 1
mixed
4 layers each
SPARTA
State averaging
Stage 11
Body 11
Pipe 0
L40S 48G
4 layers
Pipe 1
L40S 48G
4 layers
SPARTA
State averaging
Stage 12
Tail
Pipe 0
A100 40G
2 layers + lm_head
Pipe 1
A100 40G
2 layers + lm_head
SPARTA
State averaging
Coordination records →
Registration · progress · matchmaking
HTTPS, signed
Coordination Layer
Metadata Store
Central coordination service: peer registration & discovery, progress tracking, matchmaking.
Metadata Store API
Authenticated REST API
Verifies signed records
Rate limiting
CPU only
Redis
Backing store
Run-scoped keyspace
VPC-private
CPU only
Two-zone Agora architecture: Compute (Workers) and Coordination (Metadata Store). Heterogeneous example configuration; real swarms vary by participant hardware. Forward activations flow worker-to-worker Head → Body → Tail; backward activation gradients flow Tail → Body → Head along the same route. Parameter gradients never cross stage boundaries.

Compute Layer#

A worker holds the model parameters and performs the compute. It owns one pipeline stage (Head, Body, or Tail) and runs forward and backward passes for microbatches pushed to it by the previous stage. Head workers originate their own microbatches from an assigned shard of the training corpus. Workers within the same stage participate together in periodic, async SPARTA averaging rounds. The Workers section below covers the runtime structure.

Coordination Layer#

The Metadata Store holds no model parameters and touches no batch. It is a central, authenticated key/value service where every peer publishes its coordination state: stage registration and the network addresses other peers dial it on, per-stage training progress, and matchmaking records for averaging rounds. This is how peers find each other: a worker announces itself here and resolves its next hop from the same records. Every record is signed by the writing peer, and the store runs on Pluralis-owned infrastructure rather than on contributor nodes. Batch routing itself is not centralized, each worker selects a healthy worker in the next stage and pushes activations to it directly; see Batch Origination & Routing.


Component Deep Dive#

Workers#

A worker is a single process holding one stage's parameters (one or more transformer layers). It performs forward and backward on those parameters, runs its own local optimizer to apply gradients, and joins same-stage peers in periodic AllReduce rounds for state averaging. A worker has no knowledge of the rest of the pipeline beyond its own stage and the roster of the next stage it routes to.

Worker: Stage X
Connection Handlers
Listen for gRPC pushes from previous-stage workers; place batches in fwd / bwd queues. Multiplexed on the same port.
Runtime
Loops over fwd / bwd queues and dispatches batches into ModuleBackend for execution.
ModuleBackend
Stores the nn.Module for this stage. Owns the forward / backward task pools.
Stage announcements
Declares this Worker's availability in its stage (head.0.0, body1.0.1, …) for peer discovery and routing.
SPARTA Optimizer
Accumulates gradients locally, runs the local optimizer step, then matches with same-stage peers and AllReduces 5% of parameters.
Metadata Store Client
Reads and writes signed coordination records: peer discovery, expert registration, progress tracking, matchmaking.
Worker internals: six co-resident components inside a single Worker process. Runtime drives ModuleBackend for compute; the SPARTA Optimizer coordinates the parameter-averaging step with same-stage peers via the Metadata Store.

Metadata Store Client#

Agora routes all coordination state through the central Metadata Store: peer discovery, expert registration, progress tracking, and matchmaking for AllReduce. Every record is signed by the writing peer and exchanged over HTTPS; the libp2p transport carries only tensor traffic.

ModuleBackend#

The nn.Module for this stage and the forward and backward functions the Runtime invokes. Also owns the two task pools (forward and backward) where incoming batches queue up before the Runtime processes them.

Async SPARTA#

Each worker accumulates gradients from its own backward passes and runs its own optimizer step locally; there is no per-step gradient AllReduce. Same-stage replicas drift apart as a result. To re-synchronize, every 5 local steps the worker matches with same-stage peers through the Metadata Store and AllReduces 5% of its parameters. Successive rounds cover non-overlapping slices, so the full parameter set has cycled through over a 20-round window.

Connection Handlers#

gRPC listeners that receive pushes from previous-stage workers (and, on heads, the worker's own locally originated batches) and put each batch into the right queue: forward or backward. Multiple listeners share a single port.

Stage announcements#

A background thread that keeps re-announcing this worker in the Metadata Store under its stage-prefixed UID (head.0.0, body1.0.1, tail.0.0). Previous-stage workers read the announcements to pick a next hop; same-stage peers use them for matchmaking during AllReduce.

Runtime#

The main loop. Dequeues batches from the forward and backward queues and runs them through ModuleBackend. On the backward path it rebuilds the autograd graph by re-running the forward (see activation recomputation for details), calls torch.autograd.backward(), and triggers the optimizer step at the appropriate point in the batch-size accumulator.


Batch processing#

Once running, the Worker sits in an event loop processing batches:

  1. A previous-stage worker pushes a microbatch via gRPC (on a head, the worker's own data loader injects it) → Connection Handler places it in the forward queue.
  2. Runtime dequeues the batch → calls ModuleBackend.forward() → the worker pushes the output activations to a worker it selects in the next stage. On the tail there is no next stage: it computes the loss and starts the backward pass instead.
  3. The next-stage worker pushes back the activation gradients → Connection Handler places them in the backward queue.
  4. Runtime dequeues the batch → calls ModuleBackend.backward() → triggers the optimizer step and pushes grad_input to the previous hop.

See also#

  • How a new Worker joins a running swarm (state download, queue, sync mode) → Contributor Join Flow.
  • What happens at the optimizer step (ProgressTracker, Matchmaking, SPARTA AllReduce) → Communication Patterns.
  • Sync-mode entry / exit conditions and Worker-failure handling → Fault Tolerance.

Batch Origination & Routing#

There is no separate coordinator process driving the training loop. The pipeline drives itself: head workers originate microbatches from their assigned slice of the training data, every worker selects the next hop for its own output, and the tail closes the loop by computing the loss and starting the backward pass.

pretokenized data shard (pithos) loss → starts backward STAGE 0 Head embed · early layers Pipe 0 / Pipe 1 STAGE 1 Body 1 transformer block Pipe 0 / Pipe 1 STAGES 2–11 Bodies 2–11 transformer blocks Pipe 0 / Pipe 1 STAGE 12 Tail final layers · loss Pipe 0 / Pipe 1 Forward: activations · Head → Body → Tail · libp2p gRPC Backward: grad_input · Tail → Body → Head (same route, reversed)
Microbatches originate on head workers and flow directly between stages over libp2p gRPC. Forward activations flow Head → Body → Tail; backward activation gradients retrace the route Tail → Body → Head. Two pipes per stage give data-parallel redundancy.

Batch origination (heads)#

Each head worker streams a pretokenized shard of the training corpus via Pithos, assigned to it by the auth service when it joins. A background thread prefetches samples into a bounded queue, and the head injects microbatches into its own forward path. This is the same path a remote push takes, so accounting, routing, and failure handling are identical for local and remote batches. The head also collects the loss, completion, and drop reports for the batches it originated, and publishes the run's training metrics.

Next-hop selection#

Each worker maintains a live roster of the next stage's workers (refreshed from the Metadata Store every ~30s) and selects a hop per microbatch using a min-heap keyed by accumulated virtual runtime: the least-loaded worker is selected. Reserving a hop immediately charges that worker's accumulated runtime with the task's estimated duration, so concurrent sends spread across the stage rather than piling onto one worker; when the receiver's receipt arrives, the measured time updates its throughput average. New arrivals enter at the current maximum. The receiving worker acknowledges every push as accepted, busy, or dropped: a busy peer is simply routed around, while a failed or unreachable hop is short-banned from the heap. A pipeline step only stalls if a stage empties entirely: every worker in that stage unreachable or short-banned at once.

Per-microbatch flow#

  1. A head worker draws the next sample from its data shard and runs its own forward pass.
  2. Each stage pushes its output activations to the selected worker in the next stage, which runs forward in turn.
  3. The tail computes the loss (next token prediction) and reports it back to the originating head.
  4. The tail immediately starts the backward pass: activation gradients retrace the route in reverse, and each head or body worker recomputes its forward from the cached input, runs backward, and pushes grad_input to the previous hop.

The full protocol for how workers announce and refresh each other is in Communication Patterns → Periodic node announcement.