# Chutes Documents 20B-Parameter Tests Across Disconnected GPUs

By Simon Yoon

Canonical URL: https://www.tokenpost.com/news/technology/25931
Published: 2026-09-30T17:59:20.000Z
Updated: 2026-09-30T17:59:20.000Z
Section: Technology

> Its Parallax system distributes sparse mixture-of-experts training across geographically dispersed hardware, while an 8-billion-parameter run has been publicly tracked since Sept. 10.

Chutes AI is positioning distributed consumer hardware as an alternative to the tightly coupled GPU clusters normally used to train large language models.

Its Parallax system divides sparse mixture-of-experts training across machines that may be separated by geography and connected through ordinary networks. The approach assigns each participant a subset of the model’s experts while using smaller surrogate versions for experts hosted elsewhere, reducing the need for constant data exchange between GPUs.

Chutes has documented earlier Parallax experiments involving models with up to 20 billion parameters and a mix of H100, L40S, RTX 6000 Ada and RTX 4090 hardware. Parallax reduces the need for continuous communication between remote GPUs, allowing local training steps to continue with less synchronization.

Chutes began publicly tracking an 8-billion-parameter mixture-of-experts training run on Sept. 10. The estimated cost was about $11 per billion tokens processed.

Parallax’s published evidence remains primarily a systems demonstration rather than a fully replicated production benchmark. The work records validation-loss results from individual 20B experiments and does not provide measured wall-clock speedups or replicated equivalence tests.
