2 min read
Add as a preferred source on Google

GitHub Rebuilds Storage Layer as September Commits Top 7.38 Billion

The redesign will move authoritative repository data to Azure Blob Storage and separate durable storage from request-handling compute.

Archival storage containers arranged inside a cool industrial facility / TokenPost.ai
Archival storage containers arranged inside a cool industrial facility / TokenPost.ai

GitHub is rebuilding the storage infrastructure behind its code-hosting platform after repository activity surged, with 7.38 billion commits recorded in September 2026 and AI-agent workflows accelerating since late 2025.

The September total, produced by developers and AI agents combined, was more than five times the level recorded a year earlier. GitHub has not separated the portion generated by AI agents.

Overall Git activity rose from 218.2 billion events in September 2025 to 473.3 billion in August 2026. Monthly pushes increased 4.9-fold, from 690 million to 3.35 billion. Merged pull requests approached four times their year-earlier level, while GitHub Actions executions exceeded four times that level, reaching 3.26 billion in September. The busiest individual repository received about 1 billion requests in August.

The redesigned architecture will place authoritative repository data in Azure Blob Storage and separate durable storage from the computing resources that handle requests. Lightweight workers will cache data for reads and scale independently as demand changes.

The system stores five complete replicas of each repository by default across local disks on multiple file servers. Those replicas provide redundancy and support read traffic, but they also participate in writes. Updating a branch reference requires a three-phase commit in which a majority of replicas must confirm the change.

That design makes a push dependent on the slowest replica. Adding replicas can expand read capacity while slowing writes, and a failure to reach a quorum can stop a write altogether.

AI-agent workloads increase the pressure. Agents may commit or save checkpoints after nearly every action, while thousands of agents can work on separate branches in the same repository. Merges can converge on the main branch, and continuous integration and code-scanning systems can fetch the same branch thousands of times per minute after each push.

Under the proposed design, storage and compute can scale separately. A failed compute worker would be handled more like a cache miss, allowing a replacement worker to retrieve data from durable storage. Compression and garbage collection would run on separate background workers.

Internal benchmarks showed up to 35 times higher write throughput, with read capacity scaling independently. The results come from internal testing and do not represent production performance.

The storage rebuild is part of a broader infrastructure effort involving Azure migration and the separation of shared databases. April and May availability updates recorded 10 incidents and nine incidents, respectively, with several disruptions involving databases, traffic limits, deployments or load balancers rather than Git storage directly.

The redesign has no announced completion date. Further details will depend on a future engineering update and later availability reports.

Simon Yoon

Reporter

Simon Yoon reports on blockchain technology for TokenPost. Send corrections or tips to info@tokenpost.com.

Loading…