There is a particular kind of failure that senior infrastructure engineers recognize immediately. Nothing crashes. Nothing spikes. Nothing logs an error. And yet… work simply stops.
This is the story of one of those failures and why Build Forensics matters once build systems become production infrastructure.
One of the classic strategies to speed up a system is to avoid redundant work. Scalable build systems like Bazel and Buck2 heavily employ this strategy in various ways, with remote caching and content-addressable storage being two prominent examples. While remote caching prevents repeating redundant build actions, content-addressable storage (CAS) exists for the purpose of data deduplication. However, traditionally CAS operates at the granularity of a single file. When you modify a single byte in a file and store it in a CAS, the CAS stores a second, complete file. Deduplication at the file level is quite palatable for smaller files; a single build invocation typically contains many thousands of small files. Who cares if we store a couple more?
However, as file sizes increase, the cost of storing yet another slightly modified version of a file becomes more expensive. Instead of tossing a couple extra kilobytes into storage, you might be storing a few more gigabytes. Now, consider where these large files come from. These large files are often outputs of build actions, and those build actions depend on many smaller inputs. As those many inputs churn, the outputs also churn. Suddenly the cost of touching a tiny little source file isn’t just the cost of uploading a new version of the source file to CAS–it’s now also the cost of all the large outputs that are produced by the build. This dictates the growth rate of your CAS storage costs, which scales directly with the number of incoming builds. At AI-scale, these costs have become more important than ever before.
In this interview, Eugene Yokota—a software build expert who spent years maintaining Scala's sbt tool at Lightbend before working with hyperscaled Bazel monorepos at Twitter and Netflix—details his multi-year project to build a Bazel-compatible remote caching system into the newly released sbt 2.0. He explores the mechanics and benefits of Bazel, such as its robust remote caching and test cycle speed, and highlights how these modern, scalable build tools can eliminate CI bottlenecks for growing teams while protecting toolchain security.
Following a phenomenal annual EngFlow company summit, the Bazel community gathered at the Salesforce offices in Munich for an evening dedicated to expanding the boundaries of build systems, developer tooling, and software supply chain integrity.
With a keynote analyzing the complexities of SLSA 3 and technical lightning talks spanning Bazel integrations, IDE optimization, and parallelization, the event showcased how top-tier teams are solving today's hardest platform engineering problems.
If you missed the live event, we’ve aggregated the core architectural takeaways in this post.
The Build. Scale. Investigate. World Tour has officially begun! Uber and EngFlow co-hosted the February 2026 Build Meetup at Uber’s buzzing Amsterdam office. The event brought together developers and build specialists who shared "war stories" from the build community.
I spend a lot of time with engineering leaders and developer infrastructure teams. Over the past six months, a pattern has been emerging: build queues are experiencing demand in bursts, not waves.
Last week alone, conversations in Sydney, Chicago, Los Angeles, and San Francisco all surfaced the same story:
Sunday: A customer running EngFlow Bazel RBE at 100,000+ cores wants to triple capacity over the next year. PR volume is surging with no sign of slowing.
Monday: Prospective customer's CI queue is in crisis. Engineering attention diverted from product work. First-time inquiry about RBE.
Tuesday: Existing customer tried sharding load across more CI workers. Result: higher cloud costs, same bottleneck. Needs help making workloads RBE-compatible.
Thursday: A large enterprise customer skipped our customer dinner - they were occupied with urgent internal testing, preparing to double their RBE load.
Friday: Reviewed load planning with a customer preparing for an AI generation hackathon - forecasting what their infrastructure must absorb in the coming months.
Celebrating our team and customer success around the world.
2025 was a year of resilience, achieving multiple key financial milestones, record-breaking momentum, and exponential growth. EngFlow overcame significant personal and professional challenges to solidify its position as the market leader in Build Management at Scale. Today, we power the world’s most complex workloads for innovators like Arm, Asana, BMW, Canva, Databricks, Lyft, Plaid, Perplexity, Snap and Zoox.
The EngFlow team recently wrapped up a successful BazelCon 2025, which marked a significant milestone: the 10th anniversary of Bazel! Seeing the energy and innovation at the conference was especially meaningful for us, as several of us have been involved in BazelCon since its inception, both as organizers and attendees. We're reflecting on the immense growth the community has achieved in the last decade and how it spearheads developer productivity across the world.
EngFlow was a platinum sponsor of the main conference, facilitated several workshops on Training Day, and co-hosted a fun Game Night with JetBrains and VirtusLab. We loved seeing and hearing from many of our customers in-person at BazelCon, as Happy Customers are at the heart of what we do!
This blog brings you highlights from the conference and satellite events.