The Slow Part Is Between the Boxes
Draw any large, slow system as boxes and arrows, then ask each box how it’s doing. Every one will say it’s fine, and most of them will be right. The slowness lives in the arrows: the handoffs, the queues, the waiting between parts. Each part optimizes for itself, and nobody owns the space in between.
Every box is fine
Picture a small feature moving through an engineering org. An engineer builds it in a day. The pull request waits a day and a half for review. It merges, then waits for QA, then sits until the weekly release train. One config change needs a ticket to another team, and that queue is a week deep.
Every team in that story hits its own targets. The engineer was fast. Review landed inside the agreed window. QA cleared its queue. The release train left on schedule. And a one-day change took three weeks to reach a customer.
Add up the time someone was actually working on it and you get maybe two days. The rest is waiting. Lean practitioners call that ratio flow efficiency, and two days out of fifteen working days is an uncomfortable number to put on a slide.
Software has the same shape. A request crosses six services, and each team’s dashboard shows a healthy p50. The trace shows something else: spans that add up to 120 milliseconds inside a request that took 400. The missing time is in the gaps, in connection pool waits, a thread pool queue, serialization, and a retry nobody remembers adding. Every service is fast. The call graph is slow.
Local optimization makes it worse
When each part optimizes its own numbers, the whole often gets slower.
A reviewer batches code reviews into one block a day to protect focus time. Reasonable for them. Every author on the team now waits up to a day per round of feedback.
A manager keeps everyone 100% busy so no capacity goes to waste. Queueing theory has bad news for that plan: as utilization approaches 100%, wait time climbs steeply toward infinity. A fully loaded team is a team with a long line in front of it.
A service adds aggressive retries to make its own error rate look better, and multiplies the load on a dependency that was already struggling.
Nobody in those examples is doing bad work. W. Edwards Deming is often quoted as saying a bad system will beat a good person every time. These are good people inside a system that rewards each of them for making the arrows longer.
Conway’s Law makes it worse. Melvin Conway observed that organizations design systems that mirror their own communication structures. The handoff between two teams becomes an API call between two services, so the slow arrow in the org chart shows up a second time in the architecture. It works in the other direction too: decide the team boundaries first and let the services follow, which I argued for in Microservices Are a Team Decision First.
Measure the arrows
To find the gaps, measure end to end.
For a delivery process, track lead time from first commit to running in production, and split it into time spent working and time spent waiting. Most ticket trackers already record the timestamps you need. For a running system, use distributed tracing, and look at the whitespace between spans as hard as you look at the spans.
Then find the constraint. Eliyahu Goldratt’s Theory of Constraints, from The Goal, rests on the idea that a system moves only as fast as its bottleneck. Improving anything else feels productive and changes nothing end to end. Making the fastest service faster, or the fastest engineer faster, is the classic way to get busy without getting quicker.
Fix the connections
Once you can see the arrows, most of the fixes remove them or shorten them:
- Fewer handoffs. A team that can ship a change end to end beats a relay of four specialists, even when each specialist runs their leg faster.
- Smaller batches. A weekly release train makes every change wait on the slowest one in the batch. Ship each change on its own.
- Less work in progress. Finishing one thing before starting the next shrinks every queue that work passes through.
- Self-service instead of tickets. This is a big part of why a good platform raises the floor: a paved path replaces a ticket and a week-deep queue with something a team can do on its own.
- Fewer hops. A chatty sequence of six service calls can often become one.
Some of these make a single part look worse. The reviewer who drops what they’re doing to review a PR has a choppier day. A team holding spare capacity looks underused on a utilization chart. Those are good trades when the whole system gets faster, and you only see that when you measure the whole system.
Next time something big is slow, start with the arrows.