Perspectives
Bigger Networks Need Denser Compute
AI systems are getting larger, but larger does not automatically mean more efficient.
As more XPUs are added, more data has to move between them. Models are partitioned, accelerators synchronize, and intermediate data moves across the fabric.
At that point, system performance depends on more than compute.
It increasingly depends on three things:
Bandwidth. Latency. Scaling efficiency.
More XPUs do not guarantee proportional performance
In an ideal world, doubling the number of XPUs would double performance.
In reality, communication overhead grows with system size.
XPUs spend more time exchanging data, waiting for dependencies, and synchronizing with one another. The larger the system becomes, the harder it is to keep all of the compute busy.
That gap between theoretical compute and useful application performance is a scaling-efficiency problem.
And that is why the network is becoming such an important part of AI system architecture.
The industry is already building better networks
The industry clearly understands this challenge.
Scale-up fabrics are becoming faster. Switch radix is increasing. Links are getting wider. Collective operations are being accelerated in the network. Optical connectivity is moving closer to compute.
These are all important improvements.
Higher radix allows more XPUs to communicate within a smaller number of network stages. Higher bandwidth moves more data. Better switches and software improve utilization.
But there is another lever that deserves more attention:
compute density.
Higher radix and higher density solve different parts of the problem
Higher radix helps a network connect more endpoints.
But if those endpoints are still spread across boards, trays, and racks, data must still travel across those physical boundaries.
That introduces longer channels, connectors, retimers, switches, cables, and additional network stages.
At large scale, many of these components are unavoidable.
The question is not whether we still need networks.
The question is:
How much of the most frequent XPU-to-XPU communication really needs to travel through the larger network?
There are two complementary ways to improve AI system scaling:
Build a better network.
Increase radix, bandwidth, routing efficiency, and optical reach.
Build denser compute clusters.
Bring frequently communicating XPUs physically closer together.
We believe both are necessary.
Density improves usable bandwidth and latency
Raw link bandwidth is only part of the story.
What matters to the application is how much bandwidth is actually available between the XPUs that need to communicate.
The same is true for latency.
In a distributed system, data may travel through several layers:
XPU → board → switch → cable → switch → XPU
Modern fabrics are designed to make this path extremely fast.
But a denser compute domain can change the path itself:
XPU → local fabric → XPU
That does not eliminate routing, synchronization, congestion, or software complexity.
It simply gives the system architect another option:
keep more communication local.
The goal should be scaling efficiency
Imagine doubling the number of XPUs.
You have doubled theoretical compute.
But if communication overhead rises substantially, useful performance may increase by much less than 2×.
You built a larger machine.
You did not necessarily build a proportionally faster one.
So the goal should not simply be:
Connect more XPUs.
It should be:
Connect more XPUs while preserving as much of their performance as possible.
Higher radix helps.
Higher bandwidth helps.
Better software helps.
And higher compute density can help as well.
Collapse part of the scale-up network
This is the system-level question we are exploring at Plaid:
What if we could collapse part of the scale-up network by bringing the XPUs much closer together?
Not by replacing the larger network.
And not by arguing that current scale-up architectures are wrong.
Instead, by introducing a denser local compute domain underneath them.
That is the idea behind Plaid's Rack-on-Package architecture: integrate multiple complete XPUs into a system-scale fabric using short-reach, high-bandwidth connectivity.
The larger network still exists.
But more communication can happen locally before traffic needs to move into the broader rack- or data-center-scale fabric.
Bigger networks need denser compute
The AI industry will continue building faster and larger networks.
It should.
But network scale alone is not the objective.
Scaling efficiency is.
The next generation of AI systems may therefore need to advance in two directions at once:
Build networks that can reach farther.
And:
Build compute clusters that do not need to reach as far.
That combination—higher-radix networks and denser compute domains—could improve bandwidth, latency, and ultimately how effectively large AI systems use their most valuable resource: the compute itself.
