Nandan Nayampally is Chief Commercial Officer at semiconductor IP company Baya Systems.

​The IPO market has made a massive comeback recently, first with Cerebras’ IPO giving it a market cap of nearly $100 billion, followed by SpaceX securing the largest IPO in history a week after announcing its deal to sell compute to Google. It’s no coincidence that semiconductor connections strengthened both companies’ IPOs; Cerebras takes a unique approach with its wafer-scale products while SpaceX is specifically driving a strong semiconductor supply chain message with its TeraFab Initiative and selling access to its AI chip stockpile.​

Cerebras and SpaceX exemplify the two main paths being pursued. SpaceX is, of course, trying to solve the supply side by rapidly building fabs derived from Tesla’s successful GigaFactory approach to vertical integration under the same roof. On the other hand, Cerebras is taking the memory wall head-on by collocating memory on-die close to the processors themselves.

These successes reflect, in two different ways, the “memory wall” problem: how we can get to a place where compute doesn’t have to wait for memory to supply the necessary data to the right place, and what we’ve seen manifested as an acute supply capacity problem with high-bandwidth memory (HBM).

Overcoming the memory wall problem is what’s standing between us and the next generation of AI compute. And right now, the chip industry is taking a number of different approaches to solve this problem.

Wafer-Scale: The Radical Approach​

Generally, a semiconductor wafer (up to 12 inches in diameter) is diced up into chips or dies, and any defective parts of the wafer can be discarded. This often produces hundreds of high-performance chips from one wafer.

Cerebras takes this approach and doesn’t break the wafer up at all. The wafer is the chip, and the chip is the wafer.​ At this size, almost every chip design principle no longer applies. Consequently, Cerebras’ biggest breakthrough had to go all the way down to the foundational layer of a chip, the substrate. Their on-wafer 2D mesh fabric spans the entire surface of the wafer, letting any of its 900,000 AI cores talk to any other directly without data leaving the silicon.

In the more conventional approach, data has to move through interconnects (thin wires) from the memory where it’s stored to the compute chip that’s asking for it. That adds time to any computing task.

Wafer-scale chips do come with trade-offs, however. They remain exceptionally expensive to manufacture, and the technology barrier is immense, which makes them difficult to scale up. There are plenty of critical issues to address, such as how to work around defects on the wafer (which always exist) and how to deliver power from the edge of the wafer to inner computation, which magnifies what designers call the IR-drop problem many times over. They need even more cooling than the systems that generally power AI compute today.

Until these issues are addressed, it’s unlikely that widescale adoption will occur.​

CoWoS: The More Conventional Approach​

Conversely, much of the industry, led by Nvidia, has embraced the chip-on-wafer-on-substrate (CoWoS) solution to the memory wall. Its Blackwell architecture and offerings are an excellent example of this method, which arranges individually validated chips and memory cores, then layers a high-speed silicon interposer between the dies and package substrate to connect it them. This is a building-block approach that scales well, offers modular flexibility and creates bandwidth gains.​

However, CoWoS doesn’t scale the memory wall. What it does is relocate it. Higher bandwidth between memory and compute makes handling petabyte-scale AI workloads possible, but the need for data still increases. Consequently, the gap between how fast memory can supply data and how fast compute can process it re-emerges, with an added complication of the increase in energy consumption to move data around.

Until recently, architectural choices like bandwidth allocation or interconnect topology (how connections are arranged on a chip or in a system) were generally not addressed until physical integration began. Designers commit silicon, discover bottlenecks and work around them at the cost of both time and capital, which is quickly becoming unacceptable at the breakneck pace of chip development.​​

System-Level Design: The Fundamental Rethink​

CoWoS and wafer-scale are just two of the tech breakthroughs aiming to scale the memory wall. There’s simply not enough space in this article, or even several articles, to discuss all the different innovations, such as HBM stacking, die-to-die signaling or chiplet interconnects.​

However, the fact that so many possible solutions exist points to a deeper truth. The architects building the next generation of systems for AI workloads can’t look at thermal management, energy efficiency, bandwidth allocation, interconnect topology or power delivery as individual, separate concerns that can be addressed in turn or wait until physical silicon is involved.​​

Instead, each of these factors must be treated as components of a larger, unified system from the beginning. The memory wall and data movement speeds place constraints on every facet of a system. Only when they’re addressed by every facet of a system will we see real breakthroughs instead of incremental improvements.​

Cerebras bet the entire wafer. Nvidia bet on modularity. Both bets are really the same wager: that whoever rethinks the system first, rather than the chip first, has the inside track for the next decade of AI compute. The memory wall isn’t just a silicon problem anymore. It’s an architecture problem, and architecture problems don’t get solved at the transistor level. They get solved by rethinking the whole stack, while still shipping product every quarter to meet the ever-increasing demands of AI.​

Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

Share.
Exit mobile version