dispatch / sally-mckee-memory-wall-ai-infrastructure
Sally McKee Saw the Memory Wall Coming in 1994. AI Just Ran Into It at Full Speed.
A PhD student named the biggest bottleneck in modern computing three decades ago. Nobody listened. Now it's why your phone costs more and AI training runs cost billions.
# Sally McKee Saw the Memory Wall Coming in 1994. AI Just Ran Into It at Full Speed.
A post hit Hacker News yesterday with the kind of headline that usually gets scrolled past: Sally McKee, who coined the term "the memory wall", has died. She passed away in February 2025 at 61. The obituary was modest. The concept she named was anything but.
In 1994, McKee — then a PhD student at the University of Virginia — co-authored a paper with her advisor William Wulf called Hitting the Memory Wall: Implications of the Obvious. The argument was devastatingly simple: CPUs were getting faster exponentially, but memory access latency wasn't keeping up. At some point, processors would spend most of their time doing nothing — just waiting for data to arrive.
The response from the computer architecture community? Crickets. Everyone was busy optimizing cache hierarchies and branch predictors. The paper was "mostly ignored," as one HN commenter put it, because "most of the scholars in her field were concentrating on caches." McKee herself said later that she was most proud of this work. She had named a structural problem that would outlast every incremental optimization.
The Wall Is Taller Now
Here's why this matters today. The memory wall didn't go away — it got higher. Modern AI training runs aren't limited by how many math operations a GPU can perform. They're limited by how fast you can move data into the chip.
A transformer model doing inference isn't compute-bound. It's memory-bandwidth bound. Every token generation requires loading massive weight matrices from memory, doing a relatively small amount of math, then writing results back. The ratio of data movement to computation is absurd — which is exactly the imbalance McKee identified in 1994, just scaled up by a factor of a billion.
This is why Nvidia's H100 and B200 chips dedicate enormous die area not to CUDA cores, but to HBM controllers. It's why SK Hynix and Micron are printing money supplying High Bandwidth Memory stacks. It's why a single Nvidia rack-scale platform with 36 Vera CPUs and 72 Rubin GPUs consumes enough RAM for 4,600 flagship smartphones.
The memory wall is also why your phone got more expensive. When AI data centers vacuum up DRAM at that scale, the supply crunch cascades into every device with memory in it. Samsung Semiconductor just posted its best quarter ever while Samsung Mobile warned of its first-ever net loss — same company, same memory, two directions.
We Keep Rediscovering What We Already Named
McKee's biography reads like a map of computing's major transitions: DEC, Microsoft, AT&T Bell Labs, Cornell, Chalmers, Clemson. She worked on cybersecurity in her later years, but the obituary notes that she remained most proud of that 1994 paper. She had identified the bottleneck that would define not just her field, but the entire trillion-dollar AI infrastructure buildout three decades later.
The HN thread has a comment that stuck with me: "It's weird to think that all the HBM research, CXL, processing-in-memory, and model parallelism tricks are just different ways of attacking the exact problem she named in a six-page paper." He's right. We're building space elevators to get over a wall that a PhD student correctly identified while everyone else was rearranging cache lines.
What This Actually Means
There are two ways to read this. The optimistic one: computer architecture has spent thirty years developing increasingly clever workarounds — caches, prefetchers, HBM, CXL memory pooling, quantization, sparsity — and they've kept the industry moving. The pessimistic one: those are all band-aids on a structural wound that hasn't healed, and every new computing paradigm (multicore, GPUs, AI accelerators) just rediscovers the same pain.
AI is the latest paradigm to hit the wall. Training a frontier model doesn't require more FLOPs — it requires more bandwidth to feed the FLOPs you already have. That's why energy consumption, chip area, and cost all scale non-linearly with model size. You're not paying for computation. You're paying to move weights around fast enough that your $40,000 GPU doesn't spend 80% of its time idle.
Sally McKee gave us the vocabulary to understand this in 1994. The rest of the industry took thirty years to catch up to her vocabulary. And if the current AI infrastructure arms race proves anything, it's that we're still climbing the same wall she pointed at — just with more expensive ropes.
HN: Sally McKee obituary discussion · Online Tribute · Original Paper (PDF)