Crazy Wisdom
Episode #522: The Hardware Heretic: Why Everything You Think About FPGAs Is Backwards
- Robots & Atoms
- FPGA accelerated servers
- Saturn data
- vector databases
- search applications
- memory bandwidth
- DRAM and flash storage
- sparse access patterns
- distributed systems
This video isn't available here. Try YouTube, or listen to the audio below.
Watch on YouTubeAbout this episode
Timestamps
- 00:00Introduction to FPGAs and Their Role in Servers
- 02:47Understanding FPGA Limitations and Use Cases
- 05:55Exploring Different Types of Servers
- 08:47The Importance of Memory and Bandwidth
- 11:52Philosophical Insights on Search and Access Patterns
- 14:50The Relationship Between Hardware and Search Queries
- 17:45Challenges of Distributed Systems
- 20:47The CAP Theorem and Its Implications
- 23:52The Evolution of Technology and Knowledge Management
- 26:59FPGAs as IO Expanders
- 29:35The Trade-offs of FPGAs vs. ASICs and GPUs
- 32:55The Future of AI Applications with FPGAs
- 35:51Exciting Developments in Hardware and Business
Key Insights
1. FPGAs are fundamentally "crappy ASICs" with serious limitations - Despite being programmable hardware, FPGAs perform far worse than general-purpose alternatives in most cases. A $100,000 high-end FPGA might only match the memory bandwidth of a $600 gaming GPU. They're only valuable for specific niches like ultra-low latency applications or scenarios requiring massive parallel I/O operations, making them unsuitable for most computational workloads where CPUs and GPUs excel.
2. The real value of FPGAs lies in I/O expansion, not computation - Rather than using FPGAs for their processing power, Saturn Data leverages them primarily as cost-effective ways to access massive amounts of DRAM controllers and NVMe interfaces. Their server design puts 200 FPGAs in a 2U enclosure with 1.3 petabytes of flash storage and terabyte-per-second read bandwidth, essentially using FPGAs as sophisticated I/O expanders.
3. Access patterns determine hardware performance more than raw specs - The way applications access data fundamentally determines whether specialized hardware will provide benefits. Applications that do sparse reads across massive datasets (like vector databases) benefit from Saturn Data's architecture, while those requiring dense computation or frequent inter-node communication are better served by traditional hardware. Understanding these patterns is crucial for matching workloads to appropriate hardware.
4. Distributed systems complexity stems from failure tolerance requirements - The difficulty of distributed systems isn't inherent but depends on what failures you need to tolerate. Simple approaches that restart on any failure are easy but unreliable, while Byzantine fault tolerance (like Bitcoin) is extremely complex. Most practical systems, including banks, find middle ground by accepting occasional unavailability rather than trying to achieve perfect consistency, availability, and partition tolerance simultaneously.
5. Hardware specialization follows predictable cycles of generalization and re-specialization - Computing hardware consistently follows "Makimoto's Wave" - specialized hardware becomes more general over time, then gets leapfrogged by new specialized solutions. CPUs became general-purpose, GPUs evolved from fixed graphics pipelines to programmable compute, and now companies like Etched are creating transformer-specific ASICs. This cycle repeats as each generation adds programmability until someone strips it away for performance gains.
6. Memory bottlenecks are reshaping the hardware landscape - The AI boom has created severe memory shortages, doubling costs for DRAM components overnight. This affects not just GPU availability but creates opportunities for alternative architectures. When everyone faces higher memory costs, the relative premium for specialized solutions like FPGA-based systems becomes more attractive, potentially shifting the competitive landscape for memory-intensive applications.
7. Search applications represent ideal FPGA use cases due to their sparse access patterns - Vector databases and search workloads are particularly well-suited to FPGA acceleration because they involve searching through massive datasets with sparse access patterns rather than dense computation. These applications can effectively utilize the high bandwidth to flash storage and parallel I/O capabilities that FPGAs provide, making them natural early adopters for this type of specialized hardware architecture.
Episode transcript
Welcome to the Crazy Wisdom Podcast. This podcast is for you. If you have an insane drive to find the truth of things, it's not the good answers that we seek, but the good questions. I interview a range of different guests from many different fields, all with the intention to uncover the simple truths that are hidden in plain sight. Most people don't want to go there. I go there, my guests go there, and you benefit. Please let me know if you enjoy these episodes and as always, subscribe on itunes, Spotify, or wherever you listen to the podcasts.
Welcome to the Crazy Wisdom Podcast. I've got Peter Schmidt Nielsen here and he is making an FPGA accelerated server at Saturn Data. Welcome to the show.
Yeah, so it's so interesting. I did a call like a year ago with somebody who was making FPGAs and I had no idea what he was talking about, which is usually a good sign for me that I had to kind of dig in. And FPGAs have started to show up everywhere. I don't know if it's just because I've started to think about them more, so now they're starting to show up everywhere. But I saw a video of yours a few months ago talking about them and I was like, okay, well, there's another thing I don't really understand. So let's try to understand it. Why does a server Need a FPGA? What's the relationship between servers and FPGAs?
Yeah, I mean, it's a great question. A lot of people have tried to sort of make FPGA accelerators for servers. They're of course, the Alveo cards from Xilinx or one of the most popular examples of this. I think in general, probably the best starting point is some common misconceptions about FPGAs. So the first thing I'm going to say, which is probably a little bit surprising as somebody who's doing a whole startup based around FPGAs, is that in many ways they really suck. But like, I think a lot of people end up quite disappointed by what they can do. Like they think, wow, this is this, you know, custom hardware that I can, you know, customize. It's like an asic, but a little bit crappier. But the ratio there is enormous. I mean, the gulf is just huge. You'll find people who are not aware that like, you know, you take some mid end FPGA and you try to put like a little RISC V core on it and Maybe you close timing at like 2300 MHz and you compare that and say, wait a moment, this is giving me like, comparable performance to this, like, you know, dinky little $2 microcontroller. Like, they really are quite niche. So if you want to use an fpga, you probably have one. Stop me if I'm going on too long. But if you want to use an fpga, they're kind of, from my perspective, two main reasons. One is you have some sort of latency budget requirement that makes it so there's no other choice. This is sometimes the case in sort of high frequency trading applications. And the other is that you had some sort of IO requirement where you have no other choice. Like, if you want to talk to a thousand different pins and do some very intricate timing between all of them, be like, oh, I'm doing some crazy data acquisition application, I'm in some Watti lab and I'm receiving at like, you know, 800 megahertz on 1,000 wires in parallel. FPGAs are probably, if not your only choice, one of your most effective choices. So there's been a long history of massive disappointment with trying to accelerate real workloads on FPGAs because it turns out that if you just want to do something computational, general purpose CPUs and now GPUs are pretty incredible. So there are various companies and startups. There's a long series of sort of, I don't know if you want to say, like a graveyard of companies who say, we're going to put fpga, so we're going to accelerate your database, we're going to this, we're going to accelerate that. I think I tend to be, in general pretty skeptical of a lot of that, having used FPGAs, not, you know, as much as you can find. Plenty of people on Twitter have used them more than I have, but I'm not a complete beginner. You really need to find something that works to their strengths. Now, having said all of that and said that FPGAs are in many ways pretty anemic, I am putting them in servers more for that second role. As to you want to talk to a lot of IO, well, what is the IO you want to talk to? In my case, it's DRAM and Flash. So the core thesis of my company is that there are a lot of applications that want to have really large working sets that want to have a lot of memory bandwidth to, you know, a lot of memory, do some, some sort of operation on it, for which using FPGAs basically as IO expanders is pretty cost effective. They're a relatively cheap way to get a lot of memory bandwidth and in particular bandwidth to flash. So in a 2U enclosure we're putting about 1.3 petabytes of flash and most usually we're going to have about a terabyte a second of read bandwidth to that Flash. Right now we're mostly talking to sort of vector, database and search companies. There are a lot of potential applic, but I think vector database and search are really well suited to this, to our particular architecture because, you know, you get a huge amount of bandwidth to flash with gajillion iops because you have a ton of parallel NVME interfaces. So that's sort of like the high level, I guess a little bit rambled there.
Your first full transcript is free. After that, an email opens every transcript in the index — a list of readers we can write to, not a guest book.