> The weakness of GPU databases is that while they have fantastic internal bandwidth, their network to the rest of the hardware in a server system is over PCIe, which generally isn't going to be as good as what a CPU has and databases tend to be bandwidth bound. This is a real bottleneck and trying to work around it makes the entire software stack clunky.
How relevant is that when you're looking at multi-TB data sets that don't fit into computer RAM? Sure, the RAM <---> CPU bandwidth may be very wide, but the SSD connects to the computer over the same PCIe bus.
And also: when did you have this conversation? GPU performance has changed very much year by year, so what wouldn't have been worth it 2 years ago might be a huge gain now.
The difficulty with how most GPUs are connected to the rest of the system is that the data has to go RAM -> CPU -> GPU. If it could go directly RAM -> GPU, then the calculations would be better, but still not great as PCIe is still lower bandwidth and higher latency than RAM -> CPU.
It's not about GPU performance, it's about the latency and bandwidth of getting that dat to the GPU. If once you ship data to the GPU, you reuse it many times for many calculations, that cost is amortized and it doesn't matter as much. But if you ship data to the GPU and use it once, then that cost will probably not be amortized. I think of databases tending to fit in the latter category.
How relevant is that when you're looking at multi-TB data sets that don't fit into computer RAM? Sure, the RAM <---> CPU bandwidth may be very wide, but the SSD connects to the computer over the same PCIe bus.
And also: when did you have this conversation? GPU performance has changed very much year by year, so what wouldn't have been worth it 2 years ago might be a huge gain now.