For the last few months, I’ve been working on a vendor-agnostic inference library. Have 3 backends so far: legacy D3D11 for compatibility, D3D12, and Vulkan 1.3. I have reasons to believe nVidia deliberately crippling Vulkan API for their consumer GPUs. Couple examples to be specific.
nVidia driver sets quite low number for VkPhysicalDeviceLimits::maxTexelBufferElements. I don’t think that’s a hardware limit because I have D3D12 backend doing the same thing on the same hardware.
Vulkan performance is not great on nVidia. On all AMD cards I am testing, Vulkan 1.3 is the fastest backend. On nVidia however, D3D12 is faster despite tensor cores (WMMA / wave matrix multiply accumulate / cooperative matrices) are only available through Vulkan, D3D12 is using shader cores exclusively.
Yes, they could just say they're taking it towards prioritizing models that only run on one kind of equipment and making it harder to find, support or list models that run on different hardware.
The acquisition sets up corporate interests to be able to increasingly drive it.
I hope better for Nvidia and hopefully it can be cleared up, and maintained.
HF CEO keeps posting about how great local models are running in his Mac, it would be funny to see if he replaces his MacBook with a 500W Nvidia laptop after the deal gets through.
But more importantly for business, if HF fails to use the upcoming ASIC inference chips for their ZeroGPU after the acquisition, their competitors using them might have an advantage.