Then all interrupts need to be prepared for it and check for it anyways. Crucially, the code path that does the check might also get interrupted, even multiple times. No synchronization is possible, and even if it was, it'd be slower than simply saving the registers.
It doesn't matter if that code path gets interrupted. There's no need to synchronize anything. If interrupt A sees that the extra registers are idle, and then gets interrupted, that's okay. By the time the inner interrupt returns, even if the inner interrupt used them, they will be put back to idle, and A can safely switch to them. Everything else works the same as current hardware.
There's the cost of a branch, but you could remove that by making it a hardware feature. It's a branch that should predict well though.
Only when an interrupt actually does get interrupted. You could make the common case faster. Not that it matters if PCI-e is the bottleneck.