Can you elaborate on both there being 180 registers and the CPU does renaming of them? Surely the RAX register in an X86_64 CPU is always the same physical location no?
> Surely the RAX register in an X86_64 CPU is always the same physical location no?
No, it is not.
Modern CPUs all use Physical Register Files or PRFs for implementing OoO. The way they work is that there is one backing register file of ~150-200 registers, and a naming table in the frontend. Each register file in the backing table can be in one of 3 states -- waiting for data, contains data, or clean. Every time an instruction that writes data to a register is executed, during the register rename stage the renamer picks one register from the clean set, assigns that as the output of the uop and the current value of the logical register, and sets it as waiting for data.
That is, let's say you execute:
add RAX, RAX
and the current value of RAX in the rename table was PRF#1, with PRF#1 carrying data and all other physical registers clean. In the renamer, this would then turn into:
add PRF#2, PRF#1, PRF#1
if this instruction was immediately followed by another add, RAX, RAX, it would then turn into:
add, PRF#3, PRF#2, PRF2, and after renaming that instruction RAX points to PRF#3.
and so on. In PRF machines, architectural registers only exist as pointers into the PRF.
The reason for this design is that it makes OoO execution simple, and reduces unnecessary data movement. An instruction is ready to execute once all it's inputs have data in the PRF, and when it executes there is no need to move data into some special architectural register.
Registers are reaped from the "haves data" set into the clean set once no instruction or architectural register points at them. Note that multiple architectural registers can point to the same physical register: mov rax, rbx is resolved in the frontend just by pointing rax to the same register as rbx, and there is typically a special register for the value 0 that all common register-clearing operations use.
This was a fantastic explanation so thanks. A couple of questions:
PRF is Physical register file? As opposed to the logical which can point to different backing store(the actual SRAM.)
Register renaming and OoO operations are really enabled by micro-ops? In other words on a CPU with completely hard-wired control unit such things would never be possible.
> PRF is Physical register file? As opposed to the logical which can point to different backing store(the actual SRAM.)
Yep. PRF means the single backing store where all values go, and PRF systems are named for it because they are usually contrasted to the other very common way to implement OoO, ROB-backed systems, where there is a specific location for the final architectural value, and possibly multiple in-progress values in the ROB. Intel used values in the ROB in all their CPUs up to Sandy Bridge, which was their first PRF design.
(The main difference between the types is that ROB-backed is faster in circuit delay, but moves more data than a PRF design, so on the same process they potentially clock higher but produce more heat. PRF-based designs are also easier to scale wider than the ROB-based design.)
> Register renaming and OoO operations are really enabled by micro-ops? In other words on a CPU with completely hard-wired control unit such things would never be possible.
No, you can do renaming and OoO well without uops, it just requires that your instruction set is designed to be amenable to it. Before x86 had uops, it couldn't do OoO and the competing RISC CPUs were way faster because of this. PPro added uops precisely because they made it possible for an x86 cpu to do OoO, and eventually x86 beat all the competing RISC chips at their own game.
> Register renaming and OoO operations are really enabled by micro-ops? In other words on a CPU with completely hard-wired control unit such things would never be possible.
No, they're orthogonal concepts. The first OoO processor (the System 360 model 91) didn't have uops.
Wow, this is kind of blowing my mind that they were doing this with a hard-wired control unit in 1969(I think I have the date right.) Big Blue, its easy now a days to forget just how amazing they were. Do you have links you can recommend? Thanks.
Actually, the Mill uses just the forwarding network which I completely ignored in this explanation.
The really short explanation about the forwarding network is that writing/reading the PRF takes too much time for you to read it and do work on the same cycle. If you operated on it alone, all operations that were dependent on each other would have multi-cycle latencies.
Since this sucks, in addition to the PRF the machines have a forwarding network, where the result of every instruction is broadcast to all the execution units of a similar type. If the broadcasted PRF number matches the register you want, it's read from the forwarding network instead of the PRF.
AFAICT, the Mill works by completely eschewing the final register file, and exclusively using the forwarding network with some local storage at each execution unit for the past values on the network.
I am not who you responded to, but: the registers are always in the same physical location, but internally, they are virtually renamed for efficient reuse and to avoid resource contention. If the CPU knows the stream of assembly instructions to be executed, it can abstract away actual register names and show a larger number of registers than what's physically available.