AFAIK The kernel needs to perform calls to the hypervisor in order to copy the page table, but I'm not expert enough to provide you with details about this unfortunately. It is for sure not inherently due to virtualization, since for example VMware does not have this issue.
That sounds like a PV vs. HVM issue. Using hardware virtualization extensions to handle virtual memory is almost always faster than paravirtualization these days, which is why Xen introduced PVH mode in 4.4.
Yep. And EC2's "HVM"-type instances are now actually PVHVM, not pure HVM.
Since this change, there has been absolutely no reason to use anything other than HVM AMIs, and pure paravirtual instances can basically be considered a deprecated feature in EC2. New EC2 instance classes (e.g. t2) don't even support PV.
Basically, PV instances are just there to support current customers who are relying on already-built PV AMIs and have too much inertial to be nudged into switching over.
Yes, at some point I got a report about Xen 3.0 (If I remember correctly) fixing the issue, but I never see in the real world things improving much AFAIK. Here is a table that shows fork times with different environments, just to show how bad the thing is:
Linux beefy VM on VMware 6.0GB: 12.8 milliseconds per GB.
Linux running on physical machine (Unknown HW): 13.1 milliseconds per GB.
Linux running on physical machine (Xeon @ 2.27Ghz): 9 milliseconds per GB.
Linux VM on 6sync (KVM): 23.3 millisecond per GB.
Linux VM on EC2 (Xen): 239.3 milliseconds per GB.
Linux VM on Linode (Xen): 424 milliseconds per GB.
Around 30 times slower than bare metal, and I'm talking about old physical servers with slow memory compared to today's.