I can't imagine these poor video streams are useful inputs for the algorithm...
Maybe simulations are when the human is controlling the vehicle and the algorithm is only allowed to .. simulate what it's actions would be if it were in control?
When they encounter interesting situations while testing on real roads, they then take that situation over to the simulator and train their algorithms on 1000s of subtle variations of the situation, which Waymo calls 'fuzzing'.
Simulating effectively is non-trivial. In spite of testing on a comparable scale to Waymo, Uber made very little progress in 2 years because they couldn't get their simulator to work right.
They don't use the video data in the simulations, they use the processed "Where everything is" stream. That makes it more robust to sensor tech upgrades, but less to errors to do with bad sensing or sensor failure.
Computationally cheaper and simpler, no doubt, too.
If they've been recording everything, they can run the previous inputs on new algorithms and check the results, and also simulate sensor failures and what not.
Running the system on recorded inputs to see how it compares with human output is good too.
That would be entertaining to watch. Possibly also useful to check the cars don't freak in weird situations. They kind of had a real one in the Chandler crash https://www.youtube.com/watch?v=vxqBS2-4puw