They train on chat logs unless opted out. This really isn't a conspiratorial claim requiring humans to decide to steal IP if true.
His prior work predating OpenAI's interest in the problem was ingested over the last year as he made progress and used for training.
Then, with a prompting nudge from OpenAI's team who acknowledged hearing about the direction "Anthropic" (his co-collaborator) had been pursuing, they're able to point their giant amount of compute towards a known promising path to a proof and crossing the finish line first.