No, plausible given AI companies want/need session data to train their next models. Probably not someone peeking an eye to sessions directly, but probably not so hard to find the useful sessions in anonymized training data to post train a model on. As stated in the paper, OpenAI did not explicitely denied the researcher sessions were not used for training the model. So either they don't know, or don't want to tell
"I asked whether the model had been trained on, or had access to, our sessions
in Codex, into which we had been putting all our drafts for the whole of this
project. I was told the model did not look up user data. I asked again, about
training, and I did not get an answer."
Let's see what statement OpenAI will come up with for their side of the story
EDIT: precised my thought on user data vs session data
"Since August 28 we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This model’s training is ongoing and its performance continues to improve."
"When a further trained version of our internal model became available over the course of the effort"
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models "
"I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer."
Let's see what statement OpenAI will come up with for their side of the story
EDIT: precised my thought on user data vs session data