But Radford is just pretraining the decoder and qualitatively different from a s...

		kitsune_ on Feb 24, 2020 \| parent \| context \| favorite \| on: T5: The Text-to-Text Transfer Transformer But Radford is just pretraining the decoder and qualitatively different from a seq2seq approach such as MASS. If we just look at the original paper from Vaswani, than "pretraining a transformer" imho should always only have meant pretraing the encoder and decoder. Obviously that ship has sailed.