Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
stymaar
16 days ago
|
parent
|
context
|
favorite
| on:
I trained a small transformer in 1.5hrs and it bea...
AFAIK, the “large” qualifier came when transformers allowed to scale the size of language models compared to the recurrent models that where in fashion before. And although BERT isn't
large
by today's standard, it was large enough for the time.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: