Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
Philpax
72 days ago
|
parent
|
context
|
favorite
| on:
Claude Opus 4.8
It wouldn't be data distillation: instead, it would be teacher-student distillation. The teacher model has stronger representations that the student can mimic, which would give it more capability over training on the data itself.
Consider applying for YC's Fall 2026 batch!
Applications
are open till July 27.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: