Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It's very difficult to understand what the contributions are here. From what I've read so far this feels more of a proposal for future research or a press release than advancing the state of the art.

* Using large models trained on lots of data to provide the foundation for sample efficient smaller models is common.

* Transfer learning, fine tuning, character RNNs is common.

Were there any insights learned that give a deeper understanding of these phenomena?

Not knowing too much about the sentiment space, it's hard to tell how significant the resulting model is.



* advancing the state of the art

It says right at the top: "we get 91.8% accuracy versus the previous best of 90.2%" on a standard sentiment corpus. In addition, their method needs less training data than previous approaches.

* Were there any insights learned that give a deeper understanding of these phenomena?

The main appeal lies in the fact that a model trained on a (1) different and (2) very general task basically "in passing" also learned to predict sentiment (i.e., a specialized task that more or less arose from the domain the general model was trained on), and pretty much through a single neuron (out of the 4096 used). The authors speculate that this might be a general effect that could also be transferred to other prediction tasks.


If the main contribution here is the quality of the model and its interesting and powerful representation of text, I hope OpenAI does something distruptively different and releases the weights and trained model.

The accidental sentiment neuron is a function of the model, distribution of the input dataset, and the optimizer finding nice saddle points. Insight into these foundational components would make these results amazing. It sounded like training on other datasets doesn't have the same sentiment properties, which provides a lever to explore these concepts more.

At the moment it feels like the Google cat neuron. It attracted a lot of intrigue but the individual contribution from that in terms of research was more on the infrastructure side, and few people seem to refer back to that publication at this point.

That said OpenAIs mission in itself doesn't necessarily require novel research. For example, the gym is fostering a competitive atmosphere for the community to work on RL which hopefully leads to more progress in the field.

Training a model for a month is difficult and if it has captured interesting phenomena it seems in the interest of the community to release the weights and model. It would be hard for the community to reproduce this without a month of compute and 83M Amazon reviews.


Hi there, the weights and model are here: https://github.com/openai/generating-reviews-discovering-sen...


This is awesome, thanks! My apologies I must have missed it somewhere.


Also, first they write:

> We were very surprised that our model learned an interpretable feature, and that simply predicting the next character in Amazon reviews resulted in discovering the concept of sentiment.

And then they write:

> We believe the phenomenon is not specific to our model, but is instead a general property of certain large neural networks that are trained to predict the next step or dimension in their inputs.

So they can't explain why a phenomenon is occurring, but they think that it generalizes to other contexts.

I find it all very unconvincing. Is this kind of writing common in the deep learning literature?


Mind you, this is not a scientific publication but a blog post that has intentionally tried to adapt the tone to that medium, presumably to appeal to a wider audience.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: