Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

instead of thinking about what it is in practice: skip-gram negative sampling, I think it's much more intuitive to think about what it is in theory: extreme multi-class classification.

word2vec is a multi-class classification problem with a softmax output layer and cross-entropy loss. The novel part of word2vec, in my opinion, is two:

1. dataset (proximal input word & output word) generation from documents eg: skiagram, CBOW, etc 2. engineering speedup for softmax: Approximate Softmax eg Negative Sampling using NCE, hierarchal softmax, etc

If you just build word2vec w/o step 2, it's a easier to understand. Then when you get that working, add in the negative sampling speedup trick which isn't core the theoretical algorithm.



Can't really call it a speedup trick, since it actually improves the performance of the embeddings but in terms of qualitative understanding, I see where you're coming from.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: