Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

One interesting thing going forward (over the next decade) regarding deep learning and high-performance computing will be if we can circumvent the big serial bottleneck - that is, depth. The best benchmarks have generally all been set by networks with significantly increased levels of depth compared to the previous records. Gradient descent, however, requires propagating residuals back through every layer of the network (to tune the lowest layers, at any rate). This is inherently a serial task. Recurrent networks can be considered to be infinitely deep - in practice they're truncated to some finite level during training.

There are a couple of solutions to these problems that are being worked on - including various "learning-to-learn" approaches. One approach (DeepMind published, I believe) is to locally train an auxiliary network to predict the back-propagated residuals coming down from the top. It will be interesting to see if these approaches pan out, or if eventually we start to run into real serial bottlenecks again as we build ever-deeper networks.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: