If NN was able to do small data then are they better than their counter parts?
I mean if you can do it for small data and it was good then we would be seeing it dominate kaggle in all problem domains. Maybe the small data problems belong to other algorithm (such as tree base and forest, SVM).
disclaimer - I'm bias for tree base algorithm in medium and small data since it is my thesis.
I think no - SVMs are explicitly optimized to generalize the best from small data (https://en.wikipedia.org/wiki/Hinge_loss - 'The hinge loss is used for "maximum-margin" classification'), whereas NNs have more hacky regularization methods. I am not sure if the same is true for tree-based methods, but of course those are lovely due to how interpretable they are when you have few features.
I mean if you can do it for small data and it was good then we would be seeing it dominate kaggle in all problem domains. Maybe the small data problems belong to other algorithm (such as tree base and forest, SVM).
disclaimer - I'm bias for tree base algorithm in medium and small data since it is my thesis.