This is a big exaggeration. Dangerous? Because it might lead to an outbreak of bad science and poor conclusions? What's one more drop of water in that ocean?
If, in fact, Anderson's idea turns out to be half silly and half banal, the world won't even notice. Folks are too busy squaring the circle, constructing their perpetual motion machines, and printing out new labels to paste on their creationist tracts to take much notice of yet another half-baked idea.
I'm not sure what to make of this. I think Hank is taking the idea that correlation is good enough too seriously. Since I've been playing with correlations at home and learning how people find correlations in huge data sets, I've found that a correlation actually is good enough for most purposes. Think about what the correlation is used for. It makes suggestions on things to buy or things you should explore. It's making suggestions. A movie suggestion is exactly that, a suggestion. It's not an exact science and you can't have a 100% guarantee that a consumer will be into something enough to buy it. It's the same thing as trying to predict the market. The correlations can provide a useful service for getting information to people about things they might like. Hopefully the service makes it easy to buy the product and then the picture seems complete.
Gist: "My guess is that this emerging method will be one additional tool in the evolution of the scientific method. It will not replace any current methods (sorry, no end of science!) but will compliment established theory-driven science. Let's call this data intensive approach to problem solving Correlative Analytics ... In the coming world of cloud computing perfectly good answers will become a commodity. The real value of the rest of science then becomes asking good questions."
The real value of the rest of science then becomes asking good questions.
It were ever thus. Oddly enough Picasso appears to have been one of the first to recognize the relevance of this in the computer age -- "Computers are useless. They can only give you answers" -- long around 1968.
I fire up Maxima (Macsyma) all the time to do mindless derivatives, integrals, sums, and series, so that I can get frustrated as hell working on less tractable problems. The symbolic rearrangements (the ones that can be done without reference to the more complicated reality at hand) are just bookkeeping, and computers are awesome for that.
If raw data (and lots of it) were sufficient, things like dynamic programming, heuristic decomposition, and approximation algorithms (complete with correctness bounds) would be pointless. It's not and they aren't.
nb. Don't take any of the above to mean that the province of "asking interesting questions" is somehow restricted to scientists, or artists, or anyone else. But it is probably the highest hurdle towards discovering something worth pursuing. The antecedent execution thereof is rarely easy, either. Tools help -- a lot -- but by themselves they cannot somehow pull the future into the present. For the time being, at least, intelligence amplification still trumps artificial intelligence. (I hope it stays that way for a while, being human and all)
In the coming world of cloud computing perfectly good
answers will become a commodity. The real value of the
rest of science then becomes asking good questions.
Will you look at that. The same is true for software engineering. Eventually we will have machines choosing both algorithms and data structures for us based on workload and data models. But someone will still have to come up with models and make sure the relate to user's problem space.
It may be "dangerous", but basing causation just on correlation (or statistical data in general) is actually a pretty prominent idea: http://plato.stanford.edu/entries/causation-probabilistic/.
Chris Anderson obviously glosses over many tricky details, but Judea Pearl among others have written extensively on it. Ultimately, all these ideas can probably be traced back to David Hume (at least here in the West).
Hank's argument that Chris is suggesting that we don't try to find spurious correlations is rather simplistic since a probabilistic theory of causation can account for "spurious" and "real" correlations quite well. Indeed, the only reason why we know that "spurious" correlations exist is because they eventually show up in the data (i.e. the economist finds out that the stock market does not predict sun spots). Thus, a statistical definition of causation (more or less one that only uses a notion of correlation), can actually be quite robust.
I enjoyed the article Hank. As someone working on developing data-driven machine learning systems for marketing purposes, it made me think twice about some of the conclusions I am able to jump to based only on correlation. Certainly, these methods are quite useful in my field and the other less precise social sciences. In the hard sciences, however, advancing theory without a model seems preposterous save perhaps for a supportive role in special cases.
EDIT: almost feels like a republic (model) vs. democracy (data) political discussion
If, in fact, Anderson's idea turns out to be half silly and half banal, the world won't even notice. Folks are too busy squaring the circle, constructing their perpetual motion machines, and printing out new labels to paste on their creationist tracts to take much notice of yet another half-baked idea.