The article says that Google Flu Trends does worse than a "simple model of local temperatures". From a very quick, level-1 read [1], the paper you link to doesn't mention that simpler model. Instead, it compares Google flu to previous versions of itself.
I guess you can say that Google improved their flu model, but, "fixed"? Also, I don't see that the article's "thesis" is "undermined". Sorry about the scare quotes.
I mean that the article is taking small liberties to score a small point against Big Data™ (and who else best to score them against, other than Google?) but is that really enough to call bullshit on it?
I don't see anything misleading in the article is what I'm saying. So why "bullshit"?
___________
[1] Read abstract and conclusions, eyball a couple of tables, scan the rest, i.e. just enough to argue on the internet as if I know what I'm talking about.
The article says that Google Flu Trends does worse than a "simple model of local temperatures"
Indeed, that is what the article says. It's bullshit though.
You are right, though that "Google improved their flu model". No model is ever "fixed" if that means 100% correct.
Flu trends worked well (much better than a "simple model of local temperatures") except in the 2009 flu season, when it missed the A/H1N1 pandemic. It was then modified, and these modifications seem to have caused it to estimate a pandemic in the 2012/13 season which didn't occur.
"simple model[s] of local temperatures" do work quite well as a baseline, but they don't pickup pandemics either. However, in that 2012/13 season it would have done better than flu trends. [1] is a good overview.
So this is complicated topic. I have a research team working on this exact problem, and we'd love Google search data because there is no doubt that it can and does work. But like all models it breaks down when something it hasn't seen before occurs.
My bigger problem is with the thesis of the article. I'd summarize my reading of that as "big data is BS", which is a more extreme form of their title "How to Call BS on Big Data".
But the course this is based on isn't that at all. It's about understanding how big data can be used to draw wrong conclusions, NOT that big data is BS in any way at all.
I think the course is a very important and useful thing. But what it is doing is dramatically different to what this article claims, and the way that they use the implied authority of the course to support their "big data is BS" claim is what lead me to to say "bullshit".
Thanks for the clarification and it's good to hear you are speaking from experience with the kind of model being discussed (although I'd still like to know what that "simpler model" is exactly, or where it comes from anyway; but that's probably not for you -or even Google- to answer, since it's mentioned in the original article in the first place).
That said, I don't agree with you, in that I didn't read the article as saying that Big Data is BS by default. It's a short article and not terribly thorough but I didn't read a blanket condemnation in it.
Btw, I'm not sure why you and nickpsecurity are being downvoted to grey. I expected this strong disagreement to be reserved for personal attacks etc.
"No model is ever "fixed" if that means 100% correct."
One definition of broken for a proposed alternative to the status quo is if the alternative under-performs it. Kind of makes one ask why anyone would adopt it to begin with. There's a simple model using temperature that works pretty well. Google's solution is said to perform worse than that with more false positives. Google's isn't "fixed" or "working" until they show it outperforms the simple solution that works with similar error margin and cost.
In other words, it isn't good until people would want to give up existing method to get extra benefits or cost savings new one brings.
In other words, it isn't good until people would want to give up existing method to get extra benefits or cost savings new one brings.
They do.
The current state of the art methods used "in production" today all use Flu Trends data from Google[1], other forms of digital data[2], ensemble methods incorporating them all, or human-based "crowdsourced" forecasting[3]
One definition of broken for a proposed alternative to the status quo is if the alternative under-performs it.
Note that there is no mention of the "simple mean temperature" model. That's because it isn't very useful. That model predicts flu increases in winter, and picks up minor variations because of weather patterns.
To simplify even further, you can average all the CDC flu data and use that as your prediction and on an average year you'll have a decently performing model.
This isn't useful as a forecast, because the people who need forecasts already know this.
Better models (eg, SI, SEIR, Hawkes process based etc) can sometimes pick up epidemic or unusual conditions, but only after the conditions have changed. This is still useful, because there is a (best case) 2 week lag between ground conditions and CDC data being available.
Digital surveillance techniques (Flu Trends, Twitter data, etc) all push that data lag back.
This is incredibly useful for the people who need forecasts because it gives them lead time.
To understand this you need to consider the metrics. The most common metrics for flu forecasting is the "peak week", and "number of people infected at peak". Sometime the total number of people infected in a season is also reported.
Temperature-based models do really well on average at both these tasks, but they fail completely at picking the unusual seasons.
Google's solution is said to perform worse than that with more false positives.
Google flu trends picked the 2009 epidemic season really well, but failed in the 2013 season (when it falsely picked an epidemic). On average that might make it worse than a temperature based model, but that is just bad selection of metrics.
It's like reporting average income when your sample has a billionaire: the metric is misleading.
If that isn't the perfect example of "bullshit" then I don't know what is.
" From a very quick, level-1 read [1], the paper you link to doesn't mention that simpler model. Instead, it compares Google flu to previous versions of itself."
This is actually a proven method of disinformation called false/incomplete comparison that advertisers use to sell products. I'm not accusing Google of doing that so much as saying them scoring a new tech against a defective one to say something about the new tech in general should be dismissed by default since it's a broken comparison of same kind used in fraudulent advertising. Aka it's bullshit.
I guess you can say that Google improved their flu model, but, "fixed"? Also, I don't see that the article's "thesis" is "undermined". Sorry about the scare quotes.
I mean that the article is taking small liberties to score a small point against Big Data™ (and who else best to score them against, other than Google?) but is that really enough to call bullshit on it?
I don't see anything misleading in the article is what I'm saying. So why "bullshit"?
___________
[1] Read abstract and conclusions, eyball a couple of tables, scan the rest, i.e. just enough to argue on the internet as if I know what I'm talking about.