Piwik is a great project, but it tends not to work well for handling sites with millions of events per day. Your MySQL table starts to bust at the seams pretty quickly.
For big sites, you'll want that event data in GiB's of plain raw logs that you can bulk load into tools like BigQuery or Redshift for analysis.
My team has built/delivered a SaaS web content analytics platform for the past few years called Parse.ly. We instrument pageview events (like Google Analytics) automatically, and we also instrument time-on-page using heartbeat events. We collect 50 billion monthly page events for over 600 top-traffic sites, and display it all in a real-time dashboard.
To emphasize that our customers own the data they send us, we recently launched a Raw Data Pipeline product:
Basically, we host a secure S3 bucket and Kinesis stream for customers, and deliver their raw (enriched) event data there. From there, they typically load it into their own BigQuery or Redshift instance, or they analyze pieces of it directly with Python/R/Excel/etc.
Our customers tell us this strikes the right balance among data ownership, query flexibility, and hassle-free infrastructure.
I've read parse.ly tech blog a few times, it was a great so I immediately recognized your company. As someone who's working in the publishing as well and gradually moving to a more data driven approach - thanks! :)
Data ownership in GA is a "gray area" that becomes less gray if you pay $150K/yr for "GA Premium".
Google has mixed incentives in running its free analytics service. It gets web-wide analytics data, it uses data to help it sell more AdWords to customers, and it integrates GA with other services, like their display advertising products (DFP, etc.)
From a practical standpoint, you don't "own" analytics data when a) you can't easily access it in raw form and b) the SaaS provider "leaks" your data to dilute its value to you. We address (a) and (b) directly through our products and public data privacy stance. See this blog post for our public view on analytics data privacy:
While I find your stance on privacy very refreshing for an analytics company, hiding your pricing info behind a sales rep is a huge turnoff for me. If you feel your pricing is reasonable for the service that you provide, I really don't see why you can't just display it proudly on your site.
Whether to display pricing on the website is something we debated in the past, and continue to debate. (Your comment may wake up the debate for me.)
Pricing for analytics services (in the marketplace) is all over the map. Google picked $150K/year as the price for GA Premium because that's the low end of an Adobe Analytics contract, who is the market leader. We're typically cheaper than existing Adobe/GA contracts. Non-competitive "event analytics" companies like MixPanel and Heap have variable per-event pricing that would break the bank for the customers we serve. We have a bit of an aversion to per-event pricing because it feels like "punishing customers for success".
Meanwhile, per seat pricing, though attractive on the surface and popular in the SaaS segment, has several concerns in our space. First, we want customers to feel free to hand out access to our platform: part of our value proposition is democratizing access to analytics data. So we don't want "stingy seat quotas" typical with tools like Salesforce. Second, for an analytics tool, seats are a bit easier to "hack" for a pricing model -- though our dashboards can be customized per user, a single shared account can access all the data. Meanwhile, our costs don't scale with seats, but with site traffic/users instead.
For these reasons, and more, we've settled on "tiered pricing". Roughly speaking, our service is offered in three tiers. Each tier supports a larger class of site (more monthly uniques), which also bestows more features (e.g. more data retention in higher tiers). To work within the budget constraints of some companies, we will discount tiers while removing cost-affecting features, e.g. maybe you are in the highest class of site, but we disable API access and limit data retention. Because this is a tad more complex than a pricing page could express easily, and also because we think the value of the product comes through best in a guided demo, we made the decision to hide pricing and instead responsively provide demos on-demand.
So, the tl;dr is, pricing, and the display of it, is definitely something we think about, and we have (IMO valid) reasons for not displaying pricing right now, but you make a fair point: if Musk can price his rockets publicly, maybe we can figure something out, too :)
I avoid all services with prices behind sales rep when possible. I always feel I'm not an astute negotiator, the sales rep will be ergo he will see me as a mug and take me for all he can get. If I do buy, even if I'm happy I'm always left wondering if I'm paying 2x as much as other customers.
Im sure parsle.ys analytics system told them that their potentical customers were getting sticker shock; hence the sales guys need to explain the value propesition to them properly before releaving the figure ;)
Sometimes i feel Google should just blanket replace their disclosure statements across the board with this classic video; https://www.youtube.com/watch?v=8fvTxv46ano.
Less beating around the bush.
For big sites, you'll want that event data in GiB's of plain raw logs that you can bulk load into tools like BigQuery or Redshift for analysis.
My team has built/delivered a SaaS web content analytics platform for the past few years called Parse.ly. We instrument pageview events (like Google Analytics) automatically, and we also instrument time-on-page using heartbeat events. We collect 50 billion monthly page events for over 600 top-traffic sites, and display it all in a real-time dashboard.
To emphasize that our customers own the data they send us, we recently launched a Raw Data Pipeline product:
http://parse.ly/data-pipeline
Basically, we host a secure S3 bucket and Kinesis stream for customers, and deliver their raw (enriched) event data there. From there, they typically load it into their own BigQuery or Redshift instance, or they analyze pieces of it directly with Python/R/Excel/etc.
Our customers tell us this strikes the right balance among data ownership, query flexibility, and hassle-free infrastructure.