Showing posts with label predictive analytics. Show all posts
Showing posts with label predictive analytics. Show all posts

Thursday, 11 April 2013

Can you find the Big Data holy grail in the Big Data maze?

The Big Data Show and Conference in London this year is co-located with Internet World and is a live exhibition for professionals responsible for Big Data strategy, analytics and technology. Bet Buddy is excited to be presenting its latest thinking on big data research and product development in the gaming industry at 12.30pm on the 25th April at the conference. We have previously blogged about what Big Data is and the challenges organizations are facing with managing what we term Data Inflation. These challenges are manifesting tremendous growth in the Big Data software and services vendor landscape (or increasingly a Big Data vendor maze):

This increasingly complex vendor landscape reflects an increasingly complex world of big data that organizations must tackle. And as more business shifts to adopting digital and eCommerce business models, we believe that more industry specific solutions will be required to manage some of the most critical parts of the organizational value chain. The largest Big Data software vendors will continue to try to gain market share by entering new analytics and operational domains, for example Oracle will compete against not only the likes of Teradata and ParAccel (in-DBMS analytics) but also against SAP and GoodData (Business Intelligence). However, we believe it will not be possible to offer best-in-class analytics solutions across all domains without having increasingly deep industry domain expertise. Whilst databases, infrastructure, and integration will remain lucrative and large addressable markets for the biggest vendors, offering solutions tailored for your organization we believe is ultimately the holy grail of Big Data for areas such as reporting, predictive analytics and prescriptive analytics. If you want to hear more about Bet Buddy's analytics solutions and services then come and join us at the The Big Data Show and Conference this month.

Thursday, 26 April 2012

Scaling Data to Make Better Decisions

This week we have been attending a series of events held during Big Data Week. At the London Community Event, we heard from a panel discussion on big data that included (amongst others) Hilary Mason, Chief Scientist of Bit.ly, Doug Cutting, co-founder of the Apache Hadoop project, and Nick Halstead, CTO and founder of Datasift. At the same time of Big Data Week, the gaming industry have been attending GiGse 2012 in San Fransisco. During the CEO panel at GiGse, the importance of data was highlighted by Jim Ryan, co-CEO of bwin.party, who said that bwin.party now have around 70 people in their business information team analysing data and feeding back into its marketing operation. This is a big team of data analysts - bwin.party are taking data very seriously. However whilst large organisations have the resources to deploy very large analytics teams to solve big data problems, we believe that as data volumes continue to increase exponentially, adding solely staff to process and mine increasingly large data volumes will not be a scalable solution for any organisation in any industry. This view was confirmed by the expert panel at Big Data Community Event – here is a summary of some of the key discussions themes.

How important is the ‘Big’ in Big Data?
A philosophical explanation of big data centred on being able to look at data with no pre-conceived ideas. As well as an open mind, big data is about having the ability to join multiple data sets and run analytics across them, rather than taking a silo approach. Whilst the panel disagreed on the relative importance of the word ‘big’ in big data, a recurrent message was that today it’s much easier and cheaper to store and analyse very large data sets e.g. large scale data processing (i.e. map reduce) on platforms such as Amazon Web Services (AWS) has now become commoditized, and are now considered established technologies and platforms. And with the profileration of eCommerce, social media and APIs, there is a lot more volume and richness of data available to analyse today. If you are interested in reading an example as to how the combination of cloud computing (e.g. AWS) and Hadoop (map reduce) enables big data processing at scale and at a significantly reduced cost then we suggest you read this article - Big Pain or Big Profits?

What are some of the Challenges?
Arguably the biggest challenge is how do organisations find the important nuggets of information? Taking a 'boil the ocean' approach to big data is fraught with challenges, a theme we have examined in our blog on Data Inflation last month. It was stated that one of the major benefits of big data is the ability to get answers to questions back quickly, which in the past could take weeks and months - but this needs access to an increasingly important resource - the data scientist - a combination of math, computing, and domain expertise, coupled with an open and inquisitive mind. And finding these people is a major headache for most organisations. There was also discussion about academic access and use of data. It was argued commercial organisations still remain very reluctant to share information via open research projects as there is a lack of trust as to how these data sets will be used and by whom (an example of an open research project within the gaming industry is The Transparency Project). A key infrastructure challenge is that the internet has not been designed to process large data sets at low latencies. Ensuring compliance with data privacy requirements, unsurprisingly, remains a major focal point.

Should You Care?
If you want to make better decisions then the answer is yes! The end product of any data or big data project has to be focused on better decision making. In gaming this equates to supporting decision making across aspects of the business: game design, game performance, 1-2-1 marketing and consumer protection, and finance and risk management. So whilst we can argue about the importance of the word 'Big' in big data, we cannot argue about the increasing relevance of data to managing our businesses today. As the panel concluded, "we are just scratching the surface of what is possible".

Tuesday, 27 March 2012

Curbing [Data] Inflation

We are in the midst of a data deluge according to The Economist. McKinsey are telling us that data volumes are increasing at a rate of 40% per year. Whilst many extol this as a great opportunity, at Bet Buddy we believe this is causing a phenomenon that we have termed Data Inflation i.e. the risk that the value of the data an organisation captures and holds decreases as the supply of data increases.

 
How can this be so? Surely more data means more opportunities to analyse and understand consumer behaviour, therefore more opportunities to drive more personalised and targeted offers and campaigns? In theory this assumption makes sense, however in practice this is very difficult to get right.

The data available to organisations that can be utilised to support marketing, operations and customer services, risk, and compliance activities can be broadly classified as Personal, Machine Generated and Social Network Data (although data privacy laws and policies certainly restrict to what extent we can leverage much of this data). Some data generally falls neatly within these categories, for example we like Splunk’s definitions of machine generated data. However, our categories are open to interpretation e.g. whilst some may categorise click-stream data as machine generated data others argue click-stream records are personal data. Most of the data that is captured within an organisation is, however, never used - 75% remains dormant according to the Financial Times. Whilst this may sound like a lost opportunity, leaving data on the table can also make practical sense.

Knowing how to effectively capture and utilise the right data is the challenge of managing data inflation. Zynga is sending about 5TB of its data to its central store per day, which is about 10% of the data it collects - this covers game actions (personal data) and not log files (machine generated data). We estimate these daily player data volumes are >15x the core player data volumes that a medium sized online gaming operator saves per year (core player data here covers the player, game, session and transaction files). Zynga is clearly a mammoth and generates data on a magnitude alongside the world's largest 'big data' firms. It has over a quarter of a billion monthly active users on Facebook, therefore dwarfing most other gambling and gaming firms. So whilst the infrastructure they have in place is unlikely to be applicable to most gambling firms, they are however a good organisation to examine a little closer. Because of the scale of the data they generate and leverage, they had to invest early in the tools, systems and processes that allow them to better manage the risks of data inflation. As data volumes have continued to increase exponentially, we doubt very much companies such as Zynga relied purely on hiring staff to process and mine increasing data volumes, and they prioritised i.e. they didn't try to analyse everything at the same time. Whilst the magnitude of Zynga's big data challenges do not apply to most gaming and gambling operators, the principles do.

Curbing data inflation requires the organisation to think strategically across a number of key areas e.g. storage, security, transportation, and analytics. We need to step back and start thinking about data within an organisation as an ecosystem of connected data sources, internal and external customers, tools and platforms, processes, and people. Capturing data as part of everyday business-as-usual processes can be very hard, and for data that you prioritise as important for analytics, it means automating daily repetitive tasks and processes, such as:
  • Data sourcing – large gaming operators have multiple game offerings and back office account management systems, supporting multiple channels, usually from multiple vendors. This is a bigger challenge than pure data volumes for most gambling firms
  • Data cleansing – dealing with date/time stamp conversions, replacing commas in numerical values, treating missing data values, etc
  • Data transformation – core player value segmentation calculations, calculating predictor variable data for each predictive model deployed, etc
So organisations need to focus on the data that will drive the most value. Whilst this sounds obvious, we believe that for any single business area or domain, the manager should typically not be presented with more than 15 – 20 data points to manage their business on a daily business. Any more - think of very large and complex dashboards, countless MI reports and tables of data - risks a decrease in productivity and less focus on what really matters to the business i.e. an increase in data inflation. This is important as it informs what data undertakes pre-processing and when, as not every data point needs to be analysed at the same time - some data insights can be utilised after the event whilst others are best applied in real-time. There are other considerations too, such as how and when to undertaken explorative data analysis and how to best visualise data (which is often technical and complex) to enable decisions to be made (here's a good article on visualisation from Jeffrey Heer of Stanford University).

Whilst the value potential of data and analytics is large, the risk of data inflation makes the practicalities of achieving this value very difficult. 

Wednesday, 4 May 2011

Some Snippits from Westminster – The Future of Gambling


Yesterday's gambling seminar at Westminster (London, UK) brought together senior industry, parliamentary, and regulatory figures to debate how regulation, technology and international markets are shaping the future of the gambling industry.  Whilst there was much debate across a range of issues, such as EU regulatory harmonisation, taxation, US regulation and Black Friday, data privacy, sports betting and sports rights, we will share some highlights from the discussions that focused on the new opportunities that technology is providing for the industry.

Opening the ‘New technologies and platforms for gambling’ session was Mark Maydon, Commercial Director of Sporting Index.  Maydon said that the industry is witnessing an ‘arms race’ in terms of technology, in that there is a need for significant and continued investment in IT if operators are to remain competitive.  Operators need to develop more effective and scalable data processing capabilities, with Maydon suggesting that the gambling industry needs to learn from how the financial services industry has developed these capabilities.  Investment in mobile gaming was also an area of significant importance and growth for the industry.

Charles Cohen, CEO of Probability, was next and he stressed that whilst mobile gaming was a growing channel it was a very different channel to traditional online gaming.  He challenged the ‘myths of mobile gaming’  stressing mobile gaming’s objective is not about squeezing as much cash from online gamblers by encouraging them to gamble more (such as when in the pub or whilst waiting for the bus).  Rather, mobile gaming is about offering a different betting experience that appeals to a different customer i.e. not traditional PC-based online gamblers.  He stressed PC-based online games do not transfer well to the mobile channel and the best mobile games were very simple games.  He also shared some interesting insights from Probability’s own research: the average session of a mobile game is c.10 minutes and 82% of mobile gamers gamble whilst sitting at home sitting in front of the TV.  David Loveday, CEO of OpenBet, stated he felt that the TV channel, once integrated with broadband, along with online content provision, were two areas that will become increasingly prevalent in shaping the future of gambling.

Finally Martin Cruddace, Chief Legal and Regulatory Officer at Betfair, framed his views on technology in the context of social responsibility and integrity in betting.  Cruddace said the industry must lead from the ‘front foot’ on these issues and that Betfair was investing in technology to better protect the customer, including predictive analytics to identify problem gambling behaviour.  In addition, Betfair are further enhancing the current range of player protection features by offering self-exclusion via gaming vertical and are also assessing the possibility of giving players the ability to exclude themselves from gambling from certain parts of the day e.g. after 11pm.  Earlier Peter Reynolds, Head of Communications at bwin.party, emphasised the that online gaming offered a perfect audit trail of data that allowed the industry to proactively track online behaviour and act upon insights obtained.

The key messages coming from the other sessions re-iterated common themes; US regulation will inevitably happen however this will be via multi-year state projects, a lack of harmonised regulation in Europe is making it increasingly expensive for operators to compete internationally, however EU harmonisation is realistically many years away, and regulating markets is the best way to protect the consumer.  And the consumer is what the industry must focus on during these consultations and debates over the coming months as the best way to influence policy makers is by framing the debate in the context of what is best for the consumer rather than what is best for the operator.