Warning: Trying to access array offset on value of type bool in /home/sites/3b/7/74de28c56e/public_html/wp-content/plugins/related-posts-thumbnails/related-posts-thumbnails.php on line 846
 

Demystifying "Big Data" - Looking at the myths and risks. - InstaHost Solutions

8th March 2016by Insta Team

Data is now streaming from daily life: Your phone, credit cards, Smart TV’s, Tablets, Oyster Card, Supermarket Loyalty Card  and, of course, computers; then the infrastructure of cities; from sensor-equipped buildings, trains, buses, planes, bridges, and factories. The data flow so fast that the total accumulation of the past two years—a zettabyte—dwarfs the entire record of human civilisation.  Apparently we are in the midst of a Data Revolution but it is not the quantity of data that is revolutionary.  The big data revolution is that now we can do something with the data.

In marketing, familiar uses of big data include “recommendation engines” like those used by companies such as Netflix and Amazon to make purchase suggestions based on the prior interests of one customer as compared to millions of others. Target famously (or infamously) used an algorithm to detect when women were pregnant by tracking purchases of items such as unscented lotions—and offered special discounts and coupons to those valuable patrons. Credit-card companies have found unusual associations in the course of mining data to evaluate the risk of default: people who buy anti-scuff pads for their furniture, for example, are highly likely to make their payments.

In the public realm, there are all kinds of applications: allocating police resources by predicting where and when crimes are most likely to occur; finding associations between air quality and health; or using genomic analysis to speed the breeding of crops like rice for drought resistance. Google has analysed clusters of search terms by region in the United States to predict flu outbreaks faster than was possible using hospital admission records.

Big data is a collection of data from traditional and digital sources inside and outside your company that represents a source for ongoing discovery and analysis.

There is a tendency for  “big data” to be constrained to digital inputs, however by definition, we can’t exclude traditional data derived from product transaction information, financial records and interaction channels, such as the call center and point-of-sale. All of that is big data, too, even though it may be dwarfed by the volume of digital data that’s now growing at an exponential rate.

In defining big data, it’s also important to understand the mix of unstructured and multi-structured data that comprises the volume of information.

Unstructured data comes from information that is not organised or easily interpreted by traditional databases or data models, and typically, it’s text-heavy. Metadata, Twitter tweets, and other social media posts are good examples of unstructured data.

Multi-structured data refers to a variety of data formats and types and can be derived from interactions between people and machines, such as web applications or social networks. A great example is web log data, which includes a combination of text and visual images along with structured data like form or transactional information. As digital disruption transforms communication and interaction channels—and as marketers enhance the customer experience across devices, web properties, face-to-face interactions and social platforms—multi-structured data will continue to evolve.

So, Big Data is nothing new, there are simply new inputs for different data sets – so let’s explode some myths around Big Data?

Myth #1  –  Companies need “Data Scientists”

After a recent session delivering “How SME’s can harness the power of Data” We  were asked about how to effectively recruit “Data Scientists”  – this was the first time we had heard this term so we drilled down into what type of person they were looking for, this is what we were told:

  •  Needs to be a guaduate – with a good pass in Maths
  •  Knowledgable in IT with a demonstable skill set and certification
  •  Have work experience in analytics, statistics and computer programming

However , the company had no strategy on how they wanted this person to work, what the outputs of the role would be or even how the company would use the analysis!

The search for this mythical scientist is nonsensical,  most large companies have the following already:

  • Good mathematicians, statisticians  who often need the business stuff spoon-fed to them
  • Good certified IT people who understand some maths but little business application
  • Good Certified and Uncertified IT people who understand business (after working enough problems)
  • Business types who understand maths and perhaps a little IT
  • Subject matter / content / Marketing experts
  • Leaders who know how to get these people to work together

The answer is to forget the “Scientist” role and deliver through a working group from your existing cross section of experts.

Myth #2  – It’s Big
Big Data isn’t “big”. It is diverse. “Big” is a misnomer. What we’re talking about is a large volume of data points, updated at high-frequency in real-time, from various sources. It’s very granular. It’s individual transactional data; it’s a certain credit card, paying for a certain product, at a certain retailer. Big Data is actually lots and lots of very small data.

Certainly, the volume of information coming from the Web, modern call centers and other data sources can be enormous. But the main benefit of all that data isn’t in its size. It’s not even in the business insights you can get by analysing individual data sets in search of interesting patterns and relationships. To get true business intelligence from big data analytics applications, user organisations and BI and analytics vendors alike must focus on integrating and analysing a broad mix of information — in short, wide data.

Myth #3  – The more granular the data, the better
Is real-time and granular data always better? No, it’s not. The first half of a football game doesn’t predict how a whole game plays out. In fact many online betting sites use past statistical big data to predict how the football game may in fact pan out and offer a “cashing out” option to try to mitigate its losses!  Real-time can be too close to the action. Sometimes, you need to pull back for the wide-angle shot to reveal what’s really going on.

Big Data is encumbered by a huge amount of white noise. The noise as a proportion of the total signal increases with higher resolution, for example, data by minute rather than by week, or data at a town level rather than state. Do not confuse precision with accuracy. Big Data, in its raw disaggregated form, can be misleading. There needs to be an appropriate level of aggregation to cancel out all the white noise.

Myth #4  – Big Data gives you concrete answers
Ambiguity is the dominant characteristic of Big Data. Multiple sources of data (for example, transaction, customer acquisitions and media) can lead you away from what the actual, non-data based evidence is telling you. Different data, analysed incorrectly, can yield conflicting evidence. Which data do you believe? Big Data requires human judgment to intervene and resolve seemingly conflicting evidence, and that’s where the skilled (and business aware, with life experience ) analyst comes in.

The more data you have, the more likely you are to have contradictions and ambiguities that require resolution. Big Data is not all-powerful. Quite the opposite, in fact. More data gives you more witnesses but doesn’t get you closer to the truth until you leverage experienced human judgment to reconcile conflicting evidence. The future of analytics is all about combining, weighing and judging multiple sources of information and different analyses.

 

Myth #5  – Big Data is New
90% of the available data has been created in the last two years and the term big data has been around 2005, when it was launched by O’Reilly Media in 2005. However, the usage of big data and the need to understand all available data has been around much longer.

There are a lot of examples of companies that were into big data before it was called big data.  In 1995 they introduced their Clubcard, but instead of just using if for offering discounts, they understood that such a loyalty card could generate valuable insights into the shopping behaviour of their customers. Today, they receive detailed data on two-thirds of all shopping baskets thanks to the Clubcard. All this data enables Tesco to send very specific targeted emails to their customers. In 1999 they had 149.000 iterations of their newsletter and that has since then only expanded.  Currently they have over 38 million Clubcard members, of which 16 million are active users, and that offers Tesco valuable information and insights.

However, personalisation is not the only Big Data applications of Tesco. They are applying predictive data analytics to forecast how many products will be sold where. By combining weather data and sales data they know what to expect and in the past years that has resulted in over £6.3 million less food wastage in the summer, £33 million less wastage due to optimised store operations and £55 million less stock in warehouses.

As with any business initiative, a big data project involves an element of risk. Any project can fail for any number of reasons: bad management, under-budgeting, or a lack of relevant skills. However, big data projects bring their own specific risks.


The Risks

Risk #1: You need to apply it right away
Most things in life that are important and worthwhile are difficult, and the analysis of Big Data is no different. The solution is to take small steps and start with very specific objectives. Many companies still struggle with antiquated systems and processes (eg. homegrown Excel spreadsheets for critical data). Most organisations need to first focus on improving existing systems before they should even think about Big Data.

Think carefully about what you want to do with the information before you start stockpiling data.

Risk #2:  Data Security

This risk is obvious and often uppermost in our minds when we are considering the logistics of data collection and analysis. Data theft is a rampant and growing area of crime – and attacks are getting bigger and more damaging. In fact five of the six most damaging data thefts of all time (eBay, JP Morgan Chase, Adobe, Target, and Evernote) were carried out within the last two years.

The bigger your data, the bigger the target it presents to criminals with the tools to steal and sell it. In the case of Target, hackers stole credit and debit card information of 40 million customers, as well as personal identifying information such as email and geographical addresses of up to 110 million people. In March 2015 , a federal judge approved a settlement in which Target would pay $10 million into a settlement fund, from which payments would be made to everyone affected by the breach.

Risk #3:  Data Privacy

Closely related to the issue of security is privacy.  But in addition to ensuring that people’s personal data are safe from criminals, you need to be sure that the sensitive information you are storing and collecting isn’t going to be divulged through less malevolent but equally damaging misuse by yourself or by people to whom you have delegated responsibility for analysing and reporting on it.

 

Failing to follow applicable data protection laws can lead to expensive lawsuits and even prison, depending on what sort of data you are using and the jurisdiction you are in.  Big data frontrunner,  private hire and car sharing service Uber stirred up controversy when one of its executives was caught using the service’s “God mode” to track the movements of BuzzFeed journalist Johana Bhuiyan.

Risk #4:  Costs

Data collection, aggregation, storage, analysis, and reporting  – all cost money. On top of this, there will be compliancy costs – to avoid falling foul on the issues raised in the above risks 2 & 3. These costs can be mitigated by careful budgeting during the planning stages, but getting it wrong at that point can lead to spiralling costs, potentially negating any value added to your bottom line by your data-driven initiative. This is why “starting with strategy” is so vital. A well-developed strategy will clearly set out what you intend to achieve and the benefits that can be gained so they can be balanced against the resources allocated to the project.

Risk #5:  Ignoring the Human factor!

Earlier we referred to Tesco who were one of the first retail companies to use big data to drive their business forward, however, the difficulties that the retailer has been facing in the past few years have puzzled those who assume that effective use of big data protects organisations.  However the risk is that the big data is seen as the  holy grail and the human consumer / client is forgotten in the data mass,  Tesco gathered an impressive amount of data but it was not being used to improve the customer experience, for example – clubcard information was being used to identify short-term promotions which would appeal to its clubcard users, however the figures were not able to identify that consumers were becoming wary of the lack of transparency of the promotions and were looking to every day low prices on staple goods, additionally standards in stores were beginning to slide – there was big data available to support this and correlated correctly (for example – social media complaints properly aggregated would clearly discover patterns of complaints and, therefore, should shape operational goals) would have indicated key areas that needed addressing.  However, both of these issues mask the fact that with 310,000 employees – they have a valuable dataset that they clearly have failed to tap into – Data and especially digital, real-time data is extremely valuable – but, it is important to remember that effective human interaction and communication is probably the most accurate dataset out there!

 

[icegram campaigns=”423″]


Warning: Trying to access array offset on value of type bool in /home/sites/3b/7/74de28c56e/public_html/wp-content/themes/applauz/views/prev_next.php on line 10
previous
Why your business needs you to blog!

Warning: Trying to access array offset on value of type bool in /home/sites/3b/7/74de28c56e/public_html/wp-content/themes/applauz/views/prev_next.php on line 36
next
Best Practices are an organisational destination - do not rush!

Warning: Trying to access array offset on value of type bool in /home/sites/3b/7/74de28c56e/public_html/wp-content/plugins/related-posts-thumbnails/related-posts-thumbnails.php on line 846
https://www.instahost.solutions/wp-content/uploads/2018/10/logo1.png
https://www.instahost.solutions/wp-content/uploads/2017/03/logo_white.png
Insta Security
Website Secured by InstaHost.co.uk
InstaHost Solutions

Our Mission is to deliver an industry leading, comprehensive service to all of our clients regardless of client size or complexity of services required, we are committed to continually striving to develop new, innovative services and technologies in order to continue deliver cutting edge service solutions to all of our clients. We give our clients full control of their digital business without a ridiculous price tag, and our friendly team offers their expertise at all times!

Subscribe

If you wish to receive our latest news in your email box, just subscribe to our newsletter. We won’t spam you, we promise!

    Applauz

    As the pioneer of the lean startup movement, APPLAUZ has dedicated it’s time to sharing effective business strategies that help new businesses and enterpreneurs put their money to work in the right way.

    2021 Copyright by InstaHost Solutions, Powered by InstaHost.co.uk. All rights reserved.