Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts

Thursday, November 27, 2014

IBM’s Brad Becker on Watson and the ‘Humane’ Promise of Cognitive Computing


http://knowledge.wharton.upenn.edu/article/ibms-brad-becker-on-the-promise-of-cognitive-computing/

And it caused many to wonder what this technology might be able to do beyond the realm of a TV game show. Starting in 2012, IBM began to pilot uses of Watson in health care and other fields. The company has since launched a series of products based on the technology, which it calls “cognitive computing.”
Knowledge@Wharton spoke with Brad Becker, chief design officer for IBM Watson, about current and future applications of cognitive computing and how he hopes to make computers “more humane.” An edited version of the conversation follows.

Knowledge@Wharton: Your background is in user experience design. How does that play a role in IBM’s Watson Project?
Becker: [It’s based on] the idea that technology should work for people, not the other way around. Twitter  For a long time, people have worked to better understand technology. Watson is technology that works to understand us. It’s more humane, it’s helpful to humans, it speaks our language, it can deal with ambiguity, it can create hypotheses, it can learn from us. And, of course, since it’s a computer, it can scale as much as needed and has recall far beyond what humans have.
You take the traditional strength of computers, but in a way that’s more comfortable and efficient for people — more humane, I like to say. And it allows experts, or even non-experts, to do much more than they could otherwise.
Knowledge@Wharton: When you speak of a “more humane” computer, what does that mean?
Becker: Technology, traditionally, is created by technologists. That sounds like it’s a tautology, but the people who are usually creating technology love the technology and accept it as it is. Alan Cooper wrote a book called The Inmates Are Running the Asylum talking about this problem. What’s the solution?
Part of the solution is to take time to focus on who is going to be using the technology, what their needs are, how humans work, what’s the ethnography and the cognitive psychology of the people who are actually using the technology. How do we better fit the technology for humans? It’s sort of like ergonomics with furniture.
Here we use IBM Design Thinking, and we look at the business problem — both for IBM and their clients — as well as using hands-on research to understand the end users and what their specific needs are in context. We also look at things that are more horizontal: How can this technology, in general, work better for people and be at the service of people? Have you ever struggled with technology and thought, “Who came up with this?” Or felt like maybe you were dumb, because you couldn’t understand how to use this tool that was supposedly meant for you?
That is what we’re really after. We’re trying to come up with this idea of cognitive computing future today. The whole focus of this is that technology should work for people, and not the other way around. It starts with what we need and what we think is helpful for humans. How do we help augment humans? The bicycle didn’t replace legs, it augmented what they could do. That’s our goal: to take what humans are good at and then supplement the things humans aren’t good at, such as reading 50 million passages and remembering every word; making it possible for a human, with the help of Watson, to be able to do much more than they could without Watson.
The whole focus of this is that technology should work for people, and not the other way around.
Knowledge@Wharton: Can you explain in basic terms how Watson does what it does?
Becker: [Watson] is not a copy of the human brain, but it takes a similar approach [to solving problems] in that there are multiple, completely separate approaches running in parallel. We handle different kinds of queries differently, depending on the nature of that question.
We’ve also moved beyond just question and answer to discovery, where you’re looking not just for “the answer.” The answer today, if one exists, is not necessarily the same as what the answer will be tomorrow. Things change quickly; there’s ambiguity. Watson is good at dealing with that.
Discovery is an interesting application because you’re looking for weak signals in the noise. It’s a big data problem, but [one] where you’re not just looking for the most obvious things. You’re not just running a linear regression or doing typical shallow machine learning. It’s a great example of the combination of a human expert and Watson working together to sift kind of through this to find the needle in the haystack.
There are some examples in the press recently [describing how] Baylor [University] found quite a few discoveries by plugging Watson [into] all the material that’s available, just to test it. They applied it against older material to see if Watson would come up with the same discoveries that the scientific community had come up with in the last decade, and Watson found several of them within a matter of weeks.
Knowledge@Wharton: So you’re regression testing against things that have already been discovered by humans, and then seeing whether Watson can come to the same conclusion?
Becker: Right. We take 18 years to train a human to get to a level that we call “adult.” Whereas, in Watson, it can be in weeks or months to get Watson trained to provide value in a particular area or domain. In this particular case, Baylor did this to check out the Watson technology for themselves by applying it to their own data. And, sure enough, Watson was able to find those hidden connections that were out there already.
Knowledge@Wharton: You said that the kind of cognitive computing that Watson does is not quite the same as the way humans think. How are they similar and how are they different?
Becker: Well, for one thing, the human mind is interesting in that it’s a very low-power, small, portable computer attached to a stomach that can run on berries and nuts. There are all these physiological aspects of the human brain that are both limitations and strengths. There are specializations of the brain.
We’re not looking to reproduce the human brain, because quite frankly, nature’s done that just fine. We’re looking at the way the human brain works, the way people like to work, and then looking at what traditional computers do and saying: It’s not a great fit. By learning from some of the things the human brain does, we can make computers more useful to people. We don’t want to lose the perfect recall, the ability to scale and the speed at which computers can do rote tasks — but, mixed with some of the learnings from the human brain and the way that all the different pieces of the brain work together in unison to come up with a path to understanding.
We’re also adding a lot of work on natural language processing, so that computers can speak the same language we do. Since the story of the Tower of Babel, we’ve known that being able to speak a common language is important in understanding or working with someone. Until now, computers have only spoken code and you program them. Now, we’re moving to a world where you can actually use human language to work with a computer. Instead of programming a computer, you can train a computer.
Knowledge@Wharton: You mentioned the project IBM did with Baylor University. Tell us more about that.
Becker: The Baylor College of Medicine and IBM joint project used a particular toolkit with Watson to identify proteins that modified p53, which is a protein related to a lot of different cancers. They looked at 70,000 scientific articles on p53, and they were predicting proteins that turn on or off the activity of that protein. They found six potential proteins to target for new research. In general, the industry discovers one protein a year that might be interesting. By looking at 70,000 articles, Watson was able to handle at scale all of that information and came up with six promising directions. Humans will go and explore those.
Knowledge@Wharton: There are a number of different Watson-based products: the medical applications you’ve mentioned, a product called Chef Watson, etc. Can you give an overview of the products that are being spun off of the core technology?
Becker: Yes. We have Watson Engagement Advisor, [which is] a solution that allows you to have a better relationship with your customers. It takes your unstructured data and makes it available to your customers; it assists them to go in for themselves and find the information they’re looking for. You also have the ability as a business, then, to see the questions that are being asked, etc.
The second thing, that we talked about already, is Watson Discovery Advisor, which is about finding those connections and relationships – [for example,] using Watson to look through law enforcement databases and unstructured data to find potential connections between suspects, between events, etc.
We’ve talked about drug discovery where you can look for: What are the connections between different elements that have not been adequately explored? It’s one thing to find an obvious connection, but it’s another to go find a weaker, less obvious or more indirect connection that warrants exploration.
Chef Watson is another Discovery domain, where you’re discovering how there might be interesting connections between different ingredients [in cooking recipes]. The fact that people like ice cream on top of apple pie is not very interesting, but the fact that strawberries and mushrooms have a similar chemical makeup and they might pair well together is interesting — in a sense of discovery — when people are looking for something interesting and new.
Another project is Watson Explorer, which started from Enterprise Search. It’s one thing to explore the bodies of knowledge that already exist, but every institution has their own internal bodies of knowledge, as well. Watson Explorer is a way to collect all the information that’s already in your organization and make it available for people to explore and to find information and answers.
We just announced the Watson Developer Cloud. Now developers can come in, use Watson’s services and IBM Bluemix [IBM’s cloud platform] to create their own cognitive applications based on the services that we expose there. Students, universities, companies and even independent developers can come in and kick the tires and play with cognitive computing.
There are still a lot of things we don’t understand about our own brains and our own behaviors, let alone how technology can help us take them further.
Knowledge@Wharton: Are all of these hosted, cloud-based services?
Becker: No. The Watson Developer Cloud is a hosted SaaS [software as a service] based, scalable solution. There are also solutions that are based on cloud technologies but are hosted locally at the company.
Watson Explorer, for instance, works locally and finds all your information. I think the idea of cloud technologies, APIs and scalable servers are definitely a part of it always, but there are on-premises versions, based on the customer preference.
Knowledge@Wharton: When you talk about Watson working in areas like health care and crime detection, should we be concerned that people will have too much faith in its analysis? For example, we’ve seen cases in which the knowledge that when a spouse is killed the husband or wife is the most likely suspect can circumvent the exploration of other scenarios. Is there a similar concern with Watson that we’ll have too much confidence in its analysis, so that other avenues — which even though they are less likely could still be correct — may not be pursued?
Becker: That’s actually the main purpose of Discovery Advisor — to look for potential, subtle connections, not necessarily the obvious ones. The obvious connections are self-evident, so you don’t need Watson to find [them]. The fact that a spouse is an obvious suspect in a domestic murder case — you don’t need Watson for that.
The Discovery Advisor is focused on the opposite [problem]: looking for all those subtle, indirect, weak signal connections; finding the nonobvious connections and the fertile ground for investigation and for human expertise to pay attention; helping humans find where the needle in the haystack might be.
Knowledge@Wharton: This notion — that computers should work like people, rather than people working like computers — has been around for quite a while, going back to at least Apple’s Macintosh in 1984. In fact, Steve Jobs also used the bicycle analogy you mentioned. He saw the Macintosh as an “engine for the mind.” What’s taken us so long to make progress in this area?
Becker: This is a really hard, worthy challenge. There are a lot of different aspects of it. Understanding people is one challenge. There are still a lot of things we don’t understand about our own brains and our own behaviors, let alone how technology can help us take them further.
But some of this is just common sense, practical things. At IBM, we’ve hired a lot of people who are focused on building more humane computing. We have not just visual designers who make things look appropriate and refine the visual details, but also people who work through the workflows that customers are trying to do.
What is it like to be a life scientist or drug researcher? What’s it like to be a customer support representative? We work with [financial services and insurance company] USAA; what’s it like to be separated from the military? What are the questions you have? What are the concerns? What’s your state of mind? What’s that like?
We actually do ethnography. We sit with these sorts of people to understand what they care about. There are a lot of common patterns, and we’re identifying some of those, but the fact of the matter is that humans are kind of messy and complicated. We’re trying to create technology that interfaces with something that is dynamic and complicated.
Knowledge@Wharton: It seems like for many companies this issue of user experience design is often an afterthought. Is that a fair assessment?
Becker: Traditionally, I think that it was completely an afterthought. Think about a physical space, like a house, where if every time you walked in the house, there was a wall two feet in front of you and you’d run into it, or you’d hit your head on a really low thing. We set up standards for these things. For some reason, the more virtual technology has been a little bit immune to that. I think it’s catching up now. You see this across the web, across the applications — there’s an increased focus in the industry now on user experiences.
Knowledge@Wharton: If an executive came to you and said, “I want to differentiate my product by its user experience,” what advice would you give?
Humans are interesting and complex and they are hard to directly replace.
Becker: The best thing is to get out of the building and watch people use your product. That will tell you what the problems are. For the solutions, there are great ways — like IBM Design Thinking — that help you brainstorm, try out, fail fast, go in and focus and come up with what are the most important solutions for these problems. But it starts with understanding your users and, of course, understanding the capability of your technology. Most tech companies are good at doing that side of it, but understanding your users is the number one thing.
The second thing that I would say is: hire professionals; I would hire people who have experience in this, that have a passion for it. And will know how to shepherd a culture that promotes it.
Ultimately, your culture has to promote it. You mentioned Apple — it’s not that they necessarily have the most designers, but from Steve Jobs on down there was an appreciation of design and the importance that things work well for people and that you keep in mind why you’re doing it. That same culture has been growing at IBM. You have to inculcate a culture that says, “At the end of the day, we’re trying to solve a problem for somebody or provide some sort of value for someone. We’d better understand and be able to articulate what that is.”
Knowledge@Wharton: How will technologies like Watson reshape the employment landscape in the future? Won’t these technologies eliminate a lot of jobs?
Becker: It’s funny, because I was looking at some material in the history of IBM, talking about computers back in the ’60s. There were all these discussions in the ’50s and ’60s about how these new computers were going to replace humans in the office, and there would be no more office jobs. I don’t know about you, but I work in offices and there are a lot of people there — and lots of computers, too.
Humans are interesting and complex and they are hard to directly replace Twitter . But cognitive computing can help us with what our own limitations are, and help us to branch beyond those limitations.
Knowledge@Wharton: Any thoughts on when the singularity will occur? When will computer thought outpace that of human beings?
Becker: I can’t see those lines completely converging. I’m not sure we ever get there 100%. The progress in that direction will be surprisingly helpful to us — especially, as we understand ourselves better, we can make our technology work better for us. But it’s not clear to me that there’s going to be a world [in which computers surpass humans]. Because ultimately, a human has to come up with how a computer or a machine could get to the point of creating itself and creating others. It still has to be devised by humans to do this.
Knowledge@Wharton: Looking out ten or more years, then, what is the future for cognitive computing?
Becker: It’s hard to see all the fruit that this will bear. It’s really exciting already, but I think we’re going to see more of the promises fulfilled. It’s going to get easier, it’s going to be faster, it’s going to be more ubiquitous. You’re going see the fulfillment of the promise that technology will be more focused on people, more adaptable to people, more useful, more humane.
I think a lot about how the technology can serve us. Sometimes, as we have more and more technology in our lives, it feels like we’re serving the technology. We look at the carbon footprint of things, which is great, but I think there’s a responsibility in the tech industry to make sure that you also look at the cognitive footprints of the technology that you’re creating. Is it more of a burden than a blessing? We want to make sure that we understand the users and what they need, so that we can create a technology that is adapted to that, and is a net positive benefit.
That combination of making sure that technology serves us and looking, in particular, at Watson and how cognitive computing is going to be able to fit with humans a lot better than the traditional super calculators we’ve had, I think that’s exciting.
Image credit: “IBM Watson” by Clockready – Own work. Licensed under Creative Commons Attribution-Share Alike 3.0 via Wikimedia Commons.

Sunday, August 31, 2014

An architect's guide: How to use big data

This essential guide shows how organizations are using big data and offers advice for IT professionals working with the technology.


Introduction

Employees in organizations of all sizes can be faced with the daunting task of figuring out how to use big data and how to best manage it. Some IT professionals are tasked to use technology such as graph databases to crunch large volumes of data. Other developers need to be able to use tools like Hadoop to build systems capable of handling varying flows of data.
This guide brings together a range of stories that highlight examples of how to use big data, management techniques, trends with the technology, and key terms developers need to know.

1The cloud and big data

Techniques for working with big data and the cloud

What is big data? How should it be used? These are two questions people commonly ask when they begin working with large volumes of data. While use cases are evolving, there are some tips and tricks that can be gleaned from success stories.
The following is a collection of articles going over the basics of working with big data.
Feature

What is big data?

Get answers to frequently asked questions about what big data is, why it can be a problem and how big data tools will be part of the solution.Continue Reading
Feature

How big data and cloud based analytics fit together

Organizations are handling more and more data all the time, and a big problem is figuring out how to find an important piece of information in peta-bytes of big data. How can it be done? Cloud based technologies that can burst and grow are becoming the standard solution. Continue Reading
Feature

AWS Big Data Solutions overcomes common challenges

Achieving an affordable database solution that is both scalable and performant has always been a challenge, but Amazon has put scalability and performance within the reach of all sizes of business with their NoSQL solutions that have grown out of their Dynamo based big data systems.Continue Reading
Answer

How data grid technology can help wrangle big data

Many organizations are finding that current IT setups cannot meet modern demands, and in some instances, using data grid technologies can help.Continue Reading
Tip

Data persistence: What big data and cloud app developers should know

Data persistence can be problematic because it is often related to how an application is functioning. Continue Reading
Tip

How to decide between REST and SOA for big data apps

When designing big data applications, an important consideration is whether to use SOA or RESTful APIs to connect big data components and services to the rest of the application. Continue Reading

2Using Hadoop

Popular tools for working with big data

It's difficult to discuss how to use big data and managing large volumes of data without discussing Hadoop, a Java-based framework. Megacorporations like Google and IBM have capitalized on the technology, but that doesn't mean smaller companies don't stand to benefit from it as well.
Read on for technical advice about working with Hadoop.
Feature

How to reap the benefits of Hadoop with MapReduce 2.0

YARN represents the biggest architectural change in Hadoop since it's inception over seven years ago. Now, Hadoop goes beyond MapReduce to provide scheduled processing while simultaneously processing big data. Continue Reading
Feature

Where YARN fits in to Hadoop 2.0's success

MapReduce has matured, and so has Hadoop, and together under the umbrella of YARN, these powerful technologies are working together better than ever to deliver faster and more flexible big data solutions to the enterprise. Continue Reading
News

Be on the lookout: Hadoop, big data trends for 2014

Learn about the next steps in big data trends for the enterprise in 2014.Continue Reading

3Big data in marketing

Benefit from harnessing big data

Long gone are the days of marketers scratching their heads to determine ways to use big data to their advantage. Organizations of all sizes are learning the benefits of being able to analyze big data and turn that information into a powerful resource to reach consumers. With this movement comes the need for IT professionals to know how to build systems able to wrangle large volumes of data.
The following is a collection of articles that highlight the basics of using big data in marketing.
News

How big data in marketing passes control from companies to consumers

There is a shift taking place in the business world, as big data in marketing empowers customers over companies. Continue Reading
Tip

Advice for using high-performance computing to analyze big data

Researchers and business users alike analyze big data in order to glean insights as to what customers actually want and need. Continue Reading
Feature

How to use big data management tools to meet consumer demand

Organizations are taking advantage of big data management tools as applications are required to handle a growing volume of data. Continue Reading

4Graph database use cases

Make visual representations with big data

Implementing graph databases is one example of the ways organizations have learned to make meaningful use of big data. While the technology has its roots in social media, there are practical uses for graph databases that extend beyond Facebook. From dating sites to online retailers, the use cases are extensive.
Read on to learn more about big data and graph databases.
News

FAQ: Building a Facebook graph search using Neo4j database

At Big Data Techcon 2014, Software field engineer Max de Marzi makes a case for enterprise graph searches using the Neo4j database. Continue Reading
News

Graph database use cases beyond social media

While the most commonly known graph database use cases involve social media, it's not the only market to make use of the technology. Continue Reading
News

How graph searches and big data databases can be practical for enterprises

Software field engineer Max De Marzi explains why graph searches and big databases can be of practical use to the enterprise. Continue Reading

5Glossary

Common big data terms

This glossary provides common terms related to big data.
Definition

Big data

Big data is an evolving term that describes any voluminous amount of structured, semi-structured and unstructured data that has the potential to be mined for information. Although big data doesn't refer to any specific quantity, the term is often used when speaking about petabytes and exabytes of data. Continue Reading
Definition

Big data analytics

Big data analytics is the process of examining large amounts of different data types, or big data, in an effort to uncover hidden patterns, unknown correlations and other useful information. Continue Reading
Definition

Big data management

Big data management is the organization, administration and governance of large volumes of both structured and unstructured data. Continue Reading
Definition

Graph database

A graph database, also called a graph-oriented database, is a type of NoSQL database that uses graph theory to store, map and query relationships.  A graph database is essentially a collection of nodes and edges. Each node represents an entity and each edge represents a relationship between two nodes. Continue Reading
Definition

Hadoop

Hadoop is a free, Java-based programming framework that supports the processing of large data sets in a distributed computing environment. It is part of the Apache project sponsored by the Apache Software Foundation. Continue Reading
Definition

MapReduce

MapReduce is a software framework that allows developers to write programs that process massive amounts of unstructured data in parallel across a distributed cluster of processors or stand-alone computers.Continue Reading
Definition

NoSQL

NoSQL database, also called Not Only SQL, is an approach to data management and database design that's useful for very large sets of distributed data.   Continue Reading

http://searchsoa.techtarget.com/essentialguide/An-architects-guide-How-to-use-big-data?track=NL-1806&ad=895717&asrc=EM_NLN_33375111&uid=16651510&utm_medium=EM&utm_source=NLN&utm_campaign=20140829_An+architect%27s+guide+to+using+big+data_msargent

Saturday, July 19, 2014

Big Data: what is the missing link for FMCG brands?

Big Data: what is the missing link for FMCG brands?
Big Data: what is the missing link for FMCG brands?
Big Data has been extensively talked about as a strategic imperative for brands. So far, the buzz has focussed around individual purchase data collected by retailers, banks and utilities, but what about FMCG brands?
Unsurprisingly, global FMCGs such as Unilever and P&G are now committing to using data as a strategic tool to drive their businesses.
P&G’s supersavvyme community recently won the Data Strategy Award for Proximity’s work on ‘Mums on a Mission’. The promotion used both social media and coupons to track their consumers’ path to purchase.
Unilever’s chief marketing office, Keith Weed, says: “We are already able to tell a consumer when he’s walking in the park (we know his location) on a hot day (we know what the weather is like there), where the nearest place is to buy a Magnum and send him a code for a discount.”
Smaller manufacturers are catching on too. Paterson’s Shortbread and Berberana Wine are now using data to gain a competitive advantage at low cost. As first movers in their categories they have the advantage, knowing competitors will need to invest heavily to catch up.
Why does Big Data matter to FMCG?
Big Data enables brands to be more relevant and responsive to their consumers. Retailers, such as Tesco, are already masters at this with their loyalty card activity.
Combining individual consumer data with purchase data enables retailers to segment their customers in fine detail. Consequently, they can target consumers with personalised, relevant communications that really hit the mark.
While it’s clear that Unilever is embracing Big Data, I’m sure Keith Weed would agree that their example highlights the missing key element for FMCG brands: individual purchase data.
Did that Magnum discount code actually result in a sale? How do they know? Also, do they need to keep sending discount codes and continually discount the brand to keep him engaged? Or will this just annoy him and prompt him to opt out of Unilever’s database? And how does Unilever measure the return on inestment (ROI) of this campaign without accurate sales data?
Analysis companies such as Dunn Humby or Nielsen can help brands to measure overall sales performance, but not to know exactly who their consumers are and why they bought a certain item.
Brands that create Big Data that links individual purchase data to their promotional activity gain an indisputable advantage because they can accurately measure their ROI and respond intelligently. Using Big Data to drive intelligent promotions creates a massive opportunity for brands to increase sales without resorting to price discounting.
One category-leading FMCG brand is already successfully employing this tactic to sell 3.2 packs more per annum to its database; a 74 per cent increase versus their typical frequency of purchase.
Anchor Butter and its agency, whynot!, are currently using Hive’s intelligent promotional platform to gather individual purchase data from consumers and generate more active users from their database. By being more relevant and responsive, they have seen a 50 per cent increase in active consumers and recently picked up the IPM Gold Award for long-term loyalty.
How can FMCG brands link promotion data with individual purchase data?
Printing unique codes on packs unlocks the enormous potential Big Data holds for FMCG brands. Plus, thanks to modern technology, it is now cost-effective and can be implemented easily without disrupting production.
I can claim this with some confidence as Hive manages more than 2bn unique codes per year and has a 100 per cent success rate in enabling code printing.
When consumers register and enter codes, it enables brands to collect high quality behavioural and individual purchase data in a single view. They are then perfectly placed to engage consumers in a highly relevant and unique way.

Tuesday, July 1, 2014

When Big Data Isn’t an Option

http://www.strategy-business.com/article/00250?pg=all

Companies that only have access to “little data” can still use that information to improve their business.

An advertising agency met with a client—who happened to be a U.S. Marine Corps colonel—and the conversation turned to the topic of reliable data. “Look,” said the colonel, “if I’m on a battlefield trying to defend a hill and I get a piece of intelligence, even if I’m not 100 percent sure that it’s accurate, I will make decisions based on that intelligence.” He strongly believed that it’s better to have some information than none—and that you’d be a fool to disregard it just because it falls short of being definitive. One could say that the colonel was a proponent of “little data.”
There is, of course, a great deal of discussion about the potential of “big data,” the high-volume, high-velocity, high-variety information assets that require new forms of data processing to enable companies to make better decisions and operate more efficiently. Giant data sets are being created by aggregates of individuals’ behavior (on social media sites such as Twitter and Instagram, for example), by transaction logs, and by automated information-sensing devices. Companies are increasingly mining these data sources to understand more about their customers’ behavior and preferences, and even to anticipate stock market movements. Early successes by a few companies have caused others to start investing in the infrastructure, software, and talent required to mine big data.
There is, however, one important caveat. Many companies—probably most—work in relatively sparse data environments, without access to the abundant information needed for advanced analytics and data mining. For instance, point-of-sale register data is not standard in emerging markets. In most B2B industries, companies have access to their own sales and shipment data but have little visibility into overall market volumes or what their competitors are selling. Highly specialized or concentrated markets, such as parts suppliers to automakers, have only a handful of potential customers. These companies have to be content with what might be called little data—readily available information that companies can use to generate insights, even if it is sparse or of uneven quality. For these companies, the U.S. Marine colonel’s words will resonate more than the latest data-mining algorithm or social listening platform.
Several commentators have made the point that the implications of big data go beyond new data sources, analytical techniques, and technology. Rather, a paradigm shift—away from management based on gut feelings and toward data-driven decision making—is already under way, and accelerating. The shift is so profound that companies lacking complete or clean market data can no longer use this deficit as an excuse to rely on the status quo. They must make a concerted effort to use the data that is available to them (imperfect as it may be) or to explore innovative, low-cost ways to create new data.
Companies lacking complete or clean data can’t use that as an excuse to rely on the status quo.
In one example, a large beverage manufacturer wanted to improve its sales to bars, restaurants, and entertainment venues. For years, this company had been buying syndicated data from an established source, which covered more than 100,000 establishments. Unfortunately, the data was collected and structured to serve a broad set of clients and featured a standard segmentation scheme that did not provide enough insight for the beverage company into how to serve different segments. So the company decided to adopt a series of little data techniques to come up with a solution customized to its needs.
It started with observational research, visiting bars and restaurants and qualitatively cataloging the clientele and their consumption patterns. Synthesizing this information resulted in more actionable segment definitions. The next step was to quantify the segmentation—determining how many establishments were in each segment. The beverage manufacturer developed an algorithm based on observable characteristics, then asked its sales professionals to classify all the bars and restaurants in their territories based on the algorithm. (This is a classic little data technique: filling in the data gaps internally.) Finally, for each major segment, the company designed tailored product assortments, pricing, and marketing programs. Pilot projects in two large cities have shown significant lifts in total sales and share penetration, and the company is now rolling out the initiative nationwide.
Other companies have used little data successfully as well. In one case, a maker of industrial coating products had limited data on pricing broken down by customer and region. As a result, it couldn’t build robust price elasticity models using classical regression analysis. By using other analytical techniques, however, the company was able to identify specific areas in which it could improve pricing and service policies. It moved to a value-based pricing approach to ensure its most profitable customers were receiving the highest service levels. Implementation in one business unit in one region alone yielded a 4 percent increase in sales.
In another instance, a regional health insurance company trying to differentiate itself through outstanding customer experience realized that its call center was a potential source of data about customer pain points and potential solutions. The company took full transcripts of the calls—not just the summaries entered by service representatives—and applied available text-mining algorithms. From this data, the company was able to improve the format and language of its written communications, and streamline the call-center process. In addition, it uncovered an opportunity to introduce storefront locations in certain neighborhoods in order to improve its customer interactions and increase customer retention rates.
Even large companies are able to make use of little data techniques. The Chinese large-appliance giant Haier uses information gathered by service technicians to drive innovation. In the late 1990s, some technicians, for example, found that rural customers were using their washing machines to wash vegetables, leading to clogs. Haier used this information to develop a new type of washer, which the company says is “mainly for washing clothes, sweet potatoes, and peanuts.”
With the right mind-set, virtually all sources of information can be exploited to improve products, the customer experience, or a company’s profits. Little data techniques, therefore, can include just about any method that gives a company more insight into its customers without breaking the bank. As the examples above illustrate, mining little data doesn’t mean investing in expensive data acquisition, hardware, software, or technology infrastructure. Rather, companies need three things:
• The commitment to become more fact-based in their decision making.This commitment is often spurred by a sense that competition is heating up or the company is falling behind changing customer habits and preferences. But fact-based decision making can be an important source of competitive advantage for market-leading companies.
• The willingness to learn by doing. Since little data applications are not commercially available via third parties, companies have to use trial and error. However, once a few priorities have surfaced, a series of pilot projects will give the company useful experience and, with a little luck, some early successes that can inspire the rest of the organization.
• A bit of creativity. To generate richer data, companies need to get creative, in part by tapping into the customer interactions that take place naturally. For instance, retailers can intercept shoppers in store locations for quick iPad-assisted surveys. Any website with a registration form can add questions that reveal preferences beyond the basic data usually collected. Call-center conversations are another opportunity to gather data on a particular topic, and the text can be mined for greater insight into the customer. Some companies create advanced user panels of savvy customers to get input during the R&D process for new products. Others rely on their sales representatives to report trends in customer preferences and competitors’ activities. The bottom line: Companies have to put in the extra effort required to capture and interpret data that is already being generated.
Companies often start the journey by picking a product, a region, and a problem that needs attention and running one or more pilot projects. This allows executives to demonstrate to themselves and the rest of the organization that the return on effort and cost is justified. Once companies start investing in analytics, they almost never stop, because the things they learn drive improvements in the business that more than pay for the analysis. The activity becomes self-funding. In some cases, companies that start with little data end up recognizing the value of the resulting insights and expanding their investment to incorporate larger data sets and more advanced analytics. For others, little data is all that’s needed. In either case, the benefits are clear: Executives get insight into what they can do to improve their competitive position, or—to put it in terms that a Marine Corps colonel might appreciate—identify what might be charging up the hill to surprise them. It’s hard to put a price tag on that.