Google Track

Showing posts with label Predictive. Show all posts
Showing posts with label Predictive. Show all posts

Sunday, March 23, 2014

MH370 flight mystery may have an answer


Data Science perspective: MH370 flight data are crucial

It has been a while and still no answer for the missing airplane of flight MH370. This case may give many possibilities, so guessing by trying each of them is painful and may lead to wrong directions, time consuming and frustrations. So simulating both physical and mathematical model maybe an answer

Mathematical Model

Data Science is newly established profession to solve problems with data insights in high volume of data, so in this case like airplane information systems, weather condition, pattern recognition, path selections etc.
Based in the data above, Data Scientists like me can build specific models to adapt exact or similar scenarios of the missing plane in the line MH370
This may give as good picture what may happen to flight, where may landed, what weather condition we had on that time, what other circumstances occurred and what impact they had over the plane.



Picture: Possible routes

If Malaysian Airport can provide with specific technical information about MH370 airplane, historical data of all MH370 flights, we may find important patterns to solve this mystery. Period between departure and lost signal may give us the distance compared to average distance of all MH370 flights on the same period. Adding weather condition may help us to segments the other MH370 flights that had same weather condition to seek specific scenarios under same weather condition. Route lines are similar to all flights in the same airline, but segmenting them to specific conditions of the missing flight like weather, airplane type, technical conditions, air pressure etc… may lead us to better targets. Technical check before departure may give us information of what have been checked or not and if there is room for possible technical failure. If there is room for failures, we can use data from all technical failure plan crashes to predict the time occurred the failure by that distance also segmented to adapt most the missing flight model.
Big Data technology provides us with the power to analyze big and complicated data sets, and there are plenty of professionals to do so.

Physical model- Simulation

I am not the expert of the field but it can be very smart to build a flight simulation of MH370 based on the mathematical model that we provided here and including other external data that have impact on the fight itself. Sometimes visualization may bring in table other factors that may be decisive in solving the mystery.
These represent alternative approaches to what may help involved institutions to solve this case and I hope they may consider these.

Friday, January 24, 2014

Big Data and Data Science Books - A Baker's Dozen

Here are 13 informative and inspirational books on Big Data and Data Science.  This is definitely not intended to be a comprehensive list (since a complete list of such readings would itself be a form of "Big Data", and consequently the number of possibilities is a nearly uncountable number!NOTE definition of "uncountable" = an infinite set that contains too many elements to be countable.)
  1. Big Data: A Revolution That Will Transform How We Live, Work, and T..., by Viktor Mayer-Schonberger and Kenneth Cukier
  2. The Signal and the Noise: Why So Many Predictions Fail-but Some Don't, by Nate Silver
  3. Predictive Analytics: The Power to Predict Who Will Click, Buy, Lie..., by Eric Siegel 
  4. The Human Face of Big Data, by Rick Smolan and Jennifer Erwitt
  5. The Black Swan: The Impact of the Highly Improbable, by Nassim Nicholas Taleb
  6. Competing on Analytics: The New Science of Winning, by Thomas H. Davenport and Jeanne G. Harris
  7. Super Crunchers: Why Thinking-by-Numbers is the New Way to Be Smart, by Ian Ayres
  8. Big Data Marketing: Engage Your Customers More Effectively and Driv..., by Lisa Arthur
  9. Journeys to Data Mining: Experiences from 15 Renowned Researchers, by Mohamed Medhat Gaber (editor)
  10. The Fourth Paradigm: Data-Intensive Scientific Discovery, by T.Hey, S.Tansley, and K.Tolle (editors)
  11. Seven Databases in Seven Weeks: A Guide to Modern Databases and the..., by Eric Redmond and Jim Wilson
  12. Data Mining And Predictive Analysis: Intelligence Gathering And Cri..., by Colleen McCue
And here are two more, as a bonus:
 14. A Statistical Guide for the Ethically Perplexed, by Lawrence Hubert and Howard Wainer

Friday, September 21, 2012

How Big Data Brings BI and Predictive Analytics Together

Big data is breathing new life into business intelligence by putting the power of prediction into the hands of everyday decision-makers.


For as long as anyone can remember, the world of predictive analytics has been the exclusive realm of ivory-tower statisticians and data scientists who sit far away from the everyday line of business decision maker. Big data is about to change that.
As more data streams come online and are integrated into existing BI, CRM, ERP and other mission-critical business systems, the ever-elusive (and oh so profitable) single view of the customer may finally come into focus. While most customer service and field sales representatives have yet to feel the impact, companies such as IBM and MicroStrategy are working to see that they do soon.

Big Data Moves Analytics Beyond Pencil-Pushers
Imagine a world in which a CSR sitting at her console can make an independent decision on whether a problem customer is worth keeping or upgrading. Imagine, too, that a field salesman can change a retailer's wine rack on the fly based on the preferences that partiers attending the jazz festival next weekend have contributed on Facebook and Twitter.
Big data is pushing a tool more commonly used for cohort and regression analysis into the hands of line-level managers, who can then use non-transactional data to make strategic, long-term business decisions about, for example, what to put on store shelves and when to put it there.
However, big data is not about to supplant traditional BI tools, says Rita Sallam, Gartner's BI analyst. If anything, big data will make BI more valuable and useful to the business. "We're always going to need to look at the past…and when you have big data, you are going to need to do that even more. BI doesn't go away. It gets enhanced by big data."
How else you will know if what you are seeing in the initial phases of discovery will indeed bear out over time. For example, do red purses really sell better than blue ones in the Midwest? An initial pass through the data may suggest so—more red purses sold last quarter than ever before, therefore, red purses sell better.
But this is a correlation, not a cause. If you look more closely, using historical transaction data gleaned from your BI tools, you may find, say, that it is actually your latest merchandise-positioning-campaign that's paying dividends because the retailers are now putting red purses at eye level.
That's why IBM's Director of Emerging Technologies, David Barnes, is actually more inclined to refer to the resulting output from big data technologies such as Hadoop, map/reduce and R as "insights." You wouldn't want to make mission-critical business decisions based on sentiment analysis of a Twitter stream, for example.






Tuesday, August 28, 2012

Nordic Choice Business Intelligence

Here is a video where shows the job that me and BI collegues from Nordic Choice Hotels have done last 2 years.
Also we collaborated with Platon (www.platon.net) in UX and Dashboard Designing of our Business Intelligence Solution.



A Visionary Choice - Nordic Choice Hotels Business Intelligence vision from Platon on Vimeo.

Tuesday, July 24, 2012

The Future of Decision Making: Less Intuition, More Evidence



A fantastic post by Andrew McAfee

Human intuition can be astonishingly good, especially after it's improved by experience. Savvy poker players are so good at reading their opponents' cards and bluffs that they seem to have x-ray vision. Firefighters can, under extreme duress, anticipate how flames will spread through a building. And nurses in neonatal ICUs can tell if a baby has a dangerous infection even before blood test results come back from the lab.

The lexicon to describe this phenomenon is mostly mystical in nature. Poker players have a sixth sense; firefighters feel the blaze's intentions; Nurses just know what seems like an infection. They can't even tell us what data and cues they use to make their excellent judgments; their intuition springs from a deep place that can't be easily examined. . Examples like these give many people the impression that human intuition is generally reliable, and that we should rely more on the decisions and predictions that come to us in the blink of an eye.

This is deeply misguided advice. We should rely less, not more, on intuition.

A huge body of research has clarified much about how intuition works, and how it doesn't. Here's some of what we've learned:

•It takes a long time to build good intuition. Chess players, for example, need 10 years of dedicated study and competition to assemble a sufficient mental repertoire of board patterns.

•Intuition only works well in specific environments, ones that provide a person with good cues and rapid feedback . Cues are accurate indications about what's going to happen next. They exist in poker and firefighting, but not in, say, stock markets. Despite what chartists think, it's impossible to build good intuition about future market moves because no publicly available information provides good cues about later stock movements. Feedback from the environment is information about what worked and what didn't. It exists in neonatal ICUs because babies stay there for a while. It's hard, though, to build medical intuition about conditions that change after the patient has left the care environment, since there's no feedback loop.

•We apply intuition inconsistently. Even experts are inconsistent. One study determined what criteria clinical psychologists used to diagnose their patients, and then created simple models based on these criteria. Then, the researchers presented the doctors with new patients to diagnose and also diagnosed those new patients with their models. The models did a better job diagnosing the new cases than did the humans whose knowledge was used to build them. The best explanation for this is that people applied what they knew inconsistently — their intuition varied. Models, though, don't have intuition.

•It's easy to make bad judgments quickly. We have a many biases that lead us astray when making assessments. Here's just one example. If I ask a group of people "Is the average price of German cars more or less than $100,000?" and then ask them to estimate the average price of German cars, they'll "anchor" around BMWs and other high-end makes when estimating. If I ask a parallel group the same two questions but say "more or less than $30,000" instead, they'll anchor around VWs and give a much lower estimate. How much lower? About $35,000 on average, or half the difference in the two anchor prices. How information is presented affects what we think.

•We can't know tell where our ideas come from. There's no way for even an experienced person to know if a spontaneous idea is the result of legitimate expert intuition or of a pernicious bias. In other words, we have lousy intuition about our intuition.

My conclusion from all of this research and much more I've looked at is that intuition is similar to what I think of Tom Cruise's acting ability: real, but vastly overrated and deployed far too often.

So can we do better? Do we have an alternative to relying on human intuition, especially in complicated situations where there are a lot of factors at play? Sure. We have a large toolkit of statistical techniques designed to find patterns in masses of data (even big masses of messy data), and to deliver best guesses about cause-and-effect relationships. No responsible statistician would say that these techniques are perfect or guaranteed to work, but they're pretty good.

The arsenal of statistical techniques can be applied to almost any setting, including wine evaluation. Princeton economist Orley Ashenfleter predicts Bordeaux wine quality (and hence eventual price) using a model he developed that takes into account winter and harvest rainfall and growing season temperature. Massively influential wine critic Robert Parker has called Ashenfleter an "absolute total sham" and his approach "so absurd as to be laughable." But as Ian Ayres recounts in his great book Supercrunchers, Ashenfelter was right and Parker wrong about the '86 vintage, and the way-out-on-a-limb predictions Ashenfelter made about the sublime quality of the '89 and '90 wines turned out to be spot on.

Those of us who aren't wine snobs or speculators probably don't care too much about the prices of first-growth Bordeaux, but most of us would benefit from accurate predictions about such things as academic performance in college; diagnoses of throat infections and gastrointestinal disorders; occupational choice; and whether or not someone is going to stay in a job, become a juvenile delinquent, or commit suicide.

I chose those seemingly random topics because they're ones where statistically-based algorithms have demonstrated at least a 17 percent advantage over the judgments of human experts.

But aren't there at least as many areas where the humans beat the algorithms? Apparently not. A 2000 paper surveyed 136 studies in which human judgment was compared to algorithmic prediction. Sixty-five of the studies found no real difference between the two, and 63 found that the equation performed significantly better than the person. Only eight of the studies found that people were significantly better predictors of the task at hand. If you're keeping score, that's just under a 6% win rate for the people and their intuition, and a 46% rate of clear losses.

So why do we continue to place so much stock in intuition and expert judgment? I ask this question in all seriousness. Overall, we get inferior decisions and outcomes in crucial situations when we rely on human judgment and intuition instead of on hard, cold, boring data and math. This may be an uncomfortable conclusion, especially for today's intuitive experts, but so what? I can't think of a good reason for putting their interests over the interests of patients, customers, shareholders, and others affected by their judgments.

So do we just dispense with the human experts altogether, or take away all their discretion and tell them to do whatever the computer says? In a few situations, this is exactly what's been done. For most of us, our credit scores are an excellent predictor of whether we'll pay back a loan, and banks have long relied on them to make automated yes/no decisions about offering credit. (The sub-prime mortgage meltdown stemmed in part from the fact that lenders started ignoring or downplaying credit scores in their desire to keep the money flowing. This wasn't intuition as much as rank greed, but it shows another important aspect of relying on algorithms: They're not greedy, either).

In most cases, though, it's not feasible or smart to take people out of the decision-making loop entirely. When this is the case, a wise move is to follow the trail being blazed by practitioners of evidence-based medicine , and to place human decision makers in the middle of a computer-mediated process that presents an initial answer or decision generated from the best available data and knowledge. In many cases, this answer will be computer generated and statistically based. It gives the expert involved the opportunity to override the default decision. It monitors how often overrides occur, and why. it feeds back data on override frequency to both the experts and their bosses. It monitors outcomes/results of the decision (if possible) so that both algorithms and intuition can be improved.

Over time, we'll get more data, more powerful computers, and better predictive algorithms. We'll also do better at helping group-level (as opposed to individual) decision making, since many organizations require consensus for important decisions. This means that the 'market share' of computer automated or mediated decisions should go up, and intuition's market share should go down. We can feel sorry for the human experts whose roles will be diminished as this happens. I'm more inclined, however, to feel sorry for the people on the receiving end of today's intuitive decisions and judgments.

What do you think? Am I being too hard on intuitive decision making, or not hard enough? Can experts and algorithms learn to get along? Have you seen cases where they're doing so? Leave a comment, please, and let us know.



Thursday, June 28, 2012

Ten steps to predictive success

Follow these best practices to ensure a successful foray into predictive analytics.


1. Define the business proposition. What is the business problem you are trying to solve? What is the question you're trying to answer? Think like a business leader first and an analyst or IT expert second.

2. Line up a business champion. Having the support of a key executive and a stakeholder is crucial. Whenever possible help the stakeholder to become the initiator and champion of the project.

3. Start off with a quick win. Find a well-defined business problem where analytics can bring value by showing measurable results. Start small and use simple models to build credibility.

4. Know the data you have. Do you have enough data, enough history and enough granularity in the data to feed your proposed model? Getting it into the right form is the biggest part of any first-time predictive analytics project.

5. Get professional help. A statistical background and a little training aren't enough: Creating predictive models is different from traditional descriptive analytics, and is as much an art as it is a science. Get help for that first win before striking out on your own.

6. Be sure the decision maker is prepared to act. It's not enough to have a prescribed action plan. The results may dictate actions that are counterintuitive. If the business decision makers won't act or aren't in a position do so, you re wasting your time, so get a strong commitment up front.

7. Don't get ahead of yourself. Stay within the scope of the defined project, even if success breeds pressure to expand the use of your current model. Good analytics sells itself, but overextending can result in an unreliable model that will kill credibility.

8. Communicate the results in business language. Don't discuss probabilities and variances. Do talk revenue impact and fulfillment of business objectives. Use data visualization tools to hammer home the point.

9. Test, revise, repeat. Start small, test, revise and test again. Conduct A/B testing to demonstrate value. Present the results, gain critical mass, then scale out.

10. Hire me to implemment above steps with success :)

Tuesday, May 15, 2012

Effective big data strategies detailed


Businesses beginning a big data analytics program with advanced business intelligence software may be concerned about affordability. However, according to PC Advisor, there are easy steps that companies can take to make sure that their deployments are successful. These steps include deep research into the business case at hand and prudent financial planning.



Careful planning


"[Big data is] new technology solving a business problem that we often haven't proved. That's important for CIOs to keep in mind," financial consultant Jeff Muscarella told PC Advisor. "The business is going to be coming to them with all sorts of half-baked ideas for what they can do with Big Data. They have to ask: Will it really drive revenue? How and for how long?"
According to the source, carefully vetting business ideas for new big data projects is vitally important for CIOs trying to save money on their big data projects. Gathering details on each projected usage of data means less chance of failure. The source urged companies to target their big data projects, to fire "bullets" rather than "cannons" at specific problems that can provide value for the company. Muscarella told the source that companies can start small to prove that a process works before moving to the company-wide infrastructure level.



Myths vs. reality


As a widely hyped technology often presented as the future of business intelligence, advanced analytics and big data have received a large amount of press. To avoid business confusion, several publications have offered clarifications of what the technology can and cannot offer companies. The Economic Times stated that any business with a product to sell and any company hoping to help make up the potential market for big data. As companies begin to harness the power of big data, competitors could take of the systems in a bid to compete on an even level.
The source sought to puncture myths about what big data can and cannot do. It stated that big data's endgame is unknown, and that many of the features of big data analytics are still spoken of in the future tense. The source found that companies can already use the technology to provide "amazing" customer insights from vast quantities of "irrelevant stuff." While it is important to be careful when integrating big data, making the effort could become a required part of business strategy

Industry News from: http://www.panorama.com/industry-news/article-view.html?name=Effective-big-data-strategies-detailed-774397&utm_source=dlvr.it&utm_medium=facebook

Thursday, May 10, 2012

Predictive Analytics and Data Mining




Derive useful insights to make evidence-based decisions

Today's organizations accumulate huge volumes of data from a variety of sources on a daily basis. However, turning increasingly large amounts of data into useful insights and finding how to better utilize those insights in decision making remains a challenge for most.

To get answers to complex questions and gain an edge in today's marketplace requires powerful, multipurpose predictive analytic solutions so you can learn from, utilize and improve on knowledge gained from vast stores of data. BIandIT Ltd provides a wide range of software for exploring and analyzing data to help uncover unknown patterns, opportunities and insights that can drive proactive, evidence-based decision making within your organization.

Text mining applies the same analysis techniques to text-based documents. The knowledge gleaned from data and text mining can be used to fuel strategic decision making.



Components of Predictive Analytics and Data Mining

Exploratory Data Analysis – Get dynamic visualization, advanced statistical techniques and core data mining capabilities to quickly identify relationships and opportunities.

Model Development and Deployment – Streamline the data mining process to create highly accurate descriptive and predictive analytic models based on large volumes of data.

Analytics Acceleration – Generate faster results and improve data governance with in-database analytics.

Scoring Acceleration – Maximize the performance and accuracy of your analytic models.




Tuesday, May 8, 2012

Microsoft Predictive Analytics


Predictive analytics is the next step in BI: not only can you be retrospective and see what has happened in your company in the past, but now we can distill new information from the old information to actually predict what will happen in the future. Jamie MacLennan, CTO of Predixion Software, explains the difference between business intelligence and predictive analytics and shares a program that Predixion has created in Excel to review the Practice Fusion data.



Featuring Bruno Aziza

Friday, March 30, 2012

Rafal Lukawiecki in SQL Server 2012 official launch in Oslo


First, thanks for invitation from Microsoft, specially Thale Mjavatn and nice organization of such an important event. The details for the event you find in this blog:
http://biblogg.no/2012/03/15/lansering-av-microsoft-sql2012/
and the agenda of the conference is here: http://www.microsoft.com/norge/bi-sql-fagdag/index.html
Rafal was fantastic as always and did a remarkable job explaining to the audience the new technologies behind MS SQL 2012, special focus on the Business Intelligence enhancement tools included. We got introduced to 2 new technologies: PowerView and BISM and little changes on what PowerPivot is now. Also, he spoke about trends where BI and Big Data will be the most important things in the future.
Once more thank you for the opportunity and thanks to all who were part of it.

Wednesday, March 28, 2012

BI News from Panorama

Sort through data with business intelligence


With endless information accessible in today's big data sprawl, it is often difficult for businesses to sort through data that is valuable and isn't without outside assistance. Because of this, more and more organizations are turning to business intelligence in order to sort through vast reams of data to develop concrete analytics.
A new report by DSquared Media provides valuable insight into the worth of adopting business intelligence software and infrastructure. According to the study, for every $1.00 spent on business analytics, $10.66 was yielded in returns. Furthermore, 74 percent of organizatons who manually assembled data from various sources negatively affected daily operations.

The report also found that many large corporations used BI in developmental years in order to become the giants they are today. For example, Febreze used a marketing campaign aided by BI when first released, and now sales total over $1 billion a year. In addition, Target used marketing campaigns aided by BI and revenues grew from $44 billion in 2002 to $67 billion in 2010.
The study found that there was an assortment of reasons why people were influenced by BI. Ninety-five percent of respondents found that they were influenced by BI for its ability to increase insight into operations. Furthermore, 85 percent found business intelligence provided faster process and reporting cycle time, while only 48 percent were influenced by regulatory compliance.



Tuesday, March 20, 2012

Analytics in Sports

I am fan of football and my favorite team is FC Barcelona. Combining sports, specially football with Analytics is just amazing.

Thursday, March 15, 2012

BI for Customers

BI for everyone, does it sound familiar!

It is a fact that Business Intelligence was dedicated to big companies, enterprises because they have that amount of data to be considered interesting for analytics and BI. Now, Gartner started the idea of Bi for mid-size and small businesses, so they need attention too based on BI surveys. But, have you ever thought for a BI solution in Customer Level, or more detailed do you think you can handle a personal BI solution.
ELA will give you the answer.

ELA solution for Customer Intelligence

What ELA is actually?

Elegant Analytics represents the name of a general BI solution in or group of methodologies in Analytics that adapts to every Business profile. In this case, ELA will provide solution for personal finance and planning of your budget. The name of the product is PFI (Personnal Finance Intelligence). Inspired by the TV Show “Luksusfellen” here in Norway, this BI end-user tool may be a solution for all these who fail to maintain well their own economy and for those who want to perform their economy as well. The purpose of this project is to create a Customer Analytical Cube that would process data for each bank costumer using his/her history for its own benefit and then answer you most important queries that users do against their own data.

This solution will include also benchmarking against an Imaginary subject (Ola Nordman) that can be Min, Max or Avg of the customer’s measures in a certain region, for a period of time, similar age group, sex and income levels.

For having more controle and planning your own economy, will be an extra parameter as Target, so users (bank customers) will put their targets for costs and income a month, quarter or a year ahead and always will be warned when they are about to achieve the amount they targeted.

If you want to read more then follow the link where you can download the full project.

Friday, March 9, 2012

How I do Predictive Analytics without Data Mining?

This hot topic is comming soon. I will just describe a bit what will this topic will include. This topic is going to show how to do Predictive Analytics on your data without using Data Mining or DMX and just using MDX. How can Data Mining prediction Algorithms be "translated" from DMX to MDX. How accurate are they and what is the benefit of using MDX?

Wednesday, June 22, 2011

Predictive Analytics vs Data Mining

Technology Cycle:
Data warehousing is a mature technology, with approximately 70 percent of Forrester Research survey respondents indicating they have one in production. Data mining has endured significant consolidation of products since 2000, in spite of initial high-profile success stories, and has sought shelter in encapsulating its algorithms in the recommendation engines of marketing and campaign management software. Statistical inference has been transformed into predictive modelling. As we shall see, the emerging trend in predictive analytics has been enabled by the convergence of a variety of factors.

Technology Hierarchy:
In the technology hierarchy, data warehousing is generally considered an architecture for data management. Of course, when implemented, a data warehouse is a database providing information about (among many other things) what customers are buying or using which products or services and when and where are they doing so. Data mining is a process for knowledge discovery, primarily relying on generalizations of the "law of large numbers" and the principles of statistics applied to them. Predictive analytics emerges as an application that both builds on and delimits these two predecessor technologies, exploiting large volumes of data and forward-looking inference engines, by definition, providing predictions about diverse domains.

Methods:
The method of data warehousing is structured query language (SQL) and its various extensions. Data mining employs the "law of large numbers" and the principles of statistics and probability that address the issues around decision making in uncertainty. Predictive analytics carries forward the work of the two predecessor domains. Though not a silver bullet, better algorithms in operations research, risk minimization and parallel processing, when combined with hardware improvements and the lessons of usability testing, have resulted in successful new predictive applications emerging in the market. (Again, see Figure 1 on predictive analytics enabling technologies.) Widely diverging domains such as the behaviour of consumers, stocks and bonds, and fraud detection have been attacked with significant success by predictive analytics on a progressively incremental scale and scope. The work of the past decade in building the data warehouse and especially of its closely related techniques, particularly parallel processing, are key enabling factors. Statistical processing has been useful in data preparation, model construction and model validation. However, it is only with predictive analytics that the inference and knowledge are actually encoded into the model that, in turn, is encapsulated in a business application.

Definition
This results in the following definition of predictive analytics: Methods of directed and undirected knowledge discovery, relying on statistical algorithms, neural networks and optimization research to prescribe (recommend) and predict (future) actions based on discovering, verifying and applying patterns in data to predict the behavior of customers, products, services, market dynamics and other critical business transactions. In general, tools in predictive analytics employ methods to identify and relate independent and dependent variables - the independent variable being "responsible for" the dependent one and the way in which the variables "relate," providing a pattern and a model for the behavior of the downstream variables.

In data warehousing, the analyst asks a question of the data set with a predefined set of conditions and qualifications, and a known output structure. The traditional data cube addresses: What customers are buying or using which product or service and when and where are they doing so? Typically, the question is represented in a piece of SQL against a relational database. The business insight needed to craft the question to be answered by the data warehouse remains hidden in a black box - the analyst's head. Data mining gives us tools with which to engage in question formulation based primarily on the "law of large numbers" of classic statistics. Predictive analytics have introduced decision trees, neural networks and other pattern-matching algorithms constrained by data percolation. It is true that in doing so, technologies such as neural networks have themselves become a black box. However, neural networks and related technologies have enabled significant progress in automating, formulating and answering questions not previously envisioned. In science, such a practice is called "hypothesis formation," where the hypothesis is treated as a question to be defined, validated and refuted or confirmed by the data.

Tuesday, April 26, 2011

Predictive Analytics

Predictive analytics encompasses a variety of techniques from statistics, data mining and game theory that analyze current and historical facts to make predictions about future events.

In business, predictive models exploit patterns found in historical and transactional data to identify risks and opportunities. Models capture relationships among many factors to allow assessment of risk or potential associated with a particular set of conditions, guiding decision making for candidate transactions.

Predictive analytics is used in actuarial science, financial services, insurance, telecommunications, retail, travel, healthcare, pharmaceuticals and other fields.

One of the most well-known applications is credit scoring, which is used throughout financial services. Scoring models process a customer’s credit history, loan application, customer data, etc., in order to rank-order individuals by their likelihood of making future credit payments on time. A well-known example would be the FICO score.

DefinitionPredictive analytics is an area of statistical analysis that deals with extracting information from data and using it to predict future trends and behavior patterns. The core of predictive analytics relies on capturing relationships between explanatory variables and the predicted variables from past occurrences, and exploiting it to predict future outcomes. It is important to note, however, that the accuracy and usability of results will depend greatly on the level of data analysis and the quality of assumptions.

Types: Generally, the term predictive analytics is used to mean predictive modeling, "scoring" data with predictive models, and forecasting. However, people are increasingly using the term to describe related analytical disciplines, such as descriptive modeling and decision modeling or optimization. These disciplines also involve rigorous data analysis, and are widely used in business for segmentation and decision making, but have different purposes and the statistical techniques underlying them vary.

Predictive models: Predictive models analyze past performance to assess how likely a customer is to exhibit a specific behavior in the future in order to improve marketing effectiveness. This category also encompasses models that seek out subtle data patterns to answer questions about customer performance, such as fraud detection models. Predictive models often perform calculations during live transactions, for example, to evaluate the risk or opportunity of a given customer or transaction, in order to guide a decision. With advancement in computing speed, individual agent modeling systems can simulate human behavior or reaction to given stimuli or scenarios. The new term for animating data specifically linked to an individual in a simulated environment is avatar analytics.

Descriptive models: Descriptive models quantify relationships in data in a way that is often used to classify customers or prospects into groups. Unlike predictive models that focus on predicting a single customer behavior (such as credit risk), descriptive models identify many different relationships between customers or products. Descriptive models do not rank-order customers by their likelihood of taking a particular action the way predictive models do. Descriptive models can be used, for example, to categorize customers by their product preferences and life stage. Descriptive modeling tools can be utilized to develop further models that can simulate large number of individualized agents and make predictions.

Decision models: Decision models describe the relationship between all the elements of a decision — the known data (including results of predictive models), the decision and the forecast results of the decision — in order to predict the results of decisions involving many variables. These models can be used in optimization, maximizing certain outcomes while minimizing others. Decision models are generally used to develop decision logic or a set of business rules that will produce the desired action for every customer or circumstance.

Applications: Although predictive analytics can be put to use in many applications, we outline a few examples where predictive analytics has shown positive impact in recent years.

Analytical customer relationship management (CRM): Analytical Customer Relationship Management is a frequent commercial application of Predictive Analysis. Methods of predictive analysis are applied to customer data to pursue CRM objectives.

Clinical decision support systems: Experts use predictive analysis in health care primarily to determine which patients are at risk of developing certain conditions, like diabetes, asthma, heart disease and other lifetime illnesses. Additionally, sophisticated clinical decision support systems incorporate predictive analytics to support medical decision making at the point of care. A working definition has been proposed by Dr. Robert Hayward of the Centre for Health Evidence: "Clinical Decision Support systems link health observations with health knowledge to influence health choices by clinicians for improved health care."

Collection analytics: Every portfolio has a set of delinquent customers who do not make their payments on time. The financial institution has to undertake collection activities on these customers to recover the amounts due. A lot of collection resources are wasted on customers who are difficult or impossible to recover. Predictive analytics can help optimize the allocation of collection resources by identifying the most effective collection agencies, contact strategies, legal actions and other strategies to each customer, thus significantly increasing recovery at the same time reducing collection costs.

Cross-sell: Often corporate organizations collect and maintain abundant data (e.g. customer records, sale transactions) and exploiting hidden relationships in the data can provide a competitive advantage to the organization. For an organization that offers multiple products, an analysis of existing customer behavior can lead to efficient cross sell of products. This directly leads to higher profitability per customer and strengthening of the customer relationship. Predictive analytics can help analyze customers’ spending, usage and other behavior, and help cross-sell the right product at the right time.

Customer retention: With the number of competing services available, businesses need to focus efforts on maintaining continuous consumer satisfaction. In such a competitive scenario, consumer loyalty needs to be rewarded and customer attrition needs to be minimized. Businesses tend to respond to customer attrition on a reactive basis, acting only after the customer has initiated the process to terminate service. At this stage, the chance of changing the customer’s decision is almost impossible. Proper application of predictive analytics can lead to a more proactive retention strategy. By a frequent examination of a customer’s past service usage, service performance, spending and other behavior patterns, predictive models can determine the likelihood of a customer wanting to terminate service sometime in the near future. An intervention with lucrative offers can increase the chance of retaining the customer. Silent attrition is the behavior of a customer to slowly but steadily reduce usage and is another problem faced by many companies. Predictive analytics can also predict this behavior accurately and before it occurs, so that the company can take proper actions to increase customer activity.

Direct marketing: When marketing consumer products and services there is the challenge of keeping up with competing products and consumer behavior. Apart from identifying prospects, predictive analytics can also help to identify the most effective combination of product versions, marketing material, communication channels and timing that should be used to target a given consumer. The goal of predictive analytics is typically to lower the cost per order or cost per action.

Fraud detection: Fraud is a big problem for many businesses and can be of various types. Inaccurate credit applications, fraudulent transactions (both offline and online), identity thefts and false insurance claims are some examples of this problem. These problems plague firms all across the spectrum and some examples of likely victims are credit card issuers, insurance companies, retail merchants, manufacturers, business to business suppliers and even services providers. This is an area where a predictive model is often used to help weed out the “bads” and reduce a business's exposure to fraud.

Predictive modeling can also be used to detect financial statement fraud in companies, allowing auditors to gauge a company's relative risk, and to increase substantive audit procedures as needed.

The Internal Revenue Service (IRS) of the United States also uses predictive analytics to try to locate tax fraud.

Portfolio, product or economy level prediction: Often the focus of analysis is not the consumer but the product, portfolio, firm, industry or even the economy. For example a retailer might be interested in predicting store level demand for inventory management purposes. Or the Federal Reserve Board might be interested in predicting the unemployment rate for the next year. These type of problems can be addressed by predictive analytics using Time Series techniques (see below).

Underwriting: Many businesses have to account for risk exposure due to their different services and determine the cost needed to cover the risk. For example, auto insurance providers need to accurately determine the amount of premium to charge to cover each automobile and driver. A financial company needs to assess a borrower’s potential and ability to pay before granting a loan. For a health insurance provider, predictive analytics can analyze a few years of past medical claims data, as well as lab, pharmacy and other records where available, to predict how expensive an enrollee is likely to be in the future. Predictive analytics can help underwriting of these quantities by predicting the chances of illness, default, bankruptcy, etc. Predictive analytics can streamline the process of customer acquisition, by predicting the future risk behavior of a customer using application level data. Predictive analytics in the form of credit scores have reduced the amount of time it takes for loan approvals, especially in the mortgage market where lending decisions are now made in a matter of hours rather than days or even weeks. Proper predictive analytics can lead to proper pricing decisions, which can help mitigate future risk of default.

Statistical techniques: The approaches and techniques used to conduct predictive analytics can broadly be grouped into regression techniques and machine learning techniques.

Regression techniques: Regression models are the mainstay of predictive analytics. The focus lies on establishing a mathematical equation as a model to represent the interactions between the different variables in consideration. Depending on the situation, there is a wide variety of models that can be applied while performing predictive analytics. Some of them are briefly discussed below.

Linear regression model: The linear regression model analyzes the relationship between the response or dependent variable and a set of independent or predictor variables. This relationship is expressed as an equation that predicts the response variable as a linear function of the parameters. These parameters are adjusted so that a measure of fit is optimized. Much of the effort in model fitting is focused on minimizing the size of the residual, as well as ensuring that it is randomly distributed with respect to the model predictions.

The goal of regression is to select the parameters of the model so as to minimize the sum of the squared residuals. This is referred to as ordinary least squares (OLS) estimation and results in best linear unbiased estimates (BLUE) of the parameters if and only if the Gauss-Markov assumptions are satisfied.

Once the model has been estimated we would be interested to know if the predictor variables belong in the model – i.e. is the estimate of each variable’s contribution reliable? To do this we can check the statistical significance of the model’s coefficients which can be measured using the t-statistic. This amounts to testing whether the coefficient is significantly different from zero. How well the model predicts the dependent variable based on the value of the independent variables can be assessed by using the R² statistic. It measures predictive power of the model i.e. the proportion of the total variation in the dependent variable that is “explained” (accounted for) by variation in the independent variables.

Discrete choice models: Multivariate regression (above) is generally used when the response variable is continuous and has an unbounded range. Often the response variable may not be continuous but rather discrete. While mathematically it is feasible to apply multivariate regression to discrete ordered dependent variables, some of the assumptions behind the theory of multivariate linear regression no longer hold, and there are other techniques such as discrete choice models which are better suited for this type of analysis. If the dependent variable is discrete, some of those superior methods are logistic regression, multinomial logit and probit models. Logistic regression and probit models are used when the dependent variable is binary.

Logistic regression: For more details on this topic, see logistic regression.
In a classification setting, assigning outcome probabilities to observations can be achieved through the use of a logistic model, which is basically a method which transforms information about the binary dependent variable into an unbounded continuous variable and estimates a regular multivariate model (See Allison’s Logistic Regression for more information on the theory of Logistic Regression).

The Wald and likelihood-ratio test are used to test the statistical significance of each coefficient b in the model (analogous to the t tests used in OLS regression; see above). A test assessing the goodness-of-fit of a classification model is the –.

Multinomial logistic regression: An extension of the binary logit model to cases where the dependent variable has more than 2 categories is the multinomial logit model. In such cases collapsing the data into two categories might not make good sense or may lead to loss in the richness of the data. The multinomial logit model is the appropriate technique in these cases, especially when the dependent variable categories are not ordered (for examples colors like red, blue, green). Some authors have extended multinomial regression to include feature selection/importance methods such as Random multinomial logit.

Probit regressionProbit models offer an alternative to logistic regression for modeling categorical dependent variables. Even though the outcomes tend to be similar, the underlying distributions are different. Probit models are popular in social sciences like economics.

A good way to understand the key difference between probit and logit models, is to assume that there is a latent variable z.

We do not observe z but instead observe y which takes the value 0 or 1. In the logit model we assume that y follows a logistic distribution. In the probit model we assume that y follows a standard normal distribution. Note that in social sciences (example economics), probit is often used to model situations where the observed variable y is continuous but takes values between 0 and 1.

Logit versus probit: The Probit model has been around longer than the logit model. They look identical, except that the logistic distribution tends to be a little flat tailed. One of the reasons the logit model was formulated was that the probit model was difficult to compute because it involved calculating difficult integrals. Modern computing however has made this computation fairly simple. The coefficients obtained from the logit and probit model are also fairly close. However, the odds ratio makes the logit model easier to interpret.

For practical purposes the only reasons for choosing the probit model over the logistic model would be:

There is a strong belief that the underlying distribution is normal
The actual event is not a binary outcome (e.g. Bankrupt/not bankrupt) but a proportion (e.g. Proportion of population at different debt levels).
Time series models: Time series models are used for predicting or forecasting the future behavior of variables. These models account for the fact that data points taken over time may have an internal structure (such as autocorrelation, trend or seasonal variation) that should be accounted for. As a result standard regression techniques cannot be applied to time series data and methodology has been developed to decompose the trend, seasonal and cyclical component of the series. Modeling the dynamic path of a variable can improve forecasts since the predictable component of the series can be projected into the future.

Time series models estimate difference equations containing stochastic components. Two commonly used forms of these models are autoregressive models (AR) and moving average (MA) models. The Box-Jenkins methodology (1976) developed by George Box and G.M. Jenkins combines the AR and MA models to produce the ARMA (autoregressive moving average) model which is the cornerstone of stationary time series analysis. ARIMA (autoregressive integrated moving average models) on the other hand are used to describe non-stationary time series. Box and Jenkins suggest differencing a non stationary time series to obtain a stationary series to which an ARMA model can be applied. Non stationary time series have a pronounced trend and do not have a constant long-run mean or variance.

Box and Jenkins proposed a three stage methodology which includes: model identification, estimation and validation. The identification stage involves identifying if the series is stationary or not and the presence of seasonality by examining plots of the series, autocorrelation and partial autocorrelation functions. In the estimation stage, models are estimated using non-linear time series or maximum likelihood estimation procedures. Finally the validation stage involves diagnostic checking such as plotting the residuals to detect outliers and evidence of model fit.

In recent years time series models have become more sophisticated and attempt to model conditional heteroskedasticity with models such as ARCH (autoregressive conditional heteroskedasticity) and GARCH (generalized autoregressive conditional heteroskedasticity) models frequently used for financial time series. In addition time series models are also used to understand inter-relationships among economic variables represented by systems of equations using VAR (vector autoregression) and structural VAR models.

Survival or duration analysis: Survival analysis is another name for time to event analysis. These techniques were primarily developed in the medical and biological sciences, but they are also widely used in the social sciences like economics, as well as in engineering (reliability and failure time analysis).

Censoring and non-normality, which are characteristic of survival data, generate difficulty when trying to analyze the data using conventional statistical models such as multiple linear regression. The normal distribution, being a symmetric distribution, takes positive as well as negative values, but duration by its very nature cannot be negative and therefore normality cannot be assumed when dealing with duration/survival data. Hence the normality assumption of regression models is violated.

The assumption is that if the data were not censored it would be representative of the population of interest. In survival analysis, censored observations arise whenever the dependent variable of interest represents the time to a terminal event, and the duration of the study is limited in time.

An important concept in survival analysis is the hazard rate, defined as the probability that the event will occur at time t conditional on surviving until time t. Another concept related to the hazard rate is the survival function which can be defined as the probability of surviving to time t.

Most models try to model the hazard rate by choosing the underlying distribution depending on the shape of the hazard function. A distribution whose hazard function slopes upward is said to have positive duration dependence, a decreasing hazard shows negative duration dependence whereas constant hazard is a process with no memory usually characterized by the exponential distribution. Some of the distributional choices in survival models are: F, gamma, Weibull, log normal, inverse normal, exponential etc. All these distributions are for a non-negative random variable.

Duration models can be parametric, non-parametric or semi-parametric. Some of the models commonly used are Kaplan-Meier and Cox proportional hazard model (non parametric).

Classification and regression trees: Main article: decision tree learning
Classification and regression trees (CART) is a non-parametric decision tree learning technique that produces either classification or regression trees, depending on whether the dependent variable is categorical or numeric, respectively.

Decision trees are formed by a collection of rules based on variables in the modeling data set:

Rules based on variables’ values are selected to get the best split to differentiate observations based on the dependent variable
Once a rule is selected and splits a node into two, the same process is applied to each “child” node (i.e. it is a recursive procedure)
Splitting stops when CART detects no further gain can be made, or some pre-set stopping rules are met. (Alternatively, the data is split as much as possible and then the tree is later pruned.)
Each branch of the tree ends in a terminal node. Each observation falls into one and exactly one terminal node, and each terminal node is uniquely defined by a set of rules.

A very popular method for predictive analytics is Leo Breiman's Random forests or derived versions of this technique like Random multinomial logit.

Multivariate adaptive regression splines: Multivariate adaptive regression splines (MARS) is a non-parametric technique that builds flexible models by fitting piecewise linear regressions.

An important concept associated with regression splines is that of a knot. Knot is where one local regression model gives way to another and thus is the point of intersection between two splines.

In multivariate and adaptive regression splines, basis functions are the tool used for generalizing the search for knots. Basis functions are a set of functions used to represent the information contained in one or more variables. Multivariate and Adaptive Regression Splines model almost always creates the basis functions in pairs.

Multivariate and adaptive regression spline approach deliberately overfits the model and then prunes to get to the optimal model. The algorithm is computationally very intensive and in practice we are required to specify an upper limit on the number of basis functions.

Machine learning techniques. Machine learning, a branch of artificial intelligence, was originally employed to develop techniques to enable computers to learn. Today, since it includes a number of advanced statistical methods for regression and classification, it finds application in a wide variety of fields including medical diagnostics, credit card fraud detection, face and speech recognition and analysis of the stock market. In certain applications it is sufficient to directly predict the dependent variable without focusing on the underlying relationships between variables. In other cases, the underlying relationships can be very complex and the mathematical form of the dependencies unknown. For such cases, machine learning techniques emulate human cognition and learn from training examples to predict future events.

A brief discussion of some of these methods used commonly for predictive analytics is provided below. A detailed study of machine learning can be found in Mitchell (1997).

Neural networks. Neural networks are nonlinear sophisticated modeling techniques that are able to model complex functions. They can be applied to problems of prediction, classification or control in a wide spectrum of fields such as finance, cognitive psychology/neuroscience, medicine, engineering, and physics.

Neural networks are used when the exact nature of the relationship between inputs and output is not known. A key feature of neural networks is that they learn the relationship between inputs and output through training. There are two types of training in neural networks used by different networks, supervised and unsupervised training, with supervised being the most common one.

Some examples of neural network training techniques are backpropagation, quick propagation, conjugate gradient descent, projection operator, Delta-Bar-Delta etc. Some unsupervised network architectures are multilayer perceptrons, Kohonen networks, Hopfield networks, etc.

Radial basis functions: A radial basis function (RBF) is a function which has built into it a distance criterion with respect to a center. Such functions can be used very efficiently for interpolation and for smoothing of data. Radial basis functions have been applied in the area of neural networks where they are used as a replacement for the sigmoidal transfer function. Such networks have 3 layers, the input layer, the hidden layer with the RBF non-linearity and a linear output layer. The most popular choice for the non-linearity is the Gaussian. RBF networks have the advantage of not being locked into local minima as do the feed-forward networks such as the multilayer perceptron.

Support vector machines: Support Vector Machines (SVM) are used to detect and exploit complex patterns in data by clustering, classifying and ranking the data. They are learning machines that are used to perform binary classifications and regression estimations. They commonly use kernel based methods to apply linear classification techniques to non-linear classification problems. There are a number of types of SVM such as linear, polynomial, sigmoid etc.

Naïve Bayes: Naïve Bayes based on Bayes conditional probability rule is used for performing classification tasks. Naïve Bayes assumes the predictors are statistically independent which makes it an effective classification tool that is easy to interpret. It is best employed when faced with the problem of ‘curse of dimensionality’ i.e. when the number of predictors is very high.

K-nearest neighbours: The nearest neighbour algorithm (KNN) belongs to the class of pattern recognition statistical methods. The method does not impose a priori any assumptions about the distribution from which the modeling sample is drawn. It involves a training set with both positive and negative values. A new sample is classified by calculating the distance to the nearest neighbouring training case. The sign of that point will determine the classification of the sample. In the k-nearest neighbour classifier, the k nearest points are considered and the sign of the majority is used to classify the sample. The performance of the kNN algorithm is influenced by three main factors: (1) the distance measure used to locate the nearest neighbours; (2) the decision rule used to derive a classification from the k-nearest neighbours; and (3) the number of neighbours used to classify the new sample. It can be proved that, unlike other methods, this method is universally asymptotically convergent, i.e.: as the size of the training set increases, if the observations are independent and identically distributed (i.i.d.), regardless of the distribution from which the sample is drawn, the predicted class will converge to the class assignment that minimizes misclassification error.
Geospatial predictive modeling: Conceptually, geospatial predictive modeling is rooted in the principle that the occurrences of events being modeled are limited in distribution. Occurrences of events are neither uniform nor random in distribution – there are spatial environment factors (infrastructure, sociocultural, topographic, etc.) that constrain and influence where the locations of events occur. Geospatial predictive modeling attempts to describe those constraints and influences by spatially correlating occurrences of historical geospatial locations with environmental factors that represent those constraints and influences. Geospatial predictive modeling is a process for analyzing events through a geographic filter in order to make statements of likelihood for event occurrence or emergence.

Tools: There are numerous tools available in the marketplace which help with the execution of predictive analytics. These range from those which need very little user sophistication to those that are designed for the expert practitioner. The difference between these tools is often in the level of customization and heavy data lifting allowed.

In an attempt to provide a standard language for expressing predictive models, the Predictive Model Markup Language (PMML) has been proposed. Such an XML-based language provides a way for the different tools to define predictive models and to share these between PMML compliant applications. PMML 4.0 was released in June, 2009.

References:
L. Devroye, L. Györfi, G. Lugosi (1996). A Probabilistic Theory of Pattern Recognition. New York: Springer-Verlag.
John R. Davies, Stephen V. Coggeshall, Roger D. Jones, and Daniel Schutzer, "Intelligent Security Systems," in Freedman, Roy S., Flein, Robert A., and Lederman, Jess, Editors (1995). Artificial Intelligence in the Capital Markets. Chicago: Irwin. ISBN 1-55738-811-3.
Agresti, Alan (2002). Categorical Data Analysis. Hoboken: John Wiley and Sons. ISBN 0-471-36093-7.
Enders, Walter (2004). Applied Time Series Econometrics. Hoboken: John Wiley and Sons. ISBN 052183919X.
Greene, William (2000). Econometric Analysis. Prentice Hall. ISBN 0-13-013297-7.
Mitchell, Tom (1997). Machine Learning. New York: McGraw-Hill. ISBN 0-07-042807-7.
Tukey, John (1977). Exploratory Data Analysis. New York: Addison-Wesley. ISBN 0201076160.
Guidère, Mathieu; Howard N, Sh. Argamon (2009). Rich Language Analysis for Counterterrrorism. Berlin, London, New York: Springer-Verlag. ISBN 978-3-642-01140-5.

Saturday, June 19, 2010

Analyzing data

As you can imagine, the amount of data contained in a modern business is
enormous. If the data were very small, you could simply use Microsoft Excel
and perform all of the ad-hoc analysis you need with a Pivot Table. However,
when the rows of data reach into the billions, Excel is not capable of handling
the analysis on its own. For these massive databases, a concept called OnLine
Analytical Process (OLAP) is required. Microsoft’s implementation of OLAP is
called SQL Server Analysis Services (SSAS), which I cover in detail in Chapter 8.
If you’ve used Excel Pivot Tables before, think of OLAP as essentially a massive
Pivot Table with hundreds of possible pivot points and billions of rows
of data. A Pivot Table allows you to re-order and sum your data based on different
criteria. For example, you may want to see your sales broken down by
region, product, and sales rep one minute and then quickly re-order the groupings
to include product category, state, and store.
In Excel 2010 there is a new featured called PowerPivot that brings OLAP to
your desktop. PowerPivot allows you to pull in millions of rows of data and
work with it just like you would a smaller set of data. After you get your Excel
sheet how you want it, you can upload it to a SharePoint 2010 site and share
it with the rest of your organization.
With PowerPivot you are building your own Cubes right on your desktop using
Excel. If you use PowerPivot, you can brag to your friends and family that you
are an OLAP developer. Just don’t tell them you are simply using Excel and
Microsoft did some magic under the covers.
When you need a predefined and structured Cube that is already built for
you, then you turn to your IT department.