Google Track

Showing posts with label analysis. Show all posts
Showing posts with label analysis. Show all posts

Friday, January 24, 2014

Big Data and Data Science Books - A Baker's Dozen

Here are 13 informative and inspirational books on Big Data and Data Science.  This is definitely not intended to be a comprehensive list (since a complete list of such readings would itself be a form of "Big Data", and consequently the number of possibilities is a nearly uncountable number!NOTE definition of "uncountable" = an infinite set that contains too many elements to be countable.)
  1. Big Data: A Revolution That Will Transform How We Live, Work, and T..., by Viktor Mayer-Schonberger and Kenneth Cukier
  2. The Signal and the Noise: Why So Many Predictions Fail-but Some Don't, by Nate Silver
  3. Predictive Analytics: The Power to Predict Who Will Click, Buy, Lie..., by Eric Siegel 
  4. The Human Face of Big Data, by Rick Smolan and Jennifer Erwitt
  5. The Black Swan: The Impact of the Highly Improbable, by Nassim Nicholas Taleb
  6. Competing on Analytics: The New Science of Winning, by Thomas H. Davenport and Jeanne G. Harris
  7. Super Crunchers: Why Thinking-by-Numbers is the New Way to Be Smart, by Ian Ayres
  8. Big Data Marketing: Engage Your Customers More Effectively and Driv..., by Lisa Arthur
  9. Journeys to Data Mining: Experiences from 15 Renowned Researchers, by Mohamed Medhat Gaber (editor)
  10. The Fourth Paradigm: Data-Intensive Scientific Discovery, by T.Hey, S.Tansley, and K.Tolle (editors)
  11. Seven Databases in Seven Weeks: A Guide to Modern Databases and the..., by Eric Redmond and Jim Wilson
  12. Data Mining And Predictive Analysis: Intelligence Gathering And Cri..., by Colleen McCue
And here are two more, as a bonus:
 14. A Statistical Guide for the Ethically Perplexed, by Lawrence Hubert and Howard Wainer

Thursday, October 31, 2013

Your Data Analysis Takes How Long?

Reference: www.biblogg.no

andycotgreave


Andy Cotgreave, Social Content Manager at Tableau Software, looks at how analytics tools can help to save valuable time

Imagine, for a moment, that you’ve been given a task to analyse a dataset inside sixty minutes and share your results. How far do you think you would get in that time?
It’s a question I had cause to reflect on recently, after running «Fanalytics», a workshop for users of Tableau Public. In the workshop, we gave people a dataset and one hour to do something cool. Their results were astounding.
To understand why they were so impressive, let me provide a little context by considering how many of us work with data.
First, let me dispel a myth. Contrary to popular opinion, if you are using spreadsheets or traditional BI tools, it is quite possible to build beautiful charts. Unfortunately, each view of your data takes considerable time to build, Do you have that time to spare in your working life? What if the chart you take 10 minutes to build doesn’t answer your question? What if it inspires a new question? You have to go back and start again.
What if you could explore your data at the speed of thought instead? What if each mouse click changed the view instantly? This is what we call visual analytics: it allows you to find insight in your data at speeds unimaginable just a few years ago.
You’re probably wondering how all of this relates to the Fanalytics competition I mentioned earlier. Well, during the session, we gave our teams a list of every UK Number one album since 1956, downloaded from Wikipedia. The instructions they were given were to analyse the data and publish something interesting within one hour.
Did they deliver? Oh, boy, yes, and in ways that made my jaw drop. Each entry was different. The winning team analysed albums that had been to number one more than once, revealing perennially popular music, and the effects of sales on a musician’s death. Another team came up with an album explorer that found out which album was number one on your birthday. One team created a visually gorgeous dashboard, sure to engage anyone. A further team came up with a predictive model based around the likelihood of any album title to get to number one. You can see all the entrants on Tableau’s Fanalytics blog post. What was truly amazing was that they did this in one hour. Sixty minutes!
Unfortunately, many people are stuck with tools that are cumbersome or too hard to use. It often takes more than an hour just to connect to data. The simple lesson I’ve learnt from the recent session is that although some tools can make amazing charts, they are often unnecessarily complicated. With some, you need to fill in five steps in a property wizard just to draw a chart. In others, you are required to write custom scripts before you can start drawing anything.
The question we need to ask is whether we are using the right tools to answer questions quickly and in the most efficient way? If not, then perhaps it’s time to ditch these time hogs and focus on analytic tools that save you time instead!

Wednesday, May 29, 2013

TICKETS


GOBI2013 will be held in Oslo Spektrum June 10th 2013. When registering you get full access to the sessions, restaurants, cafes, expo and entertainment for the whole conference.
The web shop for conference tickets will ease the job of managing your tickets. Through the web shop you can buy any number of tickets and assign them to co-workers and/or friends who shall attend the conference. Each individual attendee is responsible for updating his or her contact information. When required information has been registered, the ticket will be sent directly to the attendee. Tickets are sent as a PDF containing a QR code on e-mail. Attendees must bring the QR code for registration at the conference site on June 10th. A conference pass will be created and handed out at registration.
Price: 5.900 NOK (EarlyBird, until 15.04.2013)
Price: 6.900 NOK (LateBird after 15.04.2013)
Tickets for GOBI2013 are administrated by Macsimum Event AS
Post Conference Seminar with Cindi Howson and Wayne Eckerson
Want to get the most out of the gurus while they’re in Oslo? Would you like to have a deep-dive seminar with BOTH keynote speakers? Attend the GOBI post-conference seminar on June 11th.
Price: 3.900 NOK (Without valid GOBI pass)
Price: 2.900 NOK (With valid GOBI pass from June 10th)

Questions about tickets? Contact: webshop.gobi@eventsystems.no

Tuesday, May 28, 2013

Post-conference seminar with Wayne Eckerson and Cindi Howson

W&C

Want to get the most out of the gurus while they’re in Oslo? Would you like to have a deep-dive seminar with BOTH keynote speakers? Attend the GOBI post-conference seminar on June 11th.
Secrets of Analytical Leaders: Insights from Information Insiders Wayne Eckerson, Principal, BI Leader Consulting
How do you bridge the worlds of business and technology? How do you harness data for business gain? How do you deliver value from BI and analytical initiatives? Based on Wayne’s book, “Secrets of Analytical Leaders: Insights from Information Insiders,” this session will unveil the secrets to success of top BI and analytical leaders from companies such as Zynga, Netflix, US Xpress, Nokia, Capital One, Kelley Blue Book and Blue KC, among others. The session will cover both the “soft stuff” of people, processes, and projects and the “hard stuff” of architecture, tools, and data required to create and sustain a successful BI and analytics program.
You Will Learn:
• How to organize a BI and analytics team for optimal performance
• How to deliver value quickly and earn credibility among business sponsors
• How to translate insights into business impact
• How to create and deploy analytical models
• How to create an agile data warehouse

BI Market Update and How to Choose a Visual Data Discovery Tool Cindi Howson, founder of BI Scorecard
As the face for the data warehouse, the BI tool is the most visible component to business users. BI tools continue to evolve to be more appealing, to reach new classes of users, and to speed the time to insight. At this session, BI tools expert Cindi Howson will offer strategies for managing your BI tool portfolio. She will highlight recent trends, the state of the market, differences in core modules, with a focus for selecting and deploying the right tool for the right user. The second half of the seminar provides an evaluation framework for evaluating dashboards and visual data discovery tools.
You Will Learn:
• State of the BI tools market and key trends
• State of BI standardization, motivations and challenges
• User segments, use cases, and tool positioning
• Differences in core modules
• Dashboard and visual data discovery use cases
• Strengths and weaknesses of leading products

The post-conference seminar is open for both GOBI-participants and others. Attend both GOBI main event and post-conference seminar and get discounts.
The post-conference seminar will take place at Dronning Eufemias gate 16, Bjørvika (Visma-bygget), on June 11th.
Agenda:
0800-0830: Registration
0830-1200: Wayne Eckerson: Secrets of Analytical Leaders: Insights from Information Insiders
1200-1230: Lunch
1300-1630: Cindi Howson: BI Market Update and How to Choose a Visual Data Discovery Tool

TICKETS AVAILABLE SOON!
Check out GOBI website www.gurusofbi.no for tickets.

Sunday, April 14, 2013

My life in Norway: Pursuing the dream (part I)

Before I start telling my dream story, it would be better to say who am I in few lines:






Born and raised in Macedonia, spent 4-5 years in Kosova and then migrated first time in Norway. My family was one of the few interested in science, specially in math, where my father was a math professor and most of my uncles studied math or engineering. I inherited the love to science and math, continued developing my self focused in math by becoming one the best in local, national and international competition of both math and physics (kind of applied mathematics).
Studied computer technology at University of Prishtina and 3 year in row won the University scholarship.
Studied with International professors from Concordia University; Vienna Institute of Technology and Institute Jean Lui Vives.

Even physically in Kosova, my dream was just to move to a more prospered countries to pursue my dream of being a great scientist. I have heard of UK, US and the big american dream, but never thought of Norway....

I moved in Norway some years ago and then I come back May 2010, pursuing my dream for a better career. I never thought that this will the time when the Revolution of my life started. I will never forget the time when I was sitting home and got a call that was actually a job opportunity to work in Norway, to work for one the best companies in the World, Nordic Choice Hotels. I answered with BIG YES and came to the first interview. It was all by the plan, the first interview was successful. Waited in Oslo for a couple of days, where I got invitation for the second round which was decisive. One day after that, I got the call of my career, saying the your job opportunity is now a job offer. Without hesitating I said YES and that was the biggest "yes" of my life, because what happened after proved this conclusion. Still not understanding in what wonderful world I was stepping in.

After signing the contract and some official paper work I started to work in June/July. I was thrilling to start with my new company and bring the successful project of Business Intelligence into live.I had time read and understand the business concept and strategy of Nordic Choice Hotels, so I was ready to dive in directly to the solution.

One of the biggest highlights of my career here is meeting the owner of Nordic Choice Hotels and bunch of other business around Norway, Mr. Petter A. Stordalen. His ability to give energy at any time in the company was special. You could feel his absence or his presence without seeing him at all.


Me and Petter Stordalen at Garden Party

During the time being at Choice, I had the opportunity to meet other important people as well, so I learned a lot from them.

Me and my department made great efforts on creating the best BI solution for the company in a given condition and situation. So we excelled by creating this solution presented in the video:



But things came to an end, sometimes without our willing, so in April I had to change my job and pursue my professional dream at Nextbridge AS. 

Why Nextbridge AS?

NextBridge is an IT-consultancy dedicated solely to business intelligence (BI). The company encompasses more than 100 years of collective experience in this field. Our services span most areas of the BI field; from DWH to reporting, from scorecards and dashboards to data mining and statistical analysis, from BI Competency Centers to maturity analysis. 
NextBridge assists the largest and most demanding clients in improving their managerial information. Our client segment is Norwegian Top 500 accounts and include Sparebank1, SG Finans, Gjensidige, BNBank, Helse Nord and NorgesEnergi . NextBridge consultants are bilingual, and speak both business and IT. Our mission is ”Bridging business and IT at the next level”. Our vision is to be the reference in the field of BI.

The blog post is public, I would not share too much of the details, so I will just jump to give a introduction of my professional profile:

""
I am an IT professional with focus on Business and Data Analytic, prefer to call myself Data Scientist. I have in depth experience using and implementing business intelligence/data analysis tools with greatest strength in the Microsoft SQL Server / Business Intelligence Studio SSIS, SSAS, SSRS. I have designed, developed, tested, debugged, and documented Analysis and Reporting processes for enterprise wide data warehouse implementations using the SQL Server / BI Suite. I also have designed/modeled OLAP cubes using SSAS and developed them using MS SQL BIDS SSAS and MDX. Served as an implementation team member where I translated source mapping documents and reporting requirements into dimensional data models. Strong ability to work closely with business and technical teams to understand, document, design and code SSAS, MDX, DMX, DAX abd ETL processes, along with the ability to effectively interact with all levels of an organization. Additional BI tool experience includes ProClarity, Microsoft Performance Point, MS Office Excel and MS SharePoint.
""

Professional highlights:

1. Worked for Capgemini Norway AS

2. Worked for Nordic Choice Hotels AS

3. Working for Nextbridge AS


Academic Honors:

MIT Honor Code Certificate: CS and Programming (04.06.2013)

Princeton University Honor Code Certificate:  Analytic Combinatorics (10.07.2013)

Stanford University Honor Code Certificate: Mathematical Thinking,, Cryptography (06.05.2013)

IIT University Honor Code Certificate: Web Intelligence and Big Data (02.06.2013)

Wesleyan University: Passion Driven Statistics (20.05.2013)


Career Highlights:

1. Over Nine years of experience in the field of Information Technology, System Analysis and Design, Data
    warehousing, Business Intelligence and Data Science in general

2. Experienced in implementing / managing large scale complex projects involving multiple stakeholders and
    leading and directing multiple project teams

3. Track record of delivering customer focused, well planned, quality products on time, while adapting to
    shifting and conflicting demands and priorities.

4. Experience in Data warehouse / Business Intelligence developments, implementation and operation setup

5. Expertise in Data Modeling, Data Analytics and Predictive Analytics SSAS, MDX and DMX

6. Strong Knowledge in Data warehouse, Data Extraction, Transformation, and Loading ETL

7. Excellent track record in developing and maintaining enterprise wide web based report systems and portals in Finance, Enterprise wide solutions and BI and Strategy Systems

8. Best new employee for 2011 of Nordic Choice Hotels AS


Achievments:

1. First place in reagional math competitions in 2 years in a row

2. First place in Physics competition in a Balkaniada (Balkan Olympics in Theoretical Physics)

3. First place in fast math competition in International Kangourou Competition

4. First place in Norway in Microsoft Virtual Academy (Microsoft Business Intelligence)

Research Work:

1. Riccati Differential Equation solution (published in printed version Research Journal)


3. Personal Finance Intelligence; published in IJSER 8 August 2012 edition



Next you will read: Life in Norway: Living the dream (part II), STAY TUNED!

Friday, September 28, 2012

The Big Data Fairy Tale


By Roel Castelein


Fairy tales usually start with ‘Once upon a time ...' and end with ‘... And they lived long and happily ever after'. But nobody explains ‘how' the heroes live long and happily ever after. Big data (analytics) promise to transform your business, but just as in fairy tale endings, big data will not explain ‘how' to transform your organization. In my view, big data might spark some behavioral change or open people's minds, but it will not transform organizations. At best, big data evolves organizations. Let's look at the concept and a concrete example to draw conclusions.


What big data analytics does is take a bunch of data, analyze and visualize it, and then derive insights that potentially can improve your organization or business. Based on these insights the actual transformation can begin, but it requires more than just big data. Let's have a look at a classic example of data analytics; the reduction of crime in New York under Mayor Giuliani with the help of CompStat.

CompStat is a data system that maps crime geographically and in terms of emerging criminal patterns, as well as charting officer performance by quantifying criminal apprehensions. The key to success was not the data or analysis, but that the organizational management that used the data and analysis was effective. Processes, structures and accountability were setup to drive the transformation. In weekly meetings, NYPD executives met with local precinct commanders from the eight boroughs in New York to discuss the problems. They devised strategies and tactics to solve problems, reduce crime, and ultimately improve quality of life in their assigned area. CompStat tracked the results of these strategies and tactics, and whether they were successful or not. Precinct commanders were held accountable for the results.

Drawing upon my own experience, I know how difficult an organizational transformation is. Even if you have the data and the analysis that shows things need to change, it requires much more than data analysis. Let's assume that the data uncovers opportunities for improvement, either in reducing cost or in increasing revenue. The next step is to design the changes in processes, in people's roles, in org charts and in the systems. This usually entails a two pronged approach; communicate the change in org charts, processes and roles, and engrain these changes in the systems to track the change results. This tracking creates a feedback loop, necessary to manage the transformation.

Another challenge in the big data transformation message is finding the right people. Ideally the team leading the transformation needs to understand an organization's data, enriched with outside data, then know how to do data analysis, and once the results are there, strategically communicate the change to get everybody on board. Next, the transformation team needs to set up a tracking and feedback process that holds participants accountable for the transformation results. And when participants do not play along, have an escalation process in place, with the possibility for punitive measures.

In the same way that Giuliani fired one of the precinct commanders when he showed up drunk at the first CompStat meeting, big data systems require a complementary management philosophy to ensure whatever transformational insights are derived get implemented and controlled.

So, when the advertisements claim that big data will transform your business, remember that big data brings the potential for transformation, not the actual transformation. That still requires commitment and hard work, just like ‘living long and happily ever after'. That's why they are called fairy tales.



Thursday, August 23, 2012

A BI Architectures Approach to Modeling and Evolving with Analytic Databases


Major shifts converging in today's BI environment bring the opportunity to discover new answers to old questions about what BI architectures are about and how they are designed.



August 21, 2012

By John O'Brien, Principal, Radiant Advisors



Whether you have been a BI architect managing a production data warehouse for many years or are embarking on building a new data warehouse, the new analytic technologies coming out today have never been so powerful and complex to understand. In fact, with so many analytic technologies available on the market, they are somewhat overwhelming as we struggle to make sense of what to do with them and which ones to use with our existing environments.



This is good for BI architects because it brings us back to BI architecture fundamentals, key data management principles, pattern recognition, and agile processes, along with BI capabilities that challenge classic BI architecture best practices to design what clearly makes sense to meet the demands of business today.



There are major shifts converging in today's BI environment, and these changes bring with them the opportunity to discover new answers to old questions about what a BI architecture is all about and how an architecture is designed. As I explore these questions in this article, I will focus on three main themes: BI architectures are strategic platforms that evolve to their full potential; good architectures are based on recognizing data management principles (patterns and so-called best practices are discovered later); and BI architecture design is purposeful at every stage of development and technology decisions follow this purpose.



Architecture Maturity and Information Capabilities



We have all seen the research that says BI architectures evolve into robust information platforms and value over time. BI teams have focused on increasing value through maturing their data warehouses from operational reporting to data marts to data warehouses and finally to enterprise data warehouses. This evolutionary approach is typical when balancing pressing tactical needs for information delivery with strategic development, and is found in many companies where business demands for information drive towards a data warehouse platform.



The classic data warehouse is the last thing to be built, if ever, because the emphasis remains on quicker delivery first and information consistency later. Unfortunately, this leads to data warehouses that reflect current information needs and doesn't foster the evolution of a mature analytics culture.



Instead, an architecture based on BI capabilities focuses on nurturing the analytic culture of the business community by first educating user communities about the BI capabilities available and then on business subject data that is delivered via BI capabilities. This approach centers the data warehouse architecture on BI capabilities such as information delivery; reporting and parameterized reporting; dimensional analytics for goals achievement; and advanced analytics for gaining insights, to name a few. These discussions recognize that the same consistent data has many usage patterns, behaviors, and roles in the decision process. A BI architecture that is designed in this way ensures that data models and chosen analytic technologies are best suited to their intended purpose.



However, this BI-capabilities approach is contrary to some BI architects' belief that there should be an all-in-one data warehouse platform in the enterprise.



Thursday, June 28, 2012

Ten steps to predictive success

Follow these best practices to ensure a successful foray into predictive analytics.


1. Define the business proposition. What is the business problem you are trying to solve? What is the question you're trying to answer? Think like a business leader first and an analyst or IT expert second.

2. Line up a business champion. Having the support of a key executive and a stakeholder is crucial. Whenever possible help the stakeholder to become the initiator and champion of the project.

3. Start off with a quick win. Find a well-defined business problem where analytics can bring value by showing measurable results. Start small and use simple models to build credibility.

4. Know the data you have. Do you have enough data, enough history and enough granularity in the data to feed your proposed model? Getting it into the right form is the biggest part of any first-time predictive analytics project.

5. Get professional help. A statistical background and a little training aren't enough: Creating predictive models is different from traditional descriptive analytics, and is as much an art as it is a science. Get help for that first win before striking out on your own.

6. Be sure the decision maker is prepared to act. It's not enough to have a prescribed action plan. The results may dictate actions that are counterintuitive. If the business decision makers won't act or aren't in a position do so, you re wasting your time, so get a strong commitment up front.

7. Don't get ahead of yourself. Stay within the scope of the defined project, even if success breeds pressure to expand the use of your current model. Good analytics sells itself, but overextending can result in an unreliable model that will kill credibility.

8. Communicate the results in business language. Don't discuss probabilities and variances. Do talk revenue impact and fulfillment of business objectives. Use data visualization tools to hammer home the point.

9. Test, revise, repeat. Start small, test, revise and test again. Conduct A/B testing to demonstrate value. Present the results, gain critical mass, then scale out.

10. Hire me to implemment above steps with success :)

Thursday, May 10, 2012

Predictive Analytics and Data Mining




Derive useful insights to make evidence-based decisions

Today's organizations accumulate huge volumes of data from a variety of sources on a daily basis. However, turning increasingly large amounts of data into useful insights and finding how to better utilize those insights in decision making remains a challenge for most.

To get answers to complex questions and gain an edge in today's marketplace requires powerful, multipurpose predictive analytic solutions so you can learn from, utilize and improve on knowledge gained from vast stores of data. BIandIT Ltd provides a wide range of software for exploring and analyzing data to help uncover unknown patterns, opportunities and insights that can drive proactive, evidence-based decision making within your organization.

Text mining applies the same analysis techniques to text-based documents. The knowledge gleaned from data and text mining can be used to fuel strategic decision making.



Components of Predictive Analytics and Data Mining

Exploratory Data Analysis – Get dynamic visualization, advanced statistical techniques and core data mining capabilities to quickly identify relationships and opportunities.

Model Development and Deployment – Streamline the data mining process to create highly accurate descriptive and predictive analytic models based on large volumes of data.

Analytics Acceleration – Generate faster results and improve data governance with in-database analytics.

Scoring Acceleration – Maximize the performance and accuracy of your analytic models.




Tuesday, May 8, 2012

Microsoft Predictive Analytics


Predictive analytics is the next step in BI: not only can you be retrospective and see what has happened in your company in the past, but now we can distill new information from the old information to actually predict what will happen in the future. Jamie MacLennan, CTO of Predixion Software, explains the difference between business intelligence and predictive analytics and shares a program that Predixion has created in Excel to review the Practice Fusion data.



Featuring Bruno Aziza

Tuesday, March 20, 2012

Analytics in Sports

I am fan of football and my favorite team is FC Barcelona. Combining sports, specially football with Analytics is just amazing.

Mike Walsh, futurist

In our VK2012 was invited Mike Walsh and he had a keynote about technology future. He instisted that the future of the World is DATA and claimed the most important profession of the future will be Data Scientist. Let the future begin and let say I am a Data Scientist.

Tuesday, April 26, 2011

Predictive Analytics

Predictive analytics encompasses a variety of techniques from statistics, data mining and game theory that analyze current and historical facts to make predictions about future events.

In business, predictive models exploit patterns found in historical and transactional data to identify risks and opportunities. Models capture relationships among many factors to allow assessment of risk or potential associated with a particular set of conditions, guiding decision making for candidate transactions.

Predictive analytics is used in actuarial science, financial services, insurance, telecommunications, retail, travel, healthcare, pharmaceuticals and other fields.

One of the most well-known applications is credit scoring, which is used throughout financial services. Scoring models process a customer’s credit history, loan application, customer data, etc., in order to rank-order individuals by their likelihood of making future credit payments on time. A well-known example would be the FICO score.

DefinitionPredictive analytics is an area of statistical analysis that deals with extracting information from data and using it to predict future trends and behavior patterns. The core of predictive analytics relies on capturing relationships between explanatory variables and the predicted variables from past occurrences, and exploiting it to predict future outcomes. It is important to note, however, that the accuracy and usability of results will depend greatly on the level of data analysis and the quality of assumptions.

Types: Generally, the term predictive analytics is used to mean predictive modeling, "scoring" data with predictive models, and forecasting. However, people are increasingly using the term to describe related analytical disciplines, such as descriptive modeling and decision modeling or optimization. These disciplines also involve rigorous data analysis, and are widely used in business for segmentation and decision making, but have different purposes and the statistical techniques underlying them vary.

Predictive models: Predictive models analyze past performance to assess how likely a customer is to exhibit a specific behavior in the future in order to improve marketing effectiveness. This category also encompasses models that seek out subtle data patterns to answer questions about customer performance, such as fraud detection models. Predictive models often perform calculations during live transactions, for example, to evaluate the risk or opportunity of a given customer or transaction, in order to guide a decision. With advancement in computing speed, individual agent modeling systems can simulate human behavior or reaction to given stimuli or scenarios. The new term for animating data specifically linked to an individual in a simulated environment is avatar analytics.

Descriptive models: Descriptive models quantify relationships in data in a way that is often used to classify customers or prospects into groups. Unlike predictive models that focus on predicting a single customer behavior (such as credit risk), descriptive models identify many different relationships between customers or products. Descriptive models do not rank-order customers by their likelihood of taking a particular action the way predictive models do. Descriptive models can be used, for example, to categorize customers by their product preferences and life stage. Descriptive modeling tools can be utilized to develop further models that can simulate large number of individualized agents and make predictions.

Decision models: Decision models describe the relationship between all the elements of a decision — the known data (including results of predictive models), the decision and the forecast results of the decision — in order to predict the results of decisions involving many variables. These models can be used in optimization, maximizing certain outcomes while minimizing others. Decision models are generally used to develop decision logic or a set of business rules that will produce the desired action for every customer or circumstance.

Applications: Although predictive analytics can be put to use in many applications, we outline a few examples where predictive analytics has shown positive impact in recent years.

Analytical customer relationship management (CRM): Analytical Customer Relationship Management is a frequent commercial application of Predictive Analysis. Methods of predictive analysis are applied to customer data to pursue CRM objectives.

Clinical decision support systems: Experts use predictive analysis in health care primarily to determine which patients are at risk of developing certain conditions, like diabetes, asthma, heart disease and other lifetime illnesses. Additionally, sophisticated clinical decision support systems incorporate predictive analytics to support medical decision making at the point of care. A working definition has been proposed by Dr. Robert Hayward of the Centre for Health Evidence: "Clinical Decision Support systems link health observations with health knowledge to influence health choices by clinicians for improved health care."

Collection analytics: Every portfolio has a set of delinquent customers who do not make their payments on time. The financial institution has to undertake collection activities on these customers to recover the amounts due. A lot of collection resources are wasted on customers who are difficult or impossible to recover. Predictive analytics can help optimize the allocation of collection resources by identifying the most effective collection agencies, contact strategies, legal actions and other strategies to each customer, thus significantly increasing recovery at the same time reducing collection costs.

Cross-sell: Often corporate organizations collect and maintain abundant data (e.g. customer records, sale transactions) and exploiting hidden relationships in the data can provide a competitive advantage to the organization. For an organization that offers multiple products, an analysis of existing customer behavior can lead to efficient cross sell of products. This directly leads to higher profitability per customer and strengthening of the customer relationship. Predictive analytics can help analyze customers’ spending, usage and other behavior, and help cross-sell the right product at the right time.

Customer retention: With the number of competing services available, businesses need to focus efforts on maintaining continuous consumer satisfaction. In such a competitive scenario, consumer loyalty needs to be rewarded and customer attrition needs to be minimized. Businesses tend to respond to customer attrition on a reactive basis, acting only after the customer has initiated the process to terminate service. At this stage, the chance of changing the customer’s decision is almost impossible. Proper application of predictive analytics can lead to a more proactive retention strategy. By a frequent examination of a customer’s past service usage, service performance, spending and other behavior patterns, predictive models can determine the likelihood of a customer wanting to terminate service sometime in the near future. An intervention with lucrative offers can increase the chance of retaining the customer. Silent attrition is the behavior of a customer to slowly but steadily reduce usage and is another problem faced by many companies. Predictive analytics can also predict this behavior accurately and before it occurs, so that the company can take proper actions to increase customer activity.

Direct marketing: When marketing consumer products and services there is the challenge of keeping up with competing products and consumer behavior. Apart from identifying prospects, predictive analytics can also help to identify the most effective combination of product versions, marketing material, communication channels and timing that should be used to target a given consumer. The goal of predictive analytics is typically to lower the cost per order or cost per action.

Fraud detection: Fraud is a big problem for many businesses and can be of various types. Inaccurate credit applications, fraudulent transactions (both offline and online), identity thefts and false insurance claims are some examples of this problem. These problems plague firms all across the spectrum and some examples of likely victims are credit card issuers, insurance companies, retail merchants, manufacturers, business to business suppliers and even services providers. This is an area where a predictive model is often used to help weed out the “bads” and reduce a business's exposure to fraud.

Predictive modeling can also be used to detect financial statement fraud in companies, allowing auditors to gauge a company's relative risk, and to increase substantive audit procedures as needed.

The Internal Revenue Service (IRS) of the United States also uses predictive analytics to try to locate tax fraud.

Portfolio, product or economy level prediction: Often the focus of analysis is not the consumer but the product, portfolio, firm, industry or even the economy. For example a retailer might be interested in predicting store level demand for inventory management purposes. Or the Federal Reserve Board might be interested in predicting the unemployment rate for the next year. These type of problems can be addressed by predictive analytics using Time Series techniques (see below).

Underwriting: Many businesses have to account for risk exposure due to their different services and determine the cost needed to cover the risk. For example, auto insurance providers need to accurately determine the amount of premium to charge to cover each automobile and driver. A financial company needs to assess a borrower’s potential and ability to pay before granting a loan. For a health insurance provider, predictive analytics can analyze a few years of past medical claims data, as well as lab, pharmacy and other records where available, to predict how expensive an enrollee is likely to be in the future. Predictive analytics can help underwriting of these quantities by predicting the chances of illness, default, bankruptcy, etc. Predictive analytics can streamline the process of customer acquisition, by predicting the future risk behavior of a customer using application level data. Predictive analytics in the form of credit scores have reduced the amount of time it takes for loan approvals, especially in the mortgage market where lending decisions are now made in a matter of hours rather than days or even weeks. Proper predictive analytics can lead to proper pricing decisions, which can help mitigate future risk of default.

Statistical techniques: The approaches and techniques used to conduct predictive analytics can broadly be grouped into regression techniques and machine learning techniques.

Regression techniques: Regression models are the mainstay of predictive analytics. The focus lies on establishing a mathematical equation as a model to represent the interactions between the different variables in consideration. Depending on the situation, there is a wide variety of models that can be applied while performing predictive analytics. Some of them are briefly discussed below.

Linear regression model: The linear regression model analyzes the relationship between the response or dependent variable and a set of independent or predictor variables. This relationship is expressed as an equation that predicts the response variable as a linear function of the parameters. These parameters are adjusted so that a measure of fit is optimized. Much of the effort in model fitting is focused on minimizing the size of the residual, as well as ensuring that it is randomly distributed with respect to the model predictions.

The goal of regression is to select the parameters of the model so as to minimize the sum of the squared residuals. This is referred to as ordinary least squares (OLS) estimation and results in best linear unbiased estimates (BLUE) of the parameters if and only if the Gauss-Markov assumptions are satisfied.

Once the model has been estimated we would be interested to know if the predictor variables belong in the model – i.e. is the estimate of each variable’s contribution reliable? To do this we can check the statistical significance of the model’s coefficients which can be measured using the t-statistic. This amounts to testing whether the coefficient is significantly different from zero. How well the model predicts the dependent variable based on the value of the independent variables can be assessed by using the R² statistic. It measures predictive power of the model i.e. the proportion of the total variation in the dependent variable that is “explained” (accounted for) by variation in the independent variables.

Discrete choice models: Multivariate regression (above) is generally used when the response variable is continuous and has an unbounded range. Often the response variable may not be continuous but rather discrete. While mathematically it is feasible to apply multivariate regression to discrete ordered dependent variables, some of the assumptions behind the theory of multivariate linear regression no longer hold, and there are other techniques such as discrete choice models which are better suited for this type of analysis. If the dependent variable is discrete, some of those superior methods are logistic regression, multinomial logit and probit models. Logistic regression and probit models are used when the dependent variable is binary.

Logistic regression: For more details on this topic, see logistic regression.
In a classification setting, assigning outcome probabilities to observations can be achieved through the use of a logistic model, which is basically a method which transforms information about the binary dependent variable into an unbounded continuous variable and estimates a regular multivariate model (See Allison’s Logistic Regression for more information on the theory of Logistic Regression).

The Wald and likelihood-ratio test are used to test the statistical significance of each coefficient b in the model (analogous to the t tests used in OLS regression; see above). A test assessing the goodness-of-fit of a classification model is the –.

Multinomial logistic regression: An extension of the binary logit model to cases where the dependent variable has more than 2 categories is the multinomial logit model. In such cases collapsing the data into two categories might not make good sense or may lead to loss in the richness of the data. The multinomial logit model is the appropriate technique in these cases, especially when the dependent variable categories are not ordered (for examples colors like red, blue, green). Some authors have extended multinomial regression to include feature selection/importance methods such as Random multinomial logit.

Probit regressionProbit models offer an alternative to logistic regression for modeling categorical dependent variables. Even though the outcomes tend to be similar, the underlying distributions are different. Probit models are popular in social sciences like economics.

A good way to understand the key difference between probit and logit models, is to assume that there is a latent variable z.

We do not observe z but instead observe y which takes the value 0 or 1. In the logit model we assume that y follows a logistic distribution. In the probit model we assume that y follows a standard normal distribution. Note that in social sciences (example economics), probit is often used to model situations where the observed variable y is continuous but takes values between 0 and 1.

Logit versus probit: The Probit model has been around longer than the logit model. They look identical, except that the logistic distribution tends to be a little flat tailed. One of the reasons the logit model was formulated was that the probit model was difficult to compute because it involved calculating difficult integrals. Modern computing however has made this computation fairly simple. The coefficients obtained from the logit and probit model are also fairly close. However, the odds ratio makes the logit model easier to interpret.

For practical purposes the only reasons for choosing the probit model over the logistic model would be:

There is a strong belief that the underlying distribution is normal
The actual event is not a binary outcome (e.g. Bankrupt/not bankrupt) but a proportion (e.g. Proportion of population at different debt levels).
Time series models: Time series models are used for predicting or forecasting the future behavior of variables. These models account for the fact that data points taken over time may have an internal structure (such as autocorrelation, trend or seasonal variation) that should be accounted for. As a result standard regression techniques cannot be applied to time series data and methodology has been developed to decompose the trend, seasonal and cyclical component of the series. Modeling the dynamic path of a variable can improve forecasts since the predictable component of the series can be projected into the future.

Time series models estimate difference equations containing stochastic components. Two commonly used forms of these models are autoregressive models (AR) and moving average (MA) models. The Box-Jenkins methodology (1976) developed by George Box and G.M. Jenkins combines the AR and MA models to produce the ARMA (autoregressive moving average) model which is the cornerstone of stationary time series analysis. ARIMA (autoregressive integrated moving average models) on the other hand are used to describe non-stationary time series. Box and Jenkins suggest differencing a non stationary time series to obtain a stationary series to which an ARMA model can be applied. Non stationary time series have a pronounced trend and do not have a constant long-run mean or variance.

Box and Jenkins proposed a three stage methodology which includes: model identification, estimation and validation. The identification stage involves identifying if the series is stationary or not and the presence of seasonality by examining plots of the series, autocorrelation and partial autocorrelation functions. In the estimation stage, models are estimated using non-linear time series or maximum likelihood estimation procedures. Finally the validation stage involves diagnostic checking such as plotting the residuals to detect outliers and evidence of model fit.

In recent years time series models have become more sophisticated and attempt to model conditional heteroskedasticity with models such as ARCH (autoregressive conditional heteroskedasticity) and GARCH (generalized autoregressive conditional heteroskedasticity) models frequently used for financial time series. In addition time series models are also used to understand inter-relationships among economic variables represented by systems of equations using VAR (vector autoregression) and structural VAR models.

Survival or duration analysis: Survival analysis is another name for time to event analysis. These techniques were primarily developed in the medical and biological sciences, but they are also widely used in the social sciences like economics, as well as in engineering (reliability and failure time analysis).

Censoring and non-normality, which are characteristic of survival data, generate difficulty when trying to analyze the data using conventional statistical models such as multiple linear regression. The normal distribution, being a symmetric distribution, takes positive as well as negative values, but duration by its very nature cannot be negative and therefore normality cannot be assumed when dealing with duration/survival data. Hence the normality assumption of regression models is violated.

The assumption is that if the data were not censored it would be representative of the population of interest. In survival analysis, censored observations arise whenever the dependent variable of interest represents the time to a terminal event, and the duration of the study is limited in time.

An important concept in survival analysis is the hazard rate, defined as the probability that the event will occur at time t conditional on surviving until time t. Another concept related to the hazard rate is the survival function which can be defined as the probability of surviving to time t.

Most models try to model the hazard rate by choosing the underlying distribution depending on the shape of the hazard function. A distribution whose hazard function slopes upward is said to have positive duration dependence, a decreasing hazard shows negative duration dependence whereas constant hazard is a process with no memory usually characterized by the exponential distribution. Some of the distributional choices in survival models are: F, gamma, Weibull, log normal, inverse normal, exponential etc. All these distributions are for a non-negative random variable.

Duration models can be parametric, non-parametric or semi-parametric. Some of the models commonly used are Kaplan-Meier and Cox proportional hazard model (non parametric).

Classification and regression trees: Main article: decision tree learning
Classification and regression trees (CART) is a non-parametric decision tree learning technique that produces either classification or regression trees, depending on whether the dependent variable is categorical or numeric, respectively.

Decision trees are formed by a collection of rules based on variables in the modeling data set:

Rules based on variables’ values are selected to get the best split to differentiate observations based on the dependent variable
Once a rule is selected and splits a node into two, the same process is applied to each “child” node (i.e. it is a recursive procedure)
Splitting stops when CART detects no further gain can be made, or some pre-set stopping rules are met. (Alternatively, the data is split as much as possible and then the tree is later pruned.)
Each branch of the tree ends in a terminal node. Each observation falls into one and exactly one terminal node, and each terminal node is uniquely defined by a set of rules.

A very popular method for predictive analytics is Leo Breiman's Random forests or derived versions of this technique like Random multinomial logit.

Multivariate adaptive regression splines: Multivariate adaptive regression splines (MARS) is a non-parametric technique that builds flexible models by fitting piecewise linear regressions.

An important concept associated with regression splines is that of a knot. Knot is where one local regression model gives way to another and thus is the point of intersection between two splines.

In multivariate and adaptive regression splines, basis functions are the tool used for generalizing the search for knots. Basis functions are a set of functions used to represent the information contained in one or more variables. Multivariate and Adaptive Regression Splines model almost always creates the basis functions in pairs.

Multivariate and adaptive regression spline approach deliberately overfits the model and then prunes to get to the optimal model. The algorithm is computationally very intensive and in practice we are required to specify an upper limit on the number of basis functions.

Machine learning techniques. Machine learning, a branch of artificial intelligence, was originally employed to develop techniques to enable computers to learn. Today, since it includes a number of advanced statistical methods for regression and classification, it finds application in a wide variety of fields including medical diagnostics, credit card fraud detection, face and speech recognition and analysis of the stock market. In certain applications it is sufficient to directly predict the dependent variable without focusing on the underlying relationships between variables. In other cases, the underlying relationships can be very complex and the mathematical form of the dependencies unknown. For such cases, machine learning techniques emulate human cognition and learn from training examples to predict future events.

A brief discussion of some of these methods used commonly for predictive analytics is provided below. A detailed study of machine learning can be found in Mitchell (1997).

Neural networks. Neural networks are nonlinear sophisticated modeling techniques that are able to model complex functions. They can be applied to problems of prediction, classification or control in a wide spectrum of fields such as finance, cognitive psychology/neuroscience, medicine, engineering, and physics.

Neural networks are used when the exact nature of the relationship between inputs and output is not known. A key feature of neural networks is that they learn the relationship between inputs and output through training. There are two types of training in neural networks used by different networks, supervised and unsupervised training, with supervised being the most common one.

Some examples of neural network training techniques are backpropagation, quick propagation, conjugate gradient descent, projection operator, Delta-Bar-Delta etc. Some unsupervised network architectures are multilayer perceptrons, Kohonen networks, Hopfield networks, etc.

Radial basis functions: A radial basis function (RBF) is a function which has built into it a distance criterion with respect to a center. Such functions can be used very efficiently for interpolation and for smoothing of data. Radial basis functions have been applied in the area of neural networks where they are used as a replacement for the sigmoidal transfer function. Such networks have 3 layers, the input layer, the hidden layer with the RBF non-linearity and a linear output layer. The most popular choice for the non-linearity is the Gaussian. RBF networks have the advantage of not being locked into local minima as do the feed-forward networks such as the multilayer perceptron.

Support vector machines: Support Vector Machines (SVM) are used to detect and exploit complex patterns in data by clustering, classifying and ranking the data. They are learning machines that are used to perform binary classifications and regression estimations. They commonly use kernel based methods to apply linear classification techniques to non-linear classification problems. There are a number of types of SVM such as linear, polynomial, sigmoid etc.

Naïve Bayes: Naïve Bayes based on Bayes conditional probability rule is used for performing classification tasks. Naïve Bayes assumes the predictors are statistically independent which makes it an effective classification tool that is easy to interpret. It is best employed when faced with the problem of ‘curse of dimensionality’ i.e. when the number of predictors is very high.

K-nearest neighbours: The nearest neighbour algorithm (KNN) belongs to the class of pattern recognition statistical methods. The method does not impose a priori any assumptions about the distribution from which the modeling sample is drawn. It involves a training set with both positive and negative values. A new sample is classified by calculating the distance to the nearest neighbouring training case. The sign of that point will determine the classification of the sample. In the k-nearest neighbour classifier, the k nearest points are considered and the sign of the majority is used to classify the sample. The performance of the kNN algorithm is influenced by three main factors: (1) the distance measure used to locate the nearest neighbours; (2) the decision rule used to derive a classification from the k-nearest neighbours; and (3) the number of neighbours used to classify the new sample. It can be proved that, unlike other methods, this method is universally asymptotically convergent, i.e.: as the size of the training set increases, if the observations are independent and identically distributed (i.i.d.), regardless of the distribution from which the sample is drawn, the predicted class will converge to the class assignment that minimizes misclassification error.
Geospatial predictive modeling: Conceptually, geospatial predictive modeling is rooted in the principle that the occurrences of events being modeled are limited in distribution. Occurrences of events are neither uniform nor random in distribution – there are spatial environment factors (infrastructure, sociocultural, topographic, etc.) that constrain and influence where the locations of events occur. Geospatial predictive modeling attempts to describe those constraints and influences by spatially correlating occurrences of historical geospatial locations with environmental factors that represent those constraints and influences. Geospatial predictive modeling is a process for analyzing events through a geographic filter in order to make statements of likelihood for event occurrence or emergence.

Tools: There are numerous tools available in the marketplace which help with the execution of predictive analytics. These range from those which need very little user sophistication to those that are designed for the expert practitioner. The difference between these tools is often in the level of customization and heavy data lifting allowed.

In an attempt to provide a standard language for expressing predictive models, the Predictive Model Markup Language (PMML) has been proposed. Such an XML-based language provides a way for the different tools to define predictive models and to share these between PMML compliant applications. PMML 4.0 was released in June, 2009.

References:
L. Devroye, L. Györfi, G. Lugosi (1996). A Probabilistic Theory of Pattern Recognition. New York: Springer-Verlag.
John R. Davies, Stephen V. Coggeshall, Roger D. Jones, and Daniel Schutzer, "Intelligent Security Systems," in Freedman, Roy S., Flein, Robert A., and Lederman, Jess, Editors (1995). Artificial Intelligence in the Capital Markets. Chicago: Irwin. ISBN 1-55738-811-3.
Agresti, Alan (2002). Categorical Data Analysis. Hoboken: John Wiley and Sons. ISBN 0-471-36093-7.
Enders, Walter (2004). Applied Time Series Econometrics. Hoboken: John Wiley and Sons. ISBN 052183919X.
Greene, William (2000). Econometric Analysis. Prentice Hall. ISBN 0-13-013297-7.
Mitchell, Tom (1997). Machine Learning. New York: McGraw-Hill. ISBN 0-07-042807-7.
Tukey, John (1977). Exploratory Data Analysis. New York: Addison-Wesley. ISBN 0201076160.
Guidère, Mathieu; Howard N, Sh. Argamon (2009). Rich Language Analysis for Counterterrrorism. Berlin, London, New York: Springer-Verlag. ISBN 978-3-642-01140-5.

Tuesday, June 9, 2009

Business Intelligence

This Blog will contain information about Business Intelligence, Data Mining, Data Modeling and Data Science including tutorials, white papers, important updates and business cases mostly based in Microsoft platform.
Our intention is open a discussion blog where experts can talk generally about Business Intelligence or can exchange views for particular problems that they experienced. We are going to talk about different BI platforms their advantages and disadvantages, against Microsoft Platform.
Analytics will be the main topic, SSAS will be the most discussed tool and SQL/MDX/DMX will be the most used scripts to explain many of the problems that BI Professionals face every day.
MDX and DMX will be part of this blog too. Advanced calculations that we can handle with MDX and problems for improving time in reporting large data warehouse calculations over dimensions.
Dimensional databases vs relational, OLAP Cubes, algorithms for time improvement will rich our Blog.
You will be updated with podcast, white papers, analysis and links that are important to our auditorium.

Best regards,
Besim Ismaili
Creator of the Blog