Google Track

Showing posts with label Gartner. Show all posts
Showing posts with label Gartner. Show all posts

Friday, November 14, 2014

22 tips for better data science


These tips are provided by Dr Granville, who brings 20 years of varied data-intensive experience working with successful start-ups, small companies across various industries, and eBay, Visa, Microsoft, GE and Wells Fargo.
1.     Leverage external data sources: tweets about your company or your competitors, or data from your vendors (for instance, customizable newsletter eBlast statistics available via vendor dashboards, or via submitting a ticket)
2.     Nuclear physicists, mechanical engineers, and bioinformatics experts can make great data scientists.
3.     State your problem correctly, and use sound metrics to measure yield provided by data science initiatives.
4.     Use the right KPIs (key metrics) and the right data from the beginning, in any project. Changes due to bad foundations are very costly. This requires careful analysis of your data to create useful databases.
5.     Fast delivery is better than extreme accuracy. All data sets are dirty anyway. Find the perfect compromise between perfection and fast return.
6.     With big data, strong signals (extremes) will usually be noise. Here's a solution.
7.     Big data has less value than useful data.
8.     Use big data from third party vendors, for competitive intelligence.
9.     You can build cheap, great, scalable, robust tools pretty fast, without using old-fashioned statistical science. Think about model-free techniques.
10.  Big data is easier and less costly than you think. Get the right tools! Here's how to get started.
11.  Correlation is not causation. This article might help you with this issue.
12.  You don't have to store all your data permanently. Use smart compression techniques, and keep statistical summaries only, for old data. Don't forget to adjust your metrics when your data changes,to keep consistency for trending purposes.
13.  A lot can be done without databases, especially for big data.
14.  Always include EDA and DOE (exploratory analysis / design of experiment) early on in any data science projects. Always create a data dictionary. And follow the traditional life cycle of any data science project.
15.  Data can be used for many purposes:
·        quality assurance
·        to find actionable patterns (stock trading, fraud detection)
·        for resale to your business clients
·        to optimize decisions and processes (operations research)
·        for investigation and discovery (IRS, litigation, fraud detection, root cause analysis)
·        machine-to-machine communication (automated bidding systems, automated driving)
·        predictions (sales forecasts, growth and financial predictions, weather)
16.  Don't dump Excel. Embrace light analytics.
17.  Data + models + gut feelings + intuition is the perfect mix. Don't remove any of these ingredients in your decision process.
18.  Leverage the power of compound metrics: KPIs derived from database fields, that have a far betterpredictive power than the original database metrics. For instance your database might include a single keyword field but does not discriminate between user query and search category (sometimes because data comes from various sources and is blended together). Detect the issue, and create a new metric called keyword type - or data source. Another example is IP address category, a fundamental metric that should be created and added to all digital analytics projects.
19.  When do you need true real time processing? When fraud detection is critical, or when processing sensitive transactional data (credit card fraud detection, 911 calls). Other than that, delayed analytics(with a latency of a few seconds to 24 hours) is good enough.
20.  Make sure your sensitive data is well protected. Make sure your algorithms can not be tampered by criminal hackers or business hackers (spying on your business and stealing everything they can, legally or illegally, and jeopardizing your algorithms - which translates in severe revenue loss). An example of business hacking can be found in section 3 in this article.
21.  Blend multiple models together to detect many types of patterns. Average these models. Here's a simple example of model blending.
22.  Ask the right questions before purchasing software.




Wednesday, May 22, 2013

Cool BI: Emerging Trends and Innovations in BI



Cindi2



Cindi Howson, founder of BI Scorecard

It’s hard to be innovative when your BI team is deluged with fixes, fighting fires, and basic data requests.  Yet to move from reactive, report-focused development to break through BI demands innovative BI teams and technologies.
In this keynote, Cindi Howson, founder of BI Scorecard and author ofSuccessful Business Intelligence: Secrets to Making BI a Killer App, highlights:
  • Being proactive when there’s no time or budget for innovation
  • Evangelizing BI in a culture resistant to change
  • Prioritizing innovations that will provide the biggest value
  • The trends most disruptive to BI including mobile, social, visual data discovery, big data, and cloud.
Check out this ketnote presentation at the Gurus Of BI (GOBI) conference on June 1oth: www.gurusofbi.no

Thursday, April 5, 2012

Big Data, the amazing thing

Big data is a term applied to data sets whose size is beyond the ability of commonly used software tools to capture, manage, and process the data within a tolerable elapsed time. Big data sizes are a constantly moving target currently ranging from a few dozen terabytes to many petabytes of data in a single data set.

In a 2001 research report[15] and related conference presentations, then META Group (now Gartner) analyst, Doug Laney, defined data growth challenges (and opportunities) as being three-dimensional, i.e. increasing volume (amount of data), velocity (speed of data in/out), and variety (range of data types, sources). Gartner continues to use this model for describing big data.






Whether through blogs, twitter, or technical articles, you’ve probably heard about Big Data, and a recognition that organizations need to look beyond the traditional databases to achieve the most cost effective storage and processing of extremely large data sets, unstructured data, and/or data that comes in too fast. As the prevalence and importance of such data increases, many organizations are looking at how to leverage technologies such as those in the Apache Hadoop ecosystem. Recognizing one size doesn’t fit all, we began detailing our approach to Big Data at the PASS Summit last October. Microsoft’s goal for Big Data is to provide insights to all users from structured or unstructured data of any size. While very scalable, accommodating, and powerful, most Big Data solutions based on Hadoop require highly trained staff to deploy and manage. In addition, the benefits are limited to few highly technical users who are as comfortable programming their requirements as they are using advanced statistical techniques to extract value. For those of us who have been around the BI industry for a few years, this may sound similar to the early 90s where the benefits of our field were limited to a few within the corporation through the Executive Information Systems.

Analysis on Hadoop for Everyone

Microsoft entered the Business Intelligence industry to enable orders of magnitude more users to make better decisions from applications they use every day. This was the motivation behind being the first DBMS vendor to include an OLAP engine with the release of SQL Server 7.0 OLAP Services that enabled Excel users to ask business questions at the speed of thought. It remained the motivation behind PowerPivot in SQL Server 2008 R2, a self-service BI offering that allowed end users to build their own solutions without dependence on IT, as well as provided IT insights on how data was being consumed within the organization. And, with the release of Power View in SQL Server 2012, that goal will bring the power of rich interactive exploration directly in the hands of every user within an organization.
Enabling end users to merge data stored in a Hadoop deployment with data from other systems or with their own personal data is a natural next step. In fact, we also introduced Hive ODBC driver, currently in Community Technology Preview, at the PASS Summit in October. This driver allows connectivity to Apache Hive, which in turn facilitates querying and managing large datasets residing in distributed storage by exposing them as a data warehouse.

Saturday, March 24, 2012

Industry News

Cloud computing changing future of BI

2012-03-23
With the onset of cloud technology, many different sectors of the business and personal world are rapidly changing. From the focus on personal computers to mobile tablets, and from legacy systems to IaaS and SaaS systems, it goes without saying that in the coming years consumers can expect a technological revolution.

All these emerging systems are impacting business intelligence as well. A recent IDC study predicts that the market for big data technology and business intelligence software will grow from $3.2 billion in 2010 to $16.9 billion in 2015.

Furthermore, Gartner predicts that companies will nearly be forced into using these new technologies or risk losing a competitive edge. According to a new Gartner study, 85 percent of Fortune 500 corporations will fail to effectively use big data to get an advantage.

According to Jeff Kaplan, managing director of THINKstrategies, businesses must adopt some semblance of business intelligence and analytics software in order to maintain an edge that may have previously been established with legacy systems.

"Without all these cloud-based resources and tools, most organizations would be unable to cope with today's explosive growth of data," he said.