Google Track

Showing posts with label dmx. Show all posts
Showing posts with label dmx. Show all posts

Tuesday, March 12, 2013

BIG Data

BIG Data: Is this about data at all?


Inspired by: Rafal Lukawiecki’s seminar about Business Analytics and Big Data, Microsoft Norway














Intro

Wikipedia defined BIG Data as a collection of data sets so large and complex that it becomes difficult to process using on-hand database management tools or traditional data processing applications. But this definition does not include the main reason why Big Data is actual nowadays and what was the purpose to re-invent this technology. Here is a visualization map for BIG Data and how is that large amount of data is generated:

















Fig1. Big Data visualization by WIPRO

As you can see from the figure, Big Data aims to represent a large set of data with a single point (a value or a expression) so that make sense for us.

BIG DATA: how big is World data?

Big Data is one of the most famous words around the world describing a new technology that will handle the big amount of data that is generated every day for analytic purposes. But, anyway handling Big Data is not any problem because we are witnessing everyday that hardware capacity expands as data volume expands and also hardware is getting cheaper day by day. So the definition BIG is not at all the case of Big Data, so the hardware capacity can hold whatever BIG Data can be. The average data set of the whole World is calculated to be 1.5 GB and that is an average memory stick, even though an average RAM (in-memory) capacity.
Anyway, if the case is not capacity and size then what is it?

BIG Data: what about data?

Data is important part of BIG Data, but is this meaning of the concept behind the BIG Data? The answer is NO and to be correct BIG Data is just a meaningless buzzword created only for masses. Behind this name does not exist any concept of the Big Data. If you look at the data you can’t say nothing than is big or small, has 1, 2, 3 … n sources and is rapidly/slowly expanding etc…, but that is not what BIG Data is interested to solve at all. If you think the size is the matter, then you are wrong again, so Big Data is not about BIGness at all.

BIG Data is interested to answer the users and not developers, is not an optimizing tool but it is actually an answering machine.

BIG Data: The real case?!

The real deal in BIG Data is that BIG Data tries to generate a single answer (I like to call: the single truth) from a huge input of data. If the answer is the only output of BIG Data processing, so logically BIG Data is dependent on QUESTION. So, the real deal behind BIG Data is the question itself. If you are in a dilemma whether to choose or not BIG Data technology over traditional database technologies you should look not in the size and not in the data itself, but simply in your queries that you are going to use on that set of data. So, I agree totally with Rafal when he says that the reason existence of Big Data technologies are in the answer that we want to get from BIG Data.

Conclusion
BIG Data is just another buzz word without having to deal with the contest itself, but better when we know before we use it. Even thou, we can’t change the trends for this buzzword; at least we can support and use the technology as much as we can.



Friday, March 9, 2012

BIandIT.com, a website that is going to have all important things in BI

My web site is Under Construction and is going to have the hotest topics and trends in Business Intelligence. You can find elegant solutions in Business Intelligence, special calculated members, Time Intelligence, Customer Intelligence, Competition Intelligence, Market Intelligence, Spatial Intelligence and Predictive Analytics. The site will have also Office Templates for Business Start-ups, for Market Analysis, ROI, First Year Costs etc... www.biandit.com is comming SOON :)

How I do Predictive Analytics without Data Mining?

This hot topic is comming soon. I will just describe a bit what will this topic will include. This topic is going to show how to do Predictive Analytics on your data without using Data Mining or DMX and just using MDX. How can Data Mining prediction Algorithms be "translated" from DMX to MDX. How accurate are they and what is the benefit of using MDX?

Wednesday, June 22, 2011

Predictive Analytics vs Data Mining

Technology Cycle:
Data warehousing is a mature technology, with approximately 70 percent of Forrester Research survey respondents indicating they have one in production. Data mining has endured significant consolidation of products since 2000, in spite of initial high-profile success stories, and has sought shelter in encapsulating its algorithms in the recommendation engines of marketing and campaign management software. Statistical inference has been transformed into predictive modelling. As we shall see, the emerging trend in predictive analytics has been enabled by the convergence of a variety of factors.

Technology Hierarchy:
In the technology hierarchy, data warehousing is generally considered an architecture for data management. Of course, when implemented, a data warehouse is a database providing information about (among many other things) what customers are buying or using which products or services and when and where are they doing so. Data mining is a process for knowledge discovery, primarily relying on generalizations of the "law of large numbers" and the principles of statistics applied to them. Predictive analytics emerges as an application that both builds on and delimits these two predecessor technologies, exploiting large volumes of data and forward-looking inference engines, by definition, providing predictions about diverse domains.

Methods:
The method of data warehousing is structured query language (SQL) and its various extensions. Data mining employs the "law of large numbers" and the principles of statistics and probability that address the issues around decision making in uncertainty. Predictive analytics carries forward the work of the two predecessor domains. Though not a silver bullet, better algorithms in operations research, risk minimization and parallel processing, when combined with hardware improvements and the lessons of usability testing, have resulted in successful new predictive applications emerging in the market. (Again, see Figure 1 on predictive analytics enabling technologies.) Widely diverging domains such as the behaviour of consumers, stocks and bonds, and fraud detection have been attacked with significant success by predictive analytics on a progressively incremental scale and scope. The work of the past decade in building the data warehouse and especially of its closely related techniques, particularly parallel processing, are key enabling factors. Statistical processing has been useful in data preparation, model construction and model validation. However, it is only with predictive analytics that the inference and knowledge are actually encoded into the model that, in turn, is encapsulated in a business application.

Definition
This results in the following definition of predictive analytics: Methods of directed and undirected knowledge discovery, relying on statistical algorithms, neural networks and optimization research to prescribe (recommend) and predict (future) actions based on discovering, verifying and applying patterns in data to predict the behavior of customers, products, services, market dynamics and other critical business transactions. In general, tools in predictive analytics employ methods to identify and relate independent and dependent variables - the independent variable being "responsible for" the dependent one and the way in which the variables "relate," providing a pattern and a model for the behavior of the downstream variables.

In data warehousing, the analyst asks a question of the data set with a predefined set of conditions and qualifications, and a known output structure. The traditional data cube addresses: What customers are buying or using which product or service and when and where are they doing so? Typically, the question is represented in a piece of SQL against a relational database. The business insight needed to craft the question to be answered by the data warehouse remains hidden in a black box - the analyst's head. Data mining gives us tools with which to engage in question formulation based primarily on the "law of large numbers" of classic statistics. Predictive analytics have introduced decision trees, neural networks and other pattern-matching algorithms constrained by data percolation. It is true that in doing so, technologies such as neural networks have themselves become a black box. However, neural networks and related technologies have enabled significant progress in automating, formulating and answering questions not previously envisioned. In science, such a practice is called "hypothesis formation," where the hypothesis is treated as a question to be defined, validated and refuted or confirmed by the data.

Tuesday, June 9, 2009

Business Intelligence

This Blog will contain information about Business Intelligence, Data Mining, Data Modeling and Data Science including tutorials, white papers, important updates and business cases mostly based in Microsoft platform.
Our intention is open a discussion blog where experts can talk generally about Business Intelligence or can exchange views for particular problems that they experienced. We are going to talk about different BI platforms their advantages and disadvantages, against Microsoft Platform.
Analytics will be the main topic, SSAS will be the most discussed tool and SQL/MDX/DMX will be the most used scripts to explain many of the problems that BI Professionals face every day.
MDX and DMX will be part of this blog too. Advanced calculations that we can handle with MDX and problems for improving time in reporting large data warehouse calculations over dimensions.
Dimensional databases vs relational, OLAP Cubes, algorithms for time improvement will rich our Blog.
You will be updated with podcast, white papers, analysis and links that are important to our auditorium.

Best regards,
Besim Ismaili
Creator of the Blog