Data Pre-processing in Data Mining

What is the meaning of Data Pre- processing?

Data Pre- processing is a very important or crucial phase in Data Mining. However, it is often neglected which should never be done.

The process of Data Pre- processing can be defined as a technique in which the raw data or the low- level data is from a set of data is transformed into an easy to understand and comprehensible form of data. It is a very beneficial step in Data Mining.

The Raw data is usually found incomplete and incompatible, due to which there are some increased chances of error and misinterpretation.

This is when the Data Pre- processing comes into play. It is a very efficient and proven method of resolving these kinds of errors.

Various Steps involved in Data Pre- processing:

Listed below are the steps that are involved in Data Pre- processing:

Data Pre-processing in Data Mining

Data Cleaning:

The first and foremost step is involved in the data pre- processing is data cleaning. As it is evident from the above section the data that is given is raw and it needs to have no irrelevant or missing parts. For that purpose, data cleaning is used.

There are some data cleaning routines through which the data should run. These routines are described below:

  • Missing Values: Sometimes, there is a situation where the data in a set of data is missing. For that, there are a couple of solutions that are listed below:
  • Ignore the tuples: This method is regarded as not a very effective method; this only comes to use when the tuple has several attributes is having missing values.
  • Fill the missing value: This approach is also not very effective and feasible. Moreover, it is a very time-consuming approach. The basic idea covering this approach is that a user has to fill the missing values. This can be done in three ways, one is doing manually, second is by using attribute mean and the third is by using the most probable value.
  • Noisy Data: Noise can be defined as a random error or variance that is found in the measured variable. It is usually caused when there is a faulty data collection, data entry error, and others. This can be handled in various ways that are listed below:
  • Binning Method: In this method, the sorted data is smoothed with the help of values around it. The data can be divided into segments of equal size and then the different methods are applied so as to complete a certain task. The data in a segment can be replaced by using the mean or boundary values in order to complete the given task.
  • Regression: The data is made smooth with the help of using a regression function. The regression can be linear or multiple. Linear regression means the regression that has only one independent variable. On the other hand, Multiple regression has various independent variables.
  • Clustering: This approach mainly groups the data in a cluster. In this approach, outliers are detected with the help of clustering. Here, similar values are arranged into a “group” or a “cluster”.

Data Integration:

In this, the data is combined from different sources into a coherent data store, as in data warehousing. Moreover, in this the conflicts are resolved.

Data Transformation:

In this step, the data given is transformed into understandable and appropriate form. There are some steps involved in data transformation. Those steps are given below:

  • Normalization: In this, the data values are scaled in a specific range for example, -1.0 to 1.0, 0.0 to 1.0 and so on. This process makes sure that there is no redundant data.
  • Smoothing: It is used to clean out the noise from the given data. This process has various techniques such as binning, clustering, and regression.
  • Aggregation: In the process of aggregation, summary or aggregation operations are performed on the given set of data.
  • Generalization of the data: In this process, the raw data or low-level data are replaced by higher level concepts with the help of using concept hierarchies.

Data Reduction:

As it has already been established that, data mining is a technique which helps the expert to handle the large amount of data. After working with large volume of data, analysis is harder in such cases. The basic aim is to increase the storage efficiency and subsequently reduce data storage and analysis costs.

The steps involved in the process of data reduction are the following:

  • Data Cube Aggregation:
    Aggregation operation is put on to data in order to construct of the data cube.
  • Attribute Subset Selection:
    It is important that the highly relevant attributes have to be used, and the remaining will be discarded. In order to perform attribute selection, one can use level of significance and p- value of the attribute.
  • Numerosity Reduction:
    Numerosity Reduction authorizes to store the model of data in place of the whole data, the example could be- Regression Models.
  • Dimensionality Reduction:
    It is used to reduce the size of data by encoding mechanisms. When reconstruction from compressed data is performed, the original data can be retrieved, this kind of reduction is also called lossless reduction. There are two effective methods of dimensionality reduction namely, Wavelet transforms and PCA (Principal Component Analysis).

Data Discretization:

Data Discretization is a process that is used in transforming continuous data attribute values to a certain finite set of intervals. In the process, there is a reduction in number of values of a continuous attribute. This is done by dividing the range of attribute intervals.

Data Sampling:

This is a technique which is used to select and work with the subset of the data set. It is possible because the subset has the similar properties of the original one.


Related Topics

Data Pre-processing in Data Mining

What is the meaning of Data Pre- processing? Data Pre- processing is a very important or crucial phase in Data Mining. However, it is often neglected which should never be done. The...

4 minutes read.

Data Mining Architecture

Data Mining Architecture: Data Mining can be defined as a process of extracting the data which is usable from huge sets of data. It is also known as KDD (Knowledge...

4 minutes read.

Outlier Analysis in Data Mining

What are Outliers? Outliers are an integral part of data analysis. An outlier can be defined as observation point that lies in a distance from other observations. An outlier is important as...

3 minutes read.

Data Mining Techniques

Data Mining Techniques Let us discuss each one of these techniques in detail: Classification: The technique used for obtaining important and relevant information about data and the metadata is called classification. As we...

4 minutes read.

Data Cleaning in Data Mining

What is data cleaning? Data mining is concerned with extracting valuable information from the data, which can help organizations make business decisions. But before performing data mining, we have to clean the...

6 minutes read.

Data Mining Applications

Data Mining Applications: As we already know, Data Mining is very useful and beneficial as we can dig deeper into the data so as to explore more about it and...

4 minutes read.

Data Mining Steps

What is Data Mining? Data Mining can be defined as the process that enables a user to find or discover various patterns and trends in a vast and large amount of...

3 minutes read.

Association Rule in Data Mining

Association Rule in Data Mining What is meant by Association Rule? One may understand Association Rule as if-then statements. It is generally used for finding and obtaining frequent patterns, correlation, and association...

4 minutes read.

Data Mining Process

Data Mining Process Data mining is a process that can be defined as a process of extracting or collecting the data that is usable from a large set of data. Data Mining...

4 minutes read.

KDD in Data Mining

What is Data Mining and why is it needed? Data Mining can be defined as the process of extraction of useful and relevant information from a set of raw data. Data Mining...

4 minutes read.

Data Mining Tasks

Data Mining Tasks Data Mining can be defined as the process of extracting important or relevant information from a set of raw data. In data mining, tasks can be categorized into two...

4 minutes read.

Data Mining Algorithms

What is Data Mining Algorithm? A data mining algorithm can be understood as a set of heuristics and calculations that are used for creating a model from a data. There are...

8 minutes read.

Clustering in Data Mining

What is meant by Clustering in Data Mining? Clustering in Data Mining can be defined as classifying or categorizing a group or set of different data objects as similar type of...

4 minutes read.

Difference between Data Warehouse and Data Mining

Difference between Data Warehouse and Data Mining Data Warehouse: Data Warehousing is a technique that is mainly used to collect and manage data from various different sources so as to give the...

8 minutes read.

Data Mining Functionalities

Data Mining Functionalities The Data Mining functionalities are basically used for specifying the different kind of patterns or trends that are usually seen in data mining tasks. Data mining is extensively...

4 minutes read.

Major Issues in Data Mining

Major Issues in Data Mining: Data Mining is not very simple to understand and implement. As it is already evident that Data Mining is a process which is very crucial...

3 minutes read.

Data Mining Tutorial

Data Mining Introduction Generally, Mining means to extract some valuable materials from the earth, for example, coal mining, diamond mining, etc. in terms of computer science, “Data Mining” is a process...

11 minutes read.

Data Mining Tools

Data Mining Tools and Techniques Data is priceless and it is not very easy to analyze. Data mining is the process in which the user searches and finds out different patterns among...

4 minutes read.

Cluster Analysis in Data Mining

Cluster Analysis in Data Mining What is meant by cluster analysis? Cluster analysis in data mining refers to the process of searching the group of objects that are similar to one and...

5 minutes read.

Regression in Data Mining

Regression in Data Mining Regression can be defined as a data mining technique that is generally used for the purpose of predicting a range of continuous values (which can also be...

3 minutes read.