Questions: 1.What is data mining? Explain its purpose in simple words. 2.List and explain the main steps in the data mining process. 3.Name and explain any three common techniques used in data mining. 4.Write any four advantages of data mining. 5.What are the main challenges of data mining? 6. Difference between Data Mining and Data Science 7. Definition of data mining, example of data mining, Importance of data mining, Applications of data mining, Steps in Data Mining, Common Data Mining Techniques, Advantages and Challenges of Data Mining
Data
Mining
•
Definition:
Data mining is the process of discovering useful patterns, relationships and
insights from large amount of data. It means digging into data to find hidden
knowledge that can help make better decisions.
•
For example ‒ Amazon collects huge
amounts of customer data (clicks, purchases, ratings). Using data mining it
finds what products are often bought together.
•
In short data mining is nothing but the knowledge discovery from data.
1)
It helps in business decision making.
2)
It find the hidden patterns which are not visible in simple reports. For
example ‒ identifying seasonal trends in product sales.
3)
It automates the process of discovering useful information.
4)
Data mining also helps in forecasting future trends.
1)
It is used in disease prediction and patient data analysis.
2)
It is used in market basket analysis and recommendation systems
3)
In banking, fraud detection, loan risk analysis.
4)
For marketing, for Customer segmentation, campaign targeting the data mining is
used.
5)
In manufacturing sector, for detecting machine faults and quality improvement
the data mining can be used.
Step 1: Data collection
•
Gather data from multiple sources such
as databases, sensors, web or files.
Step 2: Data cleaning and
preparation
•
Remove missing, duplicate or incorrect data. Convert it to suitable format so
that analysis can be done on it. For example ‒ Correct spelling errors, fill
missing age values.
Step 3: Data integration
•
Combine data from different sources into
one dataset. For example we can merger customer data with purchase_info data.
Step 4: Data selection
•
Choose the relevant variables or
features needed for mining.
Step 5: Data transformation
•
Convert data into a suitable format for
modelling.
Step 6: Data mining
•
Apply algorithms or statistical
techniques to discover patterns or relationships. For example ‒ clustering
customers into groups based on their buying habits.
Step 7: Pattern evaluation
•
Check whether the patterns found are
useful, valid and interesting. For example ‒ Customer who buy laptop also
purchase anti‒virus software. This is useful for marketing.
Step 8: Knowledge presentation
•
Present results using graphs, dashboards or reports. Make use of the tools like
powerpoint presentation, Excel or PowerBI.
•
Following are some commonly used data
mining techniques ‒
1)
Association
•
Association analysis is a method used to
find items or events that often happen together in a dataset. It helps us
discover relationships such as "when this happens, that usually happens
too." For example ‒ In a supermarket when a person buys bread he/she
usually buys butter. This is called as association.
2)
Classification
•
Classification is used to categorize
data into predefined groups. For example ‒ classifying if the emails are spam
or not spam.
3)
Clustering
•
In the clustering algorithm, the similar
items are grouped together without predefined labels. For example ‒ Grouping
the customer as registered or unregistered customers.
4)
Regression
•
This algorithm is used to predict
continuous values. For example ‒ predicting house prices based in area and
location.
5)
Artificial Neural Network (ANN)
• An Artificial
Neural Network (ANN) is a computer model that works a bit like the human brain. It has many small units
called artificial neurons, which are
connected to each other. The network learns
from data by changing the strength of these connections, just like our
brain learns from experience. ANNs are often used in image recognition, speech detection, and handwriting recognition,
where the patterns are very complex.
6)
Outlier analysis
•
It identifies unusual or unexpected data
points. For example ‒ detecting fraudulent transactions in banking data.
Advantages
1)
It helps to discover hidden trends and relationships.
2)
Use of data mining increase the business profitability.
3)
It enables personalized services to customers.
4)
It improves the decision making and forecasting.
5)
It support for automation in data driven tasks.
Challenges
1) Large volume of data: Handling
and processing huge volume of data can be time consuming.
2) Privacy concern:
In data mining sensitive data must be handled carefully.
3) Complexity of models: Some algorithms
are hard to interpret to non‒technical users.
4) Data quality issues:
Incomplete or incorrect data affects results.
•
Data mining focuses on discovering patterns from data. Data science goes beyond
that ‒ it collects, cleans, analyzes, builds models and deploys them to solve
real‒world problems.
•
Following table shows the difference
between them.

Data
mining
1.It
is the process of finding hidden
patterns and relationships from a large datasets.
2.Data
mining is used to discover usefulfor information
from data.
3.The
scope of data mining is narrow. It
only focuses on extracting knowledge from data.
4.It
is part of data science.
5.Techniques
used for data mining clustering, classification and association rules.
6.
It works mostly with structured data
like spreadsheets, databases.
7.The outcome of any data mining process is
patterns, trends, or rules found in the huge data.
8. Example ‒ A retail company finds that
customers who buy bread often buy butter.
1.
It is broader field that uses data to
analyze, build models and solve real‒ world problems.
2.
In data science process, the data is used predictions,
automation and for decision making.
3.
The scope of data science is wider.
It includes collection of raw data, cleaning, data mining, modeling,
visualization and deployment.
4.
It is a broader field that includes data
mining.
5.
It uses machine learning, data mining, visualization and big data tools.
6.It
works with both structured and
unstructured data like text, image, audio.
7.The outcome of data science is predictive
models, dashboards, reports and real‒time
applications.
8. Example Using this pattern the company
recommendation system to suggest butter when the customer buys bread.
1.What is data mining?
Explain its purpose in simple words.
2.List and explain the
main steps in the data mining process.
3.Name and explain any
three common techniques used in data mining.
4.Write any four
advantages of data mining.
5.What are the main
challenges of data mining?
Python for Data Science: Chapter 3: Foundations of Data Science : Tag: Computer Programming, Python, Data Science : Definition, example, Importance, Applications, Steps, Common Techniques, Advantages and Challenges - Data Mining
Python for Data Science
AD25201 2nd Semester AIDS Dept | 2025 Regulation | 2nd Semester 2025 Regulation
English Essentials II
EN25C02 2nd Semester | 2025 Regulation | 2nd Semester 2025 Regulation
Tamils and Technology தமிழர்களும் தொழில்நுட்பமும்
UC25H02 2nd Semester | 2025 Regulation | 2nd Semester 2025 Regulation
Linear Algebra
MA25C02 2nd Semester | 2025 Regulation
Applied Physics (CSIE) II
PH25C03 2nd Semester AIDS, CSE, IT, CSE(CY) Dept | 2025 Regulation | 2nd Semester 2025 Regulation
Digital Principles and Computer Organization
CS25C06 2nd Semester AIDS, CSE, IT, CSE(CY) Dept | 2025 Regulation | 2nd Semester 2025 Regulation
Basic Electrical and Electronics Engineering
EE25C01 2nd Semester | 2025 Regulation | 2nd Semester 2025 Regulation
Python for Data Science
AD25201 2nd Semester AIDS Dept | 2025 Regulation | 2nd Semester 2025 Regulation
Re-Engineering for Innovation
ME25C05 2nd Semester | 2025 Regulation | 2nd Semester 2025 Regulation
Python for Data Science - Laboratory
AD25201 2nd Semester AIDS Dept | 2025 Regulation | 2nd Semester 2025 Regulation