Skip to main content
Have a personal or library account? Click to login
Data Cleaning and Exploration with Machine Learning Cover

Data Cleaning and Exploration with Machine Learning

A practical guide to machine learning and data exploration with Python and Scikit-learn (English Edition)

Paid access
|Dec 2025
Product purchase options

Machine learning has become central to how organizations handle data in today’s world. With businesses generating vast amounts of information, the ability to clean, explore, and model data effectively is no longer optional, it is a critical skill for decision-making, innovation, and competitive advantage.

This book takes readers on a structured journey, starting with Python foundations and essential libraries. It discusses data cleaning, preprocessing, and exploratory analysis, and then explores text and time series data, dimensionality reduction, regression, classification, and clustering techniques. Advanced topics such as model evaluation, neural networks, deep learning, retrieval-augmented generation, and explainable AI are covered in detail, which are supported by real-world examples and case studies. Each chapter builds progressively, ensuring both theoretical grounding and practical application, and vital industry practices.

By the end of the book, readers will be equipped with the skills to handle raw datasets, uncover patterns, build and evaluate ML models, and apply advanced techniques responsibly. You will be confident in applying these methods to solve problems in their domains, making yourself a competent data practitioner, ready to deliver insights and drive impact.

WHAT YOU WILL LEARN

  • Understand Python foundations and essential data science libraries.
  • Apply data cleaning methods to handle missing or noisy data.
  • Perform exploratory data analysis using statistics and visualization.
  • Work with text, time-series, and high-dimensional datasets.
  • Build regression, classification, and clustering ML models.
  • Evaluate models with metrics, validation, and hyperparameter tuning.
  • Explore neural networks, deep learning, and explainable AI techniques.
  • Implement real-world case studies and capstone data projects.

WHO THIS BOOK IS FOR

This book is for data analysts, data scientists, ML engineers, and business professionals who want to strengthen their skills in data preparation and modeling. It is also valuable for students, researchers, and software developers aiming to apply ML techniques effectively in real-world projects.

TABLE OF CONTENTS

  1. Introduction to Data Science and Machine Learning
  2. Setting Up Your Development Environment
  3. Introduction to Integrated Development Environments
  4. Exploring Essential Python Libraries
  5. Introduction to Data Cleaning
  6. Exploratory Data Analysis Made Easy
  7. Demystifying Data Preprocessing from Raw to Refined
  8. Unraveling Insights from Text and Time Series Data
  9. Dimensionality Reduction Techniques
  10. Building Regression Models for Confident Predictions
  11. Supervised Learning for Developing Classification Models
  12. Discovering Hidden Patterns with Clustering Techniques
  13. Ensuring Model Reliability Through Evaluation
  14. Techniques and Applications of RAG Pipelines
  15. Fine-tuning and Evaluating Base LLMs
  16. Putting It All Together with Case Studies
  17. Best Practices and Tips from Industry Experts
  18. Conclusion and Further Resources
PDF ISBN: 978-93-6589-219-2 | E-Pub ISBN: 978-93-6589-110-2
Publisher: BPB Publications
Copyright owner: © 2026 BPB Publications
Publication date: 2025
Language: English
Pages: 432