
Python Feature Engineering Cookbook: A complete guide to crafting powerful features for your machine learning models 3rd Edition
Author(s): Soledad Galli (Author)
- Publisher: Packt Publishing
- Publication Date: 30 Aug. 2024
- Edition: 3rd
- Language: English
- Print length: 396 pages
- ISBN-10: B0DBQDG7SG
- ISBN-13: 9781835883587
Book Description
Leverage the power of Python to build real-world feature engineering and machine learning pipelines ready to be deployed to production
Key Features
- Craft powerful features from tabular, transactional, and time-series data
- Develop efficient and reproducible real-world feature engineering pipelines
- Optimize data transformation and save valuable time
- Purchase of the print or Kindle book includes a free PDF eBook
Book Description
Streamline data preprocessing and feature engineering in your machine learning project with this third edition of the Python Feature Engineering Cookbook to make your data preparation more efficient.
This guide addresses common challenges, such as imputing missing values and encoding categorical variables using practical solutions and open source Python libraries.
You’ll learn advanced techniques for transforming numerical variables, discretizing variables, and dealing with outliers. Each chapter offers step-by-step instructions and real-world examples, helping you understand when and how to apply various transformations for well-prepared data.
The book explores feature extraction from complex data types such as dates, times, and text. You’ll see how to create new features through mathematical operations and decision trees and use advanced tools like Featuretools and tsfresh to extract features from relational data and time series.
By the end, you’ll be ready to build reproducible feature engineering pipelines that can be easily deployed into production, optimizing data preprocessing workflows and enhancing machine learning model performance.
What you will learn
- Discover multiple methods to impute missing data effectively
- Encode categorical variables while tackling high cardinality
- Find out how to properly transform, discretize, and scale your variables
- Automate feature extraction from date and time data
- Combine variables strategically to create new and powerful features
- Extract features from transactional data and time series
- Learn methods to extract meaningful features from text data
Who this book is for
If you’re a machine learning or data science enthusiast who wants to learn more about feature engineering, data preprocessing, and how to optimize these tasks, this book is for you. If you already know the basics of feature engineering and are looking to learn more advanced methods to craft powerful features, this book will help you. You should have basic knowledge of Python programming and machine learning to get started.
Table of Contents
- Imputing Missing Data
- Encoding Categorical Variables
- Transforming Numerical Variables
- Performing Variable Discretization
- Working with Outliers
- Extracting Features from Date and Time Variables
- Performing Feature Scaling
- Creating New Features
- Extracting Features from Relational Data with Featuretools
- Creating Features from a Time Series with tsfresh
- Extracting Features from Text Variables
Editorial Reviews
Review
“Feature engineering is one of the core elements of data science, and I am excited to see a data science book that is fully dedicated to this topic. Soledad is the author and maintainer of the Feature-Engine Python library that, as the name implies, focuses on feature engineering for machine learning applications.
The book is for folks who are interested in getting started with machine learning applications and practitioners who wish to deepen their knowledge in this domain.”
Rami Krispin, Senior Manager – Data Science and Engineering at Apple
“Soledad Galli masterfully breaks down advanced concepts into digestible, actionable steps while maintaining technical depth.
What truly impresses me is how the book emphasizes not just the how but also the why of feature engineering. Understanding when to apply specific techniques is often as important as knowing how to implement them, and this book excels at providing that context.”
Serg Masís, Data Scientist and Author of Interpretable Machine Learning with Python
“In the current wave of GenAI hype, it’s easy to overlook the power of well-engineered traditional machine learning models. Soledad Galli’s book is a comprehensive and well-written guide that equips you with essential feature engineering techniques. With clear explanations and clean, well-structured code examples, this book covers everything from imputation and outlier removal to time series and NLP-specific techniques. Whether you’re starting out in machine learning or need a refresher on core feature engineering methods, this book is an invaluable resource.”
Andy McMahon, Principal AI and MLOps Engineer, author of Machine Learning Engineering with Python
About the Author
Soledad Galli is a bestselling data science instructor, author, and open-source Python developer. As the leading instructor at Train in Data, she teaches intermediate and advanced courses in machine learning that have enrolled over 64,000 students worldwide and continue to receive positive reviews. Sole is also the developer and maintainer of the Python open-source library Feature-engine, which provides an extensive array of methods for feature engineering and selection. With extensive experience as a data scientist in finance and insurance sectors, Sole has developed and deployed machine learning models for assessing insurance claims, evaluating credit risk, and preventing fraud. She is a frequent speaker at podcasts, meetups, and webinars, sharing her expertise with the broader data science community.
Wow! eBook

