The Orange Book of Machine Learning Green Edition: The essentials of making predictions using supervised regression and classification for tabular data

The Orange Book of Machine Learning Green Edition: The essentials of making predictions using supervised regression and classification for tabular data book cover

The Orange Book of Machine Learning Green Edition: The essentials of making predictions using supervised regression and classification for tabular data

Author(s): Carl McBride Ellis (Author)

  • Publisher: Packt Publishing
  • Publication Date: July 24, 2026
  • Edition: The essentials of making predictions using supervised regression and classification for tabular data
  • Language: English
  • Print length: 238 pages
  • ISBN-10: 1808081315
  • ISBN-13: 9781808081316

Book Description

Learn supervised machine learning for tabular data using Python, pandas, scikit-learn, CatBoost, LightGBM, XGBoost, TabPFN, and TabICL for regression, classification, and predictive modeling

Key Features

  • Explore, clean, and prepare tabular datasets for machine learning workflows
  • Build regression and classification models using modern machine learning tools
  • Improve predictions with calibration, conformal intervals, and optimization techniques

Book Description

Master the essential tools and techniques for supervised machine learning on tabular data with this practical guide to regression and classification. Through clear explanations, code snippets, and hands-on notebooks, you’ll learn how to use Python and leading machine learning libraries, including pandas, scikit-learn, CatBoost, LightGBM, XGBoost, TabPFN, and TabICL, to build predictive models for real-world datasets.

The book covers the complete workflow, from data exploration and cleaning to model development, evaluation, and optimization. You’ll learn how to perform regression analysis for accurate point predictions and estimate uncertainty using conformal prediction intervals. For classification tasks, you’ll explore probabilistic predictions and calibration techniques to improve model reliability. You’ll also discover practical approaches to feature engineering, feature selection, and hyperparameter optimization to enhance model performance. In addition, the book introduces tabular foundation models and in-context learning techniques, providing insight into the latest advances in machine learning for structured data.

By the end of the book, you’ll have the skills and confidence to develop, evaluate, and deploy supervised machine learning models for a wide range of tabular data applications.

What you will learn

  • Perform exploratory data analysis and data cleaning
  • Apply cross-validation for reliable model evaluation
  • Build regression models and prediction intervals
  • Develop calibrated probabilistic classification models
  • Optimize models through hyperparameter tuning
  • Engineer and select features for improved performance
  • Use ensemble learning methods effectively
  • Explore tabular foundation models and in-context learning

Who this book is for

This book is designed for motivated self-learners, university students studying applied machine learning, junior data scientists, and academic researchers looking to incorporate machine learning into their analytical workflows. Readers should have a basic familiarity with Python and data analysis concepts. Whether you’re developing predictive models for business, research, or educational purposes, this book provides the practical guidance needed to apply modern machine learning techniques to structured and tabular datasets.

Table of Contents

  1. Introduction
  2. Statistics
  3. Exploratory data analysis (EDA)
  4. Data cleaning
  5. Cross-validation
  6. Interpolation and smoothing
  7. Regression
  8. Classification
  9. GLM and GAM
  10. Ensemble estimators
  11. Hyperparameter optimization (HPO)
  12. Feature engineering and selection
  13. Tabular foundation models (TFM)

Editorial Reviews

Editorial Reviews

About the Author

Carl is a freelance data scientist based in Madrid, Spain, specializing in predictive analytics and machine learning. He is also an adjunct university lecturer in machine learning and artificial intelligence. Before transitioning into data science, Carl spent more than 20 years as an academic researcher in the physical sciences, focusing on the computer simulation of liquids using molecular dynamics and Metropolis Monte Carlo methods. He has co-authored more than 40 scientific publications and brings extensive experience in data analysis, modeling, and scientific computing. His work combines academic rigor with practical machine learning applications, helping students and professionals develop effective predictive solutions.

View on Amazon

未经允许不得转载:Wow! eBook » The Orange Book of Machine Learning Green Edition: The essentials of making predictions using supervised regression and classification for tabular data