Sam Austin AI

Building a Feature Engineering Pipeline with Scikit-learn Pipelines (2026)

September 9, 2026 13 min read Sam Austin
Contents
Feature Engineering Pipeline Scikit-learn Pipelines
Feature Engineering Pipeline Scikit-learn Pipelines

Figure 1: A scikit-learn Pipeline encapsulates preprocessing and modeling into one atomic object — the same fitted scaler runs at training and inference, eliminating an entire category of silent failures

Recall the dbt article's genuinely central insight about train-serve skew: the exact same transformation logic needs to run identically whether you're preparing training data or scoring a live prediction request. dbt solves that for warehouse-scale SQL transformations. This article solves the same problem at the model level — inside a single Python object that fits, transforms, and predicts as one atomic, leak-proof unit, at the scale of an individual model rather than an entire warehouse.

Currently at scikit-learn 1.9.0, Pipeline and ColumnTransformer remain the two objects doing almost all of the real work here, and they've genuinely stayed stable enough that code written for them years ago mostly still works today. This is deliberately the smallest-scale, most classically "data science" article in this entire arc — no Kafka, no Airflow, no warehouse — just the actual feature engineering code sitting inside the model artifact the rest of this series' MLOps tooling wraps around.

By the end of this guide, you'll understand exactly why Pipeline prevents a specific, genuinely common form of data leakage, build a real ColumnTransformer-based preprocessing setup for mixed data types, and know how this single object plugs into the model registry and CI/CD articles from earlier in this series. IMO, the "why not just preprocess with pandas first" question this article answers is one of the most practically important things a scikit-learn user can actually internalize :)

The Problem Pipeline Actually Solves

Here's the genuinely common mistake this whole article exists to prevent: fitting a scaler or an imputer on your entire dataset — training and test combined — before splitting. Doing this means information from your test set (its mean, its distribution) leaks into how your training data gets transformed, quietly inflating your evaluation metrics with information the model shouldn't have had access to at all. This is precisely the same category of failure the Great Expectations article covered, just at the feature-scaling level instead of the temporal-join level.

A Pipeline object encapsulates a chained sequence of operations — transformations, and optionally a final model — into a single object, ensuring the same transformations applied to your training set get applied consistently to new data during prediction, and critically, that any transformer's internal statistics (a scaler's mean, an imputer's fill value) get learned only from training data during cross-validation, never leaking test-fold information into the fit.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

pipeline = Pipeline([
    ("scaler", StandardScaler()),
    ("classifier", LogisticRegression())
])

pipeline.fit(X_train, y_train)
predictions = pipeline.predict(X_test)

Notice this is genuinely one object, not two separate steps you remember to run in the right order. pipeline.fit() fits the scaler on X_train only, transforms X_train using those fitted statistics, then fits the classifier on the result — and pipeline.predict() automatically applies that same fitted scaler (not a newly fitted one) before handing data to the classifier. This is the single mechanism eliminating an entire category of "I forgot to apply the same transformation at inference time" bugs — recall this being exactly the train-serve skew problem the dbt article named directly.

The Real-World Complication: Mixed Data Types

A single Pipeline step assumes every column gets the same treatment. Real datasets genuinely don't work that way — numeric columns need scaling, categorical columns need encoding, and applying StandardScaler to a categorical column is nonsensical. This is exactly the gap ColumnTransformer closes.

import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer

numeric_features = ["age", "income", "transaction_count"]
categorical_features = ["region", "account_type"]

numeric_transformer = Pipeline(steps=[
    ("imputer", SimpleImputer(strategy="median")),
    ("scaler", StandardScaler())
])

categorical_transformer = Pipeline(steps=[
    ("imputer", SimpleImputer(strategy="most_frequent")),
    ("encoder", OneHotEncoder(handle_unknown="ignore"))
])

preprocessor = ColumnTransformer(transformers=[
    ("num", numeric_transformer, numeric_features),
    ("cat", categorical_transformer, categorical_features)
])

Notice each branch is itself a small Pipeline — impute, then scale for numeric columns; impute, then encode for categorical ones. ColumnTransformer applies these two branches to their respective column subsets in parallel, then concatenates the results horizontally into one combined feature matrix. This nested structure — pipelines inside a ColumnTransformer inside a larger pipeline — is genuinely the standard shape of real-world scikit-learn feature engineering, not an advanced edge case.

That handle_unknown="ignore" parameter on the OneHotEncoder matters more than it looks — recall the ETL pipeline and Great Expectations articles' warnings about unexpected categorical values slipping through in production. Without this flag, a genuinely new category appearing at inference time (one that wasn't in the training data) would crash the pipeline outright; with it, the encoder gracefully produces all-zero indicator columns instead.

Wiring the Preprocessor Into the Full Modeling Pipeline

from sklearn.ensemble import RandomForestClassifier

full_pipeline = Pipeline(steps=[
    ("preprocessor", preprocessor),
    ("classifier", RandomForestClassifier(n_estimators=200, random_state=42))
])

full_pipeline.fit(X_train, y_train)
accuracy = full_pipeline.score(X_test, y_test)

This is genuinely the complete artifact — from raw, mixed-type input columns to a trained classifier, as one object. full_pipeline.predict(new_data) runs new, raw data through the exact same imputation, scaling, and encoding steps automatically, with zero risk of accidentally applying a differently-fitted transformer or forgetting a preprocessing step at inference time.

Hyperparameter Tuning Across the Whole Pipeline at Once

A genuinely underused capability: GridSearchCV can tune hyperparameters of both the transformers inside ColumnTransformer and the final estimator simultaneously, in one search — not two separate tuning passes.

from sklearn.model_selection import GridSearchCV

param_grid = {
    "preprocessor__num__imputer__strategy": ["mean", "median"],
    "classifier__n_estimators": [100, 200, 300],
    "classifier__max_depth": [10, 20, None]
}

grid_search = GridSearchCV(full_pipeline, param_grid, cv=5, scoring="f1")
grid_search.fit(X_train, y_train)

That double-underscore (__) syntax is doing genuinely important work — preprocessor__num__imputer__strategy navigates the nested structure: the preprocessor step, its num branch, that branch's imputer step, and finally its strategy parameter. Parameter names are constructed from step names chained together this way, letting a single GridSearchCV call search the entire pipeline's configuration space — imputation strategy and model hyperparameters — jointly, which genuinely matters since the best imputation choice can depend on which model ends up consuming its output.

Custom Transformers: When Built-In Steps Aren't Enough

Real feature engineering frequently needs domain-specific logic no built-in transformer covers. Scikit-learn's transformer interface is genuinely simple enough to extend yourself.

from sklearn.base import BaseEstimator, TransformerMixin

class RatioFeature(BaseEstimator, TransformerMixin):
    def __init__(self, numerator, denominator):
        self.numerator = numerator
        self.denominator = denominator

    def fit(self, X, y=None):
        return self

    def transform(self, X):
        X = X.copy()
        X[f"{self.numerator}_to_{self.denominator}_ratio"] = (
            X[self.numerator] / X[self.denominator].replace(0, 1)
        )
        return X

Inheriting from BaseEstimator and TransformerMixin gives your custom class the exact same .fit()/.transform()/.fit_transform() interface every built-in transformer has — meaning it slots directly into a Pipeline or ColumnTransformer step alongside StandardScaler or OneHotEncoder with zero special handling required. This is genuinely how domain-specific feature engineering (ratios, date-part extraction, text length features) integrates cleanly rather than living as a separate pandas step outside the pipeline's leak-proof boundary.

Where This Connects to the Rest of the MLOps Arc

Recall the model registry article's mlflow.sklearn.log_model() call directly — this is genuinely the object that gets logged and versioned.

import mlflow

with mlflow.start_run():
    full_pipeline.fit(X_train, y_train)
    mlflow.sklearn.log_model(
        full_pipeline,
        "model",
        registered_model_name="fraud-detector"
    )

This single line registers the entire pipeline — preprocessing and model together — as one versioned artifact. Recall the model registry article's lineage argument: loading models:/fraud-detector@champion later returns an object that preprocesses and predicts identically to what was trained, with zero risk of the serving environment applying a differently-configured scaler than training used. This is genuinely the concrete mechanism preventing train-serve skew at the code level, complementing dbt's warehouse-level solution to the same underlying problem from a different layer of the stack.

Testing Pipelines: Connecting to the CI/CD Article

Recall the CI/CD article's quality gate — a pipeline object is genuinely easy to unit-test in exactly the way the CI/CD article's evaluate.py pattern assumed was possible.

def test_pipeline_handles_unknown_category():
    sample = pd.DataFrame({
        "age": [35], "income": [50000], "transaction_count": [12],
        "region": ["never_seen_before"], "account_type": ["premium"]
    })
    result = full_pipeline.predict(sample)
    assert result is not None  # shouldn't crash on unseen category

This test directly validates the handle_unknown="ignore" behavior from earlier — confirming a genuinely realistic production scenario (a new region appearing after training) doesn't crash the deployed pipeline, exactly the kind of test that belongs in the GitHub Actions workflow from the CI/CD article, running before any model reaches the quality gate.

FeatureUnion vs. ColumnTransformer: A Genuine Distinction Worth Knowing

Worth naming since both concatenate transformer outputs and get confused for each other: FeatureUnion combines the outputs of multiple transformers each applied to the entire dataset, running each transformation independently and concatenating results — useful when you want several genuinely different feature-extraction approaches (say, both a PCA projection and the raw scaled features) applied to the same columns. ColumnTransformer applies different transformers to different, non-overlapping subsets of columns — the mixed numeric/categorical case this whole article has focused on.

A concrete technical distinction worth knowing: FeatureUnion copies the entire dataset for each parallel transformer; ColumnTransformer copies only the relevant column subset for each — meaningfully more memory-efficient when your transformers only ever need specific columns, which is genuinely the common case.

Common Mistakes People Make

Fitting a scaler or imputer on the full dataset before train/test splitting. This is precisely the data leakage problem Pipeline exists to prevent — always fit inside a Pipeline object fit only on training data, never as a separate manual step beforehand.

Forgetting handle_unknown="ignore" on OneHotEncoder. A genuinely new category value in production data will crash an otherwise-working pipeline without this flag — recall the ETL and Great Expectations articles' broader warnings about unexpected categorical values.

Applying preprocessing with pandas before the pipeline, then only wrapping the model itself. This defeats the entire leak-proofing and reproducibility guarantee — the preprocessing needs to live inside the Pipeline object, not before it.

Confusing FeatureUnion and ColumnTransformer. Use ColumnTransformer for the common "different columns need different treatment" case; reach for FeatureUnion specifically when you want multiple full-dataset transformations combined, not partitioned.

Not logging the full pipeline (preprocessing plus model) to the registry. Recall the model registry article directly — registering only the bare classifier without its preprocessing steps recreates the exact train-serve skew problem this whole arc has been trying to eliminate.

  • Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow by Aurelien Geron — the definitive practical reference for scikit-learn, including thorough coverage of Pipeline, ColumnTransformer, and the feature engineering patterns this article teaches. The chapter on training models covers exactly the data leakage problem Pipeline solves.
  • Feature Engineering and Selection by Max Kuhn and Kjell Johnson — focused specifically on feature engineering methodology, covering when and how to apply scaling, encoding, imputation, and interaction features. Directly complements the pipeline implementation patterns here with the reasoning behind each transformation choice.
  • Machine Learning Engineering by Andriy Burkov — covers the MLOps lifecycle from training to serving, including the train-serve skew problem this article's Pipeline object concretely prevents at the code level.

Wrapping This Up

Scikit-learn's Pipeline and ColumnTransformer solve feature engineering's most common, most quietly damaging mistake — data leakage from fitting transformers on data they shouldn't have seen yet — by encapsulating preprocessing and modeling into one atomic, leak-proof object that treats numeric and categorical columns differently while guaranteeing identical treatment between training and inference. GridSearchCV's double-underscore parameter syntax extends this same discipline to hyperparameter search across the entire pipeline at once, not just the final model.

Remember that handle_unknown="ignore" on categorical encoders is genuinely non-optional for any pipeline expected to survive real production data, and that logging the entire pipeline — not just the bare model — to your registry is what actually closes the train-serve skew loop this series' dbt and model registry articles both addressed from different angles. FYI, this article genuinely operates one level below everything else in this MLOps arc — while Airflow orchestrates and dbt transforms at warehouse scale, this is the actual code object living inside the model artifact those systems ultimately produce and serve :)

Now go take whatever model you trained earliest and most simply in this entire series — recall the CartPole DQN or the Snake state representation — and ask honestly whether its preprocessing logic lives inside a Pipeline object or scattered across separate script steps you'd have to remember to replicate identically at inference time. That gap, more than any benchmark in this article, is the concrete difference between a notebook experiment and something genuinely production-ready.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles