Sam Austin AI

Synthetic Data for Imbalanced Classification: SMOTE and Beyond (2026)

September 22, 2026 14 min read Sam Austin
Contents
Abstract visualization of data points showing minority class oversampling patterns
Abstract visualization of data points showing minority class oversampling patterns

Figure 1: SMOTE generates synthetic minority samples through interpolation — but knowing when to use it, and when not to, matters more than the mechanism itself

Here's the origin story worth knowing before touching a single line of SMOTE code: in the late 1990s, a graduate student named Nitesh Chawla was building a classifier to detect cancerous pixels in mammography images. He hit 97% accuracy and was thrilled — until he realized 97.6% of the pixels in his dataset were normal. A model that predicted "normal" for literally everything would have scored higher than his actual classifier. That gap between headline accuracy and genuine usefulness on the minority class is the entire reason SMOTE exists, and it's the same gap that makes "should you still use it in 2026" a genuinely live, actively debated question rather than a settled one.

Recall the scikit-learn Pipeline article's core warning about data leakage — SMOTE has an almost identical leakage trap of its own, and it's one of the most common mistakes in applied imbalanced classification. This article covers the mechanism, the honest limitations, the variant family that exists specifically to patch those limitations, and the newer GAN-based competitors (recall CTGAN directly) now challenging SMOTE's default status for genuinely hard cases.

By the end of this guide, you'll understand exactly how SMOTE generates synthetic minority samples, which variant fits which failure mode, and the specific cross-validation mistake that silently invalidates results in a genuinely large share of published SMOTE evaluations. IMO, "no pre-processing method consistently outperforms the others as imbalance severity varies" is the single most important, most consistently reproduced finding in this entire literature.

What SMOTE Actually Does

SMOTE (Synthetic Minority Oversampling Technique) generates new synthetic samples for the minority class by interpolating between existing minority samples, rather than simply duplicating them — a genuinely important distinction from naive random oversampling, which just copies existing minority rows and risks overfitting to those exact repeated examples.

  1. Identify the minority class, then for each minority sample, find its k nearest neighbors (typically k=5) within that same minority class, using distance in feature space.
  2. Generate a new synthetic point along the line segment connecting a minority sample and one of its neighbors — pick a random point somewhere between the two, not the two originals themselves.
  3. This lets a model learn broader decision-region patterns rather than memorizing repeated exact points, genuinely the mechanism-level reason SMOTE reduces overfitting compared to naive duplication.
from imblearn.over_sampling import SMOTE
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, stratify=y)

smote = SMOTE(random_state=42, k_neighbors=5)
X_resampled, y_resampled = smote.fit_resample(X_train, y_train)

imblearn (imbalanced-learn) is genuinely the standard Python package for SMOTE and virtually every variant covered below — designed to integrate directly with scikit-learn's own API conventions.

The Leakage Trap: Recall the Pipeline Article's Core Warning, Applied Here

This is genuinely worth flagging before anything else in this article, because it's a real, common, silent mistake: notice smote.fit_resample() is called on X_train only, after the train/test split — never on the full dataset before splitting.

Applying SMOTE before splitting lets synthetic minority samples derived from training data leak into your test set — your evaluation metrics then reflect how well the model does on data that's statistically entangled with what it trained on, not genuine generalization to unseen data.

The correct pattern, recall the scikit-learn Pipeline article directly: SMOTE should live inside a cross-validation loop, resampled fresh on each training fold, never touching the held-out validation or test fold at all.

from imblearn.pipeline import Pipeline as ImbPipeline
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score

pipeline = ImbPipeline([
    ("smote", SMOTE(random_state=42)),
    ("classifier", LogisticRegression())
])

scores = cross_val_score(pipeline, X, y, cv=5, scoring="f1")

Notice imblearn.pipeline.Pipeline, not sklearn.pipeline.Pipeline — this is genuinely required, not a stylistic choice. Standard scikit-learn Pipeline steps expect a 1:1 transformation (same number of rows in and out); SMOTE changes the number of rows, which only imblearn's pipeline variant correctly handles across cross-validation folds.

The Honest Limitations, Directly

SMOTE has real, well-documented weaknesses, worth taking seriously rather than treating as a universal default fix for imbalance.

  • Overlapping classes: if minority and majority class boundaries genuinely overlap, SMOTE can generate synthetic samples that actually land in majority-class territory — the interpolation logic has no awareness of where the true decision boundary sits, only of where the minority samples happen to be.
  • Not suitable for categorical data: SMOTE's core mechanism works by interpolating numerical features — averaging or blending between two points along a line segment. This genuinely doesn't make sense for a categorical column (you can't meaningfully "interpolate" between "red" and "blue"), which is exactly why the SMOTE-NC variant exists.
  • Computational cost: SMOTE relies on k-nearest-neighbors, which doesn't scale well to large datasets — genuinely worth budgeting for if your minority class itself is large, not just your overall dataset.
  • Can amplify noise: if the minority class contains noisy or mislabeled samples, SMOTE will happily interpolate around them too, potentially manufacturing more synthetic noise rather than more genuine signal.

The Variant Family: Each One Patches a Specific Weakness

Borderline-SMOTE: Focus Where It Actually Matters

Rather than generating synthetic samples uniformly across the entire minority class, Borderline-SMOTE specifically targets the borderline instances — minority samples sitting close to the actual class boundary — since these are the genuinely hardest, most decision-relevant points for a classifier to learn correctly.

from imblearn.over_sampling import BorderlineSMOTE

borderline_smote = BorderlineSMOTE(random_state=42, kind="borderline-1")
X_resampled, y_resampled = borderline_smote.fit_resample(X_train, y_train)

This generates synthetic samples near the boundary between classes, helping mitigate the impact of outliers and providing better class separation — genuinely more targeted than vanilla SMOTE's uniform interpolation across the whole minority region.

ADASYN: Adaptive Difficulty-Weighted Generation

ADASYN (Adaptive Synthetic Sampling) generates more synthetic samples specifically around hard-to-classify minority instances, using a weighted distribution based on each sample's actual local learning difficulty rather than treating every minority point equally.

from imblearn.over_sampling import ADASYN

adasyn = ADASYN(random_state=42, n_neighbors=5)
X_resampled, y_resampled = adasyn.fit_resample(X_train, y_train)

A genuinely useful empirical finding worth citing: SVM paired with SMOTE performs better than SVM paired with ADASYN as the degree of class imbalance increases — a concrete, comparative signal that the "better" variant genuinely depends on both your classifier choice and your imbalance severity, not a universal ranking.

SMOTE-NC: Handling Mixed Categorical and Continuous Data

Recall the exact limitation flagged above — SMOTE-NC (Nominal and Continuous) exists specifically to handle datasets mixing categorical and numerical features, interpolating continuous columns the normal way while handling categorical columns through majority-voting among nearest neighbors instead.

from imblearn.over_sampling import SMOTENC

smote_nc = SMOTENC(categorical_features=[1, 3, 5], random_state=42)
X_resampled, y_resampled = smote_nc.fit_resample(X_train, y_train)

Recall the scikit-learn ColumnTransformer article's mixed-type-handling discipline directly here — you're specifying exactly which column indices are categorical, the same explicit-declaration principle that article's numeric_features/categorical_features split relied on.

SVM-SMOTE: Decision-Boundary-Guided Generation

Integrates SMOTE with an SVM to identify support vectors defining the actual decision boundary, then generates synthetic samples specifically along those support vectors — ensuring new points land genuinely close to the real, learned decision boundary rather than an arbitrary interpolation region.

SMOTE-ENN and SMOTE-Tomek: Oversample, Then Clean

These combine SMOTE's oversampling with an undersampling cleaning step afterward — removing ambiguous or noisy samples near class boundaries after synthetic generation, rather than trusting the raw oversampled output.

from imblearn.combine import SMOTEENN, SMOTETomek

smote_enn = SMOTEENN(random_state=42)
X_resampled, y_resampled = smote_enn.fit_resample(X_train, y_train)

This two-stage approach genuinely addresses SMOTE's noise-amplification weakness directly — oversample to balance classes, then clean up ambiguous points the oversampling process may have introduced or exposed.

Quick Variant Comparison

VariantSolvesBest For
SMOTE (base)Naive duplication's overfitting riskGeneral-purpose starting point
Borderline-SMOTEWasted generation far from the decision boundaryClasses with a genuinely identifiable border region
ADASYNUniform treatment of easy and hard minority samplesDatasets with clearly variable per-sample difficulty
SMOTE-NCCategorical feature interpolationMixed categorical/continuous tabular data
SVM-SMOTEBoundary-blind generationWhen SVM-style boundary precision matters
SMOTE-ENN / SMOTE-TomekNoise amplification from raw oversamplingNoisy, ambiguous minority-class data

The Honest 2026 Question: Should You Still Use SMOTE At All?

Worth addressing directly rather than assuming SMOTE remains the automatic default it was a decade ago. Recall CTGAN's conditional generator and training-by-sampling from earlier in this series' CTGAN article — this is genuinely a direct competitor to SMOTE for the exact same imbalanced-tabular-data problem, using deep generative modeling instead of geometric interpolation.

  • Newer GAN-based oversampling approaches (HypoGAN, and others) have been benchmarked directly against SMOTE, ADASYN, and Borderline-SMOTE on real datasets including Wisconsin Breast Cancer and Credit Card Fraud Detection, achieving genuinely competitive F1-scores — this is an actively contested space, not a settled "SMOTE always wins" or "SMOTE is obsolete" conclusion either way.
  • The reproduced, consistent finding across the broader literature: no single rebalancing strategy consistently outperforms the others as imbalance severity varies. This mirrors the exact "no universal winner" pattern from this series' warehouse and MLOps platform comparison articles — the right choice genuinely depends on your specific data, classifier, and imbalance ratio, not a fixed ranking.
  • A genuinely practical, evidence-grounded decision rule: start with plain SMOTE as a fast, cheap baseline. If categorical features are present, move to SMOTE-NC immediately rather than applying base SMOTE incorrectly. If results are noisy or boundary-ambiguous, try SMOTE-ENN or SMOTE-Tomek. Only reach for a full GAN-based approach (recall the CTGAN article's genuine complexity and training-time cost) once simpler resampling has been validated as insufficient — the same escalation discipline running through this series' entire augmentation and synthetic-data arc.

Beyond Resampling: Class Weighting as a Genuine Alternative

Worth naming as a real alternative rather than assuming resampling is the only path: many classifiers (scikit-learn's LogisticRegression, RandomForestClassifier, and others) accept a class_weight="balanced" parameter, penalizing misclassification of the minority class more heavily during training rather than generating any synthetic data at all.

from sklearn.ensemble import RandomForestClassifier

clf = RandomForestClassifier(class_weight="balanced", random_state=42)

This genuinely avoids SMOTE's entire interpolation-and-leakage-risk machinery — worth trying as a first, cheaper baseline before reaching for any oversampling technique, since it requires zero additional pipeline complexity and no k-NN computation at all.

Common Mistakes People Make

  1. Applying SMOTE before the train/test split. Recall this being the single most important warning in this article — this causes genuine test-set leakage and inflates reported performance in a way that doesn't reflect real generalization.
  1. Using plain SMOTE on data with categorical columns. Recall the interpolation mechanism directly — averaging between "red" and "blue" produces nonsense; use SMOTE-NC and explicitly declare your categorical column indices.
  1. Assuming a "better" variant exists universally rather than for your specific data. Recall the consistently reproduced literature finding — no pre-processing method consistently outperforms the others as imbalance severity varies; validate empirically on your own data rather than trusting a general ranking.
  1. Reaching for CTGAN-based oversampling before trying SMOTE or class weighting first. Recall the escalation discipline directly — start cheap and simple, move to the more complex, slower-training generative approach only once simpler methods are validated as insufficient.
  1. Using sklearn.pipeline.Pipeline instead of imblearn.pipeline.Pipeline with a resampler in the chain. This produces a real, not cosmetic, error — standard scikit-learn pipelines assume row-count-preserving transforms, which SMOTE genuinely violates.

Want to Go Deeper?

Educative offers hands-on courses covering imbalanced classification, ML pipelines, and production model validation. If you want structured learning paths alongside the SMOTE techniques in this article, their Machine Learning with Python and Applied Machine Learning courses are solid companions.

Wrapping This Up

SMOTE and its variant family solve the exact problem Chawla's mammography classifier exposed — a model can score deceptively high on aggregate accuracy while being genuinely useless at the one class that actually matters — by generating synthetic minority samples through k-NN interpolation rather than naive duplication. Borderline-SMOTE, ADASYN, SMOTE-NC, SVM-SMOTE, and the combined SMOTE-ENN/SMOTE-Tomek variants each patch a specific, well-documented weakness in the base technique, and newer GAN-based approaches (recall CTGAN directly) now offer a genuinely competitive, more computationally expensive alternative for the hardest cases.

Remember that applying SMOTE before your train/test split causes real data leakage — always resample only within the training fold, ideally through imblearn's pipeline integration — and that no single variant or technique consistently wins across the literature, making empirical validation on your own specific data genuinely non-negotiable. FYI, this article closes the loop between the CTGAN and scikit-learn Pipeline articles from earlier in this series — SMOTE is the classical, geometric answer to the same imbalanced-data problem CTGAN's conditional generator solves through deep learning, and the leakage discipline connecting both traces directly back to the Pipeline article's original warning.

Now go check whether any classification project from earlier in this series actually applied its resampling technique inside a proper cross-validation fold, or before the split entirely. That single check, more than the choice between SMOTE and ADASYN, is what determines whether your reported performance numbers are actually trustworthy.


Keywords: SMOTE, imbalanced classification, oversampling, synthetic data, minority class, SMOTE-NC, ADASYN, Borderline-SMOTE, class imbalance, Python, imblearn, machine learning, data leakage, cross-validation

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles