Sam Austin AI

SDV Tutorial: Generate Synthetic Tabular Data with Python (2026)

September 12, 2026 14 min read Sam Austin
Contents
SDV Tutorial Generate Synthetic Tabular Data Python
SDV Tutorial Generate Synthetic Tabular Data Python

Figure 1: SDV puts GaussianCopula and CTGAN models directly in your hands — no API key, no vendor relationship, just your own infrastructure

Recall the previous two articles' repeated recommendation — start with SDV before evaluating any paid platform. This is genuinely that recommendation made concrete. Where Gretel and MOSTLY AI wrap synthesis behind a managed API, SDV puts the actual GaussianCopula and CTGAN models directly in your hands, running entirely on your own infrastructure, with zero vendor relationship required.

Maintained by DataCebo as "your one-stop shop for creating tabular synthetic data," SDV genuinely spans the full range from classical statistical methods to deep-learning-based synthesis, all behind one consistent API — single tables, multiple connected tables, or sequential time-series data. This is the hands-on tutorial the previous two articles pointed toward but didn't actually walk through.

By the end of this guide, you'll fit both a fast statistical synthesizer and a deep-learning GAN-based one, evaluate the result against your real data, and know exactly which synthesizer fits which situation. IMO, the fact that you can genuinely go from "real CSV" to "evaluated synthetic dataset" in under twenty lines of code is the whole reason this library earns the "start here" recommendation.

Installing SDV

pip install sdv

That's genuinely the entire installation — no separate model downloads, no API key, no account. SDV bundles its dependencies (Copulas, RDT for reversible data transforms, CTGAN) into one package.

Metadata: Telling SDV What Your Data Actually Is

Before fitting anything, SDV needs a Metadata object describing your table's structure — which columns are numerical, categorical, datetime, or identifiers. This is genuinely required as the first argument to every synthesizer, not optional configuration.

import pandas as pd
from sdv.metadata import Metadata

data = pd.read_csv("customer_transactions.csv")

metadata = Metadata.detect_from_dataframe(data, table_name="transactions")

detect_from_dataframe auto-infers column types from your data, but recall the Great Expectations article's core lesson directly here — always review and correct the auto-detected metadata before fitting, the same way that article insisted on explicit validation rather than blind trust in inferred schemas. An auto-detected "numerical" column that's actually a categorical ID code will genuinely produce nonsensical synthetic values if left uncorrected.

metadata.update_column(column_name="customer_id", sdtype="id")
metadata.update_column(column_name="account_type", sdtype="categorical")

Your First Synthesizer: GaussianCopula (Fast, Statistical)

The Gaussian Copula Synthesizer uses classic statistical methods rather than deep learning — genuinely the right starting point for most projects, given its speed and interpretability compared to a GAN-based alternative.

from sdv.single_table import GaussianCopulaSynthesizer

synthesizer = GaussianCopulaSynthesizer(metadata)
synthesizer.fit(data)

synthetic_data = synthesizer.sample(num_rows=1000)

Notice how closely this mirrors the scikit-learn Pipeline article's .fit()/.predict() pattern from earlier in this series — fit() learns the statistical structure of your real data, sample() generates new rows from that learned distribution. This is genuinely the entire workflow for the common case.

Customizing Column-Level Distributions

GaussianCopula lets you specify exactly which statistical distribution each column should be modeled with, rather than accepting a generic default across every field.

synthesizer = GaussianCopulaSynthesizer(
    metadata,
    enforce_min_max_values=True,
    numerical_distributions={
        "transaction_amount": "beta",
        "signup_date": "uniform"
    },
    default_distribution="norm"
)

That enforce_min_max_values=True parameter matters more than it looks — it constrains generated numerical values to stay within your real data's actual observed range. Without it, synthetic values can genuinely fall outside realistic bounds — a negative age, a transaction amount larger than any real transaction ever recorded — exactly the kind of Great Expectations-style validation failure worth catching at generation time rather than downstream.

CTGAN: When Statistical Methods Aren't Enough

Recall this being the exact deep-learning technique named in the synthetic data techniques article — CTGAN (Conditional GAN) uses generative adversarial networks specifically designed for tabular data, from the "Modeling Tabular Data using Conditional GAN" paper presented at NeurIPS 2019.

from sdv.single_table import CTGANSynthesizer

synthesizer = CTGANSynthesizer(
    metadata,
    epochs=300,
    generator_dim=(256, 256),
    discriminator_steps=1
)

synthesizer.fit(data)
synthetic_data = synthesizer.sample(num_rows=1000)

CTGAN genuinely captures more complex, non-linear relationships between columns than GaussianCopula's statistical approach can — the tradeoff is real training time (potentially minutes to hours depending on data size and epoch count, versus GaussianCopula's near-instant fitting) and a documented warning worth taking seriously: SDV itself alerts you during preprocessing if the fitting process is likely to be slow given your specific schema — a genuinely useful, honest signal built directly into the library rather than a surprise you discover after committing to a long training run.

CopulaGAN: A Genuine Hybrid Worth Knowing

CopulaGAN combines both approaches in two explicit stages: first, it learns each column's individual marginal distribution and normalizes the data using that statistical understanding (the same Gaussian-normalization approach GaussianCopula uses); then, it trains CTGAN specifically on that already-normalized data rather than the raw values.

from sdv.single_table import CopulaGANSynthesizer

synthesizer = CopulaGANSynthesizer(metadata)
synthesizer.fit(data)

This hybrid approach genuinely aims to get CTGAN's relationship-modeling strength while benefiting from the statistical normalization stabilizing the GAN's training — worth trying specifically when plain CTGAN's training feels unstable or produces poor-quality results on your particular dataset shape.

Choosing Among the Three: A Practical Decision

Start with GaussianCopulaSynthesizer regardless of your eventual choice. It's fast enough to iterate on quickly, interpretable enough to sanity-check column by column, and genuinely sufficient for a large share of tabular synthesis needs.

Move to CTGANSynthesizer specifically when your real data has genuinely complex, non-linear inter-column relationships GaussianCopula's statistical assumptions can't capture — and you've confirmed via evaluation (next section) that GaussianCopula's output isn't good enough.

Try CopulaGANSynthesizer if CTGAN alone produces unstable training or unconvincing results — the statistical pre-normalization step can genuinely help GAN training converge on data where it otherwise struggles.

Evaluating Your Synthetic Data: Don't Skip This

Recall the model monitoring article's drift-detection discipline directly — the exact same statistical comparison logic applies here, just at generation time instead of post-deployment.

from sdv.evaluation.single_table import evaluate_quality, run_diagnostic

quality_report = evaluate_quality(
    real_data=data,
    synthetic_data=synthetic_data,
    metadata=metadata
)

diagnostic_report = run_diagnostic(
    real_data=data,
    synthetic_data=synthetic_data,
    metadata=metadata
)

evaluate_quality compares column-level distributions and inter-column correlations between real and synthetic data, producing an actual quality score rather than a subjective "looks about right" judgment. run_diagnostic specifically checks for structural problems — boundary violations, category mismatches — the tabular equivalent of the Great Expectations article's schema validation, applied here to your generated output rather than incoming pipeline data.

Never skip this step regardless of which synthesizer you chose. Recall the synthetic data article's model collapse discussion directly — evaluating your generated data against the real distribution before using it for training is genuinely the concrete implementation of "validate distribution overlap on the way in" that article recommended as a collapse-prevention practice.

Multi-Table Synthesis: Preserving Referential Integrity

Recall the K2view comparison from the tools article — SDV genuinely handles the same multi-table, relational challenge natively, not just as an enterprise-tier feature.

from sdv.multi_table import HMASynthesizer
from sdv.metadata import Metadata

metadata = Metadata.detect_from_dataframes(
    data={"customers": customers_df, "orders": orders_df}
)
metadata.set_primary_key(table_name="customers", column_name="customer_id")
metadata.add_relationship(
    parent_table_name="customers",
    child_table_name="orders",
    parent_primary_key="customer_id",
    child_foreign_key="customer_id"
)

synthesizer = HMASynthesizer(metadata)
synthesizer.fit(data={"customers": customers_df, "orders": orders_df})
synthetic_tables = synthesizer.sample()

This is genuinely the open-source path to the referential-integrity capability the tools comparison article flagged as K2view's specific enterprise strength — SDV's HMASynthesizer (Hierarchical Modeling Algorithm) generates synthetic customers and synthetic orders that correctly reference each other, rather than two independently synthesized tables whose foreign keys no longer line up.

Constraints: Encoding Business Rules Directly

Real data follows rules a purely statistical model doesn't automatically know about — a shipping date can't precede an order date, a discount percentage can't exceed 100%. SDV lets you encode these directly rather than hoping the synthesizer infers them.

from sdv.cag import Inequality

constraint = Inequality(
    low_column_name="order_date",
    high_column_name="ship_date"
)

synthesizer = GaussianCopulaSynthesizer(metadata)
synthesizer.add_constraints(constraints=[constraint])
synthesizer.fit(data)

This is genuinely the SDV-native equivalent of a Great Expectations Expectation, applied at generation time rather than validation time — instead of checking after the fact that ship_date >= order_date, you're constraining the generative process itself to never produce a violation in the first place.

Anonymization: Privacy-Preserving Fields

Recall the previous article's privacy-substitution use case directly — SDV includes built-in anonymization for genuinely sensitive fields, letting you replace real names, emails, or addresses with realistic-looking fakes rather than the statistically-modeled originals.

metadata.update_column(column_name="customer_email", sdtype="email", pii=True)
metadata.update_column(column_name="full_name", sdtype="name", pii=True)

Marking a column pii=True tells SDV to generate a genuinely fake, unlinkable replacement value rather than learning and reproducing statistical patterns from the real, sensitive values — the concrete mechanism behind the "privacy substitution" use case from the synthetic data fundamentals article.

Common Mistakes People Make

Skipping metadata review after auto-detection. detect_from_dataframe is a starting point, not a final answer — an incorrectly typed column produces genuinely nonsensical synthetic values silently.

Jumping straight to CTGAN without trying GaussianCopula first. Recall the decision framework directly — GaussianCopula's speed makes it the right default to validate your approach before committing to CTGAN's genuinely longer training time.

Never running evaluate_quality or run_diagnostic. Recall the model collapse discussion from the earlier synthetic data article — validating your generated output against the real distribution is a required step, not an optional afterthought.

Synthesizing linked tables independently instead of using HMASynthesizer. This breaks referential integrity exactly the way the tools comparison article warned about for naive tabular synthesis approaches.

Forgetting pii=True on genuinely sensitive columns. Without it, SDV models and reproduces the statistical pattern of real names or emails rather than generating properly anonymized replacements.

  • Synthetic Data for Deep Learning by Salvatore Rizzello et al. — covers the theoretical foundations and practical applications of synthetic data generation across modalities, including GANs, VAEs, and the privacy-preservation use case this article addresses.
  • Generative Deep Learning by David Foster — the definitive guide to GANs, VAEs, and diffusion models from first principles, directly covering the CTGAN and CopulaGAN techniques this tutorial implements.
  • Designing Machine Learning Systems by Chip Huyen — covers data quality, data pipelines, and the train-serve skew problem, providing the broader MLOps context for where synthetic data generation fits in a production lifecycle.

Wrapping This Up

SDV genuinely delivers on the "start here before any paid platform" recommendation from the last two articles — GaussianCopula for fast, statistical synthesis; CTGAN and CopulaGAN for deep-learning-based synthesis when column relationships are genuinely complex; HMASynthesizer for multi-table referential integrity; and built-in constraints and PII anonymization for encoding business rules and privacy requirements directly into generation, all running entirely on your own infrastructure with zero vendor relationship required.

Remember that evaluating synthetic output against real data with evaluate_quality and run_diagnostic is genuinely non-optional, not a nice-to-have — it's the concrete implementation of the collapse-prevention discipline the earlier synthetic data article recommended. FYI, this tutorial closes the loop the previous two articles opened — you now have working code for the exact open-source starting point both the synthetic data techniques article and the tools comparison article pointed toward as the sensible first step before any commercial platform evaluation.

Now go run evaluate_quality on synthetic data generated from your own real dataset, and actually read the column-by-column comparison it produces before using that synthetic data for anything downstream. That evaluation step, more than the choice between GaussianCopula and CTGAN, is what separates a genuinely useful synthetic dataset from one that just looks plausible.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles