Sam Austin AI

Data Validation with Great Expectations: Catch Bad Data Early (2026)

September 8, 2026 14 min read Sam Austin
Contents
Data Validation Quality Testing
Data Validation Quality Testing

Figure 7: Data validation creates shared quality standards across teams — the foundation every reliable ML pipeline depends on

Recall the model monitoring article's churn-model story: a model degrading silently over months. This article is about catching the version of that problem that happens in seconds, not months — a null value where there shouldn't be one, a schema change nobody announced, a column that suddenly contains negative ages. The CI/CD article's quality gate blocked a worse model. Great Expectations exists to block bad data from ever reaching the pipeline that trains or feeds that model in the first place.

The tool rebranded its open-source core to GX Core, now at a stable 1.0+ release following Semantic Versioning, with a genuinely simplified API compared to the older, more verbose Great Expectations most existing tutorials still describe. This closes the last real gap in the data-and-MLOps arc running through this series — dbt transforms data, Feast serves it consistently, DVC versions it, Airflow schedules it — and Great Expectations is the thing that checks whether any of that data was actually trustworthy to begin with.

By the end of this guide, you'll understand GX Core's core components, write and run real validation checks against a DataFrame, and know exactly where this plugs into the Airflow pipelines from earlier in this series. IMO, the "Expectations as a shared language between data engineers and everyone downstream" framing is genuinely the most underrated part of this tool :)

What Great Expectations Actually Does

Great Expectations (GX) is a framework for describing data using expressive tests, then validating that the data actually meets those test criteria. The core idea is genuinely simple: instead of trusting that a pipeline's input data looks the way you assume it does, you write down explicit, checkable rules — "this column should never be null," "this value should always fall between 0 and 100" — and GX runs them automatically.

Expectations are the core unit — declarative, plain-language statements like "column age values must be between 0 and 120." Genuinely similar in spirit to assertions in a traditional Python unit test, just aimed at data instead of code.

47 built-in Expectations ship ready to use, covering the large majority of common data quality concerns, with a full gallery to browse and the ability to write custom ones for anything genuinely specific to your domain.

This gives teams a shared, common language for data quality — a data engineer, an analyst, and an ML engineer can all read an Expectation Suite and understand exactly what "valid data" means for a given dataset, without digging through someone's ad-hoc validation script.

Installing GX Core

python -m venv gx-env
source gx-env/bin/activate

pip install great_expectations

GX Core supports Python 3.10 through 3.13, with experimental support for newer versions available via an environment flag if you're on the bleeding edge. As with every tool in this series' data-and-MLOps arc, installing inside a dedicated virtual environment is genuinely recommended rather than optional.

import great_expectations as gx
print(gx.__version__)

The Core Components, in the Order You Actually Use Them

GX Core's workflow is built from a small number of Python objects, each representing a distinct piece of the validation process.

Data Context — the entry point for genuinely everything; a Python object providing access to your workflow's configuration, metadata, and validation results. Every GX workflow starts by creating one.

Data Source — tells GX how to connect to your actual data: a database, cloud storage, or a Pandas DataFrame already in memory.

Data Asset — a specific collection of records within that Data Source. If a Data Source is a database, a Data Asset is genuinely just a table, or the result of a specific query against one.

Batch Definition — tells GX how to organize the records within a Data Asset into the actual chunk you're going to validate.

import great_expectations as gx

context = gx.get_context()

data_source = context.data_sources.add_pandas("pandas_source")
data_asset = data_source.add_dataframe_asset(name="orders_asset")
batch_definition = data_asset.add_batch_definition_whole_dataframe("orders_batch")

This four-step setup should feel genuinely familiar if you've worked through the dbt or Feast articles — Context, Source, Asset, Batch is conceptually the same "declare your data's shape before doing anything with it" pattern those tools use, just applied specifically to validation rather than transformation or serving.

Writing Your First Expectations

import great_expectations.expectations as gxe

expectation_1 = gxe.ExpectColumnValuesToNotBeNull(column="order_id")
expectation_2 = gxe.ExpectColumnValuesToBeBetween(column="amount", min_value=0, max_value=100000)
expectation_3 = gxe.ExpectColumnValuesToBeInSet(
    column="order_status",
    value_set=["pending", "shipped", "delivered", "cancelled"]
)

Notice these read almost like plain English — that readability is genuinely deliberate design, not an accident. ExpectColumnValuesToBeInSet catches exactly the kind of silent schema drift the ETL pipeline article warned about — an unexpected new status value ("returned"?) appearing in production data that your downstream logic was never built to handle.

Validating a Batch

import pandas as pd

df = pd.read_csv("data/orders.csv")
batch = batch_definition.get_batch(batch_parameters={"dataframe": df})

result = batch.validate(expectation_1)
print(result)

The result object tells you whether the Expectation passed, and if not, exactly which rows failed and how many — genuinely actionable output, not just a boolean pass/fail. To reduce noise on large failing batches, GX deliberately caps how many failed values and row indexes it surfaces at once, keeping the output reviewable rather than dumping a wall of failed records.

Expectation Suites: Bundling Multiple Checks Together

Real datasets need more than one rule, and GX genuinely encourages bundling related checks into a single, named, reusable set.

suite = context.suites.add(gx.ExpectationSuite(name="orders_quality_suite"))
suite.add_expectation(gxe.ExpectColumnValuesToNotBeNull(column="order_id"))
suite.add_expectation(gxe.ExpectColumnValuesToBeBetween(column="amount", min_value=0, max_value=100000))
suite.add_expectation(gxe.ExpectColumnValuesToBeInSet(
    column="order_status",
    value_set=["pending", "shipped", "delivered", "cancelled"]
))

results = batch.validate(suite)

This is genuinely the artifact worth version-controlling — recall the DVC article's philosophy of codifying every aspect of your ML project. An Expectation Suite saved as JSON is exactly that same discipline applied to data quality rules: committed to Git, reviewed in pull requests, and evolved deliberately rather than living as tribal knowledge in someone's head.

Checkpoints: The Production Validation Unit

For anything beyond interactive, exploratory validation, GX uses Checkpoints — a bundled, repeatable unit combining a Batch and a Suite, designed specifically to run automatically as part of a real pipeline rather than a one-off notebook cell.

checkpoint = context.checkpoints.add(
    gx.Checkpoint(
        name="orders_checkpoint",
        validation_definitions=[
            gx.ValidationDefinition(
                data=batch_definition,
                suite=suite,
                name="orders_validation"
            )
        ]
    )
)

checkpoint_result = checkpoint.run()

This is the object your Airflow DAGs and CI/CD pipelines actually call — not the interactive, step-by-step Expectation-building shown above, which is genuinely more suited to exploration and initial suite design.

Wiring This Into Airflow

Recall the Airflow article's DAG pattern directly — Astronomer maintains a dedicated Great Expectations Airflow Provider specifically for running validations as a first-class task inside a DAG, using the GreatExpectationsOperator.

from great_expectations_provider.operators.great_expectations import GreatExpectationsOperator

validate_orders = GreatExpectationsOperator(
    task_id="validate_orders_data",
    conn_id="my_gx_connection",
    checkpoint_name="orders_checkpoint",
)

Placed as the task immediately after your extraction step and before your dbt transformation or model training task, this is genuinely the concrete implementation of the "validate data before it ever reaches training" principle from both the ETL and CI/CD articles — a failed validation here means the DAG halts, exactly the way a failed quality gate halted deployment in the CI/CD article.

Where Great Expectations Fits Against dbt's Own Testing

Recall the dbt article's dbt test and schema.yml-based tests — worth being explicit about the overlap and the distinction here, since both tools genuinely do "test your data."

dbt's built-in tests are SQL-native and tightly scoped to models already inside your warehouse — unique, not_null, and similar checks defined right alongside the transformation logic they validate.

Great Expectations is broader in scope and data-source-agnostic — it validates Pandas DataFrames, SQL databases, and files in cloud storage through one consistent Python API, and its Expectation library is genuinely richer than dbt's built-in test set (distributional checks, set-membership checks, complex conditional rules).

Many real pipelines genuinely use both — dbt's lightweight tests for quick, transformation-adjacent checks defined right in the model's schema file, and GX for the more comprehensive, standalone validation layer sitting at ingestion boundaries or before genuinely critical downstream steps like model training.

Automatic Documentation: Data Docs

GX automatically generates human-readable documentation from your validation results — a genuinely underused feature that pays off specifically for cross-team collaboration.

context.build_data_docs()

This produces an actual browsable HTML site showing every Expectation Suite, its most recent validation results, and historical pass/fail trends over time. This is genuinely the mechanism preserving your organization's institutional knowledge about its own data — instead of "ask the one person who remembers what this column is supposed to look like," the answer lives in a shared, always-current document nobody has to maintain by hand.

Where This Fits in the Full Series Pipeline

Extract (ETL article) → Validate (Great Expectations, this article) → Transform (dbt) →
Serve consistently (Feast) → Version (DVC) → Orchestrate all of the above (Airflow) →
Quality-gate the resulting model (CI/CD) → Monitor drift after deployment (Model Monitoring)

This genuinely completes the picture — every other article in this arc assumed the data flowing through it was basically trustworthy. Great Expectations is the article that actually earns that assumption, catching bad data at the boundary before it has a chance to silently corrupt a dbt model, skew a Feast feature, or train a model that looks fine in a notebook and quietly fails in production.

Common Mistakes People Make

Writing validation as a one-off script instead of a reusable Suite. This defeats the entire point of GX's shared-language design — a Suite committed to Git and reused across runs is genuinely the valuable artifact, not a disposable notebook cell.

Validating only after transformation, not at ingestion. Catching bad data as early as possible — right after extraction — prevents it from silently propagating through every downstream step before anyone notices.

Skipping Checkpoints and only ever running interactive validation. Interactive Expectation-building is for exploration and suite design; production pipelines need the repeatable, automatable Checkpoint object specifically.

Treating GX and dbt tests as redundant rather than complementary. Recall the distinction above — dbt's tests are lightweight and SQL-native; GX is the broader, data-source-agnostic validation layer. Many pipelines genuinely benefit from both.

Never generating Data Docs. Skipping this means your validation results live only in pipeline logs nobody reads regularly, rather than a shared, browsable record the whole team can reference.

  • Designing Data-Intensive Applications by Martin Kleppmann — the foundational text on data systems that directly informs why validation at system boundaries matters, covering consistency, integrity, and the failure modes bad data introduces.
  • Fundamentals of Data Engineering by Joe Reis & Matt Housley — covers data quality frameworks and validation patterns essential context for understanding where Great Expectations fits in the broader data stack.
  • Designing Machine Learning Systems by Chip Huyen — the definitive guide to production ML, including data validation's critical role in preventing silent training-data corruption.
  • Practical MLOps by Noah Gomez — hands-on coverage of Great Expectations, integration with Airflow pipelines, and the validation patterns that catch bad data before it reaches model training.

Wrapping This Up

Great Expectations turns "we assume the data looks right" into an actual, checkable, version-controlled claim — Expectations declare what valid data means, Suites bundle related checks together, and Checkpoints make the whole thing runnable as a repeatable step inside Airflow or any other pipeline, catching bad data at the boundary rather than letting it silently corrupt everything downstream. The GX Core 1.0+ rewrite genuinely simplified this workflow considerably compared to older tutorials' more verbose API.

Remember that GX and dbt's built-in tests solve overlapping but genuinely distinct problems — lightweight SQL-native checks versus a broader, data-source-agnostic validation framework — and that Data Docs are the underused feature turning validation results into shared institutional knowledge rather than logs nobody reads. FYI, this genuinely closes the full data-and-MLOps arc running through this series' recent articles: extraction, validation, transformation, feature consistency, versioning, orchestration, deployment gating, and post-deployment monitoring now form one complete, traceable pipeline rather than eight separate tutorials :)

Now go write a single Expectation — even just ExpectColumnValuesToNotBeNull on whatever field matters most — against the daily_sales_summary dataset from the very first ETL pipeline article in this arc. That one line is genuinely the smallest possible step from "we assume this data is fine" to "we actually check."

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles