Sam Austin AI

Batch vs Streaming Data Pipelines: Which Does Your ML Project Need? (2026)

September 8, 2026 14 min read Sam Austin
Contents
Batch vs Streaming Data Pipelines
Batch vs Streaming Data Pipelines

Figure 9: Most ML pipelines are batch-train and streaming-serve — knowing which part needs which is the expensive insight

Here's a genuinely useful real-world architecture worth holding onto through this whole article: a fraud detection system at scale, built around one deliberate split — transaction data flows continuously into Kafka, but the actual model retrains only once a night, when Spark reads the last 30 days from the data lake. The freshly retrained model then gets pushed to a Flink cluster, which does the actual real-time scoring on every incoming transaction. Training is batch. Serving is streaming. Neither piece is "better" — they're solving genuinely different problems inside the same pipeline.

This is the honest answer to this article's title: it's rarely all-or-nothing. Recall every pipeline built across this series — the ETL article's Python script, Airflow's scheduled DAGs, dbt's warehouse transformations — all of that was implicitly batch. This article names that choice explicitly and shows you when streaming genuinely earns its real operational cost instead.

By the end of this guide, you'll understand the actual architectural difference between batch and streaming, know which tool fits which half of an ML pipeline, and have a concrete framework for deciding rather than defaulting to whichever sounds more modern. IMO, "streaming because it sounds more sophisticated" is genuinely one of the most expensive mistakes a data team can make without realizing it :)

The Actual Difference, Not Just the Buzzwords

Batch processing works with static, bounded data — a fixed dataset, processed as a complete unit, on a schedule.

Stream processing works with unbounded, continuously arriving data — a series of events with no predetermined beginning or end: credit card transactions, website clicks, IoT sensor readings, arriving and needing processing as they happen.

In batch processing, time isn't a major consideration — you're working with data that's already settled, complete, sitting still.

In stream processing, time is genuinely dynamic, and events can arrive late (network delays, retries), out of order, or after the fact as backfills — this single difference is why streaming systems need entirely different mental machinery (watermarks, event-time semantics) that batch systems never have to think about at all.

Recall the Airflow article's honest caveat: Airflow is explicitly not a streaming solution — it's built for batch-oriented, scheduled or event-triggered work. This isn't a limitation to work around; it's the correct tool for the genuinely large share of ML pipelines that don't need sub-second freshness at all.

Where Each Genuinely Fits in an ML Pipeline

This is the part most comparisons skip, and it's genuinely the most useful framing: most production ML systems aren't "batch" or "streaming" — they use batch for training and streaming (or neither) for serving.

Training is almost always batch, even in systems with real-time serving. Recall the fraud detection architecture above: Spark reads a full 30-day historical window nightly. There's genuinely no reason to stream training data — a model retrain is a scheduled, bounded operation by nature, exactly the kind of job the Airflow, dbt, and DVC articles from earlier in this series were built for.

Feature computation for real-time serving is where streaming actually earns its cost. Recall the Feast article's online store — if a feature like "transactions in the last 5 minutes" needs to reflect genuinely recent activity for a fraud model scoring a live transaction, that feature has to be computed from a stream, not a nightly batch job that's already hours stale by the time it matters.

Serving/inference itself can be either — a nightly batch scoring job producing tomorrow's recommendations is genuinely fine for many use cases; a fraud model needing to approve or block a transaction in milliseconds needs a streaming-fed, low-latency serving path instead.

A common, genuinely important point of confusion: Kafka is not a processing framework. It's a transport layer — an extremely high-throughput way to move events around — but it doesn't filter, transform, or route data on its own. You still need something like Kafka Streams, Flink, or Spark Streaming layered on top to actually do anything with what Kafka is moving.

Apache Kafka: The Transport Layer

Strengths: highest-throughput event transport available, massive ecosystem, genuinely mature and battle-tested.

Weaknesses: operational complexity, no built-in processing of its own, poor support for edge deployment, and a genuinely steep learning curve — running Kafka at the edge or across distributed sites means managing separate clusters each with their own coordination quorum.

Best for: the backbone connecting your event sources to whatever processes them — genuinely the plumbing, not the destination.

Apache Spark: Batch-First, Streaming Bolted On

Spark dominates batch processing for data engineering teams — its core strength is parallel processing of massive datasets across a cluster, and its unified DataFrame API genuinely bridges batch and streaming with familiar SQL/Python/Pandas-like constructs, making it easier to learn for teams already comfortable with the tools from earlier in this series.

Spark Structured Streaming extends this into real-time, but through a micro-batch model — processing short batches of input repeatedly rather than truly event-by-event. This introduces genuine latency, typically seconds rather than milliseconds, and a heavier resource footprint than Flink for equivalent throughput.

Best for: large-scale data transformations, analytics pipelines, and lakehouse patterns — genuinely the right default when your ML pipeline's training step is the actual bottleneck, not sub-second serving freshness.

Flink offers the strongest streaming guarantees available — exactly-once semantics, genuine event-time processing (not just arrival-time), and stateful computation designed specifically for continuous, indefinitely-running applications rather than bounded jobs.

This is the framework built from the ground up for streaming, with batch capability added later — the inverse of Spark's history, and it shows in the architecture: Flink's design focuses on externalized, incremental state and dynamic scaling for genuinely low-latency, cloud-native event-driven pipelines.

The cost is real: JVM overhead, cluster dependency, a genuinely steep learning curve, and production deployment requiring real expertise in checkpointing, state management, and backpressure handling — this is not a tool to adopt casually.

Best for: complex event processing, real-time aggregations, fraud detection logic, streaming ETL, and — genuinely relevant here — ML feature pipelines specifically needing sub-second freshness.

A Concrete Decision Framework

Rather than "streaming sounds modern, let's use it," here's the sequence worth actually working through.

Does your model training step need real-time data? Almost certainly no — recall the fraud detection architecture's nightly Spark retrain. Batch is correct here in the overwhelming majority of ML projects, full stop.

Does your serving/inference step need sub-second-fresh features? If a feature computed an hour ago (or even five minutes ago) is genuinely fine for your use case — most recommendation systems, most churn models, most demand forecasting — stay in batch. Recall Feast's own explicit "who Feast is not for" guidance: if you don't need very low latency feature retrieval, don't build for it preemptively.

If yes, is the freshness requirement seconds or milliseconds? Spark Structured Streaming's micro-batch model genuinely handles "seconds" freshness fine, with a dramatically gentler operational learning curve than Flink. Don't reach for Flink's genuine complexity unless you actually need sub-second, stateful, exactly-once guarantees — fraud detection, real-time bidding, live anomaly alerting are the genuine cases, not "it would be nice if this dashboard updated faster."

Do you have the operational capacity for the tool you're considering? Recall the explicit warning: running Flink in production requires real expertise in checkpointing and backpressure handling. A small team without dedicated streaming infrastructure expertise will genuinely struggle here regardless of how well-suited Flink is to the problem on paper.

If step 3 above genuinely points you toward real streaming, here's how to pick among the actual processing options.

Kafka Streams — a lightweight library, genuinely the right choice for simpler, Kafka-native transformations where you don't want to stand up an entirely separate cluster just to process what Kafka is already moving.

Spark Streaming — the right choice if your team already knows Spark from batch work and "seconds, not milliseconds" latency is genuinely acceptable — the learning curve is dramatically gentler than Flink's.

Flink — the right choice specifically when you need to manage complex application state over time, building genuinely cloud-native, event-driven, highly time-sensitive applications — real-time fraud scoring, live feature computation for a serving model, complex windowed aggregations with exactly-once guarantees.

A genuinely common, practical progression: start with Kafka Streams for simple, Kafka-native transformations, and migrate to Flink specifically once your pipeline's complexity outgrows what Kafka Streams comfortably handles — this transition is described as natural rather than a rip-and-replace, since both frameworks connect through the same Kafka backbone.

The Architecture Pattern Worth Actually Copying

Recall the fraud detection example from the opening — this is genuinely the reusable template for combining both approaches correctly.

Ingest (Kafka) → Train (Spark, nightly, batch) → Serve (Flink, real-time)
                        ↓
              Model pushed to serving layer

Batch and streaming systems typically connect through an intermediate storage layer — Kafka itself, or a data lake (S3/HDFS) — letting both systems operate on the same underlying data at genuinely different speeds without either one blocking the other. This is precisely the same "specialized tools, each doing one job" philosophy running through this series' entire data-and-MLOps arc — dbt doesn't orchestrate, Airflow doesn't transform, and here, Spark doesn't serve in real time and Flink doesn't own your nightly retrain.

Common Mistakes People Make

Reaching for streaming because it sounds more sophisticated, not because the use case demands it. Recall the decision framework above — most ML training genuinely doesn't need real-time data, and building streaming infrastructure for a use case that would be fine with nightly batch is pure operational overhead with no real benefit.

Assuming Kafka alone solves your streaming needs. Kafka is transport, not processing — you still need Flink, Spark Streaming, or Kafka Streams layered on top to actually do anything with the events it moves.

Choosing Flink without the operational capacity to run it well. Checkpointing, state management, and backpressure handling are genuinely non-trivial — a small team without dedicated expertise here will struggle regardless of how theoretically well-suited Flink is.

Treating micro-batch and true streaming as interchangeable. Spark Structured Streaming's seconds-level latency is genuinely fine for many use cases and wrong for others (live fraud scoring, real-time bidding) — know which category your actual requirement falls into before picking.

Forgetting that training and serving can (and often should) run on genuinely different pipelines. The fraud detection architecture's batch-train/streaming-serve split isn't a compromise — it's the correct design for the large majority of real-time ML systems.

  • Streaming Systems by Tyler Akidau et al. — the definitive text on stream processing theory, covering watermarks, event-time processing, and the exact guarantees that separate Spark's micro-batch model from Flink's true streaming.
  • Designing Data-Intensive Applications by Martin Kleppmann — foundational coverage of distributed systems, event sourcing, and stream processing architecture that directly informs the Kafka/Flink/Spark decision.
  • Kafka: The Definitive Guide by Neha Narkhede et al. — essential reading for understanding Kafka's role as transport layer and when Kafka Streams is sufficient versus when you need Flink or Spark on top.
  • Fundamentals of Data Engineering by Joe Reis & Matt Housley — covers batch and streaming architecture patterns in the context of modern data platforms, including the cost tradeoffs this article emphasizes.

Wrapping This Up

The honest answer to "batch or streaming" is genuinely "probably both, for different parts of your pipeline" — training is almost always batch, even in systems needing real-time serving, and streaming earns its substantial operational cost specifically when your serving layer needs sub-second-fresh features or event-driven logic that a nightly job structurally cannot provide. Kafka moves events, Spark and its micro-batch model handle the large majority of "streaming" needs that can tolerate seconds of latency, and Flink is the right, genuinely demanding tool specifically for true low-latency, stateful, exactly-once processing.

Remember that "streaming because it sounds modern" is a real and expensive mistake — recall every batch-oriented pipeline built across this series' Airflow, dbt, and DVC articles, which remain the correct architecture for the overwhelming majority of ML training workflows regardless of how the serving side is eventually built. FYI, this genuinely closes the full data-and-MLOps arc that's run through this series' recent articles — from the first plain-Python ETL pipeline through validation, transformation, feature consistency, versioning, warehousing, and now the batch-versus-streaming architecture decision sitting underneath all of it :)

Now go look at whatever ML project you've built furthest along in this series and ask honestly: does your serving path actually need streaming, or does it just need a batch job that runs more often? For the large majority of real projects, the honest answer is genuinely the second one.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles