Sam Austin AI

Core ML Tutorial: Deploy Machine Learning Models on iPhone

September 29, 2026 14 min read Sam Austin
Contents

Core ML machine learning model running on iPhone iOS on-device inference
Core ML machine learning model running on iPhone iOS on-device inference

Figure 1: The whole point of Core ML — the model runs here, in the user's hand, with the radio switched off

Cloud inference has a cost problem and a privacy problem, and your app users notice both. Every request round-trips to a server, burns your API budget, and ships someone's photo or voice clip off their device. Core ML fixes all three by running the model right there on the iPhone. Let's get one running.

I first tried this on a small image classifier, expecting a nightmare of format conversions. It wasn't. Once you understand the pipeline, it's genuinely one of the smoother "get ML into a real app" experiences out there.

What Core ML Actually Does

Core ML is Apple's framework for running machine learning models directly inside iOS, iPadOS, macOS, watchOS, and tvOS apps. Core ML provides a unified representation for all models, and your app uses Core ML APIs and the user's data to make predictions and even fine-tune models, entirely on-device.

Why does this matter beyond the obvious privacy win? Core ML optimizes on-device performance by leveraging the CPU, GPU, and Neural Engine while minimizing memory footprint and power consumption. Running strictly on-device also removes any need for a network connection, which keeps your app responsive even with no signal.

Ever had an app feature break because someone's on a plane or in a dead zone? On-device inference just doesn't have that problem.

This is the same trade our Edge AI for beginners guide walks through for smaller hardware: push compute to where the data already lives, and you get privacy, latency, and offline capability in one move. Core ML is just the Apple-flavored version of that bet, with better silicon behind it.

The Overall Pipeline

Here's the mental model before you touch any code:

  1. Train or obtain a model in PyTorch, TensorFlow, or another supported framework
  2. Convert it to Core ML format using the coremltools Python package
  3. Drag the converted model into Xcode
  4. Write Swift code to load the model and run predictions
  5. Test on simulator or device, then ship

Two separate worlds meet here: Python for conversion, Swift for deployment. Don't expect to skip the Python step just because you're an iOS-only developer. IMO, that conversion step is where most of the actual thinking happens.

If you already ship models through a server, you have a second option worth knowing about: the same PyTorch graph can often be exported to ONNX Runtime for cross-platform coverage, or to TensorFlow Lite if you also support Android. The decision is really about how deep into Apple's stack you want to go.

Installing coremltools

Everything starts with the conversion library.

pip install -U coremltools

This Python package converts models from third-party training libraries into Core ML format. It supports TensorFlow 1.x, TensorFlow 2.x, PyTorch, and several non-neural-network frameworks like scikit-learn. Beyond conversion, coremltools also lets you read, write, and optimize existing Core ML models, and verify predictions on macOS before you ever open Xcode.

That last part matters more than it sounds. Running a prediction check in Python splits your debugging in half: if the numbers are wrong there, the bug is in conversion. If they're right there and wrong in Swift, the bug is in your app code.

Converting Your First Model

Let's convert a PyTorch model, since it's the most common starting point these days. Here's the shape of a typical conversion script:

import coremltools as ct
import torch

# Load your trained PyTorch model
model = YourModelClass()
model.load_state_dict(torch.load("weights.pth"))
model.eval()

# Trace it
example_input = torch.rand(1, 3, 224, 224)
traced_model = torch.jit.trace(model, example_input)

# Convert to Core ML
mlmodel = ct.convert(
    traced_model,
    inputs=[ct.ImageType(shape=example_input.shape)],
)

mlmodel.save("MyModel.mlpackage")

The ct.convert() call is the Unified Conversion API, and by default it produces an ML Program — a newer model type that represents computation as programmatic instructions and gives you more control over precision. This behavior targets iOS 15 and macOS 12 or newer by default, and you can override it with a minimum_deployment_target argument if you need to support older devices.

Tracing has one sharp edge: a traced model only handles the control flow it actually saw in example_input. If your model branches on data, trace with representative inputs or fall back to torch.jit.script. Don't discover this in the simulator at 11pm.

If you don't have a trained model yet, start from something that already works — our image classification with transfer learning walkthrough builds a usable classifier in one sitting, and that checkpoint is exactly what this pipeline wants.

Handling Image Inputs Correctly

If your model expects images, tell coremltools that explicitly instead of feeding it raw tensors. Declaring ct.ImageType with the right bias and scale values matches whatever normalization your model expects, for example scaling pixel values into the -1 to 1 range that many mobile classifiers use:

mlmodel = ct.convert(
    traced_model,
    inputs=[ct.ImageType(
        shape=example_input.shape,
        scale=1 / 127.5,
        bias=[-1.0, -1.0, -1.0],
        color_layout=ct.colorlayout.RGB,
    )],
)

Getting this wrong is a classic bug: the model runs without errors, but predictions look nonsensical because the input scaling doesn't match what the model trained on. If the outputs are confidently wrong rather than crashing, check preprocessing first — it is almost always preprocessing.

For anything that isn't an image — sequences, embeddings, tabular rows — use ct.TensorType in exactly the same position, and declare the dtype explicitly rather than hoping the default lines up.

Adding Metadata Before You Ship

Once conversion finishes, spend a few minutes on metadata. It shows up directly in Xcode and makes the model self-documenting for anyone else touching your project:

mlmodel.author = "Your Name"
mlmodel.short_description = "Classifies product photos into categories"
mlmodel.version = "1.0"
mlmodel.license = "MIT"

mlmodel.input_description["image"] = "A 224x224 RGB image"
mlmodel.output_description["classLabel"] = "Predicted class label"
mlmodel.save("MyModel.mlpackage")

Future-you, six months from now, will not remember what var_20 means as an output name. Label everything now.

Bringing the Model into Xcode

This part is refreshingly simple. Drag the .mlpackage (or older .mlmodel) file straight into your Xcode Project Navigator. Xcode compiles it into a resource optimized to run on-device, and auto-generates a Swift class matching your model's name and inputs.

Requirements worth checking before you start: recent workflows commonly target Xcode 16 or newer and iOS 17 or newer, though your minimum deployment target depends on what you set during conversion. Match these on both ends, or you'll get a confusing mismatch error at build time.

One practical note: Xcode only runs on macOS. If you're standing up a machine for this work, a MacBook Pro covers both the simulator and the Neural Engine profiling tools you'll want later.

Running Predictions in Swift

For image models, Apple's Vision framework pairs naturally with Core ML:

import Vision

guard let vnModel = try? VNCoreMLModel(for: MyCustomImageClassifier().model) else {
    return
}

let request = VNCoreMLRequest(model: vnModel) { request, error in
    guard let results = request.results as? [VNClassificationObservation],
          let top = results.first else { return }
    print("\(top.identifier): \(top.confidence)")
}
request.imageCropAndScaleOption = .centerCrop

let handler = VNImageRequestHandler(url: imageURL)
try? handler.perform([request])

Vision handles image resizing and cropping for you, which saves you from writing that boilerplate by hand every time. For non-image models — tabular data, text, audio features — you call the generated model class directly with its typed input struct instead of going through Vision:

import CoreML

let model = try MyTabularModel(configuration: MLModelConfiguration())
let input = MyTabularModel.Input(feature1: 0.42, feature2: 1.70)
let output = try model.prediction(input: input)
print(output.classLabel)

Expect a naming quirk here: ML Programs generate an Input struct nested inside the model type, while older Neural Network models generate a flat Input struct. If the compiler can't find MyTabularModel.Input, that's the reason.

Shrinking Your Model Size

Mobile users notice a 200MB app download. Two techniques help here, and both are one or two lines of coremltools:

import coremltools as ct
from coremltools.optimize.coreml import OpPalettizerConfig, Palettizer

# Float16: the cheapest and most common win
mlmodel = ct.convert(
    traced_model,
    inputs=[ct.ImageType(shape=example_input.shape)],
    compute_precision=ct.precision.FLOAT16,
)

# 8-bit palettization: a more aggressive cut
config = OpPalettizerConfig.global_config(nbits=8, granularity="per_channel")
mlmodel = Palettizer(config).compress(mlmodel)
mlmodel.save("MyModel_8bit.mlpackage")
  • Float16 quantization: halves weight precision from 32-bit to 16-bit, cutting size roughly in half with minimal accuracy loss
  • Palettization: clusters similar weight values together and stores an index instead of the full value, useful for more aggressive compression

Treat those argument names as a starting point — coremltools has reshaped the optimization API across 7.x and 8.x releases, so check the docs for your pinned version. Test accuracy after either step: compression is never free, and the right amount depends entirely on how sensitive your specific task is to precision loss.

If palettization still leaves you too heavy, pruning removes weights outright and knowledge distillation trades a big model for a small one that behaves like it. Our quantization explainer covers what actually happens to the numbers when you drop precision.

Core ML vs. Cloud Inference

ApproachBest ForTrade-off
Core ML (on-device)Privacy-sensitive apps, offline use, low latencyModel size limits, device compute constraints
Cloud inferenceVery large models, frequently updated modelsNetwork dependency, ongoing API costs
HybridHeavy models with a fast local fallbackTwo code paths to maintain

If your model fits comfortably in a mobile app and doesn't need constant retraining pushed to users, on-device wins on privacy and responsiveness alone. FYI, nothing stops you from doing both — cloud for heavy lifting, Core ML for fast local fallback.

A Practical Decision Framework

Let me save you some research time with straightforward guidance:

  1. Shipping to iPhone only? Use Core ML. The Neural Engine access and Vision integration are not reproducible with a cross-platform runtime
  2. Shipping to iPhone and Android? Train once, then export to both — TensorFlow Lite or ONNX Runtime covers Android, Core ML covers iOS
  3. Model over about 50MB? Convert with compute_precision=ct.precision.FLOAT16, then palettize, then re-measure accuracy before anything else
  4. Conversion produces garbage predictions? The bug is preprocessing, nine times out of ten — diff your scale and bias values against the training pipeline line by line
  5. Model needs frequent updates? Keep the weights in the cloud and push only what changes; our MLOps for beginners guide covers the release mechanics around that

The mistake isn't picking the wrong row here — it's rewriting your inference stack for on-device when a 30-line conversion would have shipped this week.

Common Mistakes People Make

Mismatched preprocessing

Recall the image inputs section directly — your Core ML input scaling must exactly match what your model saw during training, or predictions silently degrade with no error anywhere.

Wrong deployment target

Recall the conversion section directly — setting minimum_deployment_target too high locks out users on older iOS versions, and too low sacrifices newer optimizations you paid for.

Skipping on-device verification

Recall the coremltools section directly — test predictions in Python before touching Xcode, so you know conversion succeeded before you start debugging Swift.

Ignoring model size

Recall the shrinking section directly — a model that "works" in testing may still be too heavy for a real download-and-install experience once it's inside the app bundle.

Forgetting the .mlpackage recompile

Recall the Xcode section directly — editing the model file without removing the generated Swift class from your build leaves you predicting against a stale copy.

Want to Go Deeper?

If you want structured practice on mobile and edge model deployment, Educative's ML and mobile courses run through hands-on labs for shipping models to devices. The unlimited plan is useful when you're working through several deployment targets in one stretch.

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is Core ML used for?

Core ML runs trained machine learning models directly inside an iOS, iPadOS, macOS, watchOS, or tvOS app. It covers image classification, object detection, text and tabular prediction, and audio models entirely on-device, using the CPU, GPU, and Neural Engine.

Do I need to know Swift to use Core ML?

You need Swift for the app side, but not for training. The model itself is trained or obtained in PyTorch, TensorFlow, scikit-learn, or ONNX, then converted with the Python coremltools package. Xcode generates the Swift class for you, so most of what you write is loading the model and calling it.

How do I convert a PyTorch model to Core ML?

Trace the model with torch.jit.trace on an example input, then pass the traced model to ct.convert() along with either ct.ImageType or ct.TensorType describing the input. Save the result as an .mlpackage file and drag it into Xcode.

What deployment target does Core ML require?

The Unified Conversion API produces an ML Program by default, which targets iOS 15 and macOS 12 or newer, and you can change that with the minimum_deployment_target argument. Xcode 16 and iOS 17 are the common development baseline, but your real minimum is whatever you set during conversion.

Does Core ML work without an internet connection?

Yes. Once the model ships inside your app bundle, inference needs no network at all, which keeps features working in airplane mode and removes both round-trip latency and per-request API cost.

How do I make a Core ML model smaller?

Float16 quantization roughly halves weight size for almost no accuracy loss, and 8-bit palettization goes further by storing an index per weight cluster instead of the full value. Pruning and knowledge distillation are the next levers if compression alone is not enough.

What is the difference between Core ML and TensorFlow Lite?

Both run models on-device. Core ML is Apple-only but gets deep access to the Neural Engine, the Vision framework, and first-class Xcode integration. TensorFlow Lite runs on Android, iOS, embedded Linux, and more, which matters if you ship cross-platform.

Wrapping This Up

Core ML turns a trained PyTorch or TensorFlow model into something that runs privately and responsively on a user's iPhone, using coremltools for conversion and Xcode for integration. The pipeline has exactly two real steps that matter: converting correctly, and matching preprocessing between training and deployment.

Will your first conversion attempt work perfectly? Maybe not, especially around input formatting. Convert in Python first, verify the numbers, and only then open Xcode — and once you see your own model running predictions locally on a physical device with zero network calls, it's hard to go back to shipping every inference request to a server :)

When the model works and starts pulling against your app's download budget, the model quantization guide and our Edge AI hardware roundup cover what to reach for next.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles