Sam Austin AI
Practical machine learning & AI tutorials — MLOps, LLMs, local AI, deep learning and more.
-
Latest
Docker Compose for ML Projects: Multi-Container Development Setup
Docker Compose for ML projects: set up GPU-enabled multi-container development with health checks, watch mode, profiles, and resource limits.
-
SageMaker Pipelines Tutorial: End-to-End MLOps on AWS
Build reproducible, auditable ML workflows on AWS: SageMaker Pipelines tutorial covering the @step decorator, step classes, and FailStep gates.
-
Best RAM and Storage Upgrades for Local AI Workstations
Best RAM and storage upgrades for local AI workstations in 2026: how much memory 8B-14B models need, NVMe sizing, and how to shop through the DRAM shortage.
-
Best Laptops for Running Local LLMs and Edge AI Development
Best laptops for local LLMs in 2026: why memory bandwidth beats GPU name, how much memory 7B to 120B models need, and Apple vs NVIDIA vs AMD picks.
-
Speculative Decoding Explained: Speed Up Local LLM Inference
Speculative decoding explained: how draft models, n-gram lookup, Medusa, and EAGLE speed up local LLM inference 2-3x with mathematically identical output quality.
-
Offline AI Apps: Building Machine Learning Apps That Work Without Internet
Build offline AI apps: how WebLLM, Transformers.js, and ONNX Runtime Web run machine learning in the browser with WebGPU acceleration and WASM fallback.
-
Running Multiple Local LLMs: Model Switching and Routing
Run several local LLMs on one GPU: compare llama.cpp router mode, llama-swap, and LiteLLM for on-demand model switching, routing, and cloud fallback.
-
Benchmarking Local LLMs: Tokens per Second Across Hardware
Measure real local LLM speed: tokens per second and time to first token across GPUs, Apple Silicon, and CPU, plus how to benchmark your own setup.
-
Local Voice Assistants: Build a Privacy-First Alexa Alternative
Build a fully local voice assistant with Wyoming, Whisper, Piper, openWakeWord, Home Assistant, and Ollama — no cloud, no data leaving your home network.
-
Apple Neural Engine Explained: Optimizing Models for Apple Silicon
How the Apple Neural Engine works, which operations actually reach it, how to check compute unit assignment in Xcode, and how to design Core ML models it runs.