← Back to Blog
GuideOn This Site

How to Run AirLLM on a Mac (Apple Silicon): A Step-by-Step Setup Guide From an AI CTO

A tested setup for AirLLM 2.11.0 on an M-series Mac: the exact install, an ungated Llama 3.1 download to an external SSD, the four Apple Silicon fixes in one wrapper script, real speeds, and the errors you will hit.

·9 min read

To run AirLLM on a Mac, install AirLLM 2.11.0 with Apple's MLX in a Python virtual environment, download a Llama model to a drive with plenty of free space, and run it through a short wrapper script that fixes four Apple Silicon bugs. It works on any M-series Mac, including a base 16GB Mac mini. It is also slow, so read the speed section before you download 141GB of weights.

I am a fractional CTO for AI-first startups, and "can we just run the model ourselves?" comes up in almost every engagement. Last week I tested AirLLM on a 16GB M4 Mac mini and published the timings. This post is the setup behind those numbers, so you can reproduce it.

The short version:

  • Install exact versions: airllm 2.11.0, mlx 0.29.3, transformers 4.46.3, torch 2.8.0, sentencepiece.
  • Download an ungated Llama 3.1 copy to your fastest drive with the most free space.
  • Run it through the wrapper script below. It fixes two crashes, one garbled-output bug and one display bug.
  • Expect about 20 seconds per token for an 8B model and 2.5 to 3 minutes per token for 70B.

A MacBook on a desk with code on the screen

What you need before you start

An Apple Silicon Mac on macOS 14 or newer, Python 3.9 or newer, and free disk space of about twice the model's size. Memory matters far less than you would expect. Disk space and disk speed matter far more.

  • A Mac with an M-series chip. AirLLM's Mac code runs on Apple's MLX framework, which only supports Apple Silicon. My machine was a base M4 Mac mini with 16GB of memory.
  • macOS 14 (Sonoma) or newer. The MLX version below only ships for macOS 14 and up.
  • Python 3.9 or newer. I used the Python 3.9 that comes with Apple's Command Line Tools. If python3 --version fails, run xcode-select --install first.
  • Disk space of about 2x the model. AirLLM keeps the download plus a second copy split into one file per layer. Llama 3.1 8B needs about 16GB twice, so 32GB. Llama 3.1 70B needs about 141GB twice, so plan for 300GB.
  • Ideally, a fast external SSD. My internal disk had only about 35GB free, so the models lived on a 2TB Crucial X9 over USB, rated at about 1GB/s. That drive's speed sets the model's speed.

Step 1: Install AirLLM and the pinned versions

Create a virtual environment and install these exact versions; they are the ones from my working run.

python3 -m venv ~/airllm-env
source ~/airllm-env/bin/activate
pip install --upgrade pip
pip install "airllm==2.11.0" "mlx==0.29.3" "transformers==4.46.3" "torch==2.8.0" sentencepiece

Why these packages:

  • mlx and sentencepiece are not declared dependencies of AirLLM 2.11.0, but its Mac code imports both. Without them, it fails at import.
  • Pin airllm to 2.11.0. AirLLM 3.x and 4.0 came out between June and September 2026 and need Transformers 4.49 or newer. I have not tested them on a Mac. I checked the 4.0.0 Mac code, and the rope_theta bug from fix 3 below is still there, so the fixes likely still apply. Get one working run on 2.11.0 first, then experiment.

Step 2: Download a model to your external SSD

Download only the safetensors weights and the config and tokenizer files, into a folder on the drive with the most free space. Start with TinyLlama so you can test the whole pipeline in minutes.

# Smoke test: about 2GB, single-file model
hf download TinyLlama/TinyLlama-1.1B-Chat-v1.0 \
  --include "*.safetensors" "*.json" "*.model" \
  --local-dir /Volumes/YourSSD/models/tinyllama-1.1b-chat

# The real test: about 16GB
hf download unsloth/Meta-Llama-3.1-8B-Instruct \
  --include "*.safetensors" "*.json" "*.model" \
  --local-dir /Volumes/YourSSD/models/llama-3.1-8b-instruct

A few notes:

  • Replace /Volumes/YourSSD with your own drive.
  • I used Unsloth's copy of Llama 3.1, not Meta's. It holds the same weights and does not need an approval, so no Hugging Face login. Meta's licence still applies. To use Meta's gated repo (meta-llama/Llama-3.1-8B-Instruct) instead, accept the licence on its Hugging Face page, wait for approval, and run hf auth login first.
  • hf comes with the install above. On older setups the same command is huggingface-cli download.
  • The include filter matters. It skips extra files, such as Meta's original checkpoint, that would double the download.
  • Budget time. I saw about 26MB/s from Hugging Face. A 70B download takes over an hour on its own.

Step 3: Understand the four Apple Silicon fixes

AirLLM 2.11.0 runs on a Mac, but out of the box it has two crashes, one silent bug that garbles Llama 3 output, and one display problem. The wrapper script in step 4 fixes all four at runtime, so a reinstall never undoes them.

Fix 1: Create the missing index file

AirLLM only loads models that come with a model.safetensors.index.json file. Single-file models like TinyLlama do not have one, so AirLLM fails. The script builds the index from the tensor names in the file.

Fix 2: Create the layer folder first

AirLLM checks free disk space on the layer folder before it creates that folder. If the folder is new, it crashes with "No such file or directory". The script creates the folder first.

Fix 3: Use Llama 3's real rope_theta

This is the important one. AirLLM's Mac code never reads rope_theta (the base for position encoding) from the model config, so it falls back to 10,000. Llama 3.x uses 500,000. Nothing crashes; the model just writes confident nonsense. The script copies the right value from the config.

Fix 4: Stop at the end of the answer

Generation never stops on its own. It always runs to the token limit, so the answer is followed by junk. The script cuts the output at the end-of-answer marker.

Step 4: Save and run the wrapper script

Save this as airllm_mac.py. It applies the four fixes, runs one prompt, and prints the timing. It leaves the installed library untouched.

# airllm_mac.py: run AirLLM 2.11.0 on Apple Silicon with four fixes
import json
import sys
import time
from pathlib import Path

import mlx.core as mx
from safetensors import safe_open

import airllm.airllm_llama_mlx as airllm_mlx
from airllm import AutoModel

MODEL_DIR = Path(sys.argv[1])   # folder you downloaded the model into
LAYER_DIR = Path(sys.argv[2])   # folder for AirLLM's per-layer files
PROMPT = sys.argv[3] if len(sys.argv) > 3 else "What is the capital of France? Answer in one sentence."
MAX_NEW_TOKENS = 20

# Fix 1: single-file models need a generated index file
index_file = MODEL_DIR / "model.safetensors.index.json"
single_file = MODEL_DIR / "model.safetensors"
if not index_file.exists() and single_file.exists():
    with safe_open(str(single_file), framework="pt") as f:
        weight_map = {name: "model.safetensors" for name in f.keys()}
    index_file.write_text(json.dumps({"metadata": {}, "weight_map": weight_map}))

# Fix 2: the layer folder must exist before the free-space check
LAYER_DIR.mkdir(parents=True, exist_ok=True)

# Fix 3: use the model's own rope_theta (500000 for Llama 3.x)
_original_args = airllm_mlx.get_model_args_from_config

def args_with_rope_theta(config):
    args = _original_args(config)
    args.rope_theta = float(getattr(config, "rope_theta", None) or 10000)
    return args

airllm_mlx.get_model_args_from_config = args_with_rope_theta

model = AutoModel.from_pretrained(str(MODEL_DIR), layer_shards_saving_path=str(LAYER_DIR))
tokenizer = model.tokenizer

messages = [{"role": "user", "content": PROMPT}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
input_ids = tokenizer(text, return_tensors="np", add_special_tokens=False)["input_ids"]

start = time.time()
output = model.generate(mx.array(input_ids), max_new_tokens=MAX_NEW_TOKENS)
elapsed = time.time() - start

# Fix 4: cut the output at the end-of-answer marker
for stop in (tokenizer.eos_token, "<|eot_id|>", "<|end_of_text|>"):
    if stop:
        output = output.split(stop)[0]

print(output.strip())
print(f"{MAX_NEW_TOKENS} tokens in {elapsed / 60:.1f} min, {elapsed / MAX_NEW_TOKENS:.1f} s per token")

Run the smoke test first, then the real model:

python airllm_mac.py /Volumes/YourSSD/models/tinyllama-1.1b-chat /Volumes/YourSSD/airllm-layers/tinyllama
python airllm_mac.py /Volumes/YourSSD/models/llama-3.1-8b-instruct /Volumes/YourSSD/airllm-layers/llama-8b

The first run of each model splits it into per-layer files before it generates anything. For the 8B model that took 43 seconds for 35 pieces. Later runs reuse those files and start generating straight away.

If it works, both models answer "The capital of France is Paris."

What speed to expect

About 1 second per token for a 1B model, about 20 seconds per token for 8B, and 2.5 to 3 minutes per token for 70B, on a drive that reads at about 1GB/s. AirLLM reads every layer from disk for every token, so the disk sets the pace, not the chip.

Seconds per token is roughly the model's size divided by your disk's read speed. On my Mac mini, Llama 3.1 8B read about 16GB per token at an effective 0.8GB/s: 19.6 seconds per token.

My measured numbers, with the prompt "What is the capital of France? Answer in one sentence." and 20 new tokens:

  • TinyLlama 1.1B: about 1 second per token. The model is small enough that macOS kept it in its file cache.
  • Llama 3.1 8B: 19.6 seconds per token. 6.5 minutes for 20 tokens, with the correct answer.
  • Llama 3.1 70B: projected at 140 to 180 seconds per token. That is around 30 minutes for one sentence. I did not run it; the 8B numbers make the math clear.

The only lever that moves these numbers much is disk speed. A Thunderbolt or USB4 NVMe drive at about 3GB/s should roughly triple the speed of the USB drive I used. The full breakdown is in my AirLLM test results.

Troubleshooting: common AirLLM errors on a Mac

Almost every failure maps to one of the four fixes, a missing package, or disk space. Find your error message below.

  • "Found local directory … but didn't find downloaded model", then "HFValidationError: Repo id must be in the form…" Your model folder has no index file, so AirLLM treats the folder path as a Hugging Face repo name. Fix 1 creates the index. Also check that the path points at the folder that holds the .safetensors file.
  • "model.safetensors.index.json should exist." Same cause, but you passed a Hugging Face repo name instead of a local folder. Download the model yourself (step 2) and pass the folder.
  • "No such file or directory" while loading. The layer folder does not exist yet. Fix 2 creates it.
  • Fluent but wrong or garbled output from a Llama 3 model. The rope_theta bug. Make sure fix 3 runs before AutoModel.from_pretrained.
  • The right answer followed by repeated or unrelated text. Generation ran to the token limit. Fix 4 trims it.
  • NotEnoughSpaceException. The layer folder's drive has less free space than the model. Free some space or point the layer folder at a bigger drive.
  • "No module named sentencepiece" or "No module named mlx". Install them into the same virtual environment you run the script from.
  • pip pulls in AirLLM 3.x or 4.x. You installed without the version pin. Reinstall with "airllm==2.11.0" and "transformers==4.46.3".
  • An assertion about bitsandbytes when you pass compression="4bit". AirLLM's 4-bit and 8-bit options need bitsandbytes, which needs an NVIDIA GPU. On a Mac you run full-precision layers.
  • A 401 or 403 error while downloading. You are using Meta's gated repo. Accept the licence, wait for approval and log in, or use the Unsloth copy.
  • Errors with Qwen, Mistral or other non-Llama models. On a Mac, AirLLM 2.11.0 sends every model through its Llama code, whatever the architecture. Stick to Llama-architecture models.

One gap remains after all four fixes: Llama 3.1's long-context scaling is not implemented in the Mac code. Short prompts are fine; very long ones may degrade.

Should you run AirLLM on a Mac at all?

Only when you need a model that does not fit in memory and you can wait minutes per word. For anything that fits, Ollama or LM Studio is hundreds of times faster. On the same 16GB Mac mini, an 8B model at 4-bit through Ollama runs at roughly 25 to 30 tokens per second, based on published reviews of similar machines.

  • Good uses: overnight batch jobs with a large model, and checking a large model's quality on your own prompts before you pay for hardware or API capacity.
  • Bad uses: anything interactive. Chat, coding assistants and agents need tokens per second, not seconds per token.

That decision (local model or API, which model, on what hardware, at what monthly cost) is one I make with founders regularly as a fractional CTO. If your team needs a second opinion on model hosting without a full retainer, a technical advisor engagement or CTO as a service covers it. If your startup is deciding who should own AI decisions at all, read what a fractional AI officer does, and if you are raising, my notes on AI startup technical due diligence cover the hosting questions investors ask.

Frequently asked questions

Does AirLLM work on Apple Silicon? Yes. On a Mac, AirLLM runs each layer on the GPU through Apple's MLX framework. Version 2.11.0 needs four fixes, including one without which Llama 3 models produce garbled output. All four fit in a small wrapper script.

How do I install AirLLM on a Mac? Create a Python virtual environment and pip install airllm 2.11.0, mlx, transformers 4.46.3, torch and sentencepiece. MLX and sentencepiece are not declared dependencies of AirLLM 2.11.0 but are required on a Mac.

Should I use the newest AirLLM version? Not for your first run. AirLLM 3.x and 4.0 were released in 2026 and need newer Transformers. I have not tested them on a Mac, and the 4.0.0 Mac code still has the rope_theta bug. Start with 2.11.0, which is tested here.

Can I run a 70B model on a 16GB Mac with AirLLM? Yes, if you have about 300GB of free disk. Expect around 2.5 to 3 minutes per token on a 1GB/s external SSD, which means roughly half an hour for a one-sentence answer.

Can I use 4-bit compression with AirLLM on a Mac? No. AirLLM's compression depends on bitsandbytes, which needs an NVIDIA GPU. On a Mac it runs full-precision layers, which is part of why it is slow.

Which models work with AirLLM on a Mac? Llama-architecture models in safetensors format, such as Llama 3.1 and TinyLlama. On a Mac, AirLLM 2.11.0 routes every model through its Llama code, so other architectures are likely to fail.

Is AirLLM faster than Ollama on a Mac? No. For any model that fits in memory, Ollama is hundreds of times faster, because it keeps 4-bit weights in memory instead of reading full-precision weights from disk for every token.

If you are weighing self-hosted models against an API for your product, book a 30-minute call and walk me through the workload. I will tell you what I would run, where, and what it should cost. You can also get in touch if you prefer to start by email.

Written By

Kunal Vohra

Kunal Vohra

Technical Co-Founder & Fractional CTO

I've co-founded 6+ startups across India, the UAE, and the US, spanning AI, Web3, fintech, and cybersecurity. I write about the technical and strategic decisions that determine whether a startup thrives or stalls.

Comments

Loading comments…

Enjoyed this article?

More writing on AI, Web3, and building startups.