There is a persistent myth in the local AI community: to run a 70B model, you need 140GB of VRAM. A 4-bit quantized version requires ~40GB. Many consumer configurations cannot hold those weights entirely in GPU memory. Therefore, 70B at anything approaching full precision requires a multi-GPU server or an expensive cloud instance.
AirLLM breaks this assumption. Not by magic — by trading speed for accessibility in a principled way that the community has found genuinely useful (21,000+ GitHub stars).
The technique is not new. Loading model layers sequentially from disk was used in academic settings before AirLLM made it accessible. What AirLLM did was turn it into a clean open-source library that works with Hugging Face models and takes three lines of code.
The Core Insight: You Don't Need All Layers at Once
A transformer language model is a stack of attention layers. During a forward pass, the computation flows sequentially:
- Layer 1 processes the input, produces an output
- Layer 2 takes Layer 1's output, processes it, produces an output
- ... repeated through all layers
- Final layer produces the logits (predictions)
Standard inference loads every layer into GPU memory before any computation starts. This is fast — everything is resident — but requires the full model to fit in VRAM simultaneously.
AirLLM's observation: at any point in the computation, only one layer's weights are actually being used. The rest are just waiting.
So why load them all at once?
AirLLM's approach:
- Split the model into individual layer shards, save to disk
- During inference, load layer 1 → compute → free layer 1 from VRAM
- Load layer 2 → compute → free layer 2 from VRAM
- Continue through all layers
Peak VRAM usage: one layer's weights + working activations. For a 70B model with ~80 layers, each layer at FP16 is roughly 1.75GB. A 4GB GPU can handle that.
What This Costs: The Speed Trade-off
Every forward pass reads the full model from disk — sequentially. That I/O is the bottleneck.
| Hardware | 70B speed estimate | Comment |
|---|---|---|
| NVMe SSD + GPU | Measure on your machine | Storage and caching affect the result |
| SATA SSD + GPU | Measure on your machine | Compare warm and cold runs |
| HDD + GPU | Measure on your machine | Check whether the workload tolerates the delay |
The original speed estimates were not accompanied by a reproducible measurement. Use a timed run on your own checkpoint and hardware; full-VRAM and layer-loaded inference trade different memory and transfer costs.
This is the honest trade-off. AirLLM does not make 70B fast on a 4GB GPU. It makes 70B possible on a 4GB GPU.
The use cases where this trade-off makes sense:
- Batch generation tasks where latency doesn't matter (summarizing documents overnight)
- Low-frequency queries where the alternative is not running the model at all
- Development and experimentation where you need full-size model behavior
- Hardware-constrained environments (edge devices, old workstations)
Feel the trade-off in minutes, not tokens per second:
Quick Start
pip install airllm
from airllm import AutoModel
MAX_LENGTH = 128
model = AutoModel.from_pretrained("meta-llama/Llama-3-70B-Instruct")
input_text = ['Explain the trade-offs in distributed system design.']
input_tokens = model.tokenizer(
input_text,
return_tensors="pt",
return_attention_mask=False,
truncation=True,
max_length=MAX_LENGTH,
padding=False
)
generation_output = model.generate(
input_tokens['input_ids'].cuda(),
max_new_tokens=20,
use_cache=True,
return_dict_in_generate=True
)
output = model.tokenizer.decode(generation_output.sequences[0])
print(output)
Important: On first run, AirLLM downloads the model from Hugging Face and splits it into layer shards. This requires approximately 2x the model size in disk space during the splitting process (the original plus the shards). A 70B FP16 model needs ~280GB free during conversion. After conversion, set delete_original=True to reclaim half the space.
Getting 3x Faster With Compression
Layer-by-layer loading is slow because loading from disk is slow. Make the layers smaller, loading gets faster.
AirLLM's compression is block-wise quantization of weights only — not activations. This distinction matters:
Quantization schemes can target weights, activations, or both; their runtime tradeoffs differ. AirLLM's bottleneck is disk I/O, not compute. It can quantize only weights, which is easier to do accurately because you're not dealing with dynamic runtime values.
pip install -U bitsandbytes
pip install -U airllm
model = AutoModel.from_pretrained(
"meta-llama/Llama-3-70B-Instruct",
compression='4bit' # or '8bit' for a quality/speed balance
)
Reported speedup: ~3x. The maintainers report a small accuracy impact in their evaluation; compare your task before relying on that claim.
At four bits, 70B weight values alone are approximately 35GB before metadata and other storage overhead. Compression can reduce transfer work, but inspect actual shard sizes and profiling output before claiming a speedup.
Configuration Reference
model = AutoModel.from_pretrained(
"model-id-or-local-path",
compression='4bit', # '4bit', '8bit', or None
profiling_mode=False, # True to log per-layer timing
layer_shards_saving_path=None, # Custom path for shards
hf_token=None, # For gated models
prefetching=True, # Overlap load + compute (~10% speed boost)
delete_original=False, # Remove original after splitting
)
The prefetching parameter is the most impactful free optimization — it starts loading the next layer into a buffer while the current layer is computing. Enabled by default for Llama models.
Supported Models and AutoModel Detection
AutoModel.from_pretrained() detects the model class automatically. No need to import AirLLMLlama2 specifically.
Supported:
- Llama 2, 3, 3.1 (7B through 405B)
- Mistral and Mixtral
- QWen and QWen2.5
- ChatGLM (use AutoModel, not the Llama class)
- Baichuan
- InternLM
- All safetensors-format models in the Open LLM Leaderboard top 10
For gated models:
model = AutoModel.from_pretrained(
"meta-llama/Llama-2-7b-hf",
hf_token="your_hf_token"
)
MacOS: Unified Memory Changes the Math
Apple Silicon machines have unified memory — VRAM and system RAM are the same pool. A Mac Studio with 192GB unified memory can technically hold a 70B FP16 model in memory without AirLLM. A Mac with 16GB can't.
AirLLM works on Apple Silicon (M1/M2/M3 only). Install mlx and torch, ensure Python native is available (not Rosetta), and follow the project’s Apple Silicon example notebook.
The relevant consideration: a 4GB limit on Mac corresponds to using an old MacBook Air or a non-unified-memory setup. Most modern Apple Silicon Macs have 16–192GB unified memory, which changes when AirLLM is useful. On those machines, the layer-loading approach may still be worthwhile for models larger than available unified memory.
The Parametric Limit
AirLLM does not solve the problem for production serving. It solves the problem for individual inference where hardware is the constraint.
If you need:
- High throughput (dozens of concurrent users)
- Low latency (< 2 second response times)
- Reliable uptime
...AirLLM is not the answer. Those requirements need hardware that fits the model.
AirLLM's value is exactly where its trade-off is acceptable: personal use, development workflows, research environments, and batch tasks where the alternative is not running the model at all.
At 21,000+ stars, the community has clearly found that constraint useful often enough.
Test the resource budget before downloading a large checkpoint
Inventory GPU memory, system memory, free disk space, and the location where layer shards will be stored. The advertised low-VRAM path moves a constraint rather than eliminating it. Storage, temporary conversion space, activations, and transfer time still matter, and another application can consume the headroom your first test appeared to have.
Begin with a supported smaller model to confirm the environment. Then use a short prompt and a short output budget on the larger checkpoint. Record download and conversion time separately from inference time. A slow first run may include setup work that a warm run does not repeat, while a warm result can hide the cost of initializing a fresh machine.
The AirLLM repository is the reference for supported architectures, compression, and platform-specific examples. A model sharing a familiar family name is not proof that its newer architecture is supported by your installed version.
Keep weight precision and quality claims separate
Compression reduces stored weight size, but its effect on a particular task needs evaluation. Compare the same prompts and expected outputs before and after changing precision. Include tasks where exact structure or factual grounding matters, rather than judging only whether the generated text sounds fluent.
For a simple estimate, seventy billion parameters at four bits require roughly thirty-five billion bytes for weight values alone. Additional metadata and runtime storage add overhead. That arithmetic is an illustrative lower-level estimate, not a promise of a particular checkpoint's disk size or total memory requirement.
Measure completion time on your hardware instead of treating the table above as a guaranteed rate. Disk throughput, caching, model architecture, context length, and output length can change the experience substantially. Report the environment and measurement method alongside any speed result you share.
Choose batch work with a recoverable output
A document-summary job can save each completed result separately, making a slow run easier to resume. Include source identifiers so the summaries remain attached to the documents they describe. Set a bounded output length and keep failed items visible instead of silently omitting them from the final folder.
Do not assume that waiting longer fixes a quality problem. If the model invents a claim or misunderstands an input, inspect the prompt, source, and acceptance example. AirLLM changes the execution strategy; it does not make the underlying model's answer reliable by itself.
If your workload needs interactive responses, compare against a smaller model that fits comfortably in memory and an allowed hosted route. The practical choice is the path that completes your actual task within its quality, privacy, and time requirements, even if another path can technically load a larger checkpoint.
Save a known working environment before tuning
Keep the package version and checkpoint revision with your first successful output. Change one setting at a time and retain the previous configuration so you can recover from a failed experiment. If conversion runs out of disk space, inspect which artifacts were created before retrying; repeated partial downloads can consume the capacity you thought was available.
Delay cleanup of original weights until you know the transformed checkpoint loads and the output remains usable. A successful conversion message is only one part of that check. Preserve any files needed to reproduce your result, and distinguish reclaiming temporary storage from removing the checkpoint on which your saved configuration depends.
Related
- Petals resurfaces — why BitTorrent-style LLM inference still struggles
- AI tools directory — full landscape of local AI and inference tools
- AI skills registry — reusable AI workflows including LLM integrations
- Browse LLM tools — complete directory of language models and their capabilities
