Apple Machine Learning Research has published Normalizing Trajectory Models (NTM), a new way to build image generators that work in just four steps. The paper, by Jiatao Gu, Tianrong Chen, Ying Shen, David Berthelot, Shuangfei Zhai and Josh Susskind, is listed under NeurIPS on Apple's research page and posted on arXiv as 2605.08078. The arXiv record shows a first submission on May 8, 2026.
The headline claim is simple to say and hard to do. Most fast image generators gain speed by giving something up. NTM tries to keep an exact likelihood over the whole generation path and still reach good images in four steps. This post explains the idea in plain language, lists the numbers the paper reports, and says what it does not show.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is it? | A few-step image generator that models each reverse step as a normalizing flow |
| How many steps? | Four sampling steps |
| What is new? | Exact likelihood training and four-step quality together |
| How good is it? | GenEval 0.82 from scratch at 256 px; 0.76 when fine-tuned at 512 px |
| Is it state of the art? | No. The authors cite a stronger diffusion model at 0.87 GenEval |
| Can I run it? | No code link on Apple's page at the time of writing |
| Is it a product? | No. It is research |
First, a quick refresher on how images get made
A diffusion model starts with pure noise and removes it in small steps until an image appears. The big assumption is that each small step is a simple Gaussian (bell-curve) change. When the steps are small, that assumption is fine. Our guide to how diffusion image generation works walks through that loop in more detail.

The cost is time. A strong model may take 20 to 50 steps, and every step is a full pass through a large network. So the field keeps asking the same question: can we do it in far fewer steps?
Why few-step generation is hard
If you try to take only four big steps instead of fifty small ones, the Gaussian assumption breaks. The authors explain that when a step is large, the true reverse change is a mixture of Gaussians, and a single Gaussian cannot capture it. The model then has to guess the wrong shape, and quality drops.
The usual fixes are distillation (train a student to copy a slow teacher in fewer steps) and adversarial training (use a discriminator to sharpen outputs). Both work. Both also give up the clean probabilistic accounting that makes a model a true likelihood model. In plain terms, you get a fast image maker, but you lose the ability to say exactly how probable an output is.
The NTM idea in plain language
NTM keeps the likelihood by changing what each step is. Instead of a single Gaussian per step, it models every reverse step as an expressive conditional normalizing flow.
A normalizing flow is a model built from invertible transformations. Because you can run it forward and backward without losing information, the math gives you an exact likelihood through the change-of-variables formula. A flow can also represent shapes far richer than one bell curve. That is what a large step needs.
Per the paper, the design has two parts:
- A transporter. A shallow invertible flow inside each step. The authors describe two TarFlow blocks of four layers each, with causal masks that alternate direction.
- A predictor. A deep Transformer of 24 layers that attends across the whole trajectory, with full non-causal attention along the trajectory dimension.
The split matters. The flow inside a step stays shallow, so each step is cheap to invert. The depth sits in the predictor, which looks across all steps in parallel. The authors say the whole network can be trained end to end, or initialized from a pretrained flow-matching model.

Self-distillation: getting speed back
An exact-likelihood trajectory model is not automatically fast to sample. So the team adds a self-distillation step. A lightweight denoiser network, also a Transformer with non-causal attention, learns to predict the clean output directly from the predictor's outputs. The paper reports about a 9 times speedup from this denoiser while keeping fidelity, with an LPIPS of 0.121.
The point is that the distillation here is a final, separate stage. The main model was still trained with exact likelihood. That is different from methods where distillation or adversarial loss is the whole training story.
What the paper reports
The setup and numbers below come from the arXiv paper.
| Setting | Resolution | GenEval | DPG-Bench | Steps |
|---|---|---|---|---|
| Trained from scratch | 256 x 256 | 0.82 | 79.64 | 4 |
| Fine-tuned from FLUX.2-klein-base-4B | 512 x 512 | 0.76 | 83.38 | 4 |
| STARFlow baseline | not stated here | 0.56 | not given | 256 |
Other details the paper gives:
- Training data. An internal text-image set of about 70 million pairs, including CC12M.
- Compute. Batch size 1024 on 64 H100 GPUs.
- Latent space. The from-scratch model works in an FAE latent space with 16 times compression.
- ImageNet. On class-conditional ImageNet at 256 x 256, NTM reaches an FID of 3.83 in four steps. STARFlow reaches 2.67 with 256 steps. Lower is better, so STARFlow is better on that metric, but it needs 64 times the steps.
Apple's summary says the method "matches or outperforms strong image generation baselines in just four sampling steps" on text-to-image benchmarks while keeping the exact likelihood over the generative trajectory.
What the paper does not claim
Read the limits as closely as the wins. The authors state several.
- It is not the top image model. They note that exact-likelihood training alone gives competitive results, not state-of-the-art ones. They cite Qwen-Image at 0.87 GenEval, above NTM's 0.82.
- One step does not work. With a single step (T equals 1), outputs are "severely degraded," because the transporter lacks the capacity to cover so large a jump. Four steps is a working point, not a floor.
- Hard prompts stay hard. Position and attribute binding, such as which color goes with which object, remain a challenge in the fine-tuned setting.
- The ImageNet number trails the slow baseline. An FID of 3.83 against 2.67 shows the price of using four steps.
Also remember the source type. This is a research paper with the authors' own benchmarks. We have not run the model, and no code link is on Apple's page. Treat the numbers as reported, not independently reproduced.
Why it matters
Speed without losing the math
Teams that serve image models care about cost per image. Fewer steps cut that cost directly. NTM shows a route to four steps that does not throw away the likelihood. If that holds up at larger scale, you get speed and a model you can score.
What an exact likelihood buys you
An exact likelihood is useful in concrete ways:
- Evaluation. You can compare models on how well they explain held-out data, not only on sample taste.
- Anomaly detection. Low-probability inputs can be flagged as unusual.
- Principled training signals. You can build on likelihood for later steps such as reinforcement learning or search over trajectories.
The paper itself does not test all of these uses. They are the general reasons people value likelihood models, not claims from the authors.
Normalizing flows are back in the conversation
Flows were popular years ago, then diffusion took over image generation. Papers like this one, which builds on TarFlow-style blocks and compares to STARFlow, suggest researchers still see room for flows when they are combined with the trajectory view. The few-step goal is a good fit, because flows handle the "mixture of Gaussians" problem that breaks plain diffusion at large steps.
How this compares to other speed-up work
Few-step and fast generation is a busy area, and it is not limited to images. We have covered several related efforts:
- Google DiffusionGemma applies diffusion to text and reports faster generation.
- Mercury 2 from Inception is a diffusion language model aimed at very high token rates.
- NVIDIA's TwoTower diffusion LLM work explores another route to parallel decoding.
- Ideogram 4 shows where open image generation stands today in practice.
The common thread is parallelism. Autoregressive and many-step methods are serial. These newer methods try to do more work per pass. NTM's twist is to do it while keeping an exact probability model.
How to follow up
You cannot run NTM today from Apple's page, but you can do useful things now:
- Read the paper. The arXiv abstract page is arxiv.org/abs/2605.08078. The experiments section has the tables and ablations.
- Watch for code. Check Apple's page and the arXiv listing for a repository or weights.
- Try few-step models you can run. Compare a many-step and a few-step open model on your own prompts. Our list of image generation prompts is a quick test set.
- Judge by your use case. If you need the best image quality, NTM is not claimed to be that. If you need a likelihood plus speed, it is worth tracking.
The bottom line
Normalizing Trajectory Models take a clear position: you can sample in four steps without giving up an exact likelihood, if each step is a rich normalizing flow instead of a single Gaussian. The reported scores are good, not best. The honest read is that this is a promising research direction with open questions about scale, hard prompts and release. We will update this post if Apple publishes code or weights.
Related reading
- How diffusion image generation works
- Google DiffusionGemma
- Mercury 2 diffusion LLM
- NVIDIA TwoTower diffusion LLM
- Ideogram 4 open image generation
- Top AI prompts for image generation
Sources: Apple Machine Learning Research and the arXiv paper.
