explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the questions people are asking
  • First, a quick refresher on how images get made
  • Why few-step generation is hard
  • The NTM idea in plain language
  • Self-distillation: getting speed back
  • What the paper reports
  • What the paper does not claim
  • Why it matters
  • How this compares to other speed-up work
  • How to follow up
  • The bottom line
  • Related reading
← Back to blog

explainx / blog

Apple Normalizing Trajectory Models: 4-Step Image Generation

Apple, Research, Image Generation, Diffusion Models, Machine Learning

Part of Microsoft, Apple and Amazon AI

Apple researchers propose Normalizing Trajectory Models (NTM): four-step image generation with exact likelihood. What the idea is, the results, and the limits.

Oct 9, 2026·9 min read·Yash Thakker
add explainx.ai
go deep
Apple Normalizing Trajectory Models: 4-Step Image Generation

Apple Machine Learning Research has published Normalizing Trajectory Models (NTM), a new way to build image generators that work in just four steps. The paper, by Jiatao Gu, Tianrong Chen, Ying Shen, David Berthelot, Shuangfei Zhai and Josh Susskind, is listed under NeurIPS on Apple's research page and posted on arXiv as 2605.08078. The arXiv record shows a first submission on May 8, 2026.

The headline claim is simple to say and hard to do. Most fast image generators gain speed by giving something up. NTM tries to keep an exact likelihood over the whole generation path and still reach good images in four steps. This post explains the idea in plain language, lists the numbers the paper reports, and says what it does not show.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: the questions people are asking

table · 2 cols
QuestionShort answer
What is it?A few-step image generator that models each reverse step as a normalizing flow
How many steps?Four sampling steps
What is new?Exact likelihood training and four-step quality together
How good is it?GenEval 0.82 from scratch at 256 px; 0.76 when fine-tuned at 512 px
Is it state of the art?No. The authors cite a stronger diffusion model at 0.87 GenEval
Can I run it?No code link on Apple's page at the time of writing
Is it a product?No. It is research

First, a quick refresher on how images get made

A diffusion model starts with pure noise and removes it in small steps until an image appears. The big assumption is that each small step is a simple Gaussian (bell-curve) change. When the steps are small, that assumption is fine. Our guide to how diffusion image generation works walks through that loop in more detail.

Three panels showing how diffusion noise to image generation resolves speckles into a clean shape

The cost is time. A strong model may take 20 to 50 steps, and every step is a full pass through a large network. So the field keeps asking the same question: can we do it in far fewer steps?

Why few-step generation is hard

If you try to take only four big steps instead of fifty small ones, the Gaussian assumption breaks. The authors explain that when a step is large, the true reverse change is a mixture of Gaussians, and a single Gaussian cannot capture it. The model then has to guess the wrong shape, and quality drops.

The usual fixes are distillation (train a student to copy a slow teacher in fewer steps) and adversarial training (use a discriminator to sharpen outputs). Both work. Both also give up the clean probabilistic accounting that makes a model a true likelihood model. In plain terms, you get a fast image maker, but you lose the ability to say exactly how probable an output is.

The NTM idea in plain language

NTM keeps the likelihood by changing what each step is. Instead of a single Gaussian per step, it models every reverse step as an expressive conditional normalizing flow.

A normalizing flow is a model built from invertible transformations. Because you can run it forward and backward without losing information, the math gives you an exact likelihood through the change-of-variables formula. A flow can also represent shapes far richer than one bell curve. That is what a large step needs.

Per the paper, the design has two parts:

  • A transporter. A shallow invertible flow inside each step. The authors describe two TarFlow blocks of four layers each, with causal masks that alternate direction.
  • A predictor. A deep Transformer of 24 layers that attends across the whole trajectory, with full non-causal attention along the trajectory dimension.

The split matters. The flow inside a step stays shallow, so each step is cheap to invert. The depth sits in the predictor, which looks across all steps in parallel. The authors say the whole network can be trained end to end, or initialized from a pretrained flow-matching model.

Streaming ribbon of green dots flowing into a bowl, like a four-step trajectory model gathering one path

Self-distillation: getting speed back

An exact-likelihood trajectory model is not automatically fast to sample. So the team adds a self-distillation step. A lightweight denoiser network, also a Transformer with non-causal attention, learns to predict the clean output directly from the predictor's outputs. The paper reports about a 9 times speedup from this denoiser while keeping fidelity, with an LPIPS of 0.121.

The point is that the distillation here is a final, separate stage. The main model was still trained with exact likelihood. That is different from methods where distillation or adversarial loss is the whole training story.

What the paper reports

The setup and numbers below come from the arXiv paper.

table · 5 cols
SettingResolutionGenEvalDPG-BenchSteps
Trained from scratch256 x 2560.8279.644
Fine-tuned from FLUX.2-klein-base-4B512 x 5120.7683.384
STARFlow baselinenot stated here0.56not given256

Other details the paper gives:

  • Training data. An internal text-image set of about 70 million pairs, including CC12M.
  • Compute. Batch size 1024 on 64 H100 GPUs.
  • Latent space. The from-scratch model works in an FAE latent space with 16 times compression.
  • ImageNet. On class-conditional ImageNet at 256 x 256, NTM reaches an FID of 3.83 in four steps. STARFlow reaches 2.67 with 256 steps. Lower is better, so STARFlow is better on that metric, but it needs 64 times the steps.

Apple's summary says the method "matches or outperforms strong image generation baselines in just four sampling steps" on text-to-image benchmarks while keeping the exact likelihood over the generative trajectory.

What the paper does not claim

Read the limits as closely as the wins. The authors state several.

  1. It is not the top image model. They note that exact-likelihood training alone gives competitive results, not state-of-the-art ones. They cite Qwen-Image at 0.87 GenEval, above NTM's 0.82.
  2. One step does not work. With a single step (T equals 1), outputs are "severely degraded," because the transporter lacks the capacity to cover so large a jump. Four steps is a working point, not a floor.
  3. Hard prompts stay hard. Position and attribute binding, such as which color goes with which object, remain a challenge in the fine-tuned setting.
  4. The ImageNet number trails the slow baseline. An FID of 3.83 against 2.67 shows the price of using four steps.

Also remember the source type. This is a research paper with the authors' own benchmarks. We have not run the model, and no code link is on Apple's page. Treat the numbers as reported, not independently reproduced.

Why it matters

Speed without losing the math

Teams that serve image models care about cost per image. Fewer steps cut that cost directly. NTM shows a route to four steps that does not throw away the likelihood. If that holds up at larger scale, you get speed and a model you can score.

What an exact likelihood buys you

An exact likelihood is useful in concrete ways:

  • Evaluation. You can compare models on how well they explain held-out data, not only on sample taste.
  • Anomaly detection. Low-probability inputs can be flagged as unusual.
  • Principled training signals. You can build on likelihood for later steps such as reinforcement learning or search over trajectories.

The paper itself does not test all of these uses. They are the general reasons people value likelihood models, not claims from the authors.

Normalizing flows are back in the conversation

Flows were popular years ago, then diffusion took over image generation. Papers like this one, which builds on TarFlow-style blocks and compares to STARFlow, suggest researchers still see room for flows when they are combined with the trajectory view. The few-step goal is a good fit, because flows handle the "mixture of Gaussians" problem that breaks plain diffusion at large steps.

How this compares to other speed-up work

Few-step and fast generation is a busy area, and it is not limited to images. We have covered several related efforts:

  • Google DiffusionGemma applies diffusion to text and reports faster generation.
  • Mercury 2 from Inception is a diffusion language model aimed at very high token rates.
  • NVIDIA's TwoTower diffusion LLM work explores another route to parallel decoding.
  • Ideogram 4 shows where open image generation stands today in practice.

The common thread is parallelism. Autoregressive and many-step methods are serial. These newer methods try to do more work per pass. NTM's twist is to do it while keeping an exact probability model.

How to follow up

You cannot run NTM today from Apple's page, but you can do useful things now:

  1. Read the paper. The arXiv abstract page is arxiv.org/abs/2605.08078. The experiments section has the tables and ablations.
  2. Watch for code. Check Apple's page and the arXiv listing for a repository or weights.
  3. Try few-step models you can run. Compare a many-step and a few-step open model on your own prompts. Our list of image generation prompts is a quick test set.
  4. Judge by your use case. If you need the best image quality, NTM is not claimed to be that. If you need a likelihood plus speed, it is worth tracking.

The bottom line

Normalizing Trajectory Models take a clear position: you can sample in four steps without giving up an exact likelihood, if each step is a rich normalizing flow instead of a single Gaussian. The reported scores are good, not best. The honest read is that this is a promising research direction with open questions about scale, hard prompts and release. We will update this post if Apple publishes code or weights.

Related reading

  • How diffusion image generation works
  • Google DiffusionGemma
  • Mercury 2 diffusion LLM
  • NVIDIA TwoTower diffusion LLM
  • Ideogram 4 open image generation
  • Top AI prompts for image generation

Sources: Apple Machine Learning Research and the arXiv paper.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 9, 2026

Iris-3B: An Open Pixel-Space Image Model With No VAE, and an Honest Negative Result

Sperid Labs released Iris-3B on October 8, 2026, a 3B text-to-image diffusion transformer that outputs every pixel directly, with no VAE and no latent space. The paper reports the model matches Qwen-Image on OneIG. It also reports that pixel space did not beat latent models on depth or restoration.

Jun 25, 2026

Krea 2 Technical Report: Open-Weights Image Foundation Model Built for Creative Exploration

Krea 2 lands in the top 10 of the Artificial Analysis text-to-image leaderboard and 2nd among independent labs. The 58-page technical report details how they got there: no synthetic training data, a PostgreSQL-backed data warehouse they call krablets, iREPA-accelerated pretraining, a custom DPO variant called STPO to prevent policy divergence, and an RL stage with four reward signals including a dedicated artifact detector.

Jun 23, 2026

Moebius: 0.2B Parameters, 10B-Level Inpainting, 15× Faster Than FLUX

A 0.22B model matching an 11.9B industrial giant on inpainting benchmarks is not a rounding error — it is a structural claim about what task-specific specialist models can do. Moebius achieves this via a novel attention block and latent-space distillation from PixelHacker. 26ms per step. Consumer hardware. Worth understanding.