explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the questions people are asking
  • What Iris-3B is
  • How it works
  • The negative result, in the authors' words
  • What it can do
  • How to try it
  • How it compares
  • What this means for what you build or pay
  • Limits and open questions
  • Summary
  • Related reading
← Back to blog

explainx / blog

Iris-3B: An Open Pixel-Space Image Model With No VAE, and an Honest Negative Result

Image Generation, Open Source AI, Diffusion Models, Computer Vision, Sperid Labs

Part of AI Image, Video and Voice

Sperid Labs released Iris-3B, an Apache 2.0 3B pixel-space text-to-image model with no VAE. How it works, how to run it, and what its paper concedes.

Oct 9, 2026·9 min read·Yash Thakker
add explainx.ai
go deep
Iris-3B: An Open Pixel-Space Image Model With No VAE, and an Honest Negative Result

Almost every popular image model works in a latent space. A VAE squeezes the image into a small grid, the diffusion model draws in that grid, and a decoder turns the result back into pixels. On October 8, 2026, Sperid Labs released Iris-3B, an open 3B model that skips the VAE and paints every pixel itself.

The release is interesting for a second reason. The authors tested the idea that pixel-space models should help on detail-heavy tasks, and their paper reports that the gain did not show up. This post covers how Iris-3B works, what it can do today, how to run it, and how to read the paper's claims. We read the Hugging Face model card and the arXiv abstract. We did not run the model, so the quality statements are the authors'.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: the questions people are asking

table · 2 cols
QuestionShort answer
What is it?A 3B pixel-space text-to-image diffusion transformer
License?Apache 2.0 for the model. The Qwen3-VL-4B text encoder keeps its own Apache 2.0 license
Does it need a VAE?No. It outputs pixels directly
Output size?About one megapixel, 1024 by 1024 by default
Hardware?NVIDIA GPU with CUDA. About 12 GB of weights for text-to-image
Does pixel space beat latent space?The paper says no significant improvement on depth or restoration
Quality claim?Matches Qwen-Image on OneIG at 1024 by 1024 under the official evaluators
Extras?Depth estimation and image restoration and upscaling in the same repo

What Iris-3B is

The model card describes "a 3-billion-parameter model that paints every pixel directly — no VAE, no latent space — and, fine-tuned, estimates depth and restores and upscales images."

The argument for pixel space is simple. A latent encoder is lossy. It can blur fine texture and bias the model toward smooth detail. If the network outputs pixels, that loss never happens. The card says the authors "explore pixel-space generative models as an alternative to vision foundation models such as DINOv2."

The paper is "Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning," by Hanqiu Li Cai and Chema Garabito of Sperid Labs, submitted to arXiv on October 7, 2026 (arXiv 2610.09450).

How it works

The model card gives these specs:

table · 2 cols
PropertyValue
Size3B parameters, plus a frozen 4B text encoder
ArchitectureDiffusion transformer with 8 dual-stream and 16 single-stream blocks, then a 4-block pixel head that turns each 16 by 16 patch back into pixels
Text encoderQwen3-VL-4B-Instruct, frozen
TrainingRectified flow, trained from scratch at 256, then 512, then 1024 pixels, then supervised fine-tuning at 1024 pixels, 665K steps in total
DefaultsCFG scale 3, 100 steps

The arXiv abstract adds that the pixel head follows the pixel-transformer (PiT) design from PixelDiT. The authors first ablated the prediction target and representation alignment at 256 by 256 pixels to decide what to scale.

Two routes to a pixel backbone

The abstract says the team tested both routes to a pixel-space model:

  1. Pretrain from scratch. This is Iris-3B.
  2. Convert a latent model. They converted a pretrained latent model, FLUX.2 Klein base 4B, to pixel space.

They fine-tuned both families for monocular depth and for restoration and super-resolution.

The negative result, in the authors' words

The abstract is unusually frank. It says: "We find no significant improvement from using a pixel-space generative prior."

The specifics:

  • Depth. Fine-tuned with one matched direct-regression recipe, Iris-3B is "level with the latent FLUX.2 Klein." The converted pixel FLUX.2 Klein falls behind it.
  • Restoration. On 4x DIV2K restoration, neither pixel model beats a latent FLUX.2 Klein fine-tune. The converted one trails slightly.
  • Text-to-image. "Iris-3B shows that pixel-space pretraining ... scales to 3B parameters and to text-to-image quality competitive with latent models, matching Qwen-Image on OneIG under the official evaluators at 1024 by 1024."

So the release does not prove that dropping the VAE wins. It shows that it works at 3B and costs nothing in OneIG quality against one strong open model. For readers who pick models, that is the useful reading: pixel space is a viable design, not yet a clear upgrade. The authors say they release weights and code "in the hope that they help pave the way for further work on pixel-space" models.

What it can do

Text to image

The card shows examples generated at native aspect ratios of about one megapixel with the default settings. Prompts range from studio portraits and film-grain photography to a Byzantine mosaic and night landscapes.

Iris-3B pixel-space image generation of Petra's Treasury carved into pink sandstone in morning light, a 3B model without a VAE

Sample output for the prompt "The ancient city of Petra with the Treasury carved into pink sandstone, morning light." Source: the Iris-3B model card by Sperid Labs (Apache 2.0), used with credit.

Depth estimation

Fine-tuned for monocular depth, the model predicts relative log depth in a single forward pass with an empty prompt, so no text encoder is needed. The card reports AbsRel 0.071 and delta1 0.946, as the mean over NYUv2, KITTI, ETH3D, ScanNet and DIODE, using the Marigold V2 protocol zero-shot. Depth is affine-invariant, not metric. You must fit a scale and shift to compare with ground truth.

Iris-3B monocular depth estimation turning a photo into a relative depth map, an image model without a VAE used as a vision learner

Left: input photo. Right: Iris-3B depth output. Source: Iris-3B model card, Sperid Labs.

Restoration and upscaling

A second fine-tune restores and upscales images in one adversarial step, using the HYPIR recipe, 10K steps and EMA weights. The card reports LPIPS 0.292 and PSNR 20.68 on DIV2K validation at 4x. Large outputs are tiled into overlapping 1024 by 1024 windows.

Iris-3B image restoration and upscaling comparing a degraded input to the restored output from a pixel-space diffusion model

Left: degraded input. Right: Iris-3B restoration. Source: Iris-3B model card, Sperid Labs.

table · 3 cols
TaskTrainingReported result
DepthDirect regression of relative log depth, 10K stepsAbsRel 0.071, delta1 0.946 (zero-shot mean)
Restoration and upscalingOne-step adversarial restoration, 10K stepsLPIPS 0.292, PSNR 20.68 (DIV2K validation, 4x)

How to try it

bash
git clone https://github.com/speridlabs/iris-3b.git
cd iris-3b
pip install -e .          # Python 3.11+, PyTorch 2.7.1+

hf download speridlabs/iris-3b --local-dir iris-3b --exclude "depth/*" "upscaler/*"

python scripts/sample.py --checkpoint iris-3b \
    --prompt "a red fox sleeping in fresh snow, golden hour"

The weights for text-to-image are about 12 GB. The depth and upscaler folders add about 12 GB each. The Qwen3-VL-4B-Instruct text encoder downloads on the first run, and the card says you need an NVIDIA GPU with CUDA.

The card's tuning tips:

table · 3 cols
SettingDefaultEffect
--cfg-scale3Higher values follow the prompt more literally but can look harsher
--steps100Fewer steps are faster but lose some detail
--seednoneFix it to repeat an image
--negative-promptnoneThings to avoid
--txt-filenoneOne prompt per line, for batches

Write prompts as plain descriptive English sentences. For depth and upscaling, run scripts/depth.py and scripts/upscale.py. A hosted demo exists on Hugging Face Spaces for trying both without a GPU.

A cost note: 100 steps at one megapixel with a 3B transformer plus a 4B text encoder is not light. The card gives no timing or VRAM figure, so measure on your card before you plan a batch job.

How it compares

table · 4 cols
Model or routeSpaceOpen weightsNote
Iris-3BPixelYes, Apache 2.0CUDA only in the card, 1 MP output
Qwen-Image (see our Qwen-Image 3.0 post)LatentSee the postThe Iris paper says it matches Qwen-Image on OneIG at 1024 px; the paper does not name a version
FLUX 3LatentSee the postThe Iris paper tests FLUX.2 Klein base 4B, an earlier FLUX family model
Moebius 0.2B inpaintingLatentPer its postSmall inpainting model, a different job
Google Nano Banana 2.5ClosedNoA hosted model you call by API

If you need production images today, a hosted model or a mainstream latent model still has more tooling around it: LoRAs, ControlNets, ComfyUI nodes. Iris-3B is for researchers, people who want to fine-tune a pixel-space backbone, and teams who care about fine texture and want an Apache 2.0 base.

What this means for what you build or pay

  1. No licensing fee. Apache 2.0 allows commercial use of the weights. Check the terms on any data you fine-tune with.
  2. A new fine-tuning base. The same backbone supports generation, depth and restoration with no architecture change. A team that needs all three can start from one model.
  3. Do not expect a quality jump. The authors found no significant gain from pixel space on depth or restoration. Choose Iris-3B for the design, not for a promised quality edge.
  4. Plan for GPU cost. A 3B model at 100 steps and about 12 GB of weights is a datacenter or high-end desktop job, not a laptop one.

Limits and open questions

  • Text and counts. The card says text rendering inside images and exact object counts are not always reliable.
  • Safety. Like every image generator, it "can produce inaccurate, biased or unsafe content." Review outputs.
  • Resolution. About one megapixel. Use an upscaler for larger prints.
  • Evaluation scope. The OneIG match is against Qwen-Image only. We have not seen a broad head-to-head with FLUX or other models.
  • Early days. The model page showed 18 downloads on the day we looked, so there is little community testing yet.
  • Hacker News. A story on the paper appeared on October 8 with no comments when we checked, so we have no developer reports to add.

Summary

Iris-3B is a clean, open, Apache 2.0 proof that a 3B text-to-image model can work without a VAE. Its own paper says the pixel-space prior did not improve depth or restoration. Treat it as a research base and a useful second option beside latent models, and run your own prompts before you commit.

Related reading

  • Qwen-Image 3.0 and richer authentic detail
  • FLUX 3 from Black Forest Labs
  • Moebius: 0.2B inpainting versus FLUX
  • Google Nano Banana 2.5 image model
  • PixelRAG: visual retrieval over screenshots

Primary sources: the Iris-3B model card, the arXiv paper 2610.09450, the code repository and the project page. The HuggingNews summary is here.

Specs, benchmarks and download counts are accurate as of October 9, 2026. Model cards change, so re-check before you build on them.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Jun 25, 2026

Krea 2 Technical Report: Open-Weights Image Foundation Model Built for Creative Exploration

Krea 2 lands in the top 10 of the Artificial Analysis text-to-image leaderboard and 2nd among independent labs. The 58-page technical report details how they got there: no synthetic training data, a PostgreSQL-backed data warehouse they call krablets, iREPA-accelerated pretraining, a custom DPO variant called STPO to prevent policy divergence, and an RL stage with four reward signals including a dedicated artifact detector.

Jun 20, 2026

Ideogram 4.0: Open-Weight Image Generation — How to Run, API & JSON Prompts (2026)

Ideogram 4.0 is the first open-weight frontier image model built for design work — production typography, bounding-box layout, and 2K photoreal output. This guide covers what shipped, benchmark numbers, and how to run it via API, CLI, and self-hosted inference.

Oct 9, 2026

Apple Normalizing Trajectory Models: 4-Step Image Generation

Apple Machine Learning Research published Normalizing Trajectory Models, a way to generate images in four steps while keeping an exact likelihood. This post explains the idea in plain language, the reported results, and what the paper does not claim.