Short answer: GenIA is a new research release from Meta and the University of Tuebingen that turns a single image, a handful of views, or a monocular video into a complete 3D object, without retraining any model. It keeps Meta's SAM 3D Objects generative model frozen and nudges its pose and appearance predictions to agree with your inputs while it runs. The code is public under a non-commercial license, the paper is on arXiv as 2610.12388, and there is a project page with interactive comparisons. It is a research drop, not a product: the repository has a single commit and 18 stars at the time of writing, and the written results are mostly qualitative.
If you work with 3D capture, game assets, robotics perception, or just follow the open 3D stack, the idea is worth understanding, because it addresses a real weakness of generative 3D: the object looks plausible but is not quite the object in your photo.
A wireframe cube with a green corner over a grid floor, representing a reconstructed 3D object in GenIA
TL;DR: GenIA at a glance
| Question | Answer |
|---|---|
| What does it do? | Reconstructs 3D objects (Gaussians and a mesh) from an image, multi-view images, or monocular video |
| Who made it? | Meta and the University of Tuebingen; authors include Stefano Esposito, Samuel Rota Bulo, Peter Kontschieder, Andreas Geiger, and Jonathon Luiten |
| Does it retrain a model? | No. SAM 3D Objects stays frozen; alignment happens at inference |
| License | Code under CC BY-NC 4.0 (non-commercial); SAM 3D Objects license must be accepted on Hugging Face |
| What do I need? | NVIDIA GPU, CUDA 12.8, conda |
| Does it include numbers? | Runtimes yes; raw quality metrics are not listed on the pages we read |
| Maturity | One commit on the release branch, 18 stars, 0 forks |
What problem does GenIA solve?
Generative 3D foundation models can complete geometry beyond what the camera saw. Give one a photo of a mug and it will invent a plausible back side. The paper's own framing is the catch: those outputs may not faithfully match the observed geometry, appearance, or pose.
GenIA's project page shows what that looks like. On a real photo, SAM 3D's unobserved side looks smooth and synthetic. Two other recent methods, TripoSplat and CUPID, match the visible side but wash out or drift on the back. GenIA with its refinement step keeps colors consistent with the input. The pitch is not a better generator, it is a generator that is held to the evidence.
If you want the foundation, Meta released SAM 3D Objects in November 2025 as a gated model on Hugging Face, predicting geometry, texture, and layout for objects from one RGB image, alongside Meta's SAM 3D announcement. GenIA is a layer on top of it.
How does test-time input alignment work?
The method has a few moving parts. All are described on the project page and in the abstract.
Shape. GenIA uses SAM 3D's coarse shape directly, or injects an external shape into SAM 3D's shape tokens.
Pose. It denoises the object's pose over that fixed shape. Translation and scale come from geometry measured in the scene, while the learned rotation prior is kept.
Appearance. Appearance is predicted as SLAT features, the latent representation SAM 3D uses, with three guidance ingredients applied during denoising: visibility-biased attention, cross-observation fusion, and differentiable rendering guidance. In plain terms, the model attends more to the pixels that actually saw each part of the object, merges information across views, and is pushed by a renderer toward images that match the inputs.
Optional refinement. A test-time refinement step adjusts the appearance latent, a low-rank decoder adapter, and the object placement against the inputs. "Test-time" here means during your run, not during training. The ablation demo on the project page shows an appearance ladder where each added component, attention bias, rendering guidance, then refinement, improves alignment, including on held-out views.
Dependencies show where the pipeline gets its geometry: the README credits MapAnything and MoGe for depth and camera estimates, ActionMesh for temporal shapes, gsplat for Gaussian splat rendering, and nvdiffrast for differentiable rasterization.
A green ball following a curved path through three empty frame outlines, representing object tracking across video frames
What about video and moving objects?
This is the part that makes GenIA more than a single-image tool. For dynamic input, it takes externally supplied temporal shapes and recovers a shared canonical appearance plus stable placement in the world frame, with per-frame deformation as an output. Today those shapes come from ActionMesh.
That design carries its own limits, which the authors state. Dynamic geometry is not grounded directly in the observations; it depends on the external predictor. Rotation stays prior-driven and is corrected only by ICP registration, which is sensitive to noise. Failures can come from wrong placement after ICP, from non-rigid deformation errors inherited from ActionMesh, or both.
On cost, the project page gives a clear comparison. SAM 3D alone has a median runtime of 25.3 seconds per object on an A100. GenIA on dynamic sequences costs about 13 times that single-image runtime, including ActionMesh (12 times without refinement). The two video-to-4D baselines it compares against, Lift4D and HiMoR, cost about 160 times and 250 times. If those ratios hold in your hands, the practical story is that GenIA is far cheaper than optimization-heavy 4D methods while remaining slower than a single feed-forward pass.
How do I run it?
The README gives the sequence. You need conda, a CUDA 12.8 toolkit, and an NVIDIA GPU.
git clone https://github.com/facebookresearch/GenIA
cd GenIA
bash install.sh
python -m genia.core.download_weights
python demo/run_demo.py image # or: multiview, dynamic
Before the weights download works, you must accept the SAM 3D Objects license on Hugging Face, and that repository is gated behind a request form. Results land in a results/ folder. Custom data goes under demo/data/<setting>/.
The configuration page shows a Hydra setup: python -m genia.core with key=value overrides and +experiment= presets. Notable options include the depth and camera source (map_anything as default, moge, or ground_truth), the dataset mode (image, mvcustom, dyncustom, and benchmark loaders such as gso and co3d), and per-block modes for pose, appearance, and fine-tuning. Two practical knobs help on smaller GPUs: reduce dataset.frame_stride or dataset.downscale_factor, and lower appearance_init_2.inference_steps (for example to 10) to speed the appearance pass. Per-block timings are written to final/timing.json. No VRAM figure is published, so plan on trial and error.
What is the evidence, and what is missing?
Here is a plain split between what is stated and what we could verify.
| Claim | What we found |
|---|---|
| Outperforms optimization-based, per-frame image-to-3D, and video-to-4D baselines | Stated in the abstract and project page; raw metric values not listed on the pages we read |
| Highest mean held-out CLIP-I on the dynamic benchmark (with refinement) | Stated on the project page without values |
| Metrics used | PSNR (foreground-masked on training views), LPIPS and CLIP-I (held-out), co-PSNR (held-out pixels visible from an input view) |
| Runtime ratios (13x, 160x, 250x) | Given on the project page; hardware A100 for the SAM 3D reference |
| Benchmarks | The abstract says only "synthetic and real benchmarks"; configs reference GSO and CO3D loaders |
So the direction is clear and the magnitude is not. Wait for the tables in the full paper or an outside reproduction before treating any specific gain as settled. This is a normal state for a release made days after submission; the arXiv entry is version 1, dated October 8, 2026.
A play triangle with an eye above it and green motion lines, representing video-based 4D object reconstruction
The license is the real constraint
Read this before you plan a product. GenIA's code is CC BY-NC 4.0, so it is non-commercial, and third-party components carry their own terms, several also non-commercial. The SAM 3D Objects weights come under their own license, which you must accept on Hugging Face and which Meta tells commercial users to review before use. In short: fine for research, teaching, and personal projects, and a legal question for anything that ships.
Where does this fit in the 3D and world-model stack?
GenIA is one more sign that the interesting work in 3D is moving from "generate something" to "generate something faithful." We have followed the neighbors:
- LingBot-Map reconstructs scenes from streaming input, the scene-level counterpart to GenIA's object-level focus.
- Tencent HY-World 2 and World Mirror generate navigable 3D worlds.
- Tencent WorldClaw wraps agents around open-world 3D generation.
- NVIDIA Cosmos 3 is the physical-AI world-model route.
- Our explainer on what world models are gives the vocabulary.
What this means for builders
- Asset capture: if you have photos of a real object and want a mesh that matches it on all sides you can see, test-time alignment is a more promising approach than raw generation. Use the multi-view setting when you can.
- Robotics and perception research: the per-frame pose and deformation outputs for video are the interesting part. The dependence on ActionMesh and ICP tells you where failures will occur.
- Do not skip the license check. CC BY-NC code on a gated, separately licensed foundation model is not a commercial pipeline.
- Set expectations on maturity. One commit, no forks, no VRAM guidance. Expect rough edges and file issues rather than assuming production behavior.
Open questions
- What are the actual PSNR, LPIPS, and CLIP-I values, and on which benchmarks?
- How much of the dynamic-sequence quality is GenIA, and how much is the ActionMesh shapes it consumes?
- Does the approach transfer to other frozen generative 3D priors, or is it tied to SAM 3D's SLAT representation?
- Will Meta release a commercially usable variant?
Details are accurate as of October 10, 2026 and come from the GenIA repository, project page, and arXiv abstract. Star counts and commit counts will change.
