A paper posted to arXiv this week argues that you can link a vision model and a language model without ever showing them a single image-caption pair. Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data (arXiv 2610.09411, v2 submitted October 8, 2026) comes from Dominik Schnaus, Thomas Dagès, Daniel Cremers, Xi Wang and Phillip Isola. Its claim: independently trained embedding spaces share enough structure that one orthogonal map can line them up, using nothing but the two unlabeled sets of vectors.
If you have followed embeddings news, this builds on the idea behind vec2vec and universal geometry in embeddings, now pushed across modalities and across very different models, such as the self-supervised vision encoder DINOv2 and the language model Qwen3.
TL;DR
| Question | Answer from the paper |
|---|---|
| What does it need? | Two unpaired sets of embeddings. No pairs, no class labels, no shared anchors. |
| What is the method? | Wasserstein Procrustes with a coarse geometric initialization; one orthogonal map. |
| How well does it work? | Zero-shot top-1 of 47.4% on CIFAR-10 unpaired, per the authors. CLIP, trained on pairs, is far better. |
| Does a few pairs help? | Yes. With DINOv2 ViT-B/14 and Qwen3-8B on MS COCO, up to 20 pairs gives 14x to 28x lower error than the strongest baseline. |
| Can I run it? | Yes, MIT-licensed code on GitHub with a Colab notebook. |
| Biggest caveat? | Works only when geometry is shared; alignment is coarse; no way to predict success purely from unpaired data. |
Why this matters: the Platonic idea, made operational
Phillip Isola is a co-author of the Platonic Representation Hypothesis line of work, the idea that large models trained on different data and objectives drift toward similar internal representations. That is a statement about similarity. This paper turns it into a tool: if the spaces really are similar in shape, you should be able to rotate one onto the other without any supervision.
Why anyone would want that is practical. Paired data is the expensive part of multimodal AI. CLIP-style training needs hundreds of millions of image-text pairs. In science, the pairs often do not exist at all, for example matching single-cell RNA to chromatin accessibility, or brain scans to the images a person looked at. A method that aligns two modalities from unpaired samples is a way into those domains.
How it works, step by step
The recipe, as described in the paper, has three stages.
Two groups of embedding points, one dense and one sparse, showing shape structure that an embedding alignment can match
- Coarse initialization. Repeatedly take random batches from each modality and cluster them with k-means (30 centers, 30 repetitions). Match the cluster centers across modalities by maximizing linear CKA, a measure of how similar two sets of relationships are. This is a quadratic assignment problem, solved with a GRASP heuristic. Averaging the results gives a rough transport plan.
- Pseudo-pairs. Turn that plan into guessed one-to-one pairs with linear assignment on random blocks.
- Refinement. Alternate between re-assigning pairs and solving the orthogonal Procrustes problem, which finds the best rotation given the current pairs. The output is a single orthogonal matrix.
Because the map is orthogonal, it preserves distances and angles inside each space. That is a strong constraint, and it is why the method needs the two spaces to be shaped alike. If you do have a few known pairs, they feed into both the initialization and the Procrustes step, so the same algorithm covers the unpaired and few-pair regimes.
What the numbers say
All figures below are the authors' reported results, not independent replications.
- Vision-language, no pairs. Over 7 vision and 3 language models, the method beats the earlier mini-vec2vec baseline on 17 of 21 model pairs. Average FOSCTTM, an error measure where lower is better, is 0.154 versus 0.223. CLIP ViT-L/14, trained with paired data, scores 0.0006, so the gap to supervised models stays large.
- Zero-shot classification. Without any classification data, top-1 accuracy is 47.4% on CIFAR-10, with 30.1% top-5 on CIFAR-100 and 26.5% top-5 on ImageNet-100. For one model combination from prior work, CIFAR-10 top-1 rises from 37.3% to 69.1%.
- Mismatched datasets. Aligning MS COCO images with captions from a different dataset gives FOSCTTM 0.236, against 0.154 when both come from MS COCO. Alignment degrades but survives.
- Few pairs. With DINOv2 ViT-B/14 and Qwen3-8B on MS COCO, from 0 to 1,000 pairs, the method beats all baselines at every budget. At up to 20 pairs its error is 14x to 28x lower than the best baseline. The zero-pair model reaches 79.6% CIFAR-10 top-1, a level that only some baselines reach with 500 pairs.
- Beyond images and text. Across modality pairs spanning medicine, biology, chemistry, materials, astronomy, neuroscience and language, CKA explains 69% of the variance in alignment quality. Specialized comparisons include PBMC RNA-to-ATAC (0.0887 vs 0.1215 for SCOT+) and NSD fMRI (0.0332 vs 0.0353 for Platonic Brain).
The headline pattern: geometry similarity predicts success. When CKA is high, alignment works; when it is low, it does not.
The text-to-image demo, and why not to overread it
Noisy dots resolving into a clean shape, an analogy for a diffusion model conditioned on aligned embeddings
The authors also show a proof of concept: train a small diffusion model on ImageNet-1K conditioned on DINOv2 image embeddings, then feed it captions mapped into that image space by the unpaired aligner. The images have the right broad class and scene, and adding between 0 and 100 pairs sharpens details. The authors themselves call this a proof of concept, not a competitive system. The notable part is conceptual: a text-to-image pipeline whose language side was never trained on image-caption pairs.
For actual products, nothing here replaces CLIP-style retrieval or modern multimodal embedding models. If you need a production image-text embedding today, see how EmbeddingGemma 2, Perplexity's late-interaction multimodal embeddings and Cohere Embed 5 approach it with paired training.
How to try it
The code is public: dominik-schnaus/unpaired-rosetta, MIT licensed. Per its README, you clone with submodules, install with pixi (tested with Python 3.14 and PyTorch 2.14), then import WassersteinProcrustes from unpaired_rosetta. Call fit on two unpaired embedding sets, optionally with a few known pairs, then use similarity for cosine scores and transform to map images into the caption space. A Colab-ready notebook is included, and the authors also have a project page. The repo is new, with around 15 stars when we looked, so expect rough edges.
A practical test for your own data: embed 5,000 to 10,000 items from each modality, compute linear CKA between random batches first as a sanity check, run the aligner, then evaluate on a small held-out set of true pairs. If CKA is low, the paper suggests not expecting much.
How it compares with earlier unpaired methods
The paper positions itself against vec2vec, mini-vec2vec, SUE, STRUCTURE and SOTAlign. Two differences stand out. First, the earlier translation methods were mostly demonstrated between text encoders, where both sides already speak the same kind of language, while this work reports results across vision, language, single-cell, brain and other modality pairs without domain-specific design. Second, the initialization is explicit: instead of hoping an adversarial or gradient-based training loop finds a good basin, it builds a rough matching from cluster structure first.
The authors' own comparison tables are the right place to check the exact deltas. On the natural-language benchmark they list, the method matches mini-vec2vec and beats vec2vec, and on the biology and neuroscience benchmarks it is competitive with methods built specifically for those fields. That breadth, rather than any single record, is the paper's main contribution.
What a builder could do with it
Three realistic uses, all speculative until someone reproduces the results:
- Cheap zero-shot probes. If you already have a vision encoder and a text encoder, an unpaired map gives you a rough text-to-image search to bootstrap from, before you pay for any labeling.
- Few-pair fine-tuning. The few-pair results suggest that 20 labeled pairs can go much further than usual, which matters for niche domains with scarce captions.
- Diagnosing representation similarity. Even if you never deploy the map, the CKA relationship gives a quick way to ask whether two models organize your data alike.
What people are asking
Does this mean modalities are fundamentally the same?
No. It means that for the model pairs tested, the relational structure is similar enough to exploit. The authors say alignment works only when spaces share enough geometry and that noisy web text that does not describe visual scenes may share less. The mapping captures semantic clusters, often not exact sample correspondences.
Is this a privacy or security risk for vector databases?
It is a relevant question. Earlier work on translating embeddings without paired data raised exactly that concern, covered in our vec2vec post. This paper does not study attacks, so we will not claim it does, but it does widen the set of settings where unpaired alignment is feasible.
How does this relate to world-model and JEPA arguments?
It fits the broader debate on whether scaling separate models yields shared world representations. For one prominent view on what LLMs lack, see Yann LeCun on LLMs, AGI and JEPA. This paper is evidence about representation similarity, not about reasoning or AGI.
Limitations to keep in mind
- Reported by the authors. We have not rerun the experiments, and the appendices with compute costs and text-to-image metrics were not part of what we reviewed.
- Coarse alignment. Zero-pair accuracy is useful for zero-shot tasks but far from supervised models.
- No success predictor. There is no fully unsupervised way yet to know in advance if alignment will work.
- Early code. A small repo with few users.
Related reading
- vec2vec and universal geometry in embeddings
- EmbeddingGemma 2: open multimodal embeddings on device
- Perplexity pplx-embed-v2 late multimodal embeddings
- Cohere Embed 5 Pro and Fast
- Yann LeCun on LLMs, AGI and JEPA
- Tencent Youtu-Parsing Omni 5B
Paper details and repository stats reflect arXiv v2 and GitHub as of October 9, 2026.
