explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR — what people are asking
  • The problem it targets
  • How the API works
  • What the benchmark measured
  • The headline results, in context
  • What it cannot do
  • How to evaluate it in an afternoon
  • A note on licensing and terms
  • What people are asking
  • Honest limitations
  • Bottom line
  • Related on explainx.ai
← Back to blog

explainx / blog

Adaption Labs Invent API: Generate Post-Training Data From a Prompt, With No Seed Examples

Synthetic Data, Fine-Tuning, Adaption Labs, Post-Training, AI APIs

Adaption Labs' Invent API generates SFT or preference datasets from a text description. Vendor results: 17% quality, 37% diversity at 20K rows. What to verify first.

Oct 3, 2026·9 min read·Yash Thakker
add explainx.ai
go deep
Adaption Labs Invent API: Generate Post-Training Data From a Prompt, With No Seed Examples

Adaption Labs has opened API access to Invent, a system that generates post-training datasets from a text description — no seed examples required. Its technical report, "Invent a Dataset: Measuring dataset generation abilities with zero seed data," claims about 17 percent higher quality and up to 37 percent more diversity at 20,000 rows than frontier models used as data generators, with 0.0 percent exact-duplicate prompts.

If those numbers hold on your task, they change what you can do without a data team. They are also vendor-run results, so the right move is to understand what was measured and test it on your own workload.

TL;DR — what people are asking

table · 2 cols
QuestionAnswer
What is it?A hosted API that creates SFT or preference datasets from a prompt
Who built it?Adaption Labs (paper authors include Sara Hooker)
Seed data needed?No — "zero seed data" is the point
Reported quality gain?About 17% relative over the strongest baselines
Reported diversity gain?19% at 2K rows; 37% at 20K rows
Duplicates?0.0% exact-duplicate prompts at all scales; baselines 6.2–19.2%
Output formats?JSONL, JSON, CSV or Parquet; instruction pairs or preference pairs
Limits?Text only; no tool-call or agentic trajectories
Price?Not in the report; credit-based per a third-party summary
Biggest caveat?Results are from the vendor's own benchmark
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

The problem it targets

Fine-tuning a model starts with data, and teams often have none. You know the behavior you want — answer legal questions in a particular style, handle banking support tickets, write Hindi news summaries — but there is no labeled dataset. The usual workaround is to prompt a frontier model to generate examples. That works at small scale and degrades at large scale: outputs converge, prompts repeat, and diversity collapses.

Adaption's paper frames this as the zero-data regime and builds a benchmark around it: given only a natural-language description, how good is the dataset a system can generate? For the fundamentals of why data quality drives fine-tuning results, see our guide to what fine-tuning is and the decision framework in grounding, RAG and fine-tuning.

How the API works

Based on the product description, the workflow is asynchronous:

  1. Submit a request through the datasets.invent call with a prompt of up to 10,000 characters, a domain code (broad like medical or narrow like medical.symptoms_diagnosis), a target dataset size, and an output type.
  2. Receive a dataset ID.
  3. Poll the job status until it completes.
  4. Download the result as JSONL, JSON, CSV or Parquet, structured as instruction pairs for supervised fine-tuning or preference pairs for preference optimization.

Billing is credit-based, with maximum row counts set by plan tier, and language expansion costs extra based on a sample_rate setting. There is no documented self-hosted option.

What the benchmark measured

The paper evaluates eight task types:

table · 2 cols
TaskWhy it is a hard test
African QAUnderrepresented in typical training data
Legal QAPrecision matters; hallucination is costly
Serbian car adsNarrow language and genre
Customer support, bankingDomain vocabulary and policy
Hindi newsNon-English, formal register
Customer support, generalBroad intent coverage
Medical QASafety-sensitive
Hotel reviewsStyle and sentiment variation

Diversity is scored on both semantics and wording: an embedding-based score called DCScore, plus n-gram diversity, compression ratio and exact-duplicate rate. Quality is scored on prompt quality, completion quality and relevance on a 0–10 scale, and by ranking models fine-tuned on the data using an LLM as judge.

The comparison set includes proprietary reference generators — Claude Opus 5, GPT-5.6 Sol and Gemini 3.1 Pro — and open-weight generators that are permitted for training, GLM-5.3 and DeepSeek V4 Pro.

The headline results, in context

  • Quality: roughly 17 percent relative gains over the strongest baselines.
  • Diversity: 19 percent relative gain at 2,000 samples, widening to 37 percent at 20,000.
  • Duplicates: 0.0 percent exact-duplicate prompts across scales, versus 6.2 to 19.2 percent for baselines, rising with dataset size.
  • Under constraints: a reported 58 percent advantage over the next best, GLM-5.3, when restricted to training-permissible generators.
  • Post-training: models fine-tuned on Invent data were ranked first on 54 percent of prompts for Llama-3.3-70B and 41 percent for Gemma-4-31B-it.

The diversity pattern is the most plausible and useful part. It matches the known failure mode: the gap is near parity at about 200 samples and grows with scale, which is exactly where naive prompting collapses. One third-party write-up notes that Claude produced less varied data at 20,000 than at 2,000 examples.

Why you should still be skeptical

These results come from the authors, on a benchmark they designed, scored partly by an LLM judge. That does not make them wrong, but three things deserve a check:

  1. Benchmark fit. A system built to win a zero-seed-data benchmark may be tuned to its metrics. Your task may weigh things the benchmark does not.
  2. Judge bias. LLM-as-judge scoring has known biases; the paper's post-training rankings inherit them.
  3. Baseline strength. A frontier model with a well-engineered generation pipeline — seeded personas, deduplication, rejection sampling — can close part of the gap. The comparison may use simpler prompting as the baseline.

What it cannot do

The authors are direct about limits, which is a good sign:

  • No tool-call traces or agentic trajectories. If you are training an agent to use tools, this is not that.
  • Text only. Multimodal extension is future work.
  • SFT and alignment data only. It does not produce harness datasets.
  • An open question. The authors note it is unclear whether this kind of data is best absorbed through weight updates at all, versus other approaches.

Generated data also needs validation. Coverage of the product warns that synthetic data requires checking for factual errors and unsafe content before production use. In a domain like medicine or law, that is not optional.

How to evaluate it in an afternoon

You can learn more from a small, honest test than from the report.

  1. Pick one narrow behavior you already know how to judge, such as a support-ticket classifier or a style-constrained summarizer.
  2. Generate a pilot set of 1,000 to 2,000 rows with Invent, and an equivalent set with your current method (a frontier model plus your best prompt).
  3. Measure duplicates and diversity yourself. Exact-duplicate rate and a simple embedding-spread check take minutes.
  4. Hand-review 100 rows from each. Look for factual errors, repeated templates and off-topic rows.
  5. Fine-tune a small model on each and test both on a held-out set you wrote yourself.
  6. Compare cost. Include credits, review time and re-generation.
  7. Scale only if it wins. The diversity advantage is claimed to grow with size, so a second test at 10,000 rows is worth the spend if the pilot is promising.

If you are building a post-training pipeline more broadly, our overview of open-source RL-as-a-service stacks shows where data generation fits next to training and evaluation, and Prime Intellect's Prime Inference launch covers the serving side of the same loop. For the distillation angle — and the terms-of-service questions around using one model's outputs to train another — read what AI distillation is.

A note on licensing and terms

The paper describes the endpoint as permitting downstream commercial use, and its comparison focuses on generators that are permissible for training. That matters because many frontier APIs restrict using their outputs to train competing models. If you generate data with a frontier model today, check its terms. If you use Invent, confirm that the license for the generated data matches your intended use and keep a copy of the terms with the dataset.

What people are asking

Does this replace human-labeled data?

No. It reduces the cold-start problem when you have no data. For high-stakes domains you still need expert review, and for evaluation sets you should use human-written data, not generated data.

Will models fine-tuned on synthetic data collapse?

Model collapse is a risk when models train repeatedly on their own outputs. Diversity is the main protection, which is why the paper emphasizes it. Mix synthetic data with real data when you can.

Is a prompt of 10,000 characters enough to describe a behavior?

It is a generous limit for a specification. Treat the prompt as a spec: include format, tone, edge cases and examples of what to avoid.

Can I use it for non-English tasks?

The benchmark includes Hindi news, Serbian car ads and African QA, which suggests multilingual intent. Language expansion is a billed option, so check cost per language.

How does it differ from asking GPT or Claude for examples?

Per the report, the difference shows up at scale: fewer duplicates and a diversity gap that widens past a few thousand rows. At small scale the report shows near parity.

Honest limitations

  • All performance numbers are from Adaption's own technical report.
  • Pricing is not published in the report; third-party summaries may be out of date.
  • I have not run Invent myself.
  • Data-handling terms were not reviewed in full.

Bottom line

Invent offers a hosted way to create SFT and preference data from a description alone. The claimed edge — fewer duplicates and a diversity gap that grows with dataset size — targets a real failure of prompt-based generation. Test it on a narrow task with your own held-out evaluation, review the rows by hand, and check the license before scaling.

Related on explainx.ai

  • What is fine-tuning? Complete guide
  • Grounding, RAG or fine-tuning: a decision guide
  • Open-source RL-as-a-service post-training stacks
  • What is AI distillation?
  • Prime Inference launch
  • AI evals for engineers and PMs
  • How to read AI benchmarks

Sources: Invent a Dataset technical report (arXiv) · AlphaSignal summary, October 1, 2026

Results and product details reflect Adaption Labs' report and third-party summaries as of October 3, 2026.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 11, 2026

Google Research ToolGrad: Answer-First Tool-Use Dataset Generation

Generating tool-use training data by guessing a user query and searching for a working API chain fails most of the time. Google Research's ToolGrad reverses the order — verify a real tool-use chain first, write the prompt after — and the resulting Gemma-3-12B fine-tune matches Gemini 2.5 Pro and Claude 4.5 Opus on function calling.

May 14, 2026

Adaption’s AutoScientist: Automating the Frontier of Model Training and Alignment

The end of manual fine-tuning: Discover the 3,000-word technical analysis of AutoScientist. From mitigating catastrophic forgetting to gradient-free interventions, we explore how Adaption Labs is democratizing frontier-level model ownership for the 2026 enterprise.

Oct 3, 2026

Nathan Lambert Launches Trillium Labs for Open Frontier AI

On October 1–2, 2026, Nathan Lambert and Tom Zick unveiled Trillium Labs, a nonprofit built to reopen frontier post-training as science: recipes, data, code, evals, and failed runs. No public model shipped on day one. This is what Interconnects readers and open-weights builders should actually do next.