OpenAI's alignment team published LASER on October 6, 2026: a pipeline that finds the rare conversations where a safety policy is genuinely hard to apply, using about 10,000 times less grader compute than random sampling. OpenAI says it can curate evaluation data within hours.
The name is Logistic Augmented Sampling over Embeddings, Recursively. The idea is simple enough to explain without a math degree, and it is reusable well beyond OpenAI.
TL;DR: LASER in one table
| Question | Answer from OpenAI's post |
|---|---|
| What problem does it solve? | Safety evals need examples of rare, subtle policy violations; random sampling almost never finds them |
| Core trick | A cheap classifier on embeddings learns from an expensive reasoning grader, then picks the next conversations to grade |
| Compute saved | About 10,000x less grader compute than random sampling for comparable disallowed examples |
| Hit rate | About 50 percent disallowed in sampled conversations vs about 1 in 20,000 at random |
| Time to build an eval set | Within hours |
| Data used | Synthetic and de-identified conversations only, no raw user data |
| Code or paper | None linked in the post |
Why finding bad conversations is the hard part
When a lab evaluates how a chatbot handles sensitive requests, the model-side test is cheap. The expensive part is the dataset. You need realistic conversations that sit near the edge of the policy: not obviously fine, not obviously prohibited, but cases where a careful judgment call is required.
Those cases are rare. OpenAI's own figure is that about one in 20,000 randomly sampled conversations is disallowed. To find 100 examples by random sampling you would read roughly two million conversations, and if each one needs a reasoning model with long chain-of-thought to grade, the bill is enormous. The alternative, human reviewers, is slow and hard to repeat each time a policy changes.
This is the same structural problem behind many safety efforts: the failures that matter are the tail. We saw a related theme in OpenAI's post on metagaming latents, where the point was that evaluations only help if they capture behavior the model might actually exhibit.
How LASER works, step by step
OpenAI describes an iterative loop. Here is each stage in plain terms.
1. Initial sampling
The pipeline starts with a pool of synthetic conversations plus samples of de-identified production data. At this point it knows nothing about which ones are interesting.
2. A reasoning grader labels a few
An expensive reasoning model, using chain-of-thought, labels a small batch for whether each conversation violates policy. This is the gold-standard signal, and it is the cost LASER is trying to minimize.
3. A logistic classifier learns from embeddings
Every conversation can be turned into an embedding, a vector of numbers that places similar conversations near each other. LASER fits a logistic regression on those vectors to predict the grader's labels. Logistic regression is deliberately simple: it learns a single direction in embedding space along which "probably violates policy" increases.
A notable engineering detail in the post is that this regression can run inside database queries using approximate dot products on vector indices. In other words, scoring millions of stored conversations does not require pulling them out and running a model; the vector index does most of the work.
4. Sample near the decision boundary
Instead of grading the conversations the classifier is most sure about, LASER picks ones near the 50 percent line, where the classifier is least certain. This is the standard active-learning intuition: labels are most informative where the cheap model is confused.
5. Greedy diversity sampling
If you only ever grade the most uncertain points, you may grade 500 near-duplicates. The final step orders results to maximize diversity in embedding space, so the selected set covers more distinct situations.
The loop then repeats. Each round the grader's new labels improve the classifier, and the classifier steers the next round. That recursion is the R in LASER.
What the numbers mean and what they do not
OpenAI reports roughly a 50 percent disallowed rate in the conversations LASER selects, against about one in 20,000 in random samples. That ratio, combined with the cost of grading, is where the 10,000x compute figure comes from.
A few cautions when reading this:
- It is OpenAI's own measurement. The post links neither a paper nor code, so there is no outside replication yet.
- A 50 percent hit rate is by design. Sampling at the 50 percent boundary means about half the picks are positive; that does not mean half of ChatGPT traffic is unsafe.
- Coverage is not shown. LASER finds many violations cheaply, but the post does not quantify what fraction of all violation types it surfaces. Diversity sampling helps, yet a classifier trained on embeddings can only find what resembles what it has already seen.
- Grader error propagates. If the reasoning grader mislabels a policy edge case, the classifier learns the mistake and steers more sampling toward it.
None of this undermines the technique. It means the headline number describes efficiency, not completeness.
Where it fits in OpenAI's safety stack
The same alignment blog carries a steady stream of related work. Recent posts include towards safety cases for frontier AI training and a study of metagaming latents. LASER sits at the data-curation layer: it does not change the model, it improves the tests used to judge it.
That matters because OpenAI has been under pressure on how it monitors and discloses risk. Our coverage of Mark Chen's comments on moving compute toward safety monitoring and the FTC probe of OpenAI and Anthropic show the broader context. Cheaper, faster eval-set creation is one concrete way a lab can respond when policies change quickly: rebuild the test set in an afternoon.
It also lines up with OpenAI's other transparency and traceability efforts, such as the textGrain watermark plan for EU ChatGPT output, though those are separate programs.
How this differs from red-teaming and classifier filters
LASER is easy to confuse with two neighbors. Red-teaming asks people or attack models to write new adversarial prompts; LASER instead searches a large existing pool of conversations for ones that already sit near the policy line. A production safety classifier, by contrast, runs on live traffic to block or flag content; LASER's classifier is a tool for building test sets and is not described as a runtime filter.
That distinction matters for how you read the 10,000x number. It is a saving on the labeling budget for evaluation data, not a claim that ChatGPT now catches more harmful chats in real time. The practical output is a better exam for the model, written faster, and the quality of that exam depends on the policy rubric the grader applies.
There is also a historical thread. Picking the items a model is least sure about is a decades-old idea called uncertainty sampling, and embedding-based search is a staple of retrieval systems. What OpenAI adds, based on its description, is combining them with a reasoning-model grader, a diversity ordering and an in-database implementation that makes the loop cheap enough to repeat as policies change.
The privacy angle
OpenAI stresses that LASER works on synthetic and de-identified conversations and does not require raw user data access. For readers who have followed debates about training on user chats, this is the claim to scrutinize. De-identification is hard, and the post does not describe the procedure. It is reasonable to ask for an independent description of how de-identification is performed and audited.
Can you borrow the idea?
Yes, as a pattern. If you run an LLM product and need to find examples of a rare failure (a refusal that should not have happened, a jailbreak class, a hallucinated citation) you can try the same recipe at small scale:
- Embed your logged or synthetic conversations.
- Have a strong model grade a small random seed set against a written rubric.
- Fit a logistic regression on the embeddings to predict the grade.
- Grade the items closest to 0.5, plus a diverse subset.
- Retrain and repeat for a few rounds.
Keep the grading rubric versioned, and keep a random sample on the side to estimate how much the targeted set differs from real traffic. Make sure your data handling follows your own privacy commitments before pointing a grader at real conversations.
For choosing which models to grade with and what they cost, our comparison of GPT-6 Astra and Claude Fable 5.1 is a useful starting point.
What to watch next
- A paper or code release. The post is a blog write-up; a reproducible description would let others test it.
- Eval results that use LASER-curated sets. Watch for system cards that cite how their safety evals were built.
- Failure analysis. The most useful follow-up would show where LASER misses violations compared with human-curated sets.
- Adoption by other labs. The method is general, so expect similar active-learning pipelines to appear in other labs' safety work.
Figures and quotes come from OpenAI's October 6, 2026 post and may be updated by the authors.
