On August 26, 2026, Anthropic (@AnthropicAI) published Enabling independent research on how people use Claude — the first pilot giving external researchers access to real, privacy-preserved Claude usage data through Anthropic Insights (formerly Clio).
The post drew 309K+ views on X. Supporters called it rare transparency; critics asked why customer usage feeds third-party studies at all. Both reactions miss the structural point: until August 2026, only AI labs could study how frontier products are actually used at scale — everyone else got lab blog posts or skewed public datasets.
TL;DR — the pilot at a glance
| Question | Answer |
|---|---|
| Tool | Anthropic Insights (ex-Clio) — aggregate classification pipeline |
| Partners | Stanford SALT, Oxford HIP, METR |
| Sample | ~250,000 Claude.ai + Claude Code conversations (Apr–May 2026) |
| Raw access? | No — aggregated categories only, after privacy/legal review |
| Audit | Imperial College London third-party privacy audit |
| Headline finding | Over half of conversations involved consequential tasks (SALT) |
| Data released? | Yes — aggregate outputs from each project published |
| Scale next? | Expression of interest form for future researcher rounds |
What Anthropic Insights does
The pipeline Anthropic describes:
- Researcher writes a question — e.g., "What type of guidance is this person asking for?"
- Claude labels every conversation in the study sample
- Answers aggregate into categories + percentages
- Researchers see only aggregates — never underlying chats
Internal teams iterate question wording for weeks; external partners could not, because each iteration would trigger another privacy review. Anthropic had partners prototype questions on WildChat (public human–AI corpus) first — but WildChat skews casual, so some categories that looked fine on WildChat misclassified real Claude traffic.
That methodological gap matters for anyone citing the SALT numbers: they are Claude-judged aggregates, not human-labeled gold standard.
What the three labs found
Stanford SALT Lab — human–AI collaboration
Full writeup linked from Anthropic's post. Headlines:
| Finding | Detail |
|---|---|
| Consequential work | Over 50% of conversations delegate tasks that affect others or are hard to undo — vs prior research suggesting users keep high-stakes work for themselves |
| Legal/financial guidance | Highest consequential delegation rate |
| Human in charge | ~73% of conversations: human sets direction, Claude assists |
| Not copy-paste | Users usually adapt Claude output rather than use verbatim |
| Friction | Common — and often productive (clarifying intent, iterating) |
For policy readers: this supports professional AI adoption narratives — people are not only using Claude for haiku and brainstorming.
Oxford HIP Lab — emotion and behavior (ongoing)
Early results: Claude's behavior correlates with user affect — warmth with positive user states, refusal with pushback, eccentricity with intellectual engagement. Affective patterns in Claude chat resemble everyday web browsing in a separate study HIP cites.
Full paper pending; Anthropic will link when public.
METR — Claude Code productivity (ongoing)
Preliminary:
- Newer Claude models → larger self-reported speedups (Claude estimates counterfactual time without AI)
- Claude's time estimates correlate with prior developer timing studies
- Next: research acceleration — how much AI speeds researchers whose work includes AI development
Overlaps Anthropic's internal economics work on agentic coding — Anthropic connected the teams deliberately for cross-checking.
Privacy, independence, and the X backlash
Independence mechanisms Anthropic claims:
- Partners free to publish inconvenient findings
- Contractual review limited to privacy, misuse how-to, Anthropic secrets, accuracy
- Imperial College privacy audit on released aggregates
- Less than 5% of categories removed/altered — mostly safeguard-bypass methods, not misuse counts
Criticism on X (representative):
- "Giving away private usage information"
- Demands for per-account opt-out
- Skepticism that aggregate labels still leak workflow patterns
Anthropic's response frame: researchers never saw raw chats; only buckets passed audit. Whether that satisfies users who treat any usage-derived research as extraction is unsettled — and Anthropic's post does not announce a settings toggle.
Why this matters for builders
| Stakeholder | Takeaway |
|---|---|
| Product teams | Real Claude traffic is more consequential than WildChat-style public datasets suggest — eval on casual corpora undercounts professional risk |
| Agent harness designers | Human direction + iteration dominates; agents are assistants, not autopilots, in ~73% of sampled chats |
| Policy / compliance | First external study on in-product frontier usage — template for EU/US transparency debates |
| METR followers | Productivity claims will get independent Claude Code numbers — watch for divergence from vendor economics posts |
Pair with OpenAI's Hugging Face postmortem from the same week — one lab opens usage research; another documents what happens when eval agents escape sandboxes.
The research gap this pilot targets
Anthropic's framing in the official post:
| Data source | Strength | Weakness |
|---|---|---|
| Lab-published analyses | Reflect real product traffic | Answer the lab's questions, not independent hypotheses |
| Public datasets (WildChat, etc.) | Fully inspectable | Skew casual; underrepresent paid professional use |
| Surveys and polls | Cheap | Self-report bias; not behavioral |
| Anthropic Insights pilot | Real Claude traffic, external question design | Aggregate only; Claude-as-judge errors; slow iteration |
Policy debates — UK entry-level hiring cuts tied to AI, Ramp index on AI spend per employee — often mix survey data with vendor anecdotes. A reproducible external pipeline on in-product behavior is a different evidence class, even with methodological caveats.
Methodology — read before citing SALT
Claude labels Claude usage. Every category percentage is downstream of Anthropic's model applying natural-language rubrics at scale. That is scalable but not ground truth.
Known failure modes Anthropic documents:
- Question wording sensitivity — small phrasing changes shift category boundaries
- WildChat calibration drift — categories tuned on casual chat misread professional Claude Code sessions
- No human audit loop on production sample — privacy design prevents spot-checking mislabels
- Impossible-task parallel — OpenAI's Hugging Face postmortem shows eval graders can be wrong; Insights categories could inherit similar blind spots
For citations: treat SALT's "over half consequential" as directional — consistent with professional adoption stories, not a precision labor-statistic.
Policy and labor implications
If 50%+ of Claude conversations truly involve consequential delegation, several threads follow:
Professional liability. Legal and financial guidance categories topped SALT's consequential list. That aligns with Claude for Work positioning — and with risk teams asking whether outputs become advice under local regulations.
Human oversight is the norm, not the exception. ~73% direction-retention suggests users are not blindly autopiloting — they steer and edit. Friction is productive in SALT's read: iteration improves clarity. Automation fear narratives that assume one-shot delegation may misread actual behavior.
Entry-level work composition. Consequential does not mean replacing junior roles — it can mean augmenting them (draft memo, human partner reviews). Cross-read Work Foundation UK data: employers cite AI in hiring cuts, while usage research shows humans still direct most Claude work. Both can be true — cuts may target headcount budgets, not task delegation rates.
What other labs share (and do not)
| Lab | Usage research access |
|---|---|
| Anthropic (Aug 2026) | External pilot on aggregated Insights outputs |
| OpenAI | Publishes economics posts; Hugging Face incident with METR audit — eval-focused, not general usage sociology |
| Google / Meta | Occasional transparency reports; no equivalent open aggregate export in August 2026 |
| OpenRouter / aggregators | Token routing stats (Asia models 60%) — market mix, not task consequence |
Anthropic claims this is the first public independent studies on a frontier lab's own usage data. Even skeptics of aggregate privacy should track whether EU AI Act transparency expectations push competitors toward similar programs.
Scaling problems Anthropic admits
The post is unusually candid about operational cost:
- Slow vs normal lab research velocity — privacy review per dataset iteration
- Resource intensive — Anthropic runs Insights runs on partners' behalf
- WildChat workaround incomplete — external partners cannot iterate on real traffic freely
- Misuse category redaction — transparency vs safeguard evasion trade-off (less than 5% of categories altered)
Scaling means either more Anthropic staff per external study, looser privacy iteration (unlikely), or narrower research questions that pass review first pass. The expression of interest form is as much demand sensing as recruitment.
For researchers applying to the next round
Anthropic's appendix (linked from the main post) covers:
- Partner selection criteria for the first three labs
- Contractual publish even if inconvenient clause
- Imperial College privacy audit scope
- Interpretation guide for aggregate outputs — required reading before secondary analysis
If you study agent productivity, note METR overlap with Anthropic economics on agentic coding — propose different questions or holdout benchmarks to avoid redundant work. If you study affect or UX, Oxford HIP's incomplete writeup leaves room — but you compete on category design quality, not raw data access.
User privacy — what aggregate-only does and does not protect
Protects (per Anthropic + audit):
- Individual usernames, chat text, and file paths never leave the Insights pipeline to partners
- Released datasets are bucket counts, not conversation IDs
Does not fully address (per X critics):
- Workflow fingerprinting — rare category combinations may imply niche professions
- No opt-out announced for this pilot — users cannot exclude their traffic from aggregate studies via settings described in the post
- Future scale — today's 250K sample is small vs total Claude volume; expanded programs increase re-identification risk if buckets get too granular
Anthropic's threat model and audit live in the appendix — worth reading if your org's DPA with Anthropic covers research reuse of interaction metadata.
Bottom line
Anthropic Insights external pilot is a infrastructure story, not a headline statistic. The SALT finding that over half of Claude use is consequential is striking — but the bigger move is external researchers studying in-product data at all, with published aggregates and an audit trail.
Whether Anthropic can scale the program — slow privacy review, WildChat calibration gaps, resource intensity — determines if this becomes industry norm or a one-off transparency exercise.
Researchers can express interest via Anthropic's Google Form linked from the official post.
Related on explainx.ai
- David Siegel on open source AI and auditability
- METR and OpenAI Hugging Face independent assessment
- Claude vs ChatGPT for work
- Top 10 Claude Cowork use cases — real professional workflows
- AI agent monthly cost — real workflow math
- Goodhart's law and benchmark contamination
- UK employers cutting entry-level jobs — AI survey context
Official: Enabling independent research · @AnthropicAI thread · Appendix PDF (linked from post)
Findings reflect Anthropic's August 26, 2026 publication and partner preliminary results. METR and Oxford studies remain in progress.
