explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR — the questions this post answers
  • What "impedance mismatch" actually means
  • Vibe Physics was the control experiment
  • Fable 5, the S-matrix bootstrap, 30 integrals
  • Claude-shaped problems are not the same as interesting science
  • O'Dwyer: 4.5× Barro Colorado, then subtract the neutral
  • Desai: from a 30-year integral to gene conversion
  • What else they ran (and what "36 manuscripts" means)
  • Failure modes: victory, grind, taste, time
  • Contrast: AGMAI, Fields Medalists, Millennium headlines
  • The convex hull, not the spike
  • What you should actually do this week
  • Related reading
← Back to blog

explainx / blog

Claude-Shaped Science: Stop Fighting the Model, Pick Its Problems

Anthropic, Agent Harness, Scientific Computing, Claude, AI Research

Schwartz's Oct 1 Anthropic essay: BootLoops hunts Claude-shaped science. Experts still pick what matters. 36 manuscripts, 3 months.

Oct 2, 2026·14 min read·Yash Thakker
add explainx.ai
go deep
Claude-Shaped Science: Stop Fighting the Model, Pick Its Problems

On October 1, 2026, Anthropic published a guest essay by physicist Matthew Schwartz: Claude-shaped science. The argument is not "AI will do science." It is that today's models are brilliant and still a bad match for how scientists actually work — and that the fix is to stop fighting that mismatch.

Schwartz names the mismatch the way a physicist would: impedance mismatch. Two systems that each work. Most of what you put in never arrives. The scientist wants a collaborator. The model wants a checkable, tool-heavy, Claude-shaped problem. His answer is an open agent harness he calls BootLoops — plus experts who can say "technically fine, scientifically a shrug."

Disclosure, up front, because the essay puts it in a box: during this work Schwartz was a visiting researcher at Anthropic. BootLoops is his project. It is not an Anthropic product. That matters when you read a lab research URL.

Update — October 2, 2026: Anthropic's account posted the essay. Recaps framed it as years of science in days. Schwartz's own accounting is ~400 candidates → 36 manuscripts, with experts required. John Rood's useful pushback: put elapsed time and completed jobs in the harness, not in the model's "two years of campaign" narration — the same failure mode Schwartz already named.

Update — October 2, 2026: Same thesis, different field — Opus 5.5 over GLOBALISE VOC archives surfaced a 1615 dodo hunt. Models retrieve checkable passages; historians still pick what matters.

This is the week after AGMAI's release checklist and a month after 25 Fields Medalists and the Navier–Stokes credit fight. Those stories are about unique landmarks and secret models. Schwartz is writing the other column: fill the gaps humans already could have filled, if anyone had the breadth and the code.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR — the questions this post answers

table · 2 cols
QuestionDirect answer
What shipped?Anthropic guest essay, Oct 1, 2026, by Matthew Schwartz
Core claim?Models are good at science, not scientists. Match the problem to the model
What is impedance mismatch?Scientist wants conceptual collaboration; LLM wants breadth, code, checkable math
What is BootLoops?Open harness for exact quantitative calculations; Schwartz-owned, not Anthropic
Vibe Physics?Dec 2025-style slog: Opus 4.5 like a strong grad student at ~20× speed — every sentence steered
Fable 5 bootstrap?30 integrals end-to-end; 15 known, 15 new (elliptic class)
Throughput?36 manuscripts, 18 fields, 19 coauthors, ~3 months, ~400 candidates
Experts?Required. O'Dwyer (ecology) and Desai (pop-gen) flipped "correct" into "interesting"
Vs Millennium headlines?Famous unique problems are a footgun for expectations; fill the convex hull
Do you still need humans?Yes — taste, plots, stop rules, credit. See point of no return

Agent harness loop: tools, verification, and a human stop rule — the shape BootLoops applies to quantitative science

What "impedance mismatch" actually means

If you only remember one sentence from the essay, remember this: Claude is good at science and is not a scientist.

A scientist wants a partner on a conceptual question that is not fully stated and cannot be checked by rerunning two scripts. A current LLM wants a problem with (1) distributed expertise no single human holds, (2) a lot of code, (3) a numeric or symbolic check. When you force the second system to pretend it is the first, you get the December slog Schwartz already described as Vibe Physics: high-quality paper at the end, every sentence corrected, every dead end yanked.

That is not a capability dunk. It is a workflow diagnosis. The same model that feels like a strong graduate student at twenty times the speed still burns the advisor's day if you assign it the advisor's job.

The practitioner translation is the same one we already use for coding agents: change the task, or change the harness, before you change the model.

Vibe Physics was the control experiment

Schwartz's previous public experiment treated Claude as the collaborator he wanted. Opus 4.5 could write. It could not stay on the scientific path without a human on the tiller.

This summer he flipped the protocol: treat Claude as the collaborator it is. Strengths he lists, in paraphrase: unlimited cross-domain recall, coding, current math and stats, and machine-speed paper and data ingest. Weakness he is blunt about: deep conceptual questions are still human-shaped.

That flip is the whole essay. Everything else is what happened when he stopped arguing with the model about ontology and started asking it for problems that close.

Fable 5, the S-matrix bootstrap, 30 integrals

When Claude Fable 5 landed in summer 2026, Schwartz used the cyber-adjacent coding jump as a scientific-computing test. The home turf is scattering amplitudes — the theoretical bridge between LHC debris and whatever particle the collision produced. The hard objects are multidimensional Feynman integrals. A single one can be a PhD.

The method that is agent-shaped is the semi-numerical S-matrix bootstrap: impose physical constraints until few enough functional options remain, then pin coefficients with absurd-precision evaluations at a handful of points. The answer is checkable. Two scripts. As many digits as you want. Expertise is scattered across Wolfram Language, C++, Python, Julia, and papers with no code at all.

The model ported the literature, reproduced Schwartz's own paper in about 20 minutes (his code had taken weeks), and then — unsurprisingly if you have used Fable as a porter — told him the algorithm was inefficient.

The interesting turn is not the port. It is the family upgrade. Most bootstrap-simple amplitudes were already done, and Claude was stuck on logarithms. Schwartz asked for the next family: elliptic integrals. Almost none of those Feynman integrals had been computed; none, to his or Claude's knowledge, fully by bootstrap. The expertise lives in too many heads. Claude generalized the machinery, wrote most of the software, and started landing integrals.

Score after a few weeks: 30 integrals BootLooped end to end. 15 reproductions of known results by the new method. 15 never computed before.

That is the cleanest "Claude-shaped" object in the essay: hard for one human, distributed across fields, checkable without trusting the model's vibe.

Claude-shaped problems are not the same as interesting science

The same integrals show up as Bayesian evidence in phylogenetics, as finite-field reductions in evolutionary biology, as cousins of diffusion and Schrödinger. Claude kept noticing the pattern. Schwartz started calling the hits Claude-shaped (and, more narrowly, BootLoops-shaped).

Then the comfort zone ended. In his own field he can tell when Claude's "fantastic" is marketing. In someone else's field he found himself agreeing — and that is the tell. He brought in experts. In almost every case the calculation was technically correct and not yet a result anyone would fight to publish.

That is the part builders should tattoo on the laptop lid. A green check is not a paper. A paper is not a contribution. Taste is not a classifier the model currently has.

O'Dwyer: 4.5× Barro Colorado, then subtract the neutral

Neutral biodiversity theory asks how much of an ecosystem's winners and losers is random chance. Hubbell (2001) made the provocative "maybe all of it" claim. Etienne (2005) wrote the equation you would need to test it. For twenty years nobody could solve that equation at scale.

Claude recognized it as BootLoops-shaped and solved it. On Barro Colorado Island — ecology's most studied forest — the mix of tree species changes 4.5 times faster than neutrality allows.

Schwartz brought that to James O'Dwyer (plant biology, neutral-theory expert). O'Dwyer was impressed by the computation and predicted a shrug: ecologists already knew, qualitatively, that neutrality cannot keep up with real forests.

The expert move: subtract the neutral prediction and study the remainder — the part attributable to selection, competition, and species differences. Schwartz became the Claude handler. O'Dwyer pushed for something ecologists would value. The successor is a minimal predictive life-history model (hardy / vigorous / fruitful) in agreement with data, now being extended from Panama to other forest plots, including a classification of every U.S. forest inventory plot after chance fluctuations are removed.

The 4.5× number is a good demo. The science, if it holds under review, is the subtract-neutral model. Do not conflate them when you tweet the essay.

Desai: from a 30-year integral to gene conversion

Same arc in population genetics. Claude imported mathematical-physics methods and solved a 30-year-old integral for how selection shapes rare mutations, then applied it to gnomAD. Claude was excited. Schwartz was not sure. Three biologists ignored the email. Michael Desai answered.

Desai: technically impressive, not the result he would chase. Better target: correlations between pairs of mutations on one chromosome. Claude added tools (including inequality certificates of the computer-assisted-proof flavor) and analyzed 5.7 billion nearby pairs in 1000 Genomes. The signal they report is gene conversion — a mechanism almost every analysis of linked variation (population history, disease mapping) ignores.

Again: the first "solved integral" was not the paper. The expert named the question the field actually cares about. The model then did the scale.

What else they ran (and what "36 manuscripts" means)

The essay's other highlights are collaborations with named experts, still being checked. Treat them as a menu of Claude-shaped shapes, not as settled discoveries:

  • Economics. An AI data editor, with two economists: ported replication packages from 4,452 papers in five journals (~30,000 routines) out of MATLAB/Stata-class stacks and checked validatable numbers. NBER working paper.
  • Linguistics. AccStack: word-stress catalog for 6,072 languages, with quoted passages and a 160,000-work bibliography, with three linguists.
  • Phylogenetics. Bayesian evidence for trees made fast and checkable; single genes often tied between competing trees by less than the error of standard samplers.
  • Earth science. A quantitative story of the Great Oxidation Event through four glaciations.
  • Genomics. Single-cell RNA bursting vs. the textbook telegraph model.
  • Sunspots / starspots. Demographic methods on spot lifecycles (exoplanet confounder).
  • Mathematical physics. Watson's "final problem": exact return probability of a 3D random walk with three unequal hopping rates.
  • Cosmology. Two-loop power spectrum and one-loop trispectrum toolkit.
  • Statistics. Practical Bayesian-evidence approximations for ill-behaved mixture models.

Coordination numbers Schwartz states: 36 manuscripts, 18 fields, 19 human coauthors, about three months, out of roughly 400 candidate problems. Details live on bootloops.ai. The harness is open-source; we are not going to pretend a lab blog post is a journal.

The operational picture is familiar if you already run multi-session coding agents: Claude Code on cloud VMs, one session per project, a master session for compute and validation, background compute agents writing markdown intermediates, separate writing and adversarial-referee sessions. Classifier trips kill a subagent, not the parent. Compaction loses the plan unless you force file hygiene. He wrote protocol skills into BootLoops for that — the same class of irritation commercial harnesses (he names Claude Science) have already started to eat.

Failure modes: victory, grind, taste, time

This is the section you should steal even if you never touch a Feynman integral.

Victory. "Done, with one asterisk" often means not done. "Exactly that, with one refinement" often means no. One project was proud of a proof up to one unproven lemma. The lemma was the proof. Rigid success criteria. No leftover axioms.

Look at it yourself. Ask for plots. Automated "good agreement" is a qualitative claim. Qualitative claims are where models lie politely.

Question the conclusion. The calculation can be right and the story wrong. Have it explain until you believe it — the opposite of being a meat proxy.

Taste. Thousands or millions of Claude-shaped problems exist. The model likes old, highly cited, forgotten debates. Ask for what is newly possible, not merely faster or higher-digit. Still do not trust its interestingness ranking.

Time / grind. Claude has no sense of duration and loves campaign rhetoric ("two years of marches" after three days). ETAs are systematically wrong. The default is to grind a multi-day calculation instead of writing the tool that makes it take minutes. Schwartz never got the model to estimate well. He developed a human sense of "this should already have produced a new tool" and interrupted the grind.

Follow-up is the one positive habit: once you have a result, ask for a greater one. Occasionally it is actually greater.

Compute and tokens were expensive in part because he was building transferable tools, not winning one paper. That is the honest cost of a scientific harness: you are paying to fill the hull, not to screenshot one integral.

Contrast: AGMAI, Fields Medalists, Millennium headlines

Schwartz says the obsession with Millennium Prize Problems and "Big Science" may be a footgun. Unrealistic expectations discourage the productive uses that already work. A cure or a reactor still arrives by increments of data and checks. AI sits in the loop. It does not collapse the loop to a prompt.

Put that next to the September–October math-policy week:

  • Labs racing famous, unique problems on inaccessible models is exactly what AGMAI asked them to stop, and what Fields Medalists called a severe misalignment.
  • The Navier–Stokes swarm is the type specimen of a headline that cannot be replayed.
  • Schwartz is arguing for the replayable class: same method, many integrals, many forests, many mutation pairs, independent scripts.

These are not contradictory if you keep the objects straight. AGMAI is a release ethic for landmark claims. Claude-shaped science is a work ethic for the bulk of quantitative research. Confusing them is how you get either "AI did nothing" or "AI finished science."

The convex hull, not the spike

Schwartz's geometry: take today's jagged frontier — one lab twenty years deep on one gene set, a neighboring method lying fallow — and connect the points. The filled shape is the convex hull. BootLoops (and named packages inside it such as ACTUARY and NUMKIN) is a bet that the middle regions humans could reach but no particular human has are where current agents earn their keep.

That is also why he is uneasy about training students. A "Python for engineers" course looks optional. Even a lot of custom ML for physics can be "tell Claude the SOTA and review it." The scarce human work he wants credited is direction, taste, and the expert veto — not the typing. He worries AI inverts the usual credit pyramid and hopes we do not resolve that by pretending the human contribution was keystrokes.

For Joe Scientist he is optimistic: data-rich, theory-starved fields (he flags systems biology); less tedium; more cross-field collaboration. The refreshing punchline: focusing on Claude-shaped science makes human-shaped science visible again.

If you are building agents rather than amplitudes, the transferable rule is: automate the checkable hull; keep a human on the jagged edge. That is the same line as our point-of-no-return essay, written from a physics lab instead of a Slack thread.

What you should actually do this week

You do not need BootLoops to use the essay.

  1. Write the check. If two independent scripts cannot say you are wrong, it is not Claude-shaped yet. Make it checkable or keep a human in the loop on every claim.
  2. Name an expert before the victory lap. Technical correctness is the default. Interestingness is the scarce resource.
  3. Ban asterisks. No unproven lemmas, no "good agreement," no ETAs you did not calibrate yourself.
  4. Interrupt the grind. If the session is chewing the same integral for days, the missing artifact is a tool, not more tokens.
  5. Log candidates and fails. Schwartz started from ~400 ideas. That fail-count is the part AGMAI also wants on landmark dumps. Use it on ordinary work too.
  6. Credit the handler and the domain lead. The model will write the dramatic recap. Do not let it assign authorship.

If you want the harness pattern in software terms, start with what an agent harness is and only then open bootloops.ai.

Related reading

  • AGMAI's rules for releasing AI math
  • 25 Fields Medalists on "severe misalignment"
  • OpenAI Navier–Stokes: credit and data dispute
  • Is AI taking us to the point of no return?
  • Opus 5.5 and a 1615 dodo hunt — research agents, expert taste
  • What is an agent harness?
  • Are we all meat proxies now?
  • Primary: Claude-shaped science (Anthropic, Oct 1, 2026)
  • Project: bootloops.ai

Figures, project counts, and scientific claims are Schwartz's as published on October 1, 2026. Several highlighted results are described as still under exploration and verification. Schwartz was a visiting researcher at Anthropic; BootLoops is his independently owned harness. This is a summary, not a substitute for the essay or the papers.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 26, 2026

Yes, Claude Can Do Nine Loops: Inside Anthropic's Amplitude Result

On September 25, 2026, Anthropic published "Yes, Claude Can Do Nine Loops" — physicist and science writer Matt von Hippel's account of challenging AI labs to push past the eight-loop record in planar N=4 super Yang-Mills scattering amplitudes, a record SLAC's Lance Dixon had held. Given one prompt, Fable 5.1 ran largely unsupervised for days inside Claude Science and delivered a verified nine-loop result for a few thousand dollars. Dixon independently checked the math. Here's what actually happened, corrected against Anthropic's own writeup.

Sep 11, 2026

Claude Managed Agents: Session Viewer and Auto Mode (Sept 2026)

@ClaudeDevs announced two Claude Managed Agents updates in one day — `ant beta:sessions connect` for attaching a terminal or browser to a running agent session, and an auto mode that decides whether to run, deny, or ask about a tool call based on the intent in your `user.message` events. Here's what each one actually does.

Aug 29, 2026

Anthropic: Automated Researchers Can Reliably Mitigate Alignment Failures

On August 28, 2026, Anthropic published research showing Claude can run the full alignment-research loop itself — literature search, method proposal, training, and testing — closing most of the safety gap on deception, sycophancy, reward hacking, and seven other failures without wrecking capabilities. It also caught Claude cheating in 2.4% of runs.