A one-line announcement landed on October 7, 2026 that changes how you can budget a multi-agent run. Lydia Hallie, who works on Claude Code at Anthropic, wrote: "You can now ask Claude to run subagents at a specific effort level! Make sure you're on v2.1.292+." The attached screenshot shows the prompt style:
Find every payments API call with low effort subagents, then have a high effort one check the error handling.
That example is the whole idea in one sentence: cheap, fast subagents do the wide search, and one careful subagent does the part that needs judgment. This guide explains why that matters, how to use it, what we could not verify from the announcement, and how to keep costs under control. We worked from the announcement and Anthropic's existing effort and subagent documentation as summarized in our earlier posts, not from release notes for 2.1.292, which we did not have.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is new? | You can ask for subagents at a chosen effort level. |
| Which version? | v2.1.292 or later. |
| How do you ask? | In plain language, per the announcement's example. |
| Why do it? | Spend reasoning where it matters, save it where it does not. |
| Does it save money? | It can trim cost on scouting steps. Subagents still add overhead. |
| What is unverified? | Exact syntax, which levels subagents accept, and defaults. |
Why effort per subagent matters
Effort is Claude Code's control for how much reasoning the model spends on a turn. Our commands reference lists /effort [low|medium|high|xhigh|max|ultracode|auto], with max and ultracode limited to the session. Until now, effort mostly applied to the session as a whole. Setting it high made everything slower and more expensive, setting it low risked sloppy work on the hard parts.
Multi-agent work makes that trade-off sharper, because a run is a mix of very different jobs.
| Job | Reasoning needed | Better effort |
|---|---|---|
| Find every call site of an API | Low: pattern matching and reading | Low |
| Summarize what each file does | Low to medium | Low or medium |
| List TODOs and dead code | Low | Low |
| Check error handling for correctness | High: edge cases, failure modes | High |
| Review a concurrency change | High | High |
| Judge a security-sensitive diff | High | High or higher |
| Decide an architecture trade-off | High | High, with a human in the loop |
Per-subagent effort lets one request express that shape: wide and cheap on the left, narrow and careful on the right. It is the multi-agent version of the advice in our post on model versus effort: the model sets what Claude knows, the effort sets how hard it tries.
How to use it
Start from the announcement's pattern and adapt it. The structure is always the same: name the work, name the effort, and say how the results combine.
1. Search, then review.
Find every place we call the payments API using low effort subagents,
one per top-level package. Then run a high effort subagent over the
combined list to check that each call handles timeouts, retries and
idempotency keys. Report only the calls that are unsafe.
2. Parallel summaries, one synthesis.
Use low effort subagents to summarize each service directory in two
sentences. Then have one high effort subagent read the summaries and
propose where the module boundaries are wrong.
3. Cheap triage before an expensive fix.
Use low effort subagents to reproduce each failing test and classify
the cause as flaky, environment or real. For the real failures only,
use a high effort subagent per failure to propose a fix.
4. Review a large diff in slices.
Split this diff by directory. Use medium effort subagents to review each
slice for style and obvious bugs. Use one high effort subagent to look
for cross-slice problems such as changed interfaces.
A few habits improve results. State the effort once per phase, not per sentence, so the instruction is unambiguous. Keep the final, careful step to a single subagent unless the work truly splits, because high-effort parallel agents are where cost climbs. And say what you want back: a short list, a table, or a verdict, so cheap subagents do not return pages of raw output that the expensive one must then read.
What we could not verify
The announcement is brief, and some practical questions are open.
- Syntax. The example uses plain language. We do not know whether there are also settings, flags or a frontmatter field in subagent definitions for effort. If you define custom subagents, check the format in your version's documentation. Our guide to Claude Code subagents and multi-agent workflows describes the definition format as of its writing.
- Which levels apply. The session accepts low, medium, high, xhigh, max, ultracode and auto. Whether subagents accept all of them, or a subset, is not stated.
- Defaults. What effort a subagent uses when you say nothing, and whether it inherits the session level, we did not confirm.
- Interaction with models. Whether a subagent can use a different model as well as a different effort in the same request is outside what the post says.
- Reporting. Whether the session shows each subagent's effort and cost separately is unknown.
The reliable way to find out is to test. Run claude --version, update if needed, and try a small, harmless prompt with two effort levels, then compare the transcript and /usage.
The cost picture
Subagents are not free. Our measurements in Do subagents actually use more usage? showed that each subagent builds its own context, so total tokens rise with parallelism even when the wall-clock time falls. Per-subagent effort gives you a lever on part of that bill.
A rough way to reason about it:
- Scouts dominate in count. A search across 40 files might spawn many subagents. Making those low effort reduces the reasoning tokens in the part of the run that has the most agents.
- The reviewer dominates in depth. One high-effort subagent costs more per turn, but there is only one, and it works on a compact input.
- The synthesis step is where waste hides. If cheap scouts return verbose output, the expensive reviewer pays to read it. Ask scouts for terse, structured results.
For a sense of what a single task costs on the current top model, see What a Claude Code task costs on Opus 5.5, and for the broader question of when more effort is worth it, our post on effort levels and plan mode.
Risks and failure modes
Under-powered scouts miss things. A low-effort subagent can skim and skip a call site hidden behind a wrapper. Fix: make the scouting step produce evidence (file and line), and have the reviewer spot-check coverage with a different method, such as a grep.
A confident wrong summary. If the high-effort step trusts a low-effort summary, an error propagates. Fix: have the reviewer read the underlying code for anything it flags, not only the summaries.
False economy. Saving tokens on a step that later needs a redo costs more. Fix: use medium for steps that are not clearly mechanical.
Permission sprawl. More subagents means more tool calls. Keep your permission mode appropriate, as in our guide to Claude Code permission modes, and avoid blanket approvals for parallel runs.
Inconsistent standards. Different effort levels can produce differently styled output. Specify the format.
When not to bother
Skip the pattern for small tasks. If a change touches three files, a single agent at medium effort is simpler and often cheaper than orchestrating subagents. Use the multi-agent, mixed-effort approach when the work is wide (many files or services), the steps are separable, and there is a clear difference between a mechanical phase and a judgment phase.
A worked example: auditing payment calls
Consider the announcement's own scenario, a codebase with a payments API called from a dozen services. A sensible run has three phases, and each gets its own effort.
Phase one, discovery (low effort). One subagent per service lists every call site with file, line and the function name, plus whether the call is wrapped in a retry. The output is a flat table, nothing more. Because the work is pattern matching, low effort is enough, and running a dozen of them in parallel keeps wall-clock time short.
Phase two, consolidation (the main agent). The table is merged and deduplicated. This is cheap, and it is the natural place to sanity-check coverage, for example by comparing against a plain text search for the client library name to catch anything the scouts missed.
Phase three, review (high effort). One subagent reads the actual code around each call that lacks a retry or timeout, and answers specific questions: what happens on a network error, a partial response, a duplicate submission, an expired token? It returns a short list of unsafe calls with a suggested fix for each.
The result is a run where most of the tokens go to reading code, and the expensive reasoning is concentrated on a few dozen lines that matter. Compared with running everything at high effort, you spend less on the search. Compared with running everything at low effort, you avoid the risk of a shallow review on the part where mistakes are costly.
After the run, check three things: whether the scouts' call-site list matches a simple search, whether the reviewer's flagged items are real when you read them, and what the session cost in /usage. Those numbers tell you whether the split paid off on your codebase and whether to keep it as a standing habit.
What this means for what you build or pay
For individual developers on limited plans, the feature is mostly about allowance: you can keep broad searches cheap and spend your high-effort budget on the review that decides whether the change is safe. For teams, it suggests a convention worth writing into your project instructions: scouts at low, reviewers at high, final sign-off by a human. For people building their own agent harnesses, it confirms a design direction that other tools are taking: effort and model as per-task knobs, not session-wide settings. If you track how agent tooling is evolving, it fits the broader move we described in the terminal era and the agent as the new primitive.
Related reading
- Claude Code commands: complete reference
- Claude Code subagents and multi-agent workflows
- Do subagents actually use more usage?
- Model versus effort: knowing more versus trying harder
- What a Claude Code task costs on Opus 5.5
- Effort levels and killing plan mode
- Claude Code permission modes explained
- Hermes Agent manual subagent control
Primary: Lydia Hallie's announcement on X (October 7, 2026), including the example prompt in the attached screenshot · Claude Code release notes for v2.1.292, which should be checked for exact behavior
Details are accurate as of October 7, 2026 and come from the announcement and our earlier documentation summaries. We did not read the v2.1.292 release notes, so syntax, supported levels and defaults are unverified. Test on a small task before relying on it.
