explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • What should you take from an AI roll-up pitch?
  • What is Isenberg actually arguing?
  • What do the six operators report?
  • What this changes about what you build and what you pay
  • What files run the operating system?
  • Which prompts can you paste this week?
  • How does the rulebook grow in the first 60 days?
  • What does the first 100 days look like?
  • Which firms fit the pattern?
  • What breaks?
  • What people are asking
  • Related reading
← Back to blog

explainx / blog

AI Roll-Ups: How a Services Back Office Runs on Agents

AI Agents, Agent Workflows, Services, Claude Code, Codex, Operations

AI roll-ups buy a services firm, then change delivery with agents. Build shadow mode, a reviewer that can block, and tests from a corrections log.

Sep 27, 2026·21 min read·Yash Thakker
add explainx.ai
go deep
AI Roll-Ups: How a Services Back Office Runs on Agents

AI roll-ups, in the version Greg Isenberg published around 12:49 a.m. on September 27, 2026, are a purchase plus an operating change. The article drew about 175,800 views. His thesis is to buy a services firm that already has customers, licenses, and trust, then change delivery with agents. He cites a McKinsey figure of about $5 trillion in American businesses changing hands by 2035. That number is backdrop. The part a builder can use is a review architecture.

A preparer and a reviewer are different agents. Shadow mode runs before any client sees a draft. A corrections log becomes tests. That is the same shape as a careful coding-agent workflow, applied to a services back office. The acquisition thesis is the context. The operator numbers are mostly the buyers' own. This explainer is not investment advice, and it does not recommend buying a firm, a fund, or a security.

What should you take from an AI roll-up pitch?

table · 2 cols
QuestionDirect answer
What is the thesis?Buy a firm that already has clients, licenses, and trust. Change delivery with agents. He cites traditional services margins of about 5 to 10 percent EBITDA, and argues for 30 to 40 percent with the same clients, about 3 to 4x profit. That is a thesis, not a result.
Why does he say now?Models are good enough that data entry, first-draft tax returns, contract review, support tickets, and maintenance requests need checking rather than redoing. Owners are retiring. Services firms sell at low multiples because margins are assumed stuck.
What is worth copying?Three files, two agent roles, shadow mode, and a corrections log that becomes tests. The reviewer can block and cannot ship. The preparer can ship nothing on its own.
What did the buyers report?Six operators, all self-reported through Isenberg on September 27, 2026. None of the companies, he says, has been through a recession. explainx.ai did not audit them.
What do you build?Target criteria, an automation map, a rulebook, shadow-mode runs, and a preparer that cannot send client work.
What do you pay?A services-firm purchase price, his illustration is about $800,000, plus the same class of model subscription a coding agent already uses. He names Codex. Claude Code is the same class of tool. You are not paying for a new foundation model.
What should you refuse to pay for yet?Headline margin multiples. Underwrite on your own shadow-mode correction rate.
Where did the 7,000 tax returns come from?Isenberg restates a pilot explainx.ai already covered from OpenAI's Codex platform post. It is not a new measurement.
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

What is Isenberg actually arguing?

The sentence under the headline is operational. Buy the firm that already has the phone to answer, the license on the wall, and the client who renews. Then change how the work gets done. He cites traditional services margins of about 5 to 10 percent EBITDA. The thesis is 30 to 40 percent, which he frames as about 3 to 4x the profit, with the same clients. Read that as a target he is arguing for. Do not read it as a result explainx.ai measured.

He cites McKinsey for the supply of firms. About $5 trillion of American businesses change hands by 2035, many of them baby-boomer owned, often with no successor. He also says McKinsey expects about 1 million or more boomer-owned businesses to sell by 2035. explainx.ai did not re-read the McKinsey report. Those figures stay attributed to Isenberg's citation. They are not an explainx.ai finding.

His "why now" has three parts, in his words. Models are good enough that data entry, first-draft tax returns, contract review, support tickets, and maintenance requests need checking rather than redoing. Owners are retiring. Services firms sell at low multiples because buyers assume the margins are stuck. If the second and third points are right, the asset is the client list and the license, priced as if delivery cannot change. If the first point is wrong, the checking still costs as much as the doing, and the margin thesis fails.

Why buy a firm instead of building a SaaS product?

His comparison is with a software startup, not with a stock. The client list, the licenses, and the niche judgment took decades. Past engagements are the training material a startup does not have. A new product has to earn trust one signup at a time. A bookkeeping firm, a property manager, or a medical biller already has the jobs, the documents, and the exceptions that teach a rulebook.

That is also why the operating system matters more than the purchase announcement. A client list is training material only if someone turns completed jobs into rules, and only if a reviewer can reject a draft that breaks a rule. Without that, the buyer owns a services firm and a chatbot.

What do the six operators report?

Put the caveat in the same place as the numbers. Isenberg says these fund results are self-reported, and he says none of these companies has been through a recession. September 27, 2026; operator figures are the ones Isenberg reported and are not independently audited by explainx.ai.

table · 3 cols
OperatorWhat Isenberg reportsSource
Long Lake, property management18 businesses, $100 million EBITDA in under two years, margins doubledself-reported via Isenberg, Sep 27, 2026
Crescendo, contact centers90 percent of frontline tickets resolved by AI, 4x the margins of traditional operatorsself-reported via Isenberg, Sep 27, 2026
Titan MSP30 percent or more of workflows automated, targeting tripled net marginsself-reported via Isenberg, Sep 27, 2026
Dwelly, UK real estateProblem resolution from 50 days to 20, margins doubledself-reported via Isenberg, Sep 27, 2026
Thrive HoldingsNearly 50 local accounting practices in two years, and a $1 billion commitment to buy more. Larson Gross details below.self-reported via Isenberg, Sep 27, 2026
General Catalyst$1.5 billion set aside, more than $750 million into at least ten companiesself-reported via Isenberg, Sep 27, 2026

Titan's tripled net margins are a target. He does not say the target is achieved. Long Lake, Crescendo, and Dwelly are presented as results those operators claim. Young companies that are raising money have a reason to tell the generous version. That limitation is in his own piece, and it belongs next to the table.

Thrive, Larson Gross, and the tax-prep number you have seen before

Isenberg says Thrive Holdings assembled nearly 50 local accounting practices in two years and committed $1 billion to buy more. He names Larson Gross, in Bellingham, Washington, founded in 1949, with five offices and about 200 employees, as a firm where Thrive bought a stake. He says that this tax season the firm's AI processed 7,000 returns, that accountants saved 31 percent of time on average, and that one job went from 180 hours a year to 15. He says those tax agents run on OpenAI Codex.

explainx.ai already covered that pilot. OpenAI's August 2026 Codex-as-a-platform post credited Thrive Holdings and Crete with a workflow that processed 7,000 returns and cut prep time by roughly a third. That write-up is Codex as a platform. Isenberg's September 27 article restates the same pilot, adds the Larson Gross name, the 31 percent figure, and the 180-hour-to-15 example, and ties the agents to Codex. Those details are his report. They are not a new independent measurement by explainx.ai.

The deal shape he says is worth copying

General Catalyst, in his account, set aside $1.5 billion and put more than $750 million into at least ten companies. The deal shape he highlights is roughly 60 to 70 percent cash at close, with about 30 percent rolled into equity for the founder. He says that shape is worth copying. It is not a term sheet explainx.ai negotiated, and it is not a template this post tells you to sign.

His argument for a smaller, personal version

Isenberg argues that an individual can run a small version of the same idea. Funds need large checks, so they skip a $2 million bookkeeping firm. The same Codex subscription, or a Claude Code subscription in the same class of tool, is available to one person. The integration bottleneck, he says, is being in the office. Some owners prefer a person to a fund. He says SBA loans and seller notes exist at this size.

The counterweight is the part a builder has to add. Buying a firm with debt is not a side project. A license and a client's trust do not move because a prompt says they should. A wrong email to a client is a real liability. This post reports his financing remark and stops there. It does not explain how to borrow, and it does not suggest taking a loan.

He also gives two score thresholds in prose. A firm under 60 percent is usually a job, not an acquisition. A firm above 80 percent deserves a letter of intent. The detailed rubric was an image. explainx.ai does not have that image, so this post does not invent scoring categories. Use the two thresholds only as his stated cutoffs, and write target-criteria.md from the filters he already states: a license, headcount, founding year, an owner-operator, and work you can check.

What this changes about what you build and what you pay

The money in the headline does not change your stack. The review architecture does.

What you build. Five artifacts, all local:

  1. Target criteria, written down before anyone is charming on a call.
  2. An automation map of real completed jobs, classified and flagged where you are unsure.
  3. A rulebook that senior staff approve before anything is active.
  4. Shadow-mode runs, where agents work in parallel and a person still does the job the normal way.
  5. A preparer that cannot send client work. Sending, filing, and any client-visible message sit behind a reviewer that can block.

If you already run coding agents, you have built this shape on pull requests. Loop engineering is the schedule that re-prompts an agent and checks the result. Agent skills are the reusable procedures. Claude Code commands and steering with rules, hooks, and subagents are how authority gets split so one session cannot do every dangerous step. A services back office uses the same split on returns, tickets, and maintenance requests.

What you pay. Two bills, and only the first one is new. The services-firm purchase price, in his round-number illustration, is about $800,000. The model bill is the subscription a coding agent already uses. He names Codex. Claude Code is the same class of tool. You are not paying for a new foundation model. You are paying for the firm, and you are betting that review time falls faster than clients leave.

What you should not pay for yet. Headline margin multiples. A doubled margin in a two-year-old roll-up is a claim from a company that may be raising. Underwrite on your own shadow-mode correction rate. That rate is the natural metric of the corrections log: how often the approved version differs from the agent draft, and why. Isenberg did not print it as a dashboard number. It is the count his weekly review already produces. If that rate stays high, the 30 percent margin in the illustration is a wish.

His illustration, labeled as round numbers, before tooling and before anything goes wrong: buy around $800,000, change nothing about who pays, move margin from 10 percent to 30 percent, and at the same multiple the business is worth about three times the price. That is his illustration. It is not a valuation. The game is whether you can hold the margin after clients, staff, and review time are in the picture.

What files run the operating system?

He proposes an operating system for a solo operator or a tiny team. Three files do most of the work. The same idea shows up in coding setups as a small set of markdown files the agent must read, which explainx.ai covered in the agent markdown files guide.

table · 2 cols
FileJob
target-criteria.mdWhat you would buy, or what a firm you already own must still satisfy. Licenses, headcount, age, owner dependence, and whether output can be checked.
A rules directoryWhether agents can be trusted on a given task. Numbered rules, each with an example, each inactive until a named person approves it.
corrections-log.mdHow rules improve weekly. Every difference between a draft and the approved version gets a class. Repeats become proposed rules. Approved rules become tests.

Two agent sets sit on those files. One set helps you buy, or helps you decide not to. It reads public records against target-criteria.md and drafts questions. The other set helps you run the firm. It prepares work from the rules directory. The design rule covers both sets. The reviewer can block but never ship. The preparer can ship nothing on its own.

That split is the whole safety model. A single agent that drafts a client email and sends it has collapsed the two roles. A reviewer that can also "just fix it and send" has collapsed them the other way. Blocking is the reviewer's only write action. Shipping is a human act until the corrections log says otherwise, and even then only for a task the rulebook marks as safe.

Shadow mode is how you earn that mark. Agents do the work in parallel. A person does it the normal way. You compare. Clients see the person's version. This is the same discipline as a coding agent that opens a branch and cannot merge, which is the approval pattern in building useful agents with Claude Code. The agent harness is the name for the boundary that makes the block real: tools, permissions, and a stop before a consequential action.

Which prompts can you paste this week?

These are tightened from his instructions. They are not a verbatim transcript of the article. Each one tells the agent to mark inferences as inferences. Run them on public material or on files the firm has already given you. Do not point them at private personal data.

Sourcing, public records only

He lists licensing boards, secretary of state filings, association directories, marketplaces such as BizBuySell, and local journals. He prefers outreach before a firm is listed for sale. The prompt stays on those public surfaces.

text
Find owner-operated licensed services firms with 5 to 50 employees, founded before 2005, with no obvious succession. Use only public records: licensing boards, secretary of state filings, association directories, marketplaces such as BizBuySell, and local journals. Prefer firms that are not yet listed. For each firm, cite the public source. Mark every inference as an inference. Do not collect private personal data, home addresses, or personal contact details that are not already in the public filing you cite.

Automation map, after real jobs

Diligence starts with 20 to 50 anonymized completed jobs, then an automation map. His classes are: automate now, automate with human review, assist only, or keep human. Flag uncertainty. Multiply volume by hours so the map shows where the margin would have to come from.

text
Here are anonymized completed jobs. For each recurring task, classify it as automate now, automate with human review, assist only, or keep human. If you are unsure, flag the uncertainty instead of guessing. Multiply volume by hours and rank tasks by that product. Do not propose sending, filing, or emailing a client. Mark every inference as an inference.

Rule interview, nothing active until approved

The first 60 days of his rulebook start with senior staff. Answers become numbered rules with examples. Nothing is active until they approve.

text
These are interview notes from senior staff. Turn each answer into a numbered rule with one example of a correct job and one example of a wrong one. If the notes do not support a rule, say the information is missing. Do not activate any rule. Mark every inference as an inference. End with a list of rules waiting for a named person's approval.

Weekly corrections, drafts versus the approved version

Each week, compare the agent draft to the version a person accepted. Classify every change. A repeat becomes a proposed rule. An approved rule becomes a test: original input, accepted output.

text
Compare this agent draft to the approved version. Classify each change as a factual error, a client preference, missing information, or style. If the same correction appears more than once, propose one rule. For each rule a person has already approved, write a test as original input and accepted output. Do not change live client work. Mark every inference as an inference.

How does the rulebook grow in the first 60 days?

Interview senior staff before you automate their job. Ask how they know a return, a ticket, or a maintenance request is done. Ask what they refuse to send. Ask which clients have preferences that are not in the template. Turn the answers into numbered rules with examples. A rule without an example is a slogan. Nothing in the rules directory is active until those people approve it.

Then the weekly loop. Compare the agent draft to the approved version. Classify the change as a factual error, a client preference, missing information, or style. When a correction repeats, propose a rule. When a person approves that rule, add it as a test: the original input, and the output that was accepted. That is how a corrections log stops being a diary and starts being a suite. Coding agents already work this way when a failing check becomes a regression test. The services version uses the accepted workpaper as the expected output.

People who used to prepare the job move to review. That is a status change, and it is easy to skip. If the preparer is an agent and the old preparer has no review seat, you have removed the only person who can see a quiet error. His later warning is the business version of the same mistake: if margin rises while client retention falls, you sold the asset to pay for the renovation.

What does the first 100 days look like?

Clients should feel nothing for a month. The table is his sequence. The point of day 100 is a decision, not a press release.

table · 3 cols
WindowWhat clients seeWhat you do
Days 1 to 30No changeMeet every employee. Announce the plan with the former owner. Shadow mode: agents work in parallel, a person does the job the normal way, and you compare. Start the rulebook interviews.
Days 31 to 60Back office onlyIntake and document collection go live. Put the preparer on the highest-volume, lowest-risk task. Every draft is reviewed. Track minutes of human attention per job each week. People who prepared now review.
Days 61 to 100Routine status updates onlyTake the next two tasks. A client-comms agent may send routine status updates, and only those. Hold a monthly corrections review. Measure client retention and key-person retention next to margin.
Day 100The normal firm, plus whatever survived reviewYou should know how much work agents do reliably, what review costs, and whether anyone important is unhappy.

He names four measures in prose: margin, client retention, key-person retention, and minutes of human attention per job. A five-number dashboard was an image. This post does not invent the missing tile or print targets he did not write. Watch the four together. Margin without retention is the renovation problem above. Minutes of attention per job tell you whether the reviewer is actually cheaper than the old preparer.

Correction rate belongs next to those four, with a label. It is the natural metric of corrections-log.md: share of drafts that the approved version had to change, split by factual error, client preference, missing information, and style. It is not a number Isenberg printed.

How AI agents work end to end is the loop underneath a single task: gather context, act, check, stop. Shadow mode is that loop run beside a human, with the human's output as the one the client receives. The knowledge-worker guide covers the same office roles, research, drafting, and review, from the employee's side. This operating system is what changes when the person running the agent also owns the client relationship.

Which firms fit the pattern?

He lists accounting and bookkeeping, property management, insurance agencies, IT managed services, medical billing, payroll, HOA management, title and escrow, freight brokerage, and staffing. The filter under the list is more useful than the industry names. The work repeats. The output can be checked. Clients recur. Ownership is fragmented. The owner is near retirement.

A firm can sit in one of those industries and still fail the filter. A bookkeeping shop whose only judgment lives in the owner's head has nothing to put in the rules directory until you interview that owner and write the examples down. A freight desk whose "done" state is a relationship, not a document, belongs in "keep human" for the client-facing half even if the status entry can be drafted. Repeatable and checkable are the gates. Retirement and fragmentation explain why a seller might exist. They do not prove the work can be reviewed.

What breaks?

This is the section that should slow a buyer down. Isenberg's failure list is specific. Each item breaks the operating system in a different place.

Buying faster than you integrate. The scoreboard rewards speed: 18 businesses, 50 practices, ten companies. Integration is the shadow-mode month, the interviews, and the weekly corrections review. If the next close happens before the last rulebook is approved, you own a pile of firms and one tired reviewer. The margin thesis assumed delivery changed. The calendar assumed it had not.

Key people leave. Licenses, carrier appointments, and niche judgment often sit with a few seniors. Key-person retention is a named measure because the agent does not inherit their signature. A rulebook written after they leave is a guess. The first 30 days exist so you meet them while the former owner is still in the room.

Clients leave with the owner. The asset in the "why buy" argument is the client list. If the announcement is a fund, a rebrand, and a chatbot, clients who hired a person may follow that person. His deal shape, cash plus rolled equity, is one attempt to keep the founder attached. A personal buyer has the same problem on a smaller book. Announce with the former owner. Change nothing the client can see until shadow mode has a result.

Automating the relationship instead of the back office. Days 31 to 60 are intake and document collection, the highest-volume, lowest-risk task. The client-comms agent shows up in days 61 to 100, and only for routine status updates. Reversing that order means the first thing a client notices is a voice that is not the firm. Back-office drafts are checkable. A relationship is not.

Agents that can email clients without a checkpoint. This is the design rule failing in production. A preparer with send permission can ship. A reviewer who can ship will eventually ship to clear a queue. One wrong email is the liability in the counterweight above: a filing, a balance, a maintenance promise, sent to the wrong client or with the wrong number. The checkpoint is a person, or a second agent whose only power is to block, until a specific task has tests and a low correction rate.

Believing headline numbers from young companies that are raising money. The scoreboard is the temptation. $100 million of EBITDA, 4x margins, a billion-dollar buying commitment, a $1.5 billion pool. He says the figures are self-reported, and he says none of these companies has been through a recession. A downturn is when clients delay, staff leave, and the review queue gets long. Underwrite the firm in front of you. Use your shadow-mode correction rate, your minutes of attention per job, and your retention counts. Leave the headline multiple on the table until those move.

What people are asking

Does this replace the people who do the work? His 100-day plan moves preparers into review. It does not claim the firm runs unattended. Day 100 is supposed to tell you how much work agents do reliably and what review still costs. If every draft needs a full redo, the model is not "good enough to check." You are still redoing the job, and the 30 percent margin is not available.

Is Codex required? He says the Larson Gross tax agents run on OpenAI Codex. He also treats Claude Code as the same class of subscription an individual already has. The operating system is files, roles, and a block on send. The brand on the subscription is the smaller bill. Paying for a new foundation model is not the proposal.

What if the owner will not sit for a rule interview? Then you do not have a rulebook, and you should not turn agents on for client work. Past engagements are the training material only when someone who did the work will say what "correct" means. A folder of PDFs without that interview is an archive, not a test suite.

Is a letter of intent earned at 80 percent? Only in his prose, and only against a rubric this post does not have. An 80 percent score on categories we cannot see is not a reason to send paper. A letter of intent, if you are actually buying, should wait until the automation map exists and the former owner has agreed to the first 30 days of shadow mode. If you are not buying, ignore the threshold and run the shadow month inside the firm you already operate.

Related reading

The coding-agent posts below are the same review shape, pointed at repositories instead of client files.

  • Codex as a platform: the open agent harness, including the Thrive and Crete tax-prep pilot
  • Loop engineering: designing coding-agent loops that check their own work
  • What are agent skills?
  • Claude Code commands reference
  • Steering Claude Code with rules, hooks, and subagents
  • Build useful AI agents with Claude Code, including approval before a consequential send
  • Agent markdown files
  • What is an agent harness?
  • OpenAI's earlier primary post on the tax workflow: Codex as a platform

September 27, 2026; operator figures are the ones Isenberg reported and are not independently audited by explainx.ai.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 22, 2026

A Viral Agent Harness Tier List Put Claude Code in B — Does It Hold Up?

A tier-list image ranking agent harnesses — Oh My Pi alone in S-tier, Claude Code and Codex lumped into B alongside Cursor and Grok Build, GitHub Copilot and Antigravity in F — went viral on X September 21, 2026, racking up nearly 80,000 views and a comment section that disputed almost every placement. The single loudest complaint: Hermes doesn't appear on the list at all. Here's what the list actually claims, why the pushback matters more than the ranking, and how to build your own opinion instead of borrowing this one.

Sep 4, 2026

Armature Study: What Claude Code, Codex, and Cursor Actually Pick

Armature, a startup that sells "growth services to dev tools," measured 16,893 coding-agent sessions to see which tools Claude Code, Codex, and Cursor actually pick — not just mention. The findings are genuinely useful (repo language flips winners, mentions don't equal picks) and the source is a genuine conflict of interest. Here's both, held at once.

Jun 21, 2026

Why Every AI Company Wants You Using Agents: The Token Economics Nobody Talks About

A single Claude Code /loop session burns more tokens than 50 chat messages. An agentic Codex browser-use task that writes code, pushes to GitHub, and configures Vercel burns more tokens than a week of casual ChatGPT use. Anthropic, OpenAI, and every AI company building agent products has aligned incentives: the more agentic your workflow, the more they earn. This is not a conspiracy. It is business model economics. Here is how to think about it.