Short answer: Pine AI launched Pine Computer on October 9, 2026, a cloud computer built for AI agents, and reports the highest checkpoint score on UniPat AI's SaaS-Bench v1.1: 78.3%, ahead of Opus 5 with Claude Code at 74.3% and GPT-5.6 Sol with Codex at 71.1%. The same table shows Pine Computer resolving 27.4% of whole tasks, behind Claude Code at 31.1% and Codex at 29.2%. So the headline is true, and so is its opposite: Pine gets further on average, and finishes fewer jobs. The product is in private beta, and the numbers are Pine-reported.
This is a "claimed vs verified" post. We lay out what Pine says, what the benchmark's own definitions confirm, and what remains unchecked. Sources are Pine's press release, RuntimeWire's coverage, UniPat AI's SaaS-Bench page, and the SaaS-Bench paper.
A cursor arrow reaching toward a glowing button, representing an AI agent using a cloud computer
TL;DR: claims and checks
| Claim | Status |
|---|---|
| Pine Computer has the highest SaaS-Bench v1.1 checkpoint score, 78.3% | Listed on the UniPat page and in Pine materials; no independent reproduction found |
| It resolves more tasks than rivals | False by Pine's own table: 27.4% versus 31.1% (Claude Code) and 29.2% (Codex) |
| About $1.02 model cost per task | Pine-reported; excludes infrastructure; systems differ in software and budget |
| 2 to 5 times faster than AI on conventional computers | Pine's preliminary internal tests; varies by task |
| An enterprise customer took on 50 percent more work with the same people | Company-reported, unnamed customer, method not explained |
| Reads pages as structure, not screenshots | Described by Pine; bring-your-own models currently work from screenshots |
| Available now | Private beta by waitlist only |
What is Pine Computer?
Pine AI, the company behind a consumer assistant that handles phone calls, websites, and software on behalf of users, describes Pine Computer as "a computer built for AI." A developer's product uses an SDK to create a Pine Computer when a job needs one, hands it the task, and receives the finished work: files, records, and answers. Pine runs the computer, its intelligence layer, the browser, and the isolation. The developer builds the surrounding experience and keeps their own keys.
The design bet is stated in the launch quote from co-founder and chief architect Dylan Wang: "We spent years making the model smarter. Now, we're making the computer worthy of the model." In practice that means three departures from the common agent setup.
Structure instead of screenshots. The usual loop is screenshot, inspect, act. Pine says its browser and operating environment instead report page structure, the actions available, and what changed after each action, so the model reads the page as data rather than guessing from pixels. Visual input remains available for people, who can watch a live screen. For developers who bring their own model, the release says that model currently works from screenshots, which is worth noting: the structured-page advantage applies to Pine's own intelligence layer.
Isolation by default. Each Pine Computer runs in its own sandbox, separate from a user's laptop, tabs, and personal files. Pine's own explainer, Pine Computer, contrasts this with giving an agent access to your personal device.
Human takeover. When a CAPTCHA, multi-factor prompt, or approval step appears, a person can view the live desktop, complete the step, and hand control back. A live screen can be embedded in the developer's product.
Three blank browser tab outlines with a green cursor path, standing for a browser agent working across sites
If this sounds like other agent runtimes, it is partly a crowded category. Anthropic's platform now ships computer use with browser skills and a files API (Claude Platform computer use GA), OpenAI has extended Codex computer use to Windows and mobile (Codex computer use), and Cloudflare has an agent browser built on V8 isolates (Kitesurf). Pine's pitch is the environment itself, sold as a layer under someone else's product.
What is SaaS-Bench, and what do the two scores mean?
SaaS-Bench comes from UniPat AI and Peking University researchers. It is built on 23 deployable SaaS systems across six professional domains, with 106 long-horizon tasks, 74 text-only and 32 multimodal. Ninety-three percent of tasks span at least two applications. Examples named by UniPat and the paper include an expense reimbursement closeout across HR, accounting, and CRM systems, a duplicate patient-merge audit in an electronic medical record system, and a regression test audit across a code editor, a database tool, and a project tracker.
Scoring works through verification checkpoints, averaging 12.8 per task. Each checkpoint has a weight and ties to an expected final state or artifact in the system.
- Checkpoint score (CS): the weighted fraction of checkpoints passed. It rewards partial progress.
- Resolved score (RS): 1 only if every checkpoint passes, otherwise 0. It measures whether the job was finished.
That is the whole story of the headline. A system can pass 78% of checkpoints on average and still fail to pass all of them on three in four tasks. Pine's own press release says so: "By whole-task completion, Pine Computer trailed both systems."
A three-step podium with a green star, representing the SaaS-Bench v1.1 leaderboard ranking
The v1.1 table, in full
UniPat's page lists each model with its harness, because the same model scores differently under different harnesses.
| Model | Harness | Checkpoint (%) | Resolved (%) |
|---|---|---|---|
| Pine Computer (GPT-5.6 Luna) | Pine Computer Runtime | 78.3 | 27.4 |
| Opus 5 | Claude Code | 74.3 | 31.1 |
| GPT-5.6 Sol | Codex | 71.1 | 29.2 |
| GPT-5.6 Sol | Browser-Use (open source) | 69.8 | 17.9 |
| Qwen 3.8 max | Qwen Code | 65.9 | 23.6 |
| Opus 5 | Browser-Use (open source) | 64.7 | 21.7 |
| Kimi K3 | Kimi-CLI | 56.4 | 24.5 |
| Kimi K3 | Browser-Use (open source) | 56.2 | 17.0 |
| Qwen 3.8 max | Browser-Use (open source) | 31.1 | 8.5 |
Two readings stand out. First, the harness matters as much as the model: Opus 5 goes from 64.7 to 74.3 on checkpoints and from 21.7 to 31.1 on resolved tasks when moved from the open-source Browser-Use harness to Claude Code. Second, Pine's result is a whole-system result. The Pine row runs a model called GPT-5.6 Luna on Pine's runtime; the rivals run different, larger models on different harnesses. For background on those model tiers see our GPT-5.6 Sol, Terra, Luna preview, and for the harnesses see Claude Code vs Codex vs Gemini CLI and Opus 5 launch coverage.
Be careful comparing these numbers with the original paper. The May 2026 SaaS-Bench paper reported that even the strongest model resolved fewer than 4% of tasks end to end, with the best resolved rate across 14 models at 3.8%. That was an earlier evaluation round with older models; the v1.1 numbers are a different round and should not be compared directly.
Is the cost claim real?
Pine reports about $1.02 in model-token cost per task. The comparison rows are $26.50 for Opus 5 with Claude Code and $20.50 for GPT-5.6 Sol with Codex. That is roughly a 20 to 26 times gap.
Three checks apply. The cost is model tokens only, with infrastructure excluded. The systems have different software and budgets, so the gap does not isolate Pine's computer. And the numbers are Pine's.
Our own arithmetic adds a useful angle, with a loud caveat. If you divide each system's per-task cost by its resolved rate, you get a crude cost per finished task: about $3.72 for Pine ($1.02 divided by 0.274), about $85 for Claude Code, and about $70 for Codex. This assumes the average cost applies equally to finished and unfinished attempts, which is unlikely, so read it as direction, not a quote. Even so, the direction is the reason a cheaper, slightly less reliable runtime can win in a real workflow: if a failed attempt is cheap, you can retry or hand the last step to a person.
What about the speed and customer claims?
Pine says that in preliminary internal tests Pine Computer was 2 to 5 times faster than AI on conventional computers, and that results vary by task. That is plausible for a design that skips screenshot round trips, but there is no public methodology, so we label it a claim.
The press release also says an unnamed enterprise customer, using Pine's auditing automation, took on 50% more work with the same people. The release does not explain how that was measured. RuntimeWire flags the same limits. A quote from Andrew Mackenzie, co-founder of Subliminal, which evaluated the product, says Pine had "thought about all of this and built it all in"; it is an endorsement, not data.
What is still unverified?
- Independent reproduction. Pine says run data is published on Hugging Face. We found no third-party re-run of the Pine row.
- Security in practice. Isolation and permission handling are described as design features. RuntimeWire notes the beta has not yet shown how reliably they work. Agents that act inside business software are exactly where mistakes cost money; see how a browser agent's autonomy went wrong in September.
- Model dependence. The headline row uses GPT-5.6 Luna. How much of the 78.3% comes from the runtime and how much from the model is not separable with this data, and a different model on the same runtime was not published.
- Run-to-run variance. The SaaS-Bench paper lists high variance as a known failure mode: the same model can score very differently across runs on the same task. A single published number, from one team, sits inside that noise.
What does this mean for builders?
If you run agents against internal tools or SaaS apps, three takeaways hold whatever happens to Pine.
- Ask which metric a vendor quotes. Partial-progress scores are fine for debugging and misleading for buying. Your workflow either finishes or it needs a person.
- Budget for the last mile. A system that passes most checkpoints and fails the final one is useful only if you design the handoff. Pine's human-takeover feature is a bet on exactly that.
- Sandbox first, then trust. Whatever runtime you pick, keep agent access to production systems narrow, and put a policy layer between the agent and its tools. AgentBeam, the agent security platform from the explainx.ai team, stops AI agents before they take dangerous actions.
A headset beside an outlined hand receiving a green parcel, representing a human taking over an AI agent session
The pricing angle matters too. When agents hit SaaS products designed for people, vendors start charging per interaction; see headless SaaS agents and per-interaction pricing.
What to watch next
- Whether UniPat or a third party reproduces the Pine row, and with which models.
- A Pine Computer result with a different model, to separate runtime from model.
- Pine's promised open designs and specifications, which it has not dated.
- Whether the private beta turns into public pricing, which is the number that decides who can use it.
Figures are accurate as of October 10, 2026. They come from Pine's launch materials, the UniPat SaaS-Bench page, and RuntimeWire; the benchmark results have not been independently reproduced to our knowledge.
Related reading
- Claude Platform computer use, browser skills and files API
- OpenAI Codex computer use on Windows and mobile
- Cloudflare Kitesurf agent browser
- Headless SaaS agents and per-interaction pricing
- Claude Code vs Codex vs Gemini CLI vs GLM-5.2
- GPT-5.6 Sol with Claude Code setup guide
- GPT-5.6 Sol, Terra and Luna preview
