explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: claims and checks
  • What is Pine Computer?
  • What is SaaS-Bench, and what do the two scores mean?
  • The v1.1 table, in full
  • Is the cost claim real?
  • What about the speed and customer claims?
  • What is still unverified?
  • What does this mean for builders?
  • What to watch next
  • Related reading
← Back to blog

explainx / blog

Pine Computer Scores 78.3% on SaaS-Bench, but Finishes Fewer Tasks: Claimed vs Verified

AI Agents, Computer Use, Benchmarks, Pine AI, SaaS-Bench

Part of AI Agents

Pine Computer tops SaaS-Bench v1.1 at 78.3% checkpoints but resolves 27.4% of tasks, behind Claude Code at 31.1%. What is claimed and what is verified.

Oct 10, 2026·10 min read·Yash Thakker
add explainx.ai
go deep
Pine Computer Scores 78.3% on SaaS-Bench, but Finishes Fewer Tasks: Claimed vs Verified

Short answer: Pine AI launched Pine Computer on October 9, 2026, a cloud computer built for AI agents, and reports the highest checkpoint score on UniPat AI's SaaS-Bench v1.1: 78.3%, ahead of Opus 5 with Claude Code at 74.3% and GPT-5.6 Sol with Codex at 71.1%. The same table shows Pine Computer resolving 27.4% of whole tasks, behind Claude Code at 31.1% and Codex at 29.2%. So the headline is true, and so is its opposite: Pine gets further on average, and finishes fewer jobs. The product is in private beta, and the numbers are Pine-reported.

This is a "claimed vs verified" post. We lay out what Pine says, what the benchmark's own definitions confirm, and what remains unchecked. Sources are Pine's press release, RuntimeWire's coverage, UniPat AI's SaaS-Bench page, and the SaaS-Bench paper.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

A cursor arrow reaching toward a glowing button, representing an AI agent using a cloud computerA cursor arrow reaching toward a glowing button, representing an AI agent using a cloud computer

TL;DR: claims and checks

table · 2 cols
ClaimStatus
Pine Computer has the highest SaaS-Bench v1.1 checkpoint score, 78.3%Listed on the UniPat page and in Pine materials; no independent reproduction found
It resolves more tasks than rivalsFalse by Pine's own table: 27.4% versus 31.1% (Claude Code) and 29.2% (Codex)
About $1.02 model cost per taskPine-reported; excludes infrastructure; systems differ in software and budget
2 to 5 times faster than AI on conventional computersPine's preliminary internal tests; varies by task
An enterprise customer took on 50 percent more work with the same peopleCompany-reported, unnamed customer, method not explained
Reads pages as structure, not screenshotsDescribed by Pine; bring-your-own models currently work from screenshots
Available nowPrivate beta by waitlist only

What is Pine Computer?

Pine AI, the company behind a consumer assistant that handles phone calls, websites, and software on behalf of users, describes Pine Computer as "a computer built for AI." A developer's product uses an SDK to create a Pine Computer when a job needs one, hands it the task, and receives the finished work: files, records, and answers. Pine runs the computer, its intelligence layer, the browser, and the isolation. The developer builds the surrounding experience and keeps their own keys.

The design bet is stated in the launch quote from co-founder and chief architect Dylan Wang: "We spent years making the model smarter. Now, we're making the computer worthy of the model." In practice that means three departures from the common agent setup.

Structure instead of screenshots. The usual loop is screenshot, inspect, act. Pine says its browser and operating environment instead report page structure, the actions available, and what changed after each action, so the model reads the page as data rather than guessing from pixels. Visual input remains available for people, who can watch a live screen. For developers who bring their own model, the release says that model currently works from screenshots, which is worth noting: the structured-page advantage applies to Pine's own intelligence layer.

Isolation by default. Each Pine Computer runs in its own sandbox, separate from a user's laptop, tabs, and personal files. Pine's own explainer, Pine Computer, contrasts this with giving an agent access to your personal device.

Human takeover. When a CAPTCHA, multi-factor prompt, or approval step appears, a person can view the live desktop, complete the step, and hand control back. A live screen can be embedded in the developer's product.

Three blank browser tab outlines with a green cursor path, standing for a browser agent working across sitesThree blank browser tab outlines with a green cursor path, standing for a browser agent working across sites

If this sounds like other agent runtimes, it is partly a crowded category. Anthropic's platform now ships computer use with browser skills and a files API (Claude Platform computer use GA), OpenAI has extended Codex computer use to Windows and mobile (Codex computer use), and Cloudflare has an agent browser built on V8 isolates (Kitesurf). Pine's pitch is the environment itself, sold as a layer under someone else's product.

What is SaaS-Bench, and what do the two scores mean?

SaaS-Bench comes from UniPat AI and Peking University researchers. It is built on 23 deployable SaaS systems across six professional domains, with 106 long-horizon tasks, 74 text-only and 32 multimodal. Ninety-three percent of tasks span at least two applications. Examples named by UniPat and the paper include an expense reimbursement closeout across HR, accounting, and CRM systems, a duplicate patient-merge audit in an electronic medical record system, and a regression test audit across a code editor, a database tool, and a project tracker.

Scoring works through verification checkpoints, averaging 12.8 per task. Each checkpoint has a weight and ties to an expected final state or artifact in the system.

  • Checkpoint score (CS): the weighted fraction of checkpoints passed. It rewards partial progress.
  • Resolved score (RS): 1 only if every checkpoint passes, otherwise 0. It measures whether the job was finished.

That is the whole story of the headline. A system can pass 78% of checkpoints on average and still fail to pass all of them on three in four tasks. Pine's own press release says so: "By whole-task completion, Pine Computer trailed both systems."

A three-step podium with a green star, representing the SaaS-Bench v1.1 leaderboard rankingA three-step podium with a green star, representing the SaaS-Bench v1.1 leaderboard ranking

The v1.1 table, in full

UniPat's page lists each model with its harness, because the same model scores differently under different harnesses.

table · 4 cols
ModelHarnessCheckpoint (%)Resolved (%)
Pine Computer (GPT-5.6 Luna)Pine Computer Runtime78.327.4
Opus 5Claude Code74.331.1
GPT-5.6 SolCodex71.129.2
GPT-5.6 SolBrowser-Use (open source)69.817.9
Qwen 3.8 maxQwen Code65.923.6
Opus 5Browser-Use (open source)64.721.7
Kimi K3Kimi-CLI56.424.5
Kimi K3Browser-Use (open source)56.217.0
Qwen 3.8 maxBrowser-Use (open source)31.18.5

Two readings stand out. First, the harness matters as much as the model: Opus 5 goes from 64.7 to 74.3 on checkpoints and from 21.7 to 31.1 on resolved tasks when moved from the open-source Browser-Use harness to Claude Code. Second, Pine's result is a whole-system result. The Pine row runs a model called GPT-5.6 Luna on Pine's runtime; the rivals run different, larger models on different harnesses. For background on those model tiers see our GPT-5.6 Sol, Terra, Luna preview, and for the harnesses see Claude Code vs Codex vs Gemini CLI and Opus 5 launch coverage.

Be careful comparing these numbers with the original paper. The May 2026 SaaS-Bench paper reported that even the strongest model resolved fewer than 4% of tasks end to end, with the best resolved rate across 14 models at 3.8%. That was an earlier evaluation round with older models; the v1.1 numbers are a different round and should not be compared directly.

Is the cost claim real?

Pine reports about $1.02 in model-token cost per task. The comparison rows are $26.50 for Opus 5 with Claude Code and $20.50 for GPT-5.6 Sol with Codex. That is roughly a 20 to 26 times gap.

Three checks apply. The cost is model tokens only, with infrastructure excluded. The systems have different software and budgets, so the gap does not isolate Pine's computer. And the numbers are Pine's.

Our own arithmetic adds a useful angle, with a loud caveat. If you divide each system's per-task cost by its resolved rate, you get a crude cost per finished task: about $3.72 for Pine ($1.02 divided by 0.274), about $85 for Claude Code, and about $70 for Codex. This assumes the average cost applies equally to finished and unfinished attempts, which is unlikely, so read it as direction, not a quote. Even so, the direction is the reason a cheaper, slightly less reliable runtime can win in a real workflow: if a failed attempt is cheap, you can retry or hand the last step to a person.

What about the speed and customer claims?

Pine says that in preliminary internal tests Pine Computer was 2 to 5 times faster than AI on conventional computers, and that results vary by task. That is plausible for a design that skips screenshot round trips, but there is no public methodology, so we label it a claim.

The press release also says an unnamed enterprise customer, using Pine's auditing automation, took on 50% more work with the same people. The release does not explain how that was measured. RuntimeWire flags the same limits. A quote from Andrew Mackenzie, co-founder of Subliminal, which evaluated the product, says Pine had "thought about all of this and built it all in"; it is an endorsement, not data.

What is still unverified?

  • Independent reproduction. Pine says run data is published on Hugging Face. We found no third-party re-run of the Pine row.
  • Security in practice. Isolation and permission handling are described as design features. RuntimeWire notes the beta has not yet shown how reliably they work. Agents that act inside business software are exactly where mistakes cost money; see how a browser agent's autonomy went wrong in September.
  • Model dependence. The headline row uses GPT-5.6 Luna. How much of the 78.3% comes from the runtime and how much from the model is not separable with this data, and a different model on the same runtime was not published.
  • Run-to-run variance. The SaaS-Bench paper lists high variance as a known failure mode: the same model can score very differently across runs on the same task. A single published number, from one team, sits inside that noise.

What does this mean for builders?

If you run agents against internal tools or SaaS apps, three takeaways hold whatever happens to Pine.

  1. Ask which metric a vendor quotes. Partial-progress scores are fine for debugging and misleading for buying. Your workflow either finishes or it needs a person.
  2. Budget for the last mile. A system that passes most checkpoints and fails the final one is useful only if you design the handoff. Pine's human-takeover feature is a bet on exactly that.
  3. Sandbox first, then trust. Whatever runtime you pick, keep agent access to production systems narrow, and put a policy layer between the agent and its tools. AgentBeam, the agent security platform from the explainx.ai team, stops AI agents before they take dangerous actions.

A headset beside an outlined hand receiving a green parcel, representing a human taking over an AI agent sessionA headset beside an outlined hand receiving a green parcel, representing a human taking over an AI agent session

The pricing angle matters too. When agents hit SaaS products designed for people, vendors start charging per interaction; see headless SaaS agents and per-interaction pricing.

What to watch next

  1. Whether UniPat or a third party reproduces the Pine row, and with which models.
  2. A Pine Computer result with a different model, to separate runtime from model.
  3. Pine's promised open designs and specifications, which it has not dated.
  4. Whether the private beta turns into public pricing, which is the number that decides who can use it.

Figures are accurate as of October 10, 2026. They come from Pine's launch materials, the UniPat SaaS-Bench page, and RuntimeWire; the benchmark results have not been independently reproduced to our knowledge.

Related reading

  • Claude Platform computer use, browser skills and files API
  • OpenAI Codex computer use on Windows and mobile
  • Cloudflare Kitesurf agent browser
  • Headless SaaS agents and per-interaction pricing
  • Claude Code vs Codex vs Gemini CLI vs GLM-5.2
  • GPT-5.6 Sol with Claude Code setup guide
  • GPT-5.6 Sol, Terra and Luna preview
Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 8, 2026

GPT-6 Astra Clears MazeBench and Every "I'm Not a Robot" Level

Two GPT-6 Astra capability demos went viral in the same 24 hours: a reported 7x lead over Claude Fable 5.1 on MazeBench, and a full clear of all 48 levels of the "I'm Not a Robot" browser game using computer-use tools. Here's what MazeBench measures, what the CAPTCHA clear actually shows about browser control, and why "beat a human test" isn't the same claim as AGI.

Oct 7, 2026

OpenAI and Ironclad: How GPT-6 Astra Was Trained on Contracting Workflows

On October 6, 2026, OpenAI described a research partnership with Ironclad to turn contracting software workflows into training and evaluation tasks for computer-use agents. GPT-6 Astra is the first frontier model trained on them. Here is what was measured, what was not, and why software companies should care.

Oct 4, 2026

ThinkingBox: Microsoft Grades AI Agents on Database State, Not Chat

Microsoft researchers released ThinkingBox, an open sandbox and benchmark that checks what an agent actually changed in a backend instead of what it said. Across 121,680 trials, 79,853 failed attempts, and two-thirds of those ended with no tool error. Here is what the numbers show and how to apply the idea to your own agents.