The Verge has started a hands-on diary of running local AI, and its first entry is a useful reality check for anyone curious about self-hosted agents. Laptop reviewer Antonio G. Di Benedetto installed the open-source Hermes Agent on an M5 Ultra Mac Studio with 256GB of unified memory, picked a 125-billion-parameter Qwen model, and spent his first days finding out what a private "gofer" is actually good for. His verdict in the title is honest: "exciting, overwhelming, and frustrating." The original piece is on The Verge, published October 11, 2026.
This post condenses what he did, adds context from Hermes' own documentation and our earlier coverage, and turns it into a practical starting plan. Everything attributed to the diary below is the reviewer's own account, not our testing.
Quick reference: the setup and the verdict
| Question | Answer from the diary |
|---|---|
| Which agent? | Hermes Agent, a self-hosted desktop app for macOS, Windows and Linux |
| Cost? | Free if you stick to local LLMs, per The Verge |
| Which model? | Qwen 3.8 Flash Next, 125 billion parameters, about 105GB |
| Which hardware? | M5 Ultra Mac Studio, 256GB unified memory |
| How controlled? | Hermes onboarding and model picker, plus a Telegram bot from his phone |
| First tasks? | Daily briefing, Steam library sorting, financial data analysis, spec spreadsheet, benchmark automation |
| Biggest win? | Tasks he would not trust to a cloud service |
| Biggest problem? | Choosing models, knowing what to ask, and jobs that break |
Why try local AI at all?
The reviewer's motivation is privacy, not cost. He writes that giving his "most personal data to cloud services has been a major hold up," and that a helper "living solely in the box on my desk" that answers only to him feels different from subscribing to a chatbot. He still will not have it write or edit for him. The appeal is delegating trusted busy work.
That framing matches a broader mood we have tracked. Vendors are pitching local AI hardware hard: the Verge notes Apple's pitch for its new Mac desktops and a line of RTX Spark Windows machines with up to 128GB of RAM aimed at agentic AI. We covered that hardware in our posts on the Surface RTX Spark dev box and the Nvidia and Nous pairing of Hermes Desktop with RTX Spark. The privacy argument also sits against the backdrop of promises from cloud agent makers, which we examined in the Verge's analysis of AI agent privacy promises.
Local AI agent concept: a small house-shaped machine with a glowing core running models at home
What did the setup look like?
He started with Hermes because its onboarding interface and model picker are straightforward. With 256GB of unified memory, he writes, he could "run just about any model," and because local inference has no per-token bill he went big first: Qwen 3.8 Flash Next at 125 billion parameters, about 105GB on disk. Getting Hermes and Qwen running "wasn't long," and he then made them controllable from his phone using a Telegram bot.
Then came what he calls his usual nemesis: the empty text box. If you have felt that, our top 10 things you can do with Hermes Agent and the original Hermes Agent guide give starting prompts. The official Hermes Agent documentation is the reference for installation and the model picker.
A note on sizing, since it is the first thing beginners get wrong. A model's file size is the floor for the memory you need, not the ceiling: the context window and runtime overhead sit on top. The diary's 105GB model on a 256GB machine leaves generous headroom, which is why he could pick the model first and worry about fit later. On a 16GB or 32GB machine the order reverses. You choose a quantized model of a few billion parameters that fits, and you accept that it will handle simpler jobs. That trade is exactly why the reviewer plans to repeat the exercise on a Mac Mini, a MacBook Air, an AMD Strix Halo laptop and RTX Spark hardware.
Task 1: a daily briefing, and the sleep gotcha
On advice from YouTube, he had Hermes build a morning briefing as a cron job: scan email and calendar, flag anything urgent, add a short weather report, and send it daily via Telegram. He admits it is basic and "not all that useful," and that you do not need a $12,000 computer for it. The value is as a smoke test.
The failure is the instructive part. The briefing kept failing until he figured out that macOS could not be asleep when the job fired at 7:30am. Scheduled agent work inherits every limitation of the host: sleep settings, network, logged-in sessions. If your local agent runs cron jobs, set the machine to stay awake or use a small always-on box, as we describe in our Hermes remote VPS guide.
A clock placing a green bead on a track, illustrating a scheduled local AI agent cron job
Task 2: sorting 400 Steam games
His more useful task was reorganizing a Steam library of over 400 games. Steam does not auto-categorize, so doing it by hand is tedious. Hermes could see most games from the installed Steam client, proposed organization options, and sorted games by genre while he kept his own categories such as favorites, co-op and party games.
Permissions are the lesson. He had to register a Steam web API key and give it to Hermes, then revoked it right after because the heavy lifting was a one-off. That is a good pattern for any local agent: create a narrow, temporary credential, use it, delete it. A local agent can still act on your machine and accounts, and tools that supervise agent actions, such as AgentBeam from the explainx.ai team, exist for exactly that reason.
A key beside an open pocket with drops, representing revoking an API key after a local AI agent task
Task 3: data he would not upload
This is where local AI paid for itself in his telling. He had Hermes crunch financial records and build a spec comparison spreadsheet for a new laptop. Both are things he "cannot or will not trust with any cloud service": the financial data is sensitive and the laptop info was under embargo. Keeping them on his own computer "was the difference in me using this tool or not."
That is a clear decision rule. If a task involves data you could not paste into a cloud chatbot, local is justified even with a smaller model. If it does not, a cloud model may be cheaper and better. It also explains why the first useful tasks were boring ones. Privacy-driven work tends to be spreadsheets, records and drafts under embargo, not creative writing, and boring structured tasks are where current local models are most dependable.
Task 4: automating benchmarks, still unfinished
He currently runs laptop benchmarks manually, a dozen or so tests, three times each to average. He is trying to have Hermes write scripts to run them. Walking it through procedures and getting usable "vibe coded Python scripts" has been an endeavor and remains a work in progress. This is typical: agents help most when you can describe a procedure precisely, and least when you are still discovering the procedure. A practical tip is to do the task manually once while writing down every step, then hand the written procedure to the agent rather than asking it to figure the task out from a one-line goal.
What is still hard?
The diary is candid about limits:
- Model overload. "There are more models out there than I can count," some specialized, and bigger generally means more capable.
- Not a magic assistant. Even powerful local models "don't make for a magical assistant that can do anything." His daily briefing had already broken multiple times.
- Cost of the hardware. He notes Apple demoed a cluster of four Mac Studios for local AI that was nearly $50,000 worth of compute.
- Mindset. He treats Hermes as a tool, not a digital friend: no "please," no "thank you."
How to start local AI without a 256GB machine
The reviewer plans to test much smaller Qwen models on an M6 Mac Mini, an M5 MacBook Air, an Asus TUF Gaming A14 with AMD Strix Halo, and eventually RTX Spark. You can follow the same path with whatever you own:
- Pick one boring job. A daily briefing or a file-sorting task beats an open-ended "assistant."
- Match the model to memory. A model must fit in RAM or unified memory with room for context; the diary's 105GB model fits comfortably in 256GB. Start smaller than you think.
- Choose tasks for privacy. Use the sensitive-data test above.
- Give narrow permissions. Temporary keys, read-only where possible.
- Expect to debug the host. Sleep, cron and network issues will bite before the model does.
- Compare harnesses. See our Hermes Agent vs OpenClaw comparison and our post on llama.cpp performance on Apple Silicon for runtime choices.
For smaller private options, see Underdog's private personal AI launch.
What people are asking
Is local AI actually private? Inference stays on your machine, which removes the cloud provider from the picture. It does not remove the agent's own access. Anything you grant, from an API key to file access, is still exposed to whatever the agent does, so scope it narrowly.
Is it cheaper than a subscription? Not at this price point. The reviewer's machine costs around $12,000 by his own figure, so the case is privacy and control, not savings. Smaller models on hardware you already own change that math.
Which model should a beginner choose? There is no single answer in the diary. He chose the largest he could fit and says the field is overwhelming. Begin with a small model from a family you recognize, confirm your job works, and move up only if quality falls short.
What breaks first? Per the diary, the surrounding system: scheduling, permissions and procedures, not the model.
Our read
The diary is valuable because it is unglamorous. The headline capability, a 125B model on a desk, worked. The hard parts were the same ones every agent user meets: deciding what to ask, granting access safely, and keeping scheduled jobs alive. Local AI makes the privacy case strong for specific tasks and the economics case weak for casual ones. We will watch for later entries, especially the small-model tests, which will say more about what ordinary readers can do than a $12,000 desktop can.
Related reading
- Top 10 things you can do with Hermes Agent
- What is Hermes Agent and how does it work
- Hermes Agent vs OpenClaw
- Nvidia and Nous: Hermes Desktop on RTX Spark
- Surface RTX Spark dev box
- AI agent privacy promises
- The Verge local AI diary
Details are from The Verge's October 11, 2026 article and Hermes documentation, and may change.
