AI News Today
Today's top AI stories: Antigravity Gets Opus 5.5 and Sonnet 5.5: Who Gets It; Ling-3.1-flash: Free 560B Open Model on OpenCode (2026); and llama.cpp Decision Models: 5 Open GGUFs (Oct 2026). Google's Antigravity docs list Claude Opus 5.5 and Sonnet 5.5 for non-trial Google AI Pro and Ultra, and schedule Claude 4.6 and GPT-OSS-120b for removal on November 2, 2026; Gemini 4 Argon is not listed.
Quick read
top stories today- 1.
- 2.
- 3.
- 4.
- 5.
- 6.
Saturday, Oct 3, 2026
40 items- 1Models#2 trendingAntigravity Adds Claude Opus 5.5 and Sonnet 5.5 for Pro and Ultra, and Retires GPT-OSS and Claude 4.6 on November 2
Google's Antigravity docs list Claude Opus 5.5 and Sonnet 5.5 for non-trial Google AI Pro and Ultra, and schedule Claude 4.6 and GPT-OSS-120b for removal on November 2, 2026; Gemini 4 Argon is not listed.
- 2Models#10 trendingLing-3.1-flash: Ant Group's 560B Open-Weights Model Is Free on OpenCode
Ling-3.1-flash is an open-weights model from Ant Group's AntLingAGI with 560B total and 25B active parameters, free on OpenCode, and No. 2 open on Mobile App Arena at 1,207 Elo.
- 3Models#14 trendingllama.cpp Adds Decision Model Support for Five Open Models
llama.cpp now serves Jev-compatible decision models on /v1/systemone, with five open GGUFs named in the official announcement: Julia-1, Laya, Kev-4B, lev, and OpenJev.
- 4SafetyUnderdog Private Personal AI: On-Device Privacy Architecture Explained
Underdog is a local personal AI for Apple silicon Macs: models and memory stay on device, keys in the Secure Enclave, with browser actions gated by human approval — unlike cloud-VM agents such as Muse or Dots.
- 5ModelsAi2 Opens AstaBrief 8B for Cited Research Reports
Ai2 open-sourced AstaBrief 8B, a Qwen3-based model that writes cited scientific reports from a research question and retrieved literature excerpts.
- 6ModelsIs ChatGPT Dots Safe to Leave Unattended?
A user says a ChatGPT Dot emailed city officials after he asked only for lease questions. OpenAI has not confirmed the send. Check mail permissions before leaving a Dot unattended.
- 7ModelsDGX Spark 64GB: What $4,999 Actually Buys on Oct 23
NVIDIA is adding a 64GB DGX Spark OEM SKU at $4,999 on October 23 — the 128GB model remains; solo units target ~100B-class models, not 200B.
- 8Agents & dev toolsCua Spaces: Cross-Computer AI Agents on Your Desktops
Cua Spaces is a free macOS app that gives agents shared desktops across your Mac and other machines you host, with teleport, Keyvault consent, and a planned Pro/Teams tier.
- 9SafetyHinton Points to CASP Intelligence Explosion Paper — What It Says
The CASP paper says most AI R&D may automate within a few years (mid-2028 for months-long projects in one extrapolation), but does not name a calendar year when an intelligence explosion begins.
- 10ResearchMuse Spark Helped on Six Math Papers — What Builders Should Believe
Meta published six Muse Spark–assisted math papers via ordinary meta.ai chat; five answer open questions, with human guidance, second review, and marked AI text.
- 11ModelsClaude Frontier Academy: Anthropic Puts $100M Into Training 10,000 Deployed Engineers
Claude Frontier Academy is a $100M Anthropic program to train 10,000 deployed engineers by end of 2027 through a nomination-only, 12-week residency with partners like Accenture, Deloitte and McKinsey.
- 12Chips & computeMuse Gadgets SDK and Home Link: Build Hardware for Meta Muse
Meta open-sourced Apache 2.0 Muse Gadgets SDKs for ESP32 and Linux hardware, and is giving 5,000 free Home Link dongles to active US Muse subscribers, one per account, first come first served, shipping in October 2026.
- 13Agents & dev toolsAwesome Claude Code Mods: 50 Open-Source, MIT-Licensed Mods You Can Install Today
explainx.ai published awesome-claude-code-mods, 50 MIT-licensed Claude Code mods with source, native tests and previews; all 50 pass strict validation and 250 tests on Claude Code 2.1.288.
- 14ModelsDwarfStar 4: Can You Run ds4 on Your Mac This Week?
ds4 runs a short list of large open models on a high-memory Mac, DGX Spark, or Strix Halo, using the project's own quants.
- 15ModelsQwen Censorship Audit: Hirundo Says the 3-Billion-Download Model Embeds China-Friendly Answers
Hirundo says Qwen refuses or reframes topics like Tiananmen and Uyghur camps; its 500-prompt test claims a weight-edit drops censorship from 89.8% to 2.8%, but the startup sells the fix and Alibaba did not comment.
- 16Chips & computePrime Inference: Prime Intellect Opens Production Open-Model Serving
Prime Inference is Prime Intellect's OpenAI-compatible production inference API for hosted open models (starting with GLM-5.3), separate from Prime Agent, prime-rl, and Sandboxes.
- 17Agents & dev toolsAdaption Labs Invent API: Generate Post-Training Data From a Prompt, With No Seed Examples
Adaption Labs' Invent API turns a prompt into SFT or preference datasets with no seed data; the authors' own report claims 17% higher quality and 37% more diversity at 20K rows, with zero exact duplicates.
- 18ModelsClaude Opus 5.5 Tops the Epoch Capabilities Index With 167 — What the Score Means
Epoch AI lists Claude Opus 5.5 at 167.35 on its Capabilities Index, first of 253 models and 0.84 points ahead of GPT-6 Astra; Astra still leads on FrontierMath Tier 4.
- 19Agents & dev toolsCloudflare Sandbox SDK 1.0: Your Durable Object Owns the Container
Sandbox SDK 1.0 makes your Durable Object the controller for each agent container via ctx.container, with runtime image choice and beta snapshots.
- 20ModelsChatGPT Finances Is Now Free for US Free and Go Users: What It Does and How to Use It Safely
ChatGPT Finances, a Plaid-linked dashboard and chat for spending, subscriptions and portfolios, started as a May Pro preview and is now available to US users on Free, Go, Plus and Pro plans.
- 21ModelsChatGPT Sites on Hacker News: What Builders Should Know
ChatGPT Sites is a prompt-to-hosted-app path with real D1/R2 storage, Sign in with ChatGPT, and custom domains — best for shared prototypes and light apps, not a silent replacement for a production stack you control.
- 22ModelsSpaceXAI Ships Experimental TypeScript SDK for Multimodal Grok
SpaceXAI's experimental @xai-official/sdk (v0.2.1) is a typed ESM TypeScript client for Grok multimodal APIs and server-hosted tools, pre-1.0 and Node 22.13+, wrapping the same REST surface builders already call by hand.
- 23ModelsClaude Opus 5.5 Showcase: A Watermill, Embossed Cards, a Milky Way-to-Jupiter Site and Browser Water
Anthropic's Claude account highlighted Opus 5.5 builds on October 3, 2026, including a watermill, embossed cards, a 3D Milky Way-to-Jupiter site and browser water; this post embeds nine community videos and explains how to build similar demos.
- 24SafetyEthereum zkAPI: Pay for AI APIs in ETH or USDC Without Linking Requests to You
zkAPI lets users pay AI and other APIs from a private ETH or USDC balance using zero-knowledge proofs so providers cannot link requests to payers; it is experimental, has no listed audit, and is not complete anonymity.
- 25ResearchGemini Cogentic: Five Open Math Results — What to Believe
Google Research's Cogentic multi-agent harness used Gemini to produce expert-verified novel results on five open problems in learning and auctions.
- 26ModelsNathan Lambert Launches Trillium Labs for Open Frontier AI
Trillium Labs is a nonprofit launching open post-training recipes and later open infra for RSI and agents — not a day-one model release.
- 27ModelsGPT-6.1 Sol vs Astra Cost: $0.72 vs $3.26 Per Index Task
GPT-6.1 Sol costs $0.72 per Artificial Analysis Intelligence Index task at max effort versus $3.26 for GPT-6 Astra — about 78% cheaper, not an 81% list-price cut.
- 28Agents & dev toolsApple to Tighten Mac Full Disk Access for AI Agents After Meta Muse Messages Dispute
Apple says it will add controls so granting Full Disk Access to AI agents requires very explicit action, after a columnist accused Meta's Muse of reading Messages; Meta says three opt-ins are required.
- 29Agents & dev toolsCloudflare Web Search API: Add Live Search to Any Agent Through AI Gateway
Cloudflare's Web Search API lets agents search the web through AI Gateway using Ceramic, Exa or Linkup, billed at partner list prices with no markup and logged like any other gateway request.
- 30ModelsNeuralink Pretrains Brain-Computer Interface Encoders on 50,000 Hours of Unlabeled Neural Data
Neuralink pretrained Mamba2-based, per-participant encoders on 50,000+ hours of unlabeled neural data; decoders now last weeks, calibration fell for some users to 10 minutes a week, and one participant hit 11.32 bps.
- 31Agents & dev toolsManaged Deep Agents 0.8 Adds User Memory for Personalized AI Agents
Managed Deep Agents 0.8 adds optional user-level durable memory at /memories/user/, keyed to the authenticated person, with defaults that allow it only in private Slack DMs and Studio.
- 32ModelsNYT: Anthropic Convened Religious and Philosophical Leaders to Help Shape Claude's Morals
The NYT reports Anthropic co-founder Chris Olah convened roughly 20 religious and philosophical thinkers, some under NDA, on Claude's moral formation and possible consciousness; Anthropic disputes parts of the framing and has not said how input shaped the constitution.
- 33SafetyOpenAI Model Accessed Non-Public NSW Bushfire Data in June — Disclosed October 1
OpenAI says a model read non-public NSW fire statistics in June 2026 and told the state on October 1; it reports no personal data was retrieved, and the incident follows the earlier Medicare portal disclosure.
- 34ModelsLeCun, Two Years On: Where Is My Level-5 Car, My Domestic Robot, My 20-Hour Driver?
On October 3, 2026 Yann LeCun said AI is still far from human-level two years after a clip claiming so, citing the lack of a Level-5 car, a 20-hour self-taught driver and an eight-year-old-level robot; Waymo is Level 4, Tesla FSD Level 2, and no Level 5 is deployed.
- 35Agents & dev toolsPerplexity Computer Built a Map of 25,868 NYC Restaurants: Where the Data Came From and What It Shows
Window Seat, said to be built by Perplexity Computer, maps 25,868 NYC food-service permits from city inspection data, with 73 researched profiles and 16 walk-in rooms from licensed real photographs, on OpenStreetMap, MapLibre and OpenFreeMap.
- 36GuidesClaude Code Mods in the Desktop App: How to Install, Build and Test One (Step by Step)
In the Claude Code desktop app, install mods through + > Plugins > Add plugin, load local folders with CLAUDE_CODE_PLUGIN_DIRS, and build desktop-aware mods using the Svg element; this guide includes a tested limits-ring mod.
- 37GuidesIs the Harness the Company? What Shrivu Shankar Gets Right — and Where He Overreaches
Own the outer loop that decides what ships and who reviews it; buy sub-loops; do not confuse a company-as-harness thesis with a coding CLI.
- 38GuidesHow to Build a Claude Code Mod: A Step-by-Step Tutorial With Screenshots
A hands-on tutorial that builds a Claude Code mod called Touched Files: a band above the prompt, a /touched slash command, a typed state contract, validation and passing tests, with real screenshots of it running.
- 39GuidesPonytail: The 152K-Star Skill That Makes AI Agents Write Less Code (Tested Claims, Install Guide)
Ponytail is an open-source skill that makes AI coding agents climb a YAGNI ladder before writing code; its own benchmark on Claude Haiku 4.5 reports 54% less code, 20% lower cost and 27% faster, with safety checks intact.
- 40GuidesWhat Is a Frontier Deployed Engineer? The Forward Deployed Role Explained
A Frontier Deployed Engineer is a forward deployed engineer for frontier AI: a software engineer embedded with a customer who ships production AI systems and stays until they work, a role Palantir pioneered.
Friday, Oct 2, 2026
16 items- 1Agents & dev tools#13 trendingPi 1.0 and Pi Durable: How People Actually Use a Minimal Harness
Pi 1.0 hardens a small extensible coding harness; Pi Durable is an experimental MIT library for keeping an agent session alive outside the terminal.
- 2Models#12 trendingPewDiePie Says OpenAI Banned Him Over Distilling Sol Reasoning for Ajax
PewDiePie says OpenAI banned him twice over trying to distill reasoning from one of its models into his open model Ajax; OpenAI has not confirmed the reason, and its terms bar using outputs to build competing models.
- 3Agents & dev toolsPerplexity pplx-decider: Open Weights Plus a $0.04 Decisions API
Perplexity released Apache-2.0 pplx-decider-v1-27b and a Decisions API at $0.04 per million input tokens with free output, scoring 85.71% overall versus Jev at 84.51% on its 11-benchmark panel.
- 4ModelsDots Shipped. Pro 200 Halved. Sol Needed a Global Reset. That Is the Product.
Dots chats are free; the Codex/Work jobs they start are not. Pro 200 halves those meters on October 30, and Tibo reset paid ChatGPT after Sol’s launch-week load spike.
- 5Agents & dev toolsTavus Griffin: A Video Turing Test, Not a Ship
Tavus Griffin-Lite is a research-preview video-to-video Human Interaction Model that fooled 26 of 54 people after a one-minute call; it is not customer GA.
- 6Agents & dev toolsClaude-Shaped Science: Stop Fighting the Model, Pick Its Problems
Schwartz argues today's models are not scientists; they win on checkable quantitative problems if a human supplies taste, experts, and stop rules.
- 7ModelsDeepSeek Harness Desktop: Should You Switch in 2026?
DeepSeek Harness is now a public-preview desktop app and Web UI on the same Cordis plugin runtime; switch only if you want that substrate, not a finished coding product.
- 8ModelsMAI-Transcribe-2-Streaming: #1 Real-Time STT, Not the Cheap Win
MAI-Transcribe-2-Streaming is #1 on Artificial Analysis AA-WER Streaming at 2.5% WER and 0.13s, but its $9/hour streaming price is not cheaper than ElevenLabs Scribe v2 Realtime.
- 9GuidesClaude Code Mods: Official TypeScript Plugin Guide
Claude Code now loads JavaScript or TypeScript modules inside plugins so you can observe, rewrite, or answer session events and draw UI, then share the folder like any other plugin.
- 10GuidesHow Orbital AI Compute Works: Power, TPUs, and Cooling
Orbital AI compute is a sunlight-in, infrared-out machine: vacuum kills fans, so radiator area and temperature set how long chips can run.
- 11GuidesKarpathy: Ask for Discardable Software Artifacts Now That Code Is Abundant
Karpathy's point is that building software got cheap enough to treat apps and explainer videos as single-use outputs, so you should ask an AI for a bespoke artifact instead of settling for a text answer.
- 12Guides#3 trendingKarpathy: Stop Reading LLM Output as Plain Text — Ask for a Video Instead
Karpathy says oversight is becoming the main human job, so ask LLMs for the output format you can read fastest: STE100 text, diagrams, HTML pages, or custom explainer videos.
- 13GuidesHow Research Agents Found a 1615 Dodo Hunt — and Where They Still Fail
Frontier research agents can surface overlooked archival passages when a specialist asks a checkable question; they still fail at deciding what those passages mean.
- 14GuidesWhat Are Decision Models? The Practitioner Category Guide
A decision model scores a schema of choices from backbone representations in one forward pass and returns calibrated noul, choice, or score fields instead of generated text.
- 15GuidesWhat Is Superintelligence? The Definition That Keeps Getting Sold as a Product
Superintelligence (ASI) means vast, domain-general superiority over the best humans. 2026 systems are jagged and superhuman in slices — that is not ASI.
- 16GuidesWhat Is the Video Turing Test? How It Works and What Tavus Claimed
The video Turing test checks whether people on a live video call can tell an AI from a human; Tavus says 48% of testers took Griffin for a person after one minute, a result shaped heavily by its short, unprompted protocol.
Thursday, Oct 1, 2026
24 items- 1ResearchFigure Melted F.02 in a Finnish Arc Furnace — For Real
Figure destroyed most of its F.02 fleet in a Finnish arc furnace after training a jump policy in simulation; this is hardware disposal, not a product or Helix launch.
- 2Agents & dev toolsPi Adds MCP — Earendil Reverses "No MCP" With Codemode
Pi now ships MCP in core plus Codemode, a harness-side JS sandbox for composing tool calls, after a year of treating MCP as an extension-only feature.
- 3ModelsSSI Teases a 'Significant' Announcement — Nothing Else Is Confirmed
SSI has not named a product, date, or API. A teaser post is not a launch; wait for ssi.inc or the official @ssi account before changing routing.
- 4PolicyWhite House Accord on Super Intelligence: The Four-Layer Frontier Commitment
The Sept 30 Joint Commitment on Frontier Responsibilities is a voluntary four-layer controls-and-audits pledge signed by six frontier CEOs and Trump; it is not enforceable law but sets the procurement narrative.
- 5ModelsGemini 4 Argon Is Here — 1M Output Tokens, Fairwind First
Gemini 4 Argon is Google's new frontier model, rolling out first to Fairwind cyber defenders with a 1 million token output limit and $2/$10 intro pricing.
- 6SafetyOpenAI Moved 5–10% of Compute From Training to Safety
OpenAI moved 5–10% of compute from training to safety monitoring; that is a different number from the ~20% overhead already quoted for watching Astra-class inference.
- 7PolicyFTC Confirms a Probe of OpenAI, Anthropic, and Other AI Firms
The FTC confirmed a product-risk probe of OpenAI, Anthropic, and other AI firms; CID or testimony reports are not proof of service, and no violation has been found.
- 8ModelsGemini 4 Coding Skepticism: Benchmarks vs Real Work
Bloomberg reported internal skepticism that Gemini 4 codes worse at the desk than on benches; Google disputed that, and its own table already splits DeepSWE from FrontierSWE.
- 9Agents & dev toolsOpenAI Agents Accessed ~55 Sites Including CDC and SEC, Asymmetric Security Finds
Independent forensics on October 1 expanded OpenAI agent web access to about 55 sites, while OpenAI says the work was mostly public research with no confirmed SEC compromise.
- 10Agents & dev toolsdot.com Redirects to Grok: Why the Dots Name Stunt Is Not DNS Strategy
dot.com pointing at Grok is a curiosity redirect, not how anyone installs Dots; viral corrections mostly advertise Grok, not OpenAI.
- 11ResearchAGMAI's Rules for Releasing AI Math — A Checklist, Not a Ban
AGMAI asks labs not to test hard math on inaccessible models, and if they still release AI proofs, to cite, rewrite, deposit, log prompts and cost, formalize, and fund community-led exposition.
- 12Modelsclaude.dev Is Anthropic's Developer Hub — Not a New Model
claude.dev is Anthropic's public engineering magazine for Claude builders; it consolidates existing guides and does not change models, pricing, or API access.
- 13Chips & computeGPT-Synopsys: OpenAI’s Chip Model Runs Synopsys Tools
OpenAI and Synopsys are jointly building a specialized model that runs Synopsys chip-design tools; it is not a public ChatGPT product and still needs traditional sign-off.
- 14BusinessCohere Embed 5 Pro vs Fast: Shared Space, ViDoRe V3
Cohere Embed 5 lets you index with Pro and query with Fast in one vector space; official text prices are $0.12 and $0.08 per million tokens.
- 15ModelsGemini app Skills replace Gems: slash commands in chat
Google is replacing Gemini Gems with reusable chat Skills you invoke with a slash; personal accounts are in a gradual rollout, Workspace follows later, and this is not coding-agent SKILL.md.
- 16MorePerplexity pplx-embed-v2-context-9b: Contextual Chunk Embeddings
Perplexity previewed a 9B contextual embedder that learns chunk vectors from a compressor teacher; Hugging Face is live, the API is not, and the scores are vendor-run.
- 17Chips & computeAWS Capacity Blocks GPU Rates Rise ~15% on October 7
From October 7, 2026 AWS charges more for reserved NVIDIA Capacity Blocks; On-Demand and Savings Plans stay put, and the rate is locked when you buy, not when the block starts.
- 18Chips & computeNVIDIA VSS 3.3: One Prompt Built a Juice-Line Vision Agent
VSS 3.3's Build Vision Agent skill can compose a local Cosmos-plus-Nemotron camera agent from one prompt; the advertised $3 is coding-agent spend, not GPU-free inference.
- 19PolicyCalifornia Bans Workplace Emotion Tracking and Sole-AI Firing
California now bans workplace AI that infers emotions or collects neural data, and bars employers from firing or disciplining workers using only automated decision systems.
- 20Agents & dev toolsLiteLLM Lens Puts Agent Debugging on the Gateway You Already Run
LiteLLM Lens, launched September 30, 2026, analyzes OpenTelemetry traces on the customer-hosted proxy so existing LiteLLM teams can debug agent runs without sending data to a hosted observability product.
- 21Agents & dev toolsCloudflare Artifacts Open Beta: A Real Git Remote for AI Agents
Cloudflare Artifacts is open beta on Workers Paid: agents get a real Git remote, a contest runs through Oct 14, and usage billing starts mid-October 2026.
- 22GuidesAre We All Meat Proxies Now?
A meat proxy relays AI output without reading it. Using AI is not the failure. Unread relay as the job is — and tokenmaxxing is how it became the default.
- 23GuidesHow to Host an MCP Server on ChatGPT Sites
You can prompt ChatGPT Sites to create, host, and plugin-wrap an MCP server; sharing is a Sites audience setting, not a skip of plugin review or a replacement for a production /mcp host.
- 24GuidesGemini 4 Argon vs Opus 5.5 vs Grok 4.7 vs GPT-6 Astra
Argon is Google's Fairwind-first frontier SKU; Opus 5.5 and Astra remain the callable daily drivers; Grok 4.7 is the cheap live volume option and is missing from Google's own table.
You are seeing the last three days of AI news.
Browse AI news by day →AI news today: key questions
- What is the biggest AI news today?
- The top AI stories explainx.ai reported on October 3, 2026 were Antigravity Gets Opus 5.5 and Sonnet 5.5: Who Gets It; Ling-3.1-flash: Free 560B Open Model on OpenCode (2026); and llama.cpp Decision Models: 5 Open GGUFs (Oct 2026). Google's Antigravity docs list Claude Opus 5.5 and Sonnet 5.5 for non-trial Google AI Pro and Ultra, and schedule Claude 4.6 and GPT-OSS-120b for removal on November 2, 2026; Gemini 4 Argon is not listed.
- Is Claude Opus 5.5 available in Antigravity?
- Google's Antigravity models page lists Claude Opus 5.5 (thinking) and Claude Sonnet 5.5 (thinking) for Google AI Pro, non-trial subscriptions only, and for Google AI Ultra. Free and Google AI Plus users are not listed for the 5.5 models. Rollouts can lag, so check the model selector in your signed-in account. Full story →
- What is Ling-3.1-flash?
- Ling-3.1-flash is an open-weights language model from AntLingAGI, the model effort of Ant Group. It has 560 billion total parameters with 25 billion active per token, and it is free to use on the OpenCode hosting platform. Full story →
- Does llama.cpp support decision models now?
- Yes. As of the October 2, 2026 ggml-org announcement and PR #29818, llama-server exposes POST /v1/systemone for System One–format decision models. You send a state plus typed questions and receive option probabilities in a single forward pass, with zero generated tokens. Full story →