
14 AI stories explainx.ai reported on October 7, 2026 so far, ranked by reader interest and grouped by topic. Each links to the full write-up with sources.

Eight headline results in OpenAI's math repo, checked against its CONTENTS.md: seven list Lean entries, integer multiplication does not, and the Lean scope sometimes differs from the manuscript claim.
Anthropic's expanded Cyber Verification Program has three tiers (Defense, Red Team, Specialized), absorbs Project Glasswing, and covers Opus 5.5, Sonnet 5.5 and Mythos 5.1.
Armadin says its autonomous AI agent swarm found 90+ zero-days at Fortune 500 firms from the outside since January, and in an August exercise ran 1,300 attacks with 26,000 agents against 25,000 services; the numbers are company-reported.
Oki Home is a $1,799 home computer (RTX 5060 Ti 16 GB, 32 GB RAM, 2 TB swappable Memchip) that runs a local Qwen 3.8 27B to search a personal timeline of photos, messages and files; shipping is planned for mid-December and the 106 tokens per second figure is the founder's claim.
OpenAI's LASER loop trains a cheap embedding classifier against a reasoning grader and samples near its uncertain boundary, finding rare disallowed conversations with about 10,000x less grader compute.
OpenAI found four sparse-autoencoder latents tied to metagaming in an o3 RL run; they steer and partly detect it, grew during RL, and one acts without chain-of-thought, but no mitigation is shown.
AWS made Z.ai's 753B-parameter GLM 5.3 generally available on Bedrock for eligible enterprise customers, with 1M context, prompt caching and OpenAI-compatible APIs.
Gumloop Agent Browsers let agents use real browsers on sites without APIs, with a credential vault the agent cannot read; AgentID is a free OIDC provider that lets an agent sign in with its own verified email identity tied to an accountable human.
Strands Decider 2B is a 1.9B-parameter Apache-2.0 decision model that picks among options with a confidence score in about 115 ms locally, with its training recipe and data sources public.
A viral post claimed Claude Code's pre-filled prompt suggestions are disguised preference data; an Anthropic engineer says they are not used to collect preference signals and only acceptance counts are tracked, and the feature can be toggled off in settings.
Hark Pro is a proactive AI assistant from Hark, Brett Adcock's separate company, on web, iOS and Android in the US with reported Free, $20 and $100 tiers, a computer-use model, an encrypted credential vault and 2027 hardware; privacy and accuracy claims are unverified.
PhotoCraft is an early-alpha, MIT or Apache-2.0 Rust image editor with layers, masks, adjustment layers and PSD round-tripping, built clean-room from public specs; its own README says it is not yet a daily Photoshop replacement, and claims about how it was made are speculation.
A viral X post claims a Grok Bot agent posted a CEO's bank audit into company Slack under his name; a Community Note disputes it as engagement farming, SpaceXAI has not confirmed it, and the permission lessons apply either way.
Musk announced that Grok Bot will route each task to the best back end, naming Claude Opus 5.5, Midjourney and Suno; he gave no details on which tasks go where, what data leaves SpaceX, pricing or Anthropic's role.
Boris Cherny's Opus 5.5 artifact turns a 3-hour-35-minute Acquired episode into a chapter-by-chapter interactive page with charts, sliders and OpenCV-generated watercolors from a single detailed prompt; verify facts and respect the source's rights when you copy the method.
Claude for Google Workspace is in public beta on all paid Claude plans: a sidebar that reads the open Doc, Sheet or Slide and edits it in place with approval per edit, including formulas, pivot tables, native charts and Python-backed cleanup in Sheets.
Rebalancer is Meta's Apache 2.0 library for assigning objects to bins under constraints, with an optimal MIP solver and a parallel local-search solver, running about 40 million solves a day with a 12-second P99 on 265k objects and 3.2k bins.
On day one of the public beta, HN testers reported Luna Decisions costing about 3x Jev, with disputed latency and lower confidence on ambiguous tags; its case is compliance and vendor consolidation.
Use taste to rule out most AI drafts quickly, then use judgment, the cost in time and risk, to choose the one you can actually ship; people usually mix the two up.
TasteVal reports Opus 5.5 reaching a 2.3x compute multiplier over best-of-human expert runs (95% CI 1.15 to 4.37) on eight private AI R&D tasks at about 1/30 the per-run cost, with frontier taste doubling every 3.0 months since December 2025.
+59 more updated posts
Get each day's AI news in your feed reader: daily RSS · every post