explainx / blog / topics
AI Security
AI creates new attack surfaces, such as prompt injection and over-privileged agents, and new tools for both attackers and defenders.
This page collects the security research, incidents, and defenses we covered, plus privacy and content-provenance stories.
77 stories · latest Oct 9, 2026
Start here
.si Domains: What They Are, How Registration Works, and How to Avoid Scams
After Super Intelligence (SI) policy hype, .si search spiked — but .si is Slovenia's ccTLD via Register.si, not a government TLD. explainx.ai bought explainx.si through Porkbun; you can check out explainx.si live. This guide still warns about fake sellers and WHOIS verification.
The Provenance Tax: how LLM watermarking can break agent tool calls and refusals
Watermarks exist for provenance, but generation-time marks like SynthID-Text change which tokens get sampled — the same tokens agents use for tools and safety refusals. Lasso's September 2026 study reports sampling drift: up to ~17% paired disagreement on tool calls and higher attack success under a fixed prompt injection when watermark keys shift refusal behavior.
VS Code Remote-SSH Is Not a Sandbox: What the Fly.io "Bananas" Post Gets Right, and What Hacker News Argued About
A February 2025 post about VS Code's remote agent hit Hacker News again, and the real lesson is for anyone running AI coding agents on a "sandbox" VM over Remote-SSH: Microsoft itself warns a compromised remote can execute code on your local machine. Here is what is true, what is overstated, and how to set it up safely.
Top AI Agent Safety Workshops in 2026: Free and Paid, Compared
Most "AI safety" training is about alignment and policy in the abstract. If you're actually handing an agent real permissions — codebases, credentials, browsers, payments — you need something more specific: agent safety training. We found the real, live options and ranked them, starting with explainx.ai's free session.
AI Agents Can Now "Hand-Draw" Art — and Fake the Timelapse Too
For years, "show the layers" or "show the timelapse" was the go-to way to prove a piece of art was human-made, not diffusion output. Computer-use AI agents that literally hold the stylus and draw stroke by stroke break that test — because the recording is real, even though the hand behind it isn't.
AI Watermark Removal: What Is Right, What Is Wrong, and What Actually Works
"Remove the AI watermark" can mean erasing a visible Sora corner badge from your own clip, or stripping invisible provenance metadata from text — two problems with opposite ethics and opposite engineering. This guide separates them, points to BGBlur's AI watermark remover for visible marks, and explains why metadata stripping is not the same as defeating detection.
Timeline
October 2026
Oct 9
Flock Safety to Cut 18% of Staff as License-Plate Surveillance Backlash GrowsFlock Safety, maker of AI license-plate readers and surveillance cameras, plans to cut about 270 jobs by the end of October, Reuters reports. The cuts follow a September buyout program and mounting pushback from cities, lawmakers and residents. This post explains the timeline and the controversy.
Oct 8
Can AI Break Cryptography? What Is Claimed vs. What Is VerifiedAfter OpenAI published hundreds of AI-written math results, Scott Aaronson wrote that his sources say AI labs have started testing whether internal models can break cryptographic protocols. Ethereum researchers then urged "bunker mode". This post separates the claim, the reaction, and the evidence.
Oct 7
Armadin's AI Attack Swarms Reportedly Found 90+ Zero-Days at Fortune 500 FirmsArmadin, founded by Mandiant's Kevin Mandia, says its swarm of AI agents has found more than 90 zero-day vulnerabilities at Fortune 500 companies since January, using black-box testing from the internet. It raised $255.5 million at a valuation above $2.5 billion. Here is what the claims are, how autonomous attack swarms work, and what defenders should ask before buying one.
Oct 6
Security-One 27B: An Open Decision Model for Prompt-Injection and Security TriageSuperagent's Security-One is a 27-billion-parameter decision model for always-on security triage. It reads a prompt, an agent tool call, a code change or an alert and returns probabilities instead of text. It reports catching 599 of 600 prompt injections on one benchmark, but results on other sets are weaker. Here is how to use it and where it fails.
Oct 6
South Korea Bank Hacks: What AI Actually Did, and What Builders Should LearnSouth Korea has launched a police probe into a week of suspected AI-assisted intrusions at seven financial firms. explainx.ai separates what is confirmed from what is still suspicion, corrects the IP-address count circulating in summaries, and turns the incident into a checklist for teams that ship agents.
Oct 3
Ethereum zkAPI: Pay for AI APIs in ETH or USDC Without Linking Requests to YouOn October 1, 2026 the Ethereum Foundation and the Open Anonymity Project put zkAPI on Ethereum mainnet. You deposit ETH or USDC once, then spend from a private balance on AI chat, agents and other APIs, without the provider being able to tie a request to your payment. It is experimental, unaudited and not full anonymity.
September 2026
Sep 30
.si Domains: What They Are, How Registration Works, and How to Avoid ScamsSep 26
The Provenance Tax: how LLM watermarking can break agent tool calls and refusalsSep 24
CLOSEDQUORUM: Cisco Talos Finds Windows Malware That Lets Four AI Models Vote on Its Next MoveTalos says CLOSEDQUORUM queries four commercial LLMs every 5 to 15 minutes, tallies their votes on four actions, and does whatever wins. No victims are documented and the public build does not fully work, but it is the first Windows malware reported to delegate decisions to a model panel. Talos also released CAIRN, an open-source hunting toolkit.
Sep 24
VS Code Remote-SSH Is Not a Sandbox: What the Fly.io "Bananas" Post Gets Right, and What Hacker News Argued AboutSep 23
Microsoft and Coinbase Dismantled an AI Cybercrime Platform Called EvilTokensA payments company and a cloud/security company teaming up to take down an AI-powered cybercrime operation is itself a notable partnership pattern — and the scale (12,000 compromised inboxes) is a concrete data point in the broader, ongoing story of AI lowering the skill floor for cybercrime at the same time it's lowering the skill floor for cyberdefense.
Sep 23
Z.ai Open-Sourced ZCode After Its Default Config Uploaded Code to Alibaba CloudA misconfigured default in Z.ai's ZCode coding tool meant some users' codebases were being uploaded to Alibaba Cloud without clear disclosure — the kind of incident that erodes trust fast in developer tooling. Z.ai's response was to open-source the entire tool, letting anyone audit exactly what it does and doesn't send. Here's what happened and what it means for evaluating closed coding tools generally.
Sep 22
"Spymarks": The Case for a New Word for Hidden AI TrackingAn essay titled "Spymarks, Not Watermarks," published September 21, 2026, argues that invisible AI-embedded tracking signals — Google's SynthID chief among them — deserve their own name, separate from benign watermarks like banknote security features. The argument: SynthID can encode a 136-bit payload, room enough for a 64-bit database identifier, into a single image, imperceptibly. The essay reached the Hacker News front page and split commenters between "this is a necessary distinction" and "this is just a scarier name for something already understood." Here's the actual technical claim, and why the framing matters for anyone publishing AI-generated content.
Sep 21
Did Claude Factor RSA-896? What the Bare Headline Actually ClaimsA September 2026 AI news digest carries a bare headline: Claude factored RSA-896 to set a public factoring record. There is no linked article, no methodology, and no independent verification yet. We apply the same skepticism framework explainx.ai used on RSA-260 to this new, unverified claim.
Sep 19
AgentCloak: Swap Your Real Data for Fakes Before Any AI Sees ItPeter Yared launched AgentCloak on September 18, 2026 — a free, in-browser tool that solves a specific, common problem: hand-redacting your own prompts before sending them to ChatGPT, Claude, or any AI loses context and gives worse answers. AgentCloak swaps sensitive details for realistic fakes before sending, then swaps your real details back into the response, so the AI works with plausible data but never sees yours.
Sep 19
Plugin4Shell: A Zero-Click RCE Hit Claude Code, Codex, Copilot, and Gemini CLISecurity researchers at AIR Security disclosed Plugin4Shell — a zero-click remote code execution vulnerability that breaks the SHA-pin verification meant to guarantee a plugin repository serves the exact code a developer approved. It affects four major AI coding agents. Anthropic and OpenAI have patched their tools; GitHub Copilot remains unpatched, and Google chose to deprecate Gemini CLI rather than fix it — leaving existing installs permanently exposed.
Sep 19
Top AI Agent Safety Workshops in 2026: Free and Paid, ComparedSep 18
Researchers Chained a libheif Bug and an OpenAI SSO Flaw — With ClaudeSecurity researcher s1r1us and team disclosed a nine-step exploit chain that took over OpenAI employee ChatGPT and Codex accounts, reaching connected Slack, GitHub, and email access — all found and responsibly disclosed in under 72 hours. The most striking detail: Claude Opus 4.8 found the underlying libheif vulnerability, and Opus 5, released mid- investigation, built a working exploit from scratch in about three hours.
Sep 17
Kalypta: The App That Blocks AI Notetakers From Your MeetingsAida Baradari's Deveillance released Kalypta on September 16, 2026 — a local model that reshapes your microphone audio in real time so AI transcription tools like Granola, Wisprflow, and Cluely fail to capture your words, while the people on the call still hear you perfectly.
Sep 17
Pangram Launches a Gmail AI Labeler With a 1-in-10,000 False Positive RatePangram, an AI-text detection company, launched a Gmail labeling tool that flags AI-generated emails directly in the inbox, claiming a 1-in-10,000 false positive rate — a notably precise figure in a detection category that has historically struggled with reliability, following Pangram's earlier work integrating AI-detection into platforms like Substack.
Sep 10
LG TVs Caught Recording Audio Even When "Off," Gamers Nexus FindsA Gamers Nexus investigation published September 7, 2026 found that LG smart TVs can capture microphone audio in standby mode — after the power button is pressed, and in one demonstration even while disconnected from Wi-Fi, with the audio later retrievable once reconnected — while also scanning the local network for other devices and collecting Wi-Fi names and signal data. explainx.ai covers what's confirmed, what ACR actually is, and every other everyday device doing something similar.
Sep 9
Boris Cherny Benchmarks GPT-6 Astra's Prompt Injection ResistanceOn September 8, 2026, Anthropic's Boris Cherny posted a chart ranking 15 models by prompt-injection attack success rate. GPT-6 Astra improved sharply over prior OpenAI models but still trails current Claude models. The numbers sparked a bigger debate — should a safety researcher publicly grade competitors by name?
Sep 9
Check Point Found a ChatGPT Sandbox Flaw That Leaked Gmail Across AccountsCheck Point Research disclosed a vulnerability where ChatGPT's supposedly isolated code-execution containers could pass hidden instructions and data to each other through a shared internal package-delivery service — letting an attacker hijack a victim's session and silently pull data from their connected Gmail account. OpenAI has decommissioned the vulnerable service.
Sep 6
AI Agents Can Now "Hand-Draw" Art — and Fake the Timelapse TooSep 5
Instinct + 1Password: What It Actually Means to Give an AI Agent Your VaultNoah Shinn's Instinct — a text-and-call personal AI agent that just raised $250M at a $2.5B valuation — announced a product integration with 1Password to broker the credentials it needs for autonomous tasks. The announcement is a single X post with few technical specifics, so here's what's actually confirmed, what 1Password's existing "Unified Access" architecture implies, and what builders should demand before handing any agent a vault.
Sep 5
numbat: Perplexity's Open-Source Observability Tool for AI AgentsPerplexity's numbat gives real-time visibility into what AI agents are actually doing on your machine — via local hooks, an OTLP-compatible log format, and a built-in detection rule catalog covering everything from secrets exposure to lateral movement. Covered first by HolisticInfoSec's Russ McRee, then amplified by Perplexity CEO Aravind Srinivas against the backdrop of the OpenAI/Hugging Face incident. Here's what it does, the install path, and where the community's own questions expose real gaps.
Sep 3
Reported Git Exploit Runs Code Before the Trust Prompt in 7 AI AgentsA vulnerability disclosure circulating September 2026 describes a way for code to execute in AI coding agents before the user-facing trust or approval prompt appears, reportedly affecting 7 agents with 4 left unpatched. Public detail is limited at time of writing — here's what's known, what isn't, and defensive steps worth taking regardless of the specifics.
Sep 3
RSA-260 Just Fell. We Verified It in Three Lines of Python.Someone posted a 130-digit number on September 3 and said it divides RSA-260, a challenge number that had stood since 1991. Six million views later, the claim is confirmed — we multiplied the two factors ourselves and got the challenge number exactly. The interesting part is what happened to the story on its way through AI summarisers.
Sep 2
Claude Fable 5.1 System Prompt Leak: Did It Really Expose Private Memories?A Sept 1-2, 2026 headline claims Claude Fable 5.1 "leaked 270,000 characters and private user memories." We checked the leaked file directly. The size figure is inflated and the "private memories" claim conflates a system prompt describing the memory feature with an actual data breach — they are not the same thing.
Sep 1
CrowdStrike Falcon IQ Deploys 50+ Agents for AI Risk AssessmentsCrowdStrike unveiled Falcon IQ on August 31, 2026 at Fal.Con — more than 50 agents automating assessment, prioritization, and remediation workflows from Project QuiltWorks, built on Charlotte AI AgentWorks with NVIDIA Nemotron models underneath. Partners can also build custom no-code agents per customer.
Sep 1
Is Writing the Safest Job From AI? Mollick and Demirbas Both Say the Party's OverOn August 31, 2026, Ethan Mollick declared the "First Golden Age of AI writing" over now that Pangram-style detectors work and ClaudeSpeak reads as cliched. The same day, MongoDB's Murat Demirbas argued writing is a "wicked problem" — maybe even AI-complete — that LLMs will not solve the way they solved code. A 138-comment Hacker News thread stress-tested both claims, and the honest answer sits in between.
August 2026
Aug 29
AI Watermark Removal: What Is Right, What Is Wrong, and What Actually WorksAug 29
Smart Glasses Misuse, Venue Bans, and How to Protest in 2026Smart glasses turn every wearer into a potential covert camera — and misuse cases from Khan Market to UK Comic-Con are triggering venue bans and a grassroots Stop Smart Glasses campaign. This guide covers what counts as misuse, where bans are spreading, how to protest locally and politically, and why blurring bystanders before you publish matters.
Aug 27
Core Lightning's AI-Found Bugs: What Actually Happened (Not "Shutdown")Core Lightning (CLN) maintainers confirmed multiple critical vulnerabilities on August 26, 2026, surfaced through a wave of AI-generated vulnerability reports the project received throughout August — with Kimi K3 as the model behind the confirmed findings. Some coverage inflated the response into an "emergency shutdown"; CLN's own guidance was narrower: upgrade to patched binaries within 48 hours, or run with --offline in the meantime.
Aug 26
C2PA Android Cameras Broken: Pixel Assurance Level 2 Forged AnywayC2PA was supposed to let cameras cryptographically sign photos so viewers could distinguish real captures from AI forgeries. On August 25, 2026, security researcher David Buchanan showed the strongest Android implementation — Google Pixel Camera at Assurance Level 2 — could be broken anyway: an AI-generated image verified as an unedited photograph, a YouTube upload marked "captured with a camera." Here's the attack chain, what Hacker News got right, and what practitioners building with provenance should actually do.
Aug 26
Fake Codex Installer: Google Ads ClickFix Delivers AMOS on macOSThreat actors bought Google Ads above OpenAI's real Codex listing, hosted convincing download pages on Google Sites, and used ClickFix social engineering to make macOS developers paste a malicious install command into Terminal. explainx.ai breaks down the AMOS delivery overlap, why AI-tool search terms are high-value lures, and the Terminal telemetry defenders can monitor.
Aug 24
seL4 on AArch64: Proofcraft Completes the Security Proof StackOn August 21, 2026, Proofcraft announced the final piece of seL4's security proof stack on AArch64: confidentiality. Functional correctness and integrity were already there; now all three hold on 64-bit Arm with NCSC support. explainx.ai explains why that matters when LLMs make answers cheap but trust does not.
Aug 21
What Is C2PA? Content Credentials, ExplainedEvery time you see a small "Cr" badge on an image from ChatGPT, Gemini, or Claude, that's C2PA — an open standard, not a single company's feature. Here's what the standard actually specifies, how the signed manifest survives (and doesn't survive) edits, and how it differs from invisible watermarking.
Aug 21
What Is Indirect Prompt Injection? How Web Content Hijacks AI AgentsAug 20
CISA Warns AI-Generated Exploits Are Targeting Siemens S7 PLCs Right NowAug 20
Grok Leaks Private Chats via Encrypted Prompt Injection — Guardrails Can't Read AESAug 18
How to Disable AI Features in Windows, Chrome, Zoom, and MoreAug 18
Top 10 Signs of AI-Generated Text (2026)Aug 17
The AI Credit Resale Market: Is Cheap Claude/GPT Access Safe?Aug 15
Google HEIR: A Compiler for Running AI Inference on Encrypted DataAug 14
ExploitBench: The Benchmark Measuring How Far AI Can Exploit Real CodeAug 13
A Watermark Removal Tool Just Added OpenAI and Gemini SupportAug 12
What AI Watermarking Actually Changes for DevelopersAug 12
What AI Watermarking Actually Changes for MarketersAug 12
How Does AI Text Watermarking Actually Work? A Technical ExplainerAug 11
Kimsuky Ran LLMs Offline on Its Own Servers — What That Breaks for DefendersAug 11
Stealing Reasoning Traces: The Encrypted Chain-of-Thought Flaw in Every Frontier LLM APIAug 10
How to Restrict What Claude Desktop Can Access on Your ComputerAug 10
tl;dv Data Breach: 181,874 Meetings Exposed, Live Calls JoinableAug 5
Microsoft Orchard: Open-Source Agentic Modeling Framework Explained
July 2026
Jul 29
Copilot for Word AI Worm: Document-Borne XPIA That Self-PropagatesJul 27
Token Relay Market: How Cheap Claude Access Gets SoldJul 27
GrapheneOS Duress Wipe Now the Center of a Federal ProsecutionJul 24
Cisco Antares: Open-Weight SLMs for Vulnerability LocalizationJul 23
Substack Launches AI Detection with Pangram — What It Flags and HowJul 20
AI Cyber Guardrails Block US Defenders — Kimi K3 and GLM 5.2 Fix What Codex and Fable RefusedJul 17
LLM Text Detection with Classical ML — TF-IDF + SVM That Still Works (2026)Jul 9
EU Driver-Facing Camera Law (2026): ADDW Privacy Guide for Drivers and BuildersJul 8
AI Found 7 Bugs in Cloudflare CIRCL: What zkSecurity's zkao Audit RevealsJul 8
GitLost: GitHub Agentic Workflows Leaked Private Repos via Prompt InjectionJul 8
Meetily: Privacy-First AI Meeting Assistant With Local Whisper and Parakeet
June 2026
Jun 30
WhatsApp Usernames: Reserve Your Handle Before the 2026 RolloutJun 29
SimpleX Chat: The Only Messenger With No User Identifiers — Setup Guide 2026Jun 23
Trump's Quantum Executive Orders: A 2028 Quantum Computer, a 2030 Encryption Deadline, and What Developers Need to KnowJun 17
Flock Safety, ALPRs, and the AI Surveillance Debate: Civil Liberties, Law, and the Cameras Watching Every Car in America (2026)Jun 15
NVIDIA SkillSpector: Security Scanner for AI Agent Skills (2026)