explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • What people are asking after the clip
  • How 3.3 is different from the May blueprint
  • Adaptive EVS is the other half of 3.3
  • What this means for people who build with skills
  • Honest limitations
  • What to do this week
  • Related reading
← Back to blog

explainx / blog

NVIDIA VSS 3.3: One Prompt Built a Juice-Line Vision Agent

NVIDIA, VSS Blueprint, Cosmos, Agent Skills, Vision Agents

NVIDIA's VSS 3.3 Build Vision Agent skill composed a Cosmos + Nemotron factory watcher from one prompt. About $3 of coding-agent tokens — plus your GPUs.

Oct 1, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
NVIDIA VSS 3.3: One Prompt Built a Juice-Line Vision Agent

Cameras that see everything and tell you nothing are not a poetry problem. They are a composition problem: ingest, detect, verify, store, search, report — usually five teams and a graveyard of YAML.

NVIDIA's answer in VSS Blueprint 3.3 is a single agent skill: vss-build-vision-ai. The tutorial now circulating as https://youtu.be/PPXoPSCdGyU walks that skill through an orange-juice line. One English prompt. A plan you can read. Four questions. Then a live stack.

The official engineering write-up landed September 29, 2026: Lower the Cost of Building and Running Visual AI Agents with NVIDIA VSS Blueprint 3.3. The film is the same claim with a leaky filler. This is not a new Cosmos 3 world model. It is the May VSS blueprint growing a composer.

NVIDIA tutorial: one prompt, Cosmos watching the line, Nemotron writing the report. The overflow in the official write-up is synthetically generated.
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR

table · 2 cols
QuestionAnswer
What shipped?VSS 3.3 — Build Vision Agent skill + Adaptive EVS
Skill idvss-build-vision-ai
WatcherCosmos reasoning VLM on camera chunks
ReporterNemotron (docs default: Nemotron Nano 9B v2)
Demo promptTell me if the juice ends up outside a bottle
Video numbers~20 services, 26 minutes, 7.5M tokens, about $3 compose
Official hardwareTwo RTX PRO 6000 Blackwell GPUs for the bottling write-up
Runtime billLocal GPU, "pennies per task" — not a cloud token meter
Is the spill real?Stack is real; NVIDIA says the overflow footage is synthetic

What people are asking after the clip

"Did one prompt really deploy 20 services?"

Yes, in the sense NVIDIA means: the skill reads the VSS Blueprint repo, maps the ask onto four validated foundations (base, alerts, lvs, search), computes a delta, shows an architecture, then writes a self-contained _builds/<name>/ tree. It does not rewrite the repo's deploy/docker/ tree.

The video's sequence matches the blog:

  1. You describe cameras, alerts, and a web UI.
  2. The skill proposes Cosmos to watch and Nemotron to write.
  3. It asks where to build, which GPUs, host address, cameras.
  4. Only after you confirm does Compose come up.

That is progressive disclosure with a human gate, not "the model SSHed into prod." If you want the same pattern in software agents, it is closer to a deploy skill than to a chat widget.

"Is $3 the price of a factory AI?"

No. Split the invoice the way NVIDIA does.

table · 3 cols
CostWhat it isWhat it is not
~$3 / 7.5M tokensCoding-agent spend to plan, generate, and verify the stackPer-frame inference
26 minutesWall clock for that compose on their machineYour laptop
Pennies per reportLocal Nemotron writing from stored incidentsA hosted API with no GPU
Two Blackwell PRO 6000sThe published bottling-line hostIncluded in the $3

If you do not already own the GPUs, the honest first number is capex + power, then the $3. The useful claim is narrower: composition got cheap enough to demo in half an hour. Runtime VLM cost is a different knob — Adaptive EVS.

"Which Cosmos is watching the line?"

The video says Cosmos is NVIDIA's open reasoning VLM. Docs for the default VSS agent pair nvidia/cosmos3-nano-reasoner on port 30082 with nvidia/nvidia-nemotron-nano-9b-v2 on 30081. The bottling write-up puts FP8 Cosmos 3 Nano on the same GPU as the detector.

That is the Reasoner surface — text out about video — not the Generator that rolls futures for robots. Keep them separate the way explainx.ai already does in the Cosmos 3 guide and the SIGGRAPH physical-AI recap.

The video's runtime loop: Cosmos looks at the stream in 10-second chunks and asks the prompt. A leak hits the dashboard "within seconds." Chat: generate a report. Nemotron writes from the incident store. Under a minute, presenter says, no cloud token bill.

"Can I paste this onto my cameras tonight?"

Only if you already have:

  • A coding agent that loads agentskills.io packages (NVIDIA names Claude Code, Codex, NemoClaw).
  • The VSS 3.3 skills installed (symlink from the blueprint repo, not a one-off copy).
  • GPUs the profiles actually fit.
  • RTSP (or recorded) sources, and a trusted network. NVIDIA's getting-started list ends with authentication, TLS, rate limits, and isolation.

The sample prompt from the official post is the one to reuse — not the one-liner from the film:

text
Build a VSS vision agent for an orange juice bottling line.
Use two RTSP cameras on the filler and capper.
Detect bottle overflows and juice spills, verify each alert with the VLM,
make alert clips searchable, and generate a shift report for the line supervisor.

The film compresses that to "Tell me if the juice ends up outside a bottle." Same intent. The long prompt is what the skill is specified against.

How 3.3 is different from the May blueprint

explainx.ai's May VSS post described the Metropolis Video Search and Summarization stack: VLMs, RAG, skills, UI. 3.3 does not replace that. It adds a composer and a runtime pruner.

table · 3 cols
LayerMay-era VSSVSS 3.3
How you stand it upPick skills and Compose by handvss-build-vision-ai starts from a Foundation profile
Shared infraEasy to duplicate Kafka / Redis / ESSkill converges shared roles onto one instance
ChangeRebuild or edit YAMLSmaller delta on a running deploy
VLM tokensFull frames into the reasonerAdaptive EVS drops unchanged patches
MCPIn the blueprintStill in the graph — tools, not a new protocol

MCP here is the same Model Context Protocol already in the VSS agent: tools for search, sensors, clips. The new skill is the thing that decides which services exist so those tools have something to call.

NVIDIA also says earlier skills still exist for deploy, cameras, search, reports. 3.3 organizes them. vss-build-vision-ai sits on top.

Adaptive EVS is the other half of 3.3

A juice filler is mostly the same pixels, frame after frame. Sending all of them to a VLM is how visual agents get expensive.

Adaptive Efficient Video Sampling compares patches to the prior frame (cosine similarity), drops the still ones, and batches work around activity. NVIDIA's published RTX PRO 6000 Blackwell + Cosmos 3 Super FP8 numbers:

table · 2 cols
MetricClaimed change
Alert contextualization latency1,021 ms → 844 ms (17% faster)
Concurrent real-time VLM streams13 → 19 (46% more)
60-minute summaryAbout half the time, 80% fewer VLM input tokens

Those are NVIDIA's own benches. They vary with motion, chunk length, and similarity threshold. EVS already existed at a fixed prune rate in vLLM and Cosmos NIM. 3.3's adaptive path lives in the real-time VLM microservice, not on a remote API.

If your scene is a busy intersection, do not expect juice-line savings. Benchmark your footage. NVIDIA says to set VIA_EVS_SESSION, VLM_VIDEO_PRUNING_RATE, and VLLM_EVS_SIMILARITY_THRESHOLD in override.env.

What this means for people who build with skills

Most explainx.ai readers will not own a bottling line. They will own a coding agent and a pile of cameras, or they will evaluate whether a SKILL.md is doing real work.

Steal three habits from this demo:

  1. Foundation + delta, not greenfield YAML. Four tested profiles beat a model inventing Kafka from memory. That is the same instinct as ACES / SkillEvaluator: measure whether the skill helped, do not trust the README.
  2. Show the architecture before up. The video is explicit — nothing deploys until the plan is readable and four questions are answered.
  3. Separate compose tokens from runtime tokens. A cheap plan that stands up 20 services can still light two GPUs forever. Adaptive EVS is the runtime half; the $3 is not.

If you already run Pi + MCP + Codemode, treat VSS skills as another harness-side package, not a Pi feature. Same skill-file idea, different host.

Honest limitations

  • Vendor demo. Times, token counts, and $3 are the presenter's and NVIDIA's. Reproduce on your harness before you quote them in a budget.
  • Synthetic spill. Good for showing the alert path. Bad as evidence your line will page correctly.
  • Twenty services is still twenty services. The skill hides the first compose. You still operate Kafka, Redis, Elasticsearch, VIOS, HAProxy, and two NIMs.
  • Hardware floor. Published path is a two-GPU Blackwell workstation, not a MacBook.
  • Security is your job. NVIDIA lists TLS, auth, isolation. A juice-line agent with open RTSP on a plant VLAN is a different product than the film.
  • Not a robot policy. Cosmos-the-watcher does not replace Helix or a VLA. It writes tickets.

What to do this week

  1. Read the 3.3 blog and the VSS skills docs.
  2. If you already cloned the May-era blueprint, check out the 3.3 skills branch and symlink skills so git pull stays current.
  3. Run the long sample prompt against recorded video first. Do not point it at production RTSP until the architecture diagram and resolved.yml look boring.
  4. Enable Adaptive EVS only on RT-VLM workloads and measure miss rate on your own quiet-vs-event footage.
  5. NVIDIA also scheduled a live build October 1, 2026, 9 a.m. PT — same one-prompt claim. Treat the recording as a second source, not a second product.

Related reading

  • NVIDIA VSS Blueprint (May 2026 primer)
  • NVIDIA Cosmos 3: Reasoner vs Generator
  • What are agent skills?
  • What is MCP?
  • NVIDIA ACES: does this skill actually lift?
  • NVIDIA SIGGRAPH 2026: Cosmos, MCP, physical AI
  • Pi adds MCP and Codemode
  • Official: VSS 3.3 technical blog
  • Official: VSS agent skills
  • Film: YouTube tutorial

Compose times, token counts, dollar figures, Adaptive EVS percentages, and GPU placements are NVIDIA's September 29, 2026 post and the accompanying tutorial. The overflow clip is labeled synthetic. Verify on your own host before treating any of those numbers as a quote for procurement.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Aug 24, 2026

NVIDIA ACES: 27% of Agent Skill Runs Don’t Beat Baseline

NVIDIA's ACES framework (August 2026) pairs with-skill vs baseline agent runs on 947 tasks. Composite lift averages +0.21 — but 27% of cases show zero or negative lift. Document scans barely predict runtime value.

Jun 15, 2026

NVIDIA SkillSpector: Security Scanner for AI Agent Skills (2026)

Research shows 26.1% of agent skills contain vulnerabilities and 5.2% show likely malicious intent. NVIDIA's SkillSpector is an open-source scanner that catches them before they reach your agent.

Sep 29, 2026

Claude Code Can Build Evals and Hillclimb Them: /claude-api build-eval

On September 28, 2026, Lance Martin published Anthropic’s playbook for eval design and hillclimbing in the claude-api skill. /claude-api build-eval interviews you, samples production traces, and validates graders. /claude-api hillclimb then patches one change per round against a train/test split. One internal support benchmark cut cost to about one-fifth while lifting held-out accuracy from 78.6% to 90.5%.