Cameras that see everything and tell you nothing are not a poetry problem. They are a composition problem: ingest, detect, verify, store, search, report — usually five teams and a graveyard of YAML.
NVIDIA's answer in VSS Blueprint 3.3 is a single agent skill: vss-build-vision-ai. The tutorial now circulating as https://youtu.be/PPXoPSCdGyU walks that skill through an orange-juice line. One English prompt. A plan you can read. Four questions. Then a live stack.
The official engineering write-up landed September 29, 2026: Lower the Cost of Building and Running Visual AI Agents with NVIDIA VSS Blueprint 3.3. The film is the same claim with a leaky filler. This is not a new Cosmos 3 world model. It is the May VSS blueprint growing a composer.
TL;DR
| Question | Answer |
|---|---|
| What shipped? | VSS 3.3 — Build Vision Agent skill + Adaptive EVS |
| Skill id | vss-build-vision-ai |
| Watcher | Cosmos reasoning VLM on camera chunks |
| Reporter | Nemotron (docs default: Nemotron Nano 9B v2) |
| Demo prompt | Tell me if the juice ends up outside a bottle |
| Video numbers | ~20 services, 26 minutes, 7.5M tokens, about $3 compose |
| Official hardware | Two RTX PRO 6000 Blackwell GPUs for the bottling write-up |
| Runtime bill | Local GPU, "pennies per task" — not a cloud token meter |
| Is the spill real? | Stack is real; NVIDIA says the overflow footage is synthetic |
What people are asking after the clip
"Did one prompt really deploy 20 services?"
Yes, in the sense NVIDIA means: the skill reads the VSS Blueprint repo, maps the ask onto four validated foundations (base, alerts, lvs, search), computes a delta, shows an architecture, then writes a self-contained _builds/<name>/ tree. It does not rewrite the repo's deploy/docker/ tree.
The video's sequence matches the blog:
- You describe cameras, alerts, and a web UI.
- The skill proposes Cosmos to watch and Nemotron to write.
- It asks where to build, which GPUs, host address, cameras.
- Only after you confirm does Compose come up.
That is progressive disclosure with a human gate, not "the model SSHed into prod." If you want the same pattern in software agents, it is closer to a deploy skill than to a chat widget.
"Is $3 the price of a factory AI?"
No. Split the invoice the way NVIDIA does.
| Cost | What it is | What it is not |
|---|---|---|
| ~$3 / 7.5M tokens | Coding-agent spend to plan, generate, and verify the stack | Per-frame inference |
| 26 minutes | Wall clock for that compose on their machine | Your laptop |
| Pennies per report | Local Nemotron writing from stored incidents | A hosted API with no GPU |
| Two Blackwell PRO 6000s | The published bottling-line host | Included in the $3 |
If you do not already own the GPUs, the honest first number is capex + power, then the $3. The useful claim is narrower: composition got cheap enough to demo in half an hour. Runtime VLM cost is a different knob — Adaptive EVS.
"Which Cosmos is watching the line?"
The video says Cosmos is NVIDIA's open reasoning VLM. Docs for the default VSS agent pair nvidia/cosmos3-nano-reasoner on port 30082 with nvidia/nvidia-nemotron-nano-9b-v2 on 30081. The bottling write-up puts FP8 Cosmos 3 Nano on the same GPU as the detector.
That is the Reasoner surface — text out about video — not the Generator that rolls futures for robots. Keep them separate the way explainx.ai already does in the Cosmos 3 guide and the SIGGRAPH physical-AI recap.
The video's runtime loop: Cosmos looks at the stream in 10-second chunks and asks the prompt. A leak hits the dashboard "within seconds." Chat: generate a report. Nemotron writes from the incident store. Under a minute, presenter says, no cloud token bill.
"Can I paste this onto my cameras tonight?"
Only if you already have:
- A coding agent that loads agentskills.io packages (NVIDIA names Claude Code, Codex, NemoClaw).
- The VSS 3.3 skills installed (symlink from the blueprint repo, not a one-off copy).
- GPUs the profiles actually fit.
- RTSP (or recorded) sources, and a trusted network. NVIDIA's getting-started list ends with authentication, TLS, rate limits, and isolation.
The sample prompt from the official post is the one to reuse — not the one-liner from the film:
Build a VSS vision agent for an orange juice bottling line.
Use two RTSP cameras on the filler and capper.
Detect bottle overflows and juice spills, verify each alert with the VLM,
make alert clips searchable, and generate a shift report for the line supervisor.
The film compresses that to "Tell me if the juice ends up outside a bottle." Same intent. The long prompt is what the skill is specified against.
How 3.3 is different from the May blueprint
explainx.ai's May VSS post described the Metropolis Video Search and Summarization stack: VLMs, RAG, skills, UI. 3.3 does not replace that. It adds a composer and a runtime pruner.
| Layer | May-era VSS | VSS 3.3 |
|---|---|---|
| How you stand it up | Pick skills and Compose by hand | vss-build-vision-ai starts from a Foundation profile |
| Shared infra | Easy to duplicate Kafka / Redis / ES | Skill converges shared roles onto one instance |
| Change | Rebuild or edit YAML | Smaller delta on a running deploy |
| VLM tokens | Full frames into the reasoner | Adaptive EVS drops unchanged patches |
| MCP | In the blueprint | Still in the graph — tools, not a new protocol |
MCP here is the same Model Context Protocol already in the VSS agent: tools for search, sensors, clips. The new skill is the thing that decides which services exist so those tools have something to call.
NVIDIA also says earlier skills still exist for deploy, cameras, search, reports. 3.3 organizes them. vss-build-vision-ai sits on top.
Adaptive EVS is the other half of 3.3
A juice filler is mostly the same pixels, frame after frame. Sending all of them to a VLM is how visual agents get expensive.
Adaptive Efficient Video Sampling compares patches to the prior frame (cosine similarity), drops the still ones, and batches work around activity. NVIDIA's published RTX PRO 6000 Blackwell + Cosmos 3 Super FP8 numbers:
| Metric | Claimed change |
|---|---|
| Alert contextualization latency | 1,021 ms → 844 ms (17% faster) |
| Concurrent real-time VLM streams | 13 → 19 (46% more) |
| 60-minute summary | About half the time, 80% fewer VLM input tokens |
Those are NVIDIA's own benches. They vary with motion, chunk length, and similarity threshold. EVS already existed at a fixed prune rate in vLLM and Cosmos NIM. 3.3's adaptive path lives in the real-time VLM microservice, not on a remote API.
If your scene is a busy intersection, do not expect juice-line savings. Benchmark your footage. NVIDIA says to set VIA_EVS_SESSION, VLM_VIDEO_PRUNING_RATE, and VLLM_EVS_SIMILARITY_THRESHOLD in override.env.
What this means for people who build with skills
Most explainx.ai readers will not own a bottling line. They will own a coding agent and a pile of cameras, or they will evaluate whether a SKILL.md is doing real work.
Steal three habits from this demo:
- Foundation + delta, not greenfield YAML. Four tested profiles beat a model inventing Kafka from memory. That is the same instinct as ACES / SkillEvaluator: measure whether the skill helped, do not trust the README.
- Show the architecture before
up. The video is explicit — nothing deploys until the plan is readable and four questions are answered. - Separate compose tokens from runtime tokens. A cheap plan that stands up 20 services can still light two GPUs forever. Adaptive EVS is the runtime half; the $3 is not.
If you already run Pi + MCP + Codemode, treat VSS skills as another harness-side package, not a Pi feature. Same skill-file idea, different host.
Honest limitations
- Vendor demo. Times, token counts, and $3 are the presenter's and NVIDIA's. Reproduce on your harness before you quote them in a budget.
- Synthetic spill. Good for showing the alert path. Bad as evidence your line will page correctly.
- Twenty services is still twenty services. The skill hides the first compose. You still operate Kafka, Redis, Elasticsearch, VIOS, HAProxy, and two NIMs.
- Hardware floor. Published path is a two-GPU Blackwell workstation, not a MacBook.
- Security is your job. NVIDIA lists TLS, auth, isolation. A juice-line agent with open RTSP on a plant VLAN is a different product than the film.
- Not a robot policy. Cosmos-the-watcher does not replace Helix or a VLA. It writes tickets.
What to do this week
- Read the 3.3 blog and the VSS skills docs.
- If you already cloned the May-era blueprint, check out the 3.3 skills branch and symlink skills so
git pullstays current. - Run the long sample prompt against recorded video first. Do not point it at production RTSP until the architecture diagram and
resolved.ymllook boring. - Enable Adaptive EVS only on RT-VLM workloads and measure miss rate on your own quiet-vs-event footage.
- NVIDIA also scheduled a live build October 1, 2026, 9 a.m. PT — same one-prompt claim. Treat the recording as a second source, not a second product.
Related reading
- NVIDIA VSS Blueprint (May 2026 primer)
- NVIDIA Cosmos 3: Reasoner vs Generator
- What are agent skills?
- What is MCP?
- NVIDIA ACES: does this skill actually lift?
- NVIDIA SIGGRAPH 2026: Cosmos, MCP, physical AI
- Pi adds MCP and Codemode
- Official: VSS 3.3 technical blog
- Official: VSS agent skills
- Film: YouTube tutorial
Compose times, token counts, dollar figures, Adaptive EVS percentages, and GPU placements are NVIDIA's September 29, 2026 post and the accompanying tutorial. The overflow clip is labeled synthetic. Verify on your own host before treating any of those numbers as a quote for procurement.
