NVIDIA's VSS 3.3 Unveils Cost-Cutting AI Innovations with Live Demo on October 1
September 29, 2026
The latest update introduces two cost-reduction innovations: Build Vision Agent skill (vss-build-vision-ai) to accelerate development by composing and extending deployments, and Adaptive Efficient Video Sampling (EVS) to cut runtime VLM processing by pruning unchanged video regions and batching work around events.
A live learning event is scheduled for October 1 at 9 a.m. PT to demonstrate building a visual AI agent from a single prompt, using an orange juice bottling line scenario as the workflow example.
To get started with VSS 3.3, users should clone the VSS Blueprint repo, install skills, describe the desired agent, review architecture and override.env, enable Adaptive EVS for RT-VLM workloads, and benchmark by running tests on representative footage.
Performance results with Adaptive EVS on an RTX PRO 6000 Blackwell and Cosmos 3 Super FP8 show a 17% reduction in alert contextualization latency, 46% more concurrent real-time VLM streams, and 80% fewer VLM input tokens for a 60‑minute video, with results varying by scene motion and configuration.
The Build Vision Agent skill enables starting from one of four validated developer profiles, computing the smallest delta to customize a deployment, reusing shared infrastructure to minimize duplicates, and generating deployment plans and architecture diagrams for review before deployment.
NVIDIA’s VSS Blueprint 3.3 aims to lower the cost of building and running visual AI agents by integrating vision-language models, retrieval-augmented generation, and MCP tools to deliver natural-language search, visual Q&A, verified alerts, and automated reporting from live and recorded video.
The content highlights reducing three main cost drivers—development, operating, and change—by reusing infrastructure, lowering token usage, and enabling incremental deployment changes through the Build Vision Agent skill and Adaptive EVS.
Adaptive EVS dynamically prunes video patches based on cosine similarity to the previous frame and uses event-aware batching to retain tokens only for active moments, reducing GPU usage and token counts while increasing the number of concurrent streams.
Summary based on 1 source
Get a daily email with more AI stories
Source

NVIDIA Technical Blog • Sep 29, 2026
Lower the Cost of Building and Running Visual AI Agents with NVIDIA VSS Blueprint 3.3