NVIDIA Launches Alpamayo 2 Super: Revolutionizing Autonomous Vehicle AI with 34-Billion-Parameter Model

August 4, 2026
NVIDIA Launches Alpamayo 2 Super: Revolutionizing Autonomous Vehicle AI with 34-Billion-Parameter Model
  • NVIDIA unveils Alpamayo 2 Super, a 34‑billion‑parameter reasoning vision‑language‑action model designed to unify trajectory generation, high‑level intent reasoning, scene understanding, and data labeling for autonomous vehicles.

  • The system provides five synchronized outputs per driving scenario: a planned trajectory, a Chain‑of‑Causation reasoning trace, a high‑level meta‑action indicator, auto‑labels, and visual question‑answering responses, enabling traceability for safety auditing and compliance.

  • Alpamayo 2 Super outputs include future ego trajectories, CoC traces, meta‑actions, grounded scene answers via VQA, and auto‑labels, creating a common foundation across offline policy teaching, evaluation, data engines, and task customization.

  • A standardized evaluation framework uses open‑loop, closed‑loop, and IoU‑based metrics for meta‑actions and VQA grounding, with internal judge‑model similarity scores for CoC auto‑labeling (0.652) and 2D grounding IoU of 0.71.

  • The article describes four workflows: generating trajectories with CoC traces, predicting meta‑actions, answering natural‑language questions about multi‑camera scenes, and generating CoC auto‑labels with 2D grounding on user clips.

  • Meta‑actions summarize plans (yield, lane changes, stop, accelerate) to bridge end‑to‑end models with modular AV stacks, with a notebook showing IoU accuracy across lateral, longitudinal, and lane components (about 74.6, 61.9, and 73.6 on 94k clips).

  • Training data comprises roughly 115,000 hours of multi‑camera driving video with egomotion and trajectory annotations, plus about 3.7 million CoC reasoning traces collected from cameras, inertial sensors, and GPS.

  • VQA and grounding enable natural‑language scene descriptions and 2D bounding‑box localization across surround cameras, supporting interactive debugging and tighter links between evidence, reasoning, and action.

  • The article cautions that benchmarks are not roads, notes that the full model is memory‑intensive requiring tens of gigabytes of GPU memory, and warns that tools alone do not guarantee success against competitors.

  • The system uses input from up to seven cameras for full 360‑degree perception and targets Level 4 autonomy, handling driving within a defined area without a human driver.

  • Architecturally, Alpamayo 2 Super combines a 32‑billion‑parameter vision‑language backbone with a 2.3‑billion‑parameter diffusion‑based action decoder, ingesting multi‑camera video, text, and motion history to output trajectories, CoC reasoning, meta‑actions, training labels, and VQA grounded to image regions.

  • Auto‑labeling enables offline generation of structured CoC labels at scale, shortening annotation cycles from months to days, and can incorporate predicted future trajectories as observed data.

Summary based on 5 sources


Get a daily email with more Startups stories

More Stories