Vision Zero Agent · team-4 demo

Download video (MP4) · Slides (PDF) · Slides (HTML) · Narration (MP3) · War-footage agent

Narration by scene

Voiced with Piper TTS (en_US-lessac-medium, research/non-commercial voice).

  1. 01_title
    A news desk gets hours of raw footage from remote cameras in conflict zones and disaster areas. Somewhere inside are five seconds that matter. Nobody can sit and watch it all live. The Vision Zero Agent watches it for them. Today it looks for explosions; the event is just a prompt, so it can watch for fires, floods or collapses too.
  2. 10_war
    We start with the hardest footage: five public domain archival films. NVIDIA Cosmos Reason, on CoreWeave GPUs, watched every five second clip, and we hand checked two hundred twenty two of them. It found fifty one of fifty six real explosions, at eighty five percent precision. Bomb releases are its known limit: parachutes and cargo drops fool it, and we show that too.
  3. 10b_live
    This is how it works live. As footage arrives, the agent cuts it into five second segments and checks each one as it lands. The moment it sees an explosion, it saves the clip and alerts the news desk on screen and by email. On three films it had never seen, every blast raised an alert, with no false alarms in thirty six segments. Its misses are the first flash, before the cloud forms.
  4. 02_problem
    The same agent protects people on everyday cameras. Our archive on VAST holds over three thousand five hundred clips from thirteen cameras: street cameras where cars cut close to pedestrians and cyclists, and warehouse floors where workers step into the path of forklifts and robots.
  5. 03_overview
    The agent scans the whole archive, ranks every camera by risk, and keeps a live feed of the worst moments first.
  6. 04_warehouse
    On the warehouse floor, NVIDIA Cosmos Reason, running on CoreWeave GPUs, watches every five second clip. Here a worker walks right up to a moving forklift as it turns. YOLO boxes from the VAST pipeline mark the people, in red when the clip is risky. Cosmos rates it risk three, names the machine, and counts the workers nearby.
  7. 06_street
    The same agent covers city streets. On the cyclist's helmet camera it flags close passes, blocked bike lanes and drivers cutting across the rider's path.
  8. 07_rule
    Safety teams can write a watch rule in plain English. A model on Weights and Biases inference turns it into a filter, and matching clips appear instantly.
  9. 09_wandb
    Every Cosmos and language model call is traced in Weights and Biases Weave, and every number is measured against human checks: fifty one of fifty six explosions in archival film, eight of ten street clips, and six of six warehouse clips.
  10. 11_stack
    Built in one day with Cursor, on VAST for video search and storage, NVIDIA Cosmos on CoreWeave GPUs, and Weights and Biases for inference, tracing and evaluation.
  11. 12_close
    Vision Zero Agent. Any camera, any archive, every finding checked on the video. Thank you.