« All posts

SAGA Framework Pinpoints Which AI Model Generated a Video

SAGA is a new framework that attributes AI-generated videos to their source model across five levels, using minimal labeled data and interpretable signatures.

As synthetic videos become nearly indistinguishable from real footage, binary real-or-fake detectors are no longer enough. SAGA (Source Attribution of Generative AI videos) is introduced as the first comprehensive framework designed to identify not just whether a video is synthetic, but exactly which generative model produced it.

The system provides multi-granular attribution across five levels: authenticity, generation task (text-to-video or image-to-video), model version, development team, and the specific generator used. A novel video transformer architecture, built on features from a robust vision foundation model, captures spatio-temporal artifacts that make different generators distinguishable.

A key contribution is a data-efficient pretrain-and-attribute strategy that matches fully supervised performance using only 0.5% of source-labeled data per class. The paper also introduces Temporal Attention Signatures (T-Sigs), an interpretability method that visualizes learned temporal patterns, offering the first explanation of why generators differ in their signatures.

Extensive experiments, including cross-domain scenarios on public datasets, show SAGA sets a new benchmark for synthetic video provenance, providing interpretable insights valuable for forensic investigation and regulatory oversight.

This synthesis was produced from its source by AI; there is no human editor or manual review step. How we work