LAION releases 10-million-hour open video dataset with 1.3B URLs

Laion Big Video Dataset

LAION releases 10-million-hour open video dataset with 1.3B URLs

LAION-BVD (Big Video Dataset) is a massive open resource for multimodal learning, built from 1.3 billion platform-specific video URLs mined from CommonCrawl. The team downloaded 80 million videos totaling 10 million hours, then used content-aware scene detection to extract clips and generate synthetic captions for video and audio. Models trained on this data—including ViCLIP, CLAP, and CLIP—perform competitively on standard benchmarks, with gains scaling up to 2.1% over InternVid-trained models. The dataset also includes 300 million frames for image-text pre-training, offering a visual distribution distinct from typical web corpora. Released exclusively for research, it aims to counter the concentration of large-scale video data in proprietary companies.

Large-scale video datasets and the models trained on them are increasingly concentrated within a small number of predominantly proprietary technology companies, limiting independent scientific investigation and reproducibility.
  1. voidUpdate

    Damn, that's a lot of videos for them to contact the creators and ask for permission to use their content as machine learning training content. Unless of course, they didn't, and just went ahead with it anyway...

  2. vivzkestrel

    - can someone with expertise give us an overview of the architecture involved doing this

    - let us say you ran yt-dlp inside python aiohttp

    - surely your ll run a limit soon as your ip address will be flagged

    - what solutions do we have to auto rotate proxies in python

    - are there better, faster and more reliable ways to go about downloading a 100 million videos without getting your ip address blocked?

  3. topwalktown

    "Overall, we attempted to download 130M videos and achieved a link success rate of approximately 60%, resulting in 80M successfully retrieved videos with a total duration of 10M hours."

    I am astonished that the success rate is so high. How Youtube didn't block them, I don't know. But I think that this URL list won't age well because youtube will very quickly block any researcher trying to download these videos themselves.

More from this day

2026-08-27