Laion Big Video Dataset
projects.laion.aiDamn, that's a lot of videos for them to contact the creators and ask for permission to use their content as machine learning training content. Unless of course, they didn't, and just went ahead with it anyway...
Under the EU's AI Act non-profits and universities can just bypass these constents to make non-commercial datasets btw
But I've been told that the AI act was a terrible anti-innovation legislation made by clueless and corrupt bureaucrats. How is it possible that they made a sensible decision?
Does that apply to everyone who doesn't live in the EU as well? And what happens when a major AI company uses this dataset as training data, ignoring consent as usual?
- can someone with expertise give us an overview of the architecture involved doing this
- let us say you ran yt-dlp inside python aiohttp
- surely your ll run a limit soon as your ip address will be flagged
- what solutions do we have to auto rotate proxies in python
- are there better, faster and more reliable ways to go about downloading a 100 million videos without getting your ip address blocked?
Google may supply cache boxes to ISPs which can be manipulated to let download any video you want
"Overall, we attempted to download 130M videos and achieved a link success rate of approximately 60%, resulting in 80M successfully retrieved videos with a total duration of 10M hours."
I am astonished that the success rate is so high. How Youtube didn't block them, I don't know. But I think that this URL list won't age well because youtube will very quickly block any researcher trying to download these videos themselves.
I believe they have the actual videos to download, "exclusively for academic and non-commercial research." and you have to request access, https://projects.laion.ai/bvd/bvd-raw-access.html
They say they used yt-dlp and "employ a residential proxy network". But, yeah, youtube seems to now more aggressively block yt-dlp.
Ha, is it just me or does a “residential proxy network” sound like a fancy way of saying “botnet”?
Basically all of them are botnets, yes. They generally claim to have consent but I'm pretty sure 99% of it is "some app the user uses has it buried in a 200 page ToS"-style consent.
No, they are not. Just ask any AI.
Nope, they’re different things. Proxy networks like Proxybase [0] use open-source clients and ask for the user’s consent before allowing them to join the network.
That’s a pretty penny on the bandwidth bill alone!
> youtube seems to now more aggressively block yt-dlp.
Another thing to thank AI for! /s
I was wondering if youtube blocked them a lot and they got a lot of videos from other sources - but no, you can see on this HF page that 93.1% of URls are youtube.com, https://huggingface.co/datasets/laion/BVD-URLs