Live data from Hacker News

Yark: Advanced and easy YouTube archiver now stable

github.com

111–120 of 148 posts

Re: Yark: Advanced and easy YouTube archiver now stable

#111
post #2

I've been working on polishing my YouTube archiver project for the last while and I've finally released a solid version of it, it has an offline web viewer/visualiser for archived channels and it's managed using a cli. Most importantly, its easy to use :)

From the readme, it is not clear to me what is meant by metadata.

Will this archive subtitles?

Will it archive comments? If is can, can comments be updated?

Also, from the readme it looks like all metadata is kept centralized instead of each video having its own metadata file.

Are containers other than mp4 supported?

Re: Yark: Advanced and easy YouTube archiver now stable

#112

Earlier quoted context omitted.

YT is not subsidized by Google, it makes a profit. The scale of storage and serving is not of any real comparable scale.

I wonder how many orders of magnitude separate the amount of data stored in YouTube vs Wikipedia? 5? 6? How about data served by them?

About 8 for storage, wikipedia is gigabytes, youtube is around exabytes.

Re: Yark: Advanced and easy YouTube archiver now stable

#114
post #16
post #13

How cool would it be if everyone had IPFS running in their browser, and everyone dedicated some time to filling it with a backup of the internet, including YouTube.

I did some napkin math. If 1 billion people each backed up 10gb, we’d almost have enough to store a copy of YouTube with zero data redundancy. Google is massive.

I’m not aware that Google make public the total size of YouTube. Your estimate is 10EB. That seems quite low.

Re: Yark: Advanced and easy YouTube archiver now stable

#115
post #103
post #87

Earlier quoted context omitted.

That's impressive indeed, but if we boil it down to just the things that are worth saving, removing duplicates, long and pointless livestreams, long videos that are just endless loops, etc., it could be done with much less space. If we also ignore auto generated spam[1] and harmful content (Elsagate, etc.), it would require even less. It's a nice thought experiment, but we really don't need to archive all of YT. That…

I would argue that backing up selectively instead of just grabbing everything makes the job harder, not easier. You might need less storage, but who will decide (and how) what gets stored and what doesn't?

In a world where backing up to IPFS would be as easy and widespread as hitting Ctrl+D to bookmark and sharing backups would not be dangerous legally, your answer gets an easy answer.

Each person gets to decide what they find important, backup and commit resources to.

We already have the sharing technology (BitTorrent, IPFS), and the backuping tech (ArchiveBox, TubeArchivist). But they are not integrated and they are not easy to configure and use for a nontechnical person. And they are unlikely to become mainstream thanks to the copyright cartel.

Many years ago there was the dream of every home having a home computer giving people ownership over their digital life. A world where everyone has their own email server, their own diaspora pod for their family, their own blog, etc. etc. and all they had to do was buy or build a box and plug it in.

Projects like freedomplug and sandstorm.io and YunoHost and others. There are even newer projects like Umbrel. And of course NASs that now have apps on them.

Arguably, NASs and Umbrel are really easy to configure and use. OMV is somewhat harder but not unreasonable.

Alas, a self sovereign world is not what people desire. People are content with letting themselves be controlled by a few megacorps. Everything is in the cloud (someone else's computer).

Also, with the rise of botnets and DDOSs, it became unfeasible to self host something public without something like Cloudflare in front.

I miss the old internet.

Re: Yark: Advanced and easy YouTube archiver now stable

#116
post #78
post #32

Earlier quoted context omitted.

Normally in modern adaptive streaming, every video variant is muxed into a separate stream without audio, and different audio variants are muxed into their own individual streams.

wow, I wonder if that's why it always feels so frequent an experience of mine where the audio and video feel subtly out of sync with one another. It's very minute but detectable. Feels like that experience has increased in the last 6 months or so.

That kind of issue on the encoder OR player side would've been easily caught in testing, including through automated tests, so as crazygringo says, it's most likely your OS/hardware.

If you need to troubleshoot sync issues, this is a helpful tool: https://www.youtube.com/watch?v=HD4emXqHCsE

Re: Yark: Advanced and easy YouTube archiver now stable

#117
post #41

Earlier quoted context omitted.

It's unbelievable that Wikipedia is free and survives on donations. Youtube sells ads and is subsidized by one of the biggest ad companies in the world that happens to have a lot of cheap cloud storage available.

My understanding is that Wikipedia does not survive on, or even need donations, at all: they are entirely funded through their endowment. The donations go to the Wikimedia Foundation which spends the money on a bunch of other social stuff that's not related to running Wikipedia at all. So unless you want to support all those other causes, you're effectively wasting your money by donating to Wikipedia's frequent donat…

While that is true and it is much more worthwhile to donate to the Internet Archive instead of Wikipedia (many Wikipedia references depend on being archived by IA), I would not want to start a anti-Wikipedia-donations movement. Sure they have plenty of money now and are wasting some of it and are begging for more, but if they would stop receiving any, they can eventually burn through everything they have. And having reserves and a steady inflow are important for planning future projects.

However, at this point, IA needs your donations much much much more than Wikipedia does.

Re: Yark: Advanced and easy YouTube archiver now stable

#118
post #103
post #87

Earlier quoted context omitted.

That's impressive indeed, but if we boil it down to just the things that are worth saving, removing duplicates, long and pointless livestreams, long videos that are just endless loops, etc., it could be done with much less space. If we also ignore auto generated spam[1] and harmful content (Elsagate, etc.), it would require even less. It's a nice thought experiment, but we really don't need to archive all of YT. That…

I would argue that backing up selectively instead of just grabbing everything makes the job harder, not easier. You might need less storage, but who will decide (and how) what gets stored and what doesn't?

On the contrary. Instead of everyone grabbing everything, each person would only archive what's important to them. This not only distributes the workload naturally, but serves as an implicit filter of content people find enjoyable and would actually watch, rather than archiving content nobody cares about.

But this is all hypothetical. We don't need a global YouTube archive. We need to stop using it altogether, and replace it with decentralized services. In the meantime, the existing personal archiving solutions work well.

Re: Yark: Advanced and easy YouTube archiver now stable

#119
post #103

Earlier quoted context omitted.

I would argue that backing up selectively instead of just grabbing everything makes the job harder, not easier. You might need less storage, but who will decide (and how) what gets stored and what doesn't?

In a world where backing up to IPFS would be as easy and widespread as hitting Ctrl+D to bookmark and sharing backups would not be dangerous legally, your answer gets an easy answer. Each person gets to decide what they find important, backup and commit resources to. We already have the sharing technology (BitTorrent, IPFS), and the backuping tech (ArchiveBox, TubeArchivist). But they are not integrated and they are…

The current decentralized trend ("web3", etc.) becoming mainstream is a pipe dream only tech enthusiasts care about. The sad reality is that it has very slim chances of ever gaining mass adoption. The general public and non-technical users couldn't care less about owning and managing their data, running their own services, paying for services, and everything that entails. Even if they're aware that their privacy is being violated and that their personal data is sold on shady adtech markets, they see it as a cost worth paying for in exchange for the services they get for "free".

So even if all these technical solutions to problems only technical users care about become as easy to use for laypeople as modern web browsers are, the general public just won't care about it.

I've long believed that the blame for this lies mostly on early WWW architects. If the focus from the very start had been on sharing content as much as it was on consuming it, and user-friendly tools analogous to the web browser had been built, then the general public would be educated that the web works by being in control of your data, and sharing it selectively with specific people, companies, or the world. ISPs would be forced to deliver symmetrical connections to enable this, centralized services would be much less influential, and the web landscape would look much different today.

This was actually planned as a second phase in the original HyperText proposal[1], but was never completed for some reason. I'd be very interested to know what happened to this effort. If someone has insider knowledge, or can contact TBL, I'd be very grateful.

Alas, it's too late for this now. The centralized web is how most people experience the "internet", and that train has no chance of stopping.

[1]: https://www.w3.org/Proposal.html

Re: Yark: Advanced and easy YouTube archiver now stable

#120
I have a youtube archiver script I'm using right now that I pulled from a thread on the data hoarder subreddit. My main issue is that I need the downloader to remove emojis from the filename because I sometimes sync them to Dropbox. Can this project do that?
Post reply on HN