Live data from Hacker News

Yark: Advanced and easy YouTube archiver now stable

github.com

121–130 of 148 posts

Re: Yark: Advanced and easy YouTube archiver now stable

#121
post #119

Earlier quoted context omitted.

In a world where backing up to IPFS would be as easy and widespread as hitting Ctrl+D to bookmark and sharing backups would not be dangerous legally, your answer gets an easy answer. Each person gets to decide what they find important, backup and commit resources to. We already have the sharing technology (BitTorrent, IPFS), and the backuping tech (ArchiveBox, TubeArchivist). But they are not integrated and they are…

The current decentralized trend ("web3", etc.) becoming mainstream is a pipe dream only tech enthusiasts care about. The sad reality is that it has very slim chances of ever gaining mass adoption. The general public and non-technical users couldn't care less about owning and managing their data, running their own services, paying for services, and everything that entails. Even if they're aware that their privacy is b…

I fully agree with what you wrote but I just want to mention that, as you also imply, the ideas behind "the current decentralized trend ("web3", etc.)", really aren't new. They just build on older ideas with regards to decentralization.

There are many ideas that fit under decentralization: torrents, the fediverse, crypto, even the old idea of the semantic web (because it was about standardized formats for metadata and carrying that metadata with the data instead of having it siloed in a central entity).

All of the hype around web3 really is only about crypto, because web3 is a marketable term for speculators.

Currently I am very cautiously hopeful about what the hype surrounding Mastodon (caused by Twitters self-immolation) will lead to.

Re: Yark: Advanced and easy YouTube archiver now stable

#122
A while ago I did something similar. I'm already downloading various playlists via yt-dlp and wrote a web interface [1] to view/play/search them.

My biggest annoyance at the time was importing my existing videos (and converting them to a streamable format, generating thumbnails and hover-previews etc). Do you have any plans of allowing to import existing yt-dlp folder (in the standard layout with a bunch of mkv files, the info.json, the subtitles etc). Because my current archive contains a lot of already-deleted videos :(

[1] https://github.com/Mikescher/youtube-dl-viewer

Re: Yark: Advanced and easy YouTube archiver now stable

#123
post #43
post #8

Earlier quoted context omitted.

ytdl is one of its dependencies You clearly didn't even skim the readme to see what this does

From the HN guidelines: > Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".

There's a difference between not reading an entire article or missing reading a paragraph versus not even clicking the link or skimming a couple screenshots to figure out what you're commenting on.

Re: Yark: Advanced and easy YouTube archiver now stable

#124
post #41

Earlier quoted context omitted.

It's unbelievable that Wikipedia is free and survives on donations. Youtube sells ads and is subsidized by one of the biggest ad companies in the world that happens to have a lot of cheap cloud storage available.

YT is not subsidized by Google, it makes a profit. The scale of storage and serving is not of any real comparable scale.

Not for the first 15 years it didn't.

Re: Yark: Advanced and easy YouTube archiver now stable

#125

Does this have the ability to bet set to "grab highest available resolution" instead of specifying one? A lot of the material I'd like to archive has material from well before and after Youtube started supporting HD resolutions.

I would actually like the opposite: cap the resolution. I don't really need 4k videos.

Re: Yark: Advanced and easy YouTube archiver now stable

#126

Earlier quoted context omitted.

Thank you. I have a couple like that and didn’t know how confident to be with replicating It seems Apple is fine with adblocking by default for web content, but inconsistently as with youtube. Hard to predict what's risky to invest dev time into

Apparently they don't like when your app does something in a way that breaks ToS or EULA, even if they are not worth the pixels they are written on. Apple being the trillion dollar company it is, is naturally averse to the legal risk posed by having a ToS-breaking app on their store. This is why having a single company controlling your entire device is bad -- they do what's best for them, rather than whatever you mig…

Agree it's bad, it's just where the money is for b2c indie dev. I hope their app store monopoly gets broken up.

I get confused about their stance on this because they allow products like AdGuard or browsers that block ads by default

Re: Yark: Advanced and easy YouTube archiver now stable

#127
post #94

Earlier quoted context omitted.

Thank you. I have a couple like that and didn’t know how confident to be with replicating It seems Apple is fine with adblocking by default for web content, but inconsistently as with youtube. Hard to predict what's risky to invest dev time into

I'm not an expert on Apple's rules but my understanding is the download thing is circumvention of the terms under which Youtube licenses/re-licenses content so it's treated fairly strictly.

Then why do they allow other ad-blocking products or browsers that block ads by default? I don't understand the consistency of their position except that YouTube has a powerful lobby

Re: Yark: Advanced and easy YouTube archiver now stable

#128

Earlier quoted context omitted.

Long ago I had my podcast downloader keep all files it downloads and recently I've been using OpenAI's Whisper to go through and create transcripts of the 8000 or so hours of data I have downloaded over the years. It's very cool to be able to search through and remind myself of something I heard once. Not exactly life changing, but still, nice to be able to quickly drill down and find audio for something when a curio…

8000 hours? Napkin math time, that's 20 years of 10+ hours daily. I call BS.

It's about an hour or two or a day.

This includes data from 1995 on. The early data is backfill of radio shows that transitioned to podcasts and dumped old episodes in their feed at some point. My reader itself started in 2012, I downloaded around 7000 hours of new podcasts, which works out to 1.7 hours per day. So, around 2 hours per day, since I don't listen every day, and to be fair, I haven't listened to every podcast I've downloaded, some don't interest me. But 1-2 hours of listening a day is the sweet spot for me.

Re: Yark: Advanced and easy YouTube archiver now stable

#129
If you're interested in this kind of thing, you may also want to check out:

- the Distributed YouTube Archive Discord: https://discord.com/invite/PQqks7eSKc

- ArchiveTeam also do a significant amount of YT archiving: https://wiki.archiveteam.org/index.php/YouTube

- a similar, but private effort: https://reddit.com/r/Archivists/comments/5uvfpw/youtube_arch...

Re: Yark: Advanced and easy YouTube archiver now stable

#130

Earlier quoted context omitted.

Long ago I had my podcast downloader keep all files it downloads and recently I've been using OpenAI's Whisper to go through and create transcripts of the 8000 or so hours of data I have downloaded over the years. It's very cool to be able to search through and remind myself of something I heard once. Not exactly life changing, but still, nice to be able to quickly drill down and find audio for something when a curio…

What kind of hardware do you have that makes it feasible to process thousands of hours of podcasts? I want to do the same but I’ve heard that Whisper requires some serious GPU might for decent accuracy (Linux Unplugged podcast specifically).

Yep, it takes a bit of GPU RAM. I'm using 3 machines with NVidia 3080 or better. I let them go for a few weeks over the winter break when I was mostly disconnected from the tech world. The workers prioritized podcasts I'm personally likely to want to search, and got through almost a third of my archive.

Now it's down to 1 or 2 machines depending on what's going on, so it'll take much longer to finish up, but I'm in no rush.

Post reply on HN