Live data from Hacker News

Internet Archive as a default host-of-record for startups

twitter.com

181–188 of 188 posts

Re: Internet Archive as a default host-of-record for startups

#181

Earlier quoted context omitted.

No you are right of course. I was imagining that a full text Wayback Machine search engine would mostly be useful to look for words unique enough that sifting through a lot of results (even if those were not ranked "well") could still be useful..? I was very naively "back of the envelope" prototyping a search engine. I both realize that this is not the way that these things are built, and both would really like to ha…

I may have given the wrong impression: it sounds like a good idea to me. > I was imagining that a full text Wayback Machine search engine would mostly be useful to look for words unique enough that sifting through a lot of results (even if those were not ranked "well") could still be useful..? If we think about use cases, users may often search specific domains. In that case, results ordered by frequency and/or date…

I agree & appreciate your response.

I have 0 time for this, but I also can't easily let go ;-)

Want to collab on this?

Re: Internet Archive as a default host-of-record for startups

#182

Earlier quoted context omitted.

I may have given the wrong impression: it sounds like a good idea to me. > I was imagining that a full text Wayback Machine search engine would mostly be useful to look for words unique enough that sifting through a lot of results (even if those were not ranked "well") could still be useful..? If we think about use cases, users may often search specific domains. In that case, results ordered by frequency and/or date…

I agree & appreciate your response. I have 0 time for this, but I also can't easily let go ;-) Want to collab on this?

Wow, great offer and a great project. I also have 0 time and I'm dealing with extra obligations anyway. I'm afraid I need to be disciplined to keep my primary life goals - things that take a decade or more - on course. I am really thrilled by the energy you are showing, though, and hate to contribute anything negative. Please don't stop for me!

You made my day.

Re: Internet Archive as a default host-of-record for startups

#183

Earlier quoted context omitted.

I agree & appreciate your response. I have 0 time for this, but I also can't easily let go ;-) Want to collab on this?

Wow, great offer and a great project. I also have 0 time and I'm dealing with extra obligations anyway. I'm afraid I need to be disciplined to keep my primary life goals - things that take a decade or more - on course. I am really thrilled by the energy you are showing, though, and hate to contribute anything negative. Please don't stop for me! You made my day.

Three cheers, thank you for the nice exchange! Happy upcoming Holidays. I'll (obviously) post on HN if anything comes out of this.

Re: Internet Archive as a default host-of-record for startups

#184
post #16

Earlier quoted context omitted.

You're consenting by posting it on public internet in the first place.

Posting something on the public internet is not consent for you to scrape it and post it on your own site forever. And requiring an explicit opt-in would basically mean no IA. To be clear, the IA is a positive, maybe even a great one. But it skirts by because most people don't care. (They did as you say post whatever on the public internet.) Add the facts that they're a non-profit, aren't trying to monetize their hos…

I strongly believe that theres no freedom of speech without the freedom to replicate that speech.

Re: Internet Archive as a default host-of-record for startups

#185

Hi, I manage the Wayback Machine at the Internet Archive. Very happy so many people here care about preserving, and making available, our cultural heritage! Please know a dedicated, and talented, team of engineers works every day to do a better job of archiving more of the public Web, and making it available via the Wayback Machine. As noted the Internet Archive is experimenting with filecoin.io and storj.io and is a…

Thank you!

Re: Internet Archive as a default host-of-record for startups

#186
post #47
post #43

Is it really worth archiving ?

It's difficult to know beforehand! A startup might be a total flop that delivers nothing of value to us now, but it might be a valuable datapoint for future historians to understand how startup culture changed and evolved. Or it may be interesting to future founders - I've seen some startups with perfectly fine ideas fail, and then a few years later, someone succeeds doing something very similar. This may not be the…

that's optimistic, and I could see the potential. But I've become a bit suspicious of the content on the internet and the costs in archiving it. Maybe Internet Archive could have a curation threshold .

I would rather revert to the internet where quality content was published openly on the web. But presently web content, at least that upranked by Google, is very low quality (seo, clickbait, biased, trivial)

Re: Internet Archive as a default host-of-record for startups

#187
post #155

What compression does IA use to store websites? Using a 2x better compression will allow them to store 2x more websites/content. I am doing some compression research, and would love to help IA in any way I can. There are some amazing SOTA compression algorithms available now. And if IA ignores images/video, and focuses only on text, they can store an insane amount of websites at a very low cost.

right now, most of it is coming in as ZSTD with a central index. older stuff was gzip

I am working on a ML-based compression which will be slow to compress/decompress but will give much better compression than zstd/gzip (even 2-3x better). Do you think that it's a useful algorithm for archival reasons or for storing huge amounts of data in cloud?

Re: Internet Archive as a default host-of-record for startups

#188
post #157
post #3

Personally, I think eternally archiving everything and infinitely available public data has been not-so-great. If this was an "archive with consent" sort of system, then sure. My response may be better summarized as, "Does IA support robots.txt, and if not why?"

They do respect robots.txt matter of fact, it will remove the website outright and most of its archives will be hidden until that website's robots.txt is offline. Their documentation about it is rather crap right now, but its in several of their FAQs about it

Respecting robots.txt would mean not saving pages blocked by robots.txt. Not just hiding.
Post reply on HN