Live data from Hacker News

The sum of all knowledge and the sorry state of the web

christianheilmann.com

141–150 of 187 posts

Re: The sum of all knowledge and the sorry state of the web

#141

Earlier quoted context omitted.

I second the call that WP is web done right. Yes, of course there's bias; it's not possible to produce bias-free content, and WP's particularly bad in the fields of politics and history, and really any field where facts aren't settled and feelings are strong. Enter critical thinking. If you dig just a little (e.g. read the talk pages and the edit histories), you can soon learn that some topic has been taken over by P…

> Anything to do with Israel/Palestine/West Bank is unreliable; the boss is a zionist, and so are a lot of the senior staff, so it's not surprising. But WP is a million times better than the web that search engines expose. Is there good evidence for this claim? Evidence is useful, especially for these claims of general bias. On the contrary, in general Wikipedia's content is likely biased against Israel when consider…

I find it interesting that your way of proving anti-Zionist sentiment is to demonstrate that there are quantifiably more leftist editors on Wikipedia. That does not demonstrate your thesis of anti-Israeli bias. Plenty of people who at least claim to be leftists are profuse in their support of Israel.

Neither should it surprise that a scholarly endeavor for bored young people on the internet tends to tilt leftward: there is a big education differential between political poles, and a strong demographic tilt based on age. To put it in a nutshell: if boomers spent their time writing citations instead of pounding out all-caps screeds about ivermectin in the comments sections of local newspapers, there would be a more robust wikipedia contingent.

Re: The sum of all knowledge and the sorry state of the web

#142
post #129

Earlier quoted context omitted.

I’m not sure you and the author are talking about the same thing. He mentions the fact you can’t even read the news from 10 years ago, the content has simply disappeared. No amount of YouTube videos can replace that. The problem is not “everything, everywhere” or a lack of filters but the extreme commercialization of all content available, closed networks, the short life of URLs…

You absolutely can read news from 10 even 20 even 100 years ago. One example: https://news.google.com/newspapers I think you forget or may not have grown up with the microfiche. The reality of the internet is that everyone has a voice and things will only be archived if someone gives a damn to archive them. And that's fine. Some information deserves to be transient. Hell, we've survived millennia without this level o…

This is very interesting. Do you know how are these licensed?

Re: The sum of all knowledge and the sorry state of the web

#143
post #105
post #69

Earlier quoted context omitted.

Homophobia generally spread throughout the world with Christianity, carried by colonialism. Again, China is a good example. And fixing homophobia in western societies strongly corresponds to decreasing importance of religion. It’s not limited to homophobia of course - pretty much every single Catholic claim about human sexuality is antiscientific bull - but I think homophobia is a good enough example of a harmful, fa…

For one thing, you may be underselling the homophobia in Islam a bit. And there's Tacitus, who says that the Germans punished homosexuals. That's a bit before the Christians, and it comes from someone from a culture not opposed to the practice. Second, I find it hard to believe that the Chinese were turned homophobic, and remained so, by a few Christians despite millenia of tolerance, when the vast majority of the co…

I guess his take on germanic tribes might have been accurate. The first recorded "law" in the Nordics is monetarily punnishment for accusing someone of being homosexual. (Quoted from distant memory, don't take my word for it).

Re: The sum of all knowledge and the sorry state of the web

#144

Earlier quoted context omitted.

From my admittedly limited understanding, the failure of the semantic web is one of mankind's biggest missed opportunities. Now the knowledge graph is just locked behind Google's neural network layers and only being used for ads.

The idea behind semantic web was inspiring and great, however it required considerable work on the part of people creating stuff for the web and that was never going to happen. Maybe it could have happened in some things like academia based or knowledge based websites, but on the larger scale it was doomed.

(Warning: Personal plug incoming)

I fully agree, especially when it comes to the "semantic" part of the semantic web. Reusing and publishing ontologies that define those semantics always seemed like an afterthought of the semantic web, when it should be part of the foundation that things on the semantic web are built on.

In most other parts that make up a website (JS and HTML) we figured out how to make reuse (mostly) work by replacing flimsy web references with package management. Ontologies never had something like that, and thus were stuck in an early 00s era of software/ontology development.

Where I work, we are building Plow, a package manager for ontologies (https://github.com/field33/plow) as part of our tech stack to improve that situation and allow people to build applications with large-scale stable semantics at the core.

As part of building Plow we are aiming to make the process of creating and sharing ontologies easier and with that also lowering the barrier of entry to that domain.

Re: The sum of all knowledge and the sorry state of the web

#145

>Ever tried to look up some news from 12 years ago? Internet Archive which is a lot easier to search in than going manually through microfilm.

One reason Archive.Today is far more frequently used and referenced on HN and elsewhere, including by myself, is that it provides actual useful site access in many, many cases where the Internet Archive does not.

I'm a big fan of the Internet Archive and its mission. I'm concerned over Archive Today's severe lack of transparency (I've had a few exchanges with the site's operators, I've no idea who they are or what their motives are). I find the site useful but troubling.

I've also conducted microfilm (and -fiche) searches, as well as cataloguing of same. One affordance of a microform archive is that it is indexed, in ways that many online archives, including The Internet Archive ... aren't, usefully. The exchange is one of access-from-anywhere (yes, useful) against usable search and curation. It's ... an uncomfortable trade-off.

(The history of usefully cataloguing and indexing archives is itself a very old one. The US Librarian of Congress's annual letters to Congress, available through the otherwise almost wholly useless Hathi Trust, though, come to think of it, downloads of the entire letter rather than one single page at a time require using a different service ... are one interesting view to that process and the creation of the Library of Congress Classification and Subject Headings.)

Re: The sum of all knowledge and the sorry state of the web

#146
post #50

Earlier quoted context omitted.

I’m not sure you and the author are talking about the same thing. He mentions the fact you can’t even read the news from 10 years ago, the content has simply disappeared. No amount of YouTube videos can replace that. The problem is not “everything, everywhere” or a lack of filters but the extreme commercialization of all content available, closed networks, the short life of URLs…

> No amount of YouTube videos can replace that. The irony here is that Youtube videos from ten years ago are still alive and well. As Youtube makes a much better places for publishing and archiving content than the rest of the Web. With Youtube you don't have to worry about URLs changing or domain names expiring or anything like that. You just publish your video once, get a unique video-id and don't have to worry abo…

> There is no ISBN when you write a blog, no library were you could look up that ISBN.

An ISBN is a string of characters, much like a URI/URL, and offers no more "protection" for long term access and the latter. Books get mangled and lost too; their only benefit is that it is harder to mangle and lose them, and there is likely to be more than one of them.

Re: The sum of all knowledge and the sorry state of the web

#147
post #50

Earlier quoted context omitted.

> No amount of YouTube videos can replace that. The irony here is that Youtube videos from ten years ago are still alive and well. As Youtube makes a much better places for publishing and archiving content than the rest of the Web. With Youtube you don't have to worry about URLs changing or domain names expiring or anything like that. You just publish your video once, get a unique video-id and don't have to worry abo…

Unlisted youtube videos from 10 years ago are all gone... Google made the decision to delete every video 'shared by URL' because of the possibility that the URL generation algorithm had leaked. It was legally less risky to delete all the content than to risk leaking all the content to the open web. IMO, they made the wrong call - it would have been better for the internet as a whole to notify all users that "We have…

Do you have a source for this?

Re: The sum of all knowledge and the sorry state of the web

#148
I'm all for openness. Open standards for example. But 'open information' never worked; and never will. Powerful interests want to control information. They did it in Communist states. They do it in Capitalist states. They even do it on 'Anarchist' Wikipedia. That doesn't mean to say I favour closed off information. I just think people need to understand that - in opening up information and discourse - we are up against powerful interests = other people. Many of them are fithy rich and they desire to tell us what to think, read and say - because they "care" about us and "know what's best" for us.

Re: The sum of all knowledge and the sorry state of the web

#149

>Ever tried to look up some news from 12 years ago? I have a better one for you. Ever wondered why it's so hard? Why web protocols have nothing related to archiving? Why web browsers are a hellscape for aggregating information over time in a meaningful way? Why this continues to be true, despite countless Microsoft and Google engineers writing all these heartfelt posts about knowledge? If your answer is "because it's…

From my admittedly limited understanding, the failure of the semantic web is one of mankind's biggest missed opportunities. Now the knowledge graph is just locked behind Google's neural network layers and only being used for ads.

Maybe something like a 'WikiInfo' (or another better name :) ), that contains a hierarchy of (potentially all) known pages and topics? I think the only way to tackle this problem is collaboratively and distributedly.

You could add for example a 'Newspapers' topic, and then say 'The Springfield Times' and then have 'Articles by date', 'Articles by topic', etc. like a huge database (browsed hierarchically like "WHERE dates BETWEEN '20121211' and '20121213'", etc.). The primary datastructure could be a database, and users can add hierarchies as queries to the underlying database -- a collaborative index (in the literal sense, like a Homepage of the internet) is shown. Any unique 'object' (like a specific newspaper) gets an UUID and a row in the db. I don't know how modern dbs handle sparse data, but that'd definitely be a requirement (i.e. each object can have a handful of millions of possible properties, like publication date, location, author, colour, etc.).

Re: The sum of all knowledge and the sorry state of the web

#150

I definitely worry about the future of scholarship as our print media becomes more and more fragmented and hidden. It's not necessarily about ensuring essential knowledge is carried forward, but rather a "sense of the past". In researching history, the further back you go the more fragmented and unreliable your sources usually become. So it becomes harder and harder to figure out the broad sweep of these past culture…

while I have the same general concerns, both wikipedia and archive.org have offline backup and running options for free. and, they are surprisingly small. all one has to do is write a simple script to auto download these backups daily/weekly/whatever and you can access all that info of your solar powered raspberry pi.

while I think there could be a short period where this info is largely unavailable, I don't believe it will be lost forever. if the Internet goes away, once some new Internet like technology comes around to replace it, these data repos will likely get out back up very quickly. just a matter of how long that blip is

Post reply on HN