Live data from Hacker News

The semantic web is now widely adopted

csvbase.com

11–20 of 269 posts

Re: The semantic web is now widely adopted

#11
The argument about LLMs is wrong, not because of reasons stated but because semantic meaning shouldn't solely be defined by the publisher.

The real question is whether the average publisher is better than an LLM at accurately classifying their content. My guess is, when it comes to categorization and summarization, an LLM is going to handily win. An easy test is: are publishers experts on topics they talk about? The truth of the internet is no, they're not usually.

The entire world of SEO hacks, blogspam, etc exists because publishers were the only source of truth that the search engine used to determine meaning and quality, which has created all the sorts of misaligned incentives that we've lived with for the past 25 years. At best there are some things publishers can provide as guidance for an LLM, social card, etc, but it can't be the only truth of the content.

Perhaps we will only really reach the promise of 'the semantic web' when we've adequately overcome the principal-agent problem of who gets to define the meaning of things on the web. My sense is that requires classifiers that are controlled by users.

Re: The semantic web is now widely adopted

#12
post #8

The author gives two reasons why AI won't replace the need for metadata: 1: LLMs "routinely get stuff wrong" 2: "pricy GPU time" 1: I make a lot of tests on how well LLMs get categorization and data extraction right or wrong for my Product Chart ( https://www.productchart.com ) project. And they get pretty hard stuff right 99% of the time already. This will only improve. 2: Loading the frontpage of Reddit takes hundr…

GPU compute price is dropping fast and will continue to do so.

Re: The semantic web is now widely adopted

#13
post #8

The author gives two reasons why AI won't replace the need for metadata: 1: LLMs "routinely get stuff wrong" 2: "pricy GPU time" 1: I make a lot of tests on how well LLMs get categorization and data extraction right or wrong for my Product Chart ( https://www.productchart.com ) project. And they get pretty hard stuff right 99% of the time already. This will only improve. 2: Loading the frontpage of Reddit takes hundr…

LLMs have no soul, so I like content and curation from real people

The main problem is that the incentive for well-intentioned people to add detailed and accurate metadata is much lower than the incentive for SEO dudes to abuse the system if the metadata is used for anything of consequence. There's a reason why search engines that trusted website metadata went extinct.

That's the whole benefit of using LLMs for categorization: they work for you, not for the SEO guy... well, prompt injection tricks aside.

Re: The semantic web is now widely adopted

#14
Metadata in PDFs is also typically based on semantic web standards.

https://www.meridiandiscovery.com/articles/pdf-forensic-anal...

Instead of using JSON-LD it uses RDF written as XML. Still uses the same concept of common vocabularies, but instead of schema.org it uses a collection of various vocabularies including Dublin Core.

Re: The semantic web is now widely adopted

#15
Ehm... The semantic web as an idea was/is a totally different thing: the idea is the old libraries of Babel/Bibliotheca Universalis by Conrad Gessner (~1545) [1] or the ability to "narrow"|"select"|"find" just "the small bit of information I want". Observing that a book it's excellent to develop and share a specific topic, it have some indexes to help directly find specific information but that's not enough, a library of books can't be traversed quick enough to find a very specific bit of information like when John Smith was born and where.

The semantic web original idea was the interconnection of every bit of information in a format a machine can travel for a human, so the human can find any specific bit ever written with little to no effort without having to humanly scan pages of moderately related stuff.

We never achieve such goal. Some have tried to be more on the machine side, like WikiData, some have pushed to the extreme the library science SGML idea of universal classification not ended to JSON but all are failures because they are not universal nor easy to "select and assemble specific bit of information" on human queries.

LLMs are a, failed, tentative of achieve such result from another way, their hallucinations and slow formation of a model prove their substantial failure, they SEEMS to succeed for a distracted eye perceiving just the wow effect, but they practically fails.

Aside the issue with ALL test done on the metadata side of the spectrum so far is simple: in theory we can all be good citizens and carefully label anything, even classify following Dublin Core at al any single page, in practice very few do so, all the rest do not care, or ignoring the classification at all or badly implemented it, and as a result is like an archive with some missing documents, you'll always have holes in information breaking the credibility/practical usefulness of the tool.

Essentially that's why we keep using search engines every day, with classic keyword based matches and some extras around. Words are the common denominator for textual information and the larger slice of our information is textual.

[1] https://en.wikipedia.org/wiki/Bibliotheca_universalis

Re: The semantic web is now widely adopted

#16

Earlier quoted context omitted.

LLMs have no soul, so I like content and curation from real people

The main problem is that the incentive for well-intentioned people to add detailed and accurate metadata is much lower than the incentive for SEO dudes to abuse the system if the metadata is used for anything of consequence. There's a reason why search engines that trusted website metadata went extinct. That's the whole benefit of using LLMs for categorization: they work for you, not for the SEO guy... well, prompt i…

There is value-add if you can prove whatever content you are producing is from an authentic human, because I dislike LLM produced garbage

Re: The semantic web is now widely adopted

#17
post #9

> Before JSON-LD there was a nest of other, more XMLy, standards emitted by the various web steering groups. These actually have very, very deep support in many places (for example in library and archival systems) but on the open web they are not a goer. If archival systems and library's are using XML, wouldn't it be preferable to follow their lead and whatever standards they are using? Since they are the ones who ar…

The format really isn’t much of an issue. From an information point of view, the content of the different formats are identical, and translation among them is straightforward.

Promoting JSON-LD potentially makes it more palatable to the modern web creators, perhaps increasing adoption. The bots have already adapted.

Re: The semantic web is now widely adopted

#18
post #8

The author gives two reasons why AI won't replace the need for metadata: 1: LLMs "routinely get stuff wrong" 2: "pricy GPU time" 1: I make a lot of tests on how well LLMs get categorization and data extraction right or wrong for my Product Chart ( https://www.productchart.com ) project. And they get pretty hard stuff right 99% of the time already. This will only improve. 2: Loading the frontpage of Reddit takes hundr…

How does Product Chart use LLMs?

Re: The semantic web is now widely adopted

#19
post #8

The author gives two reasons why AI won't replace the need for metadata: 1: LLMs "routinely get stuff wrong" 2: "pricy GPU time" 1: I make a lot of tests on how well LLMs get categorization and data extraction right or wrong for my Product Chart ( https://www.productchart.com ) project. And they get pretty hard stuff right 99% of the time already. This will only improve. 2: Loading the frontpage of Reddit takes hundr…

For my part, I stopped reading at the free bashing of blockchain•.

Reminded me of the angst and negativity of these original "Web3" people, already bashing everything that was not in their mood back then.

• The crypto ecosystem is shady, I know, but the tech is great

Re: The semantic web is now widely adopted

#20
So much jumping to defend llms as the future. I'd like to point that llms hallucinate, could be injected, and often lack context which well structured metadata can provide. At least, I don't want for an llm to hollucinate the author's picture and bio based on hints in the article, thank you very much.

I don't think that one is necessarily better than the other, but imagining that llms are a silver bullet when another trending story in the front pages is about prompt injection used against the slack ai bots sounds a bit over optimistic.

Post reply on HN