Live data from Hacker News

The semantic web is now widely adopted

csvbase.com

21–30 of 269 posts

Re: The semantic web is now widely adopted

#21
post #4

In all honesty, llms are probably going to make all this entirely redundant. As such semantic web was not a natural follower to what we had before, and not web 3.0.

The article addresses this point with the following: > It would of course be possible to sic Chatty-Jeeps on the raw markup and have it extract all of this stuff automatically. But there are some good reasons why not. > > The first is that large language models (LLMs) routinely get stuff wrong. If you want bots to get it right, provide the metadata to ensure that they do. > > The second is that requiring an LLM to re…

The first point is moot, because human annotation would also have some amount of error, either through mistakes (interns being paid nothing to add it) or maliciously (SEO). Plus, human annotation would be multi-lingual, which leads to a host of other problems that LLMs don't have to the same extent.

The second point is silly, because there is no reason for everyone to train their own LLMs on the raw web. You'd have a few companies or projects that handle the LLM training, and everyone else uses those LLMs.

I'm not a big fan of LLMs, and not even a big believer in their future, but I still think they have a much better chance of being useful for these types of tasks than the semantic web. Semantic web is a dead idea, people should really allow it to rest.

Re: The semantic web is now widely adopted

#22
post #8

The author gives two reasons why AI won't replace the need for metadata: 1: LLMs "routinely get stuff wrong" 2: "pricy GPU time" 1: I make a lot of tests on how well LLMs get categorization and data extraction right or wrong for my Product Chart ( https://www.productchart.com ) project. And they get pretty hard stuff right 99% of the time already. This will only improve. 2: Loading the frontpage of Reddit takes hundr…

> Reddit takes hundreds of http requests, parses megabytes of text, image and JavaScript code [...] to show some links to articles

Yes, and I hate it. I closed Reddit many times because the wait time wasn't worth it.

Re: The semantic web is now widely adopted

#23
post #8

The author gives two reasons why AI won't replace the need for metadata: 1: LLMs "routinely get stuff wrong" 2: "pricy GPU time" 1: I make a lot of tests on how well LLMs get categorization and data extraction right or wrong for my Product Chart ( https://www.productchart.com ) project. And they get pretty hard stuff right 99% of the time already. This will only improve. 2: Loading the frontpage of Reddit takes hundr…

For my part, I stopped reading at the free bashing of blockchain•. Reminded me of the angst and negativity of these original "Web3" people, already bashing everything that was not in their mood back then. • The crypto ecosystem is shady, I know, but the tech is great

As someone who stopped getting involved in blockchain "tech" 12 years ago because of the prevalence of scams and bad actors and lack of interesting tech beyond the merkle tree, what's great about it?

FWIW I am genuinely asking. I don't know anything about the current tech. There's something about "zero knowledge proofs" but I don't understand how much of that is used in practice for real blockchain things vs just being research.

As far as I know, the throughput of blockchain transactions at scale is miserably slow and expensive and their usual solution is some kind of side channel that skips the full validation.

Distributed computation on the blockchain isn't really used for anything other than converting between currencies and minting new ones mostly AFAIK as well.

What is the great tech that we got from the blockchain revolution?

Re: The semantic web is now widely adopted

#24
post #8

The author gives two reasons why AI won't replace the need for metadata: 1: LLMs "routinely get stuff wrong" 2: "pricy GPU time" 1: I make a lot of tests on how well LLMs get categorization and data extraction right or wrong for my Product Chart ( https://www.productchart.com ) project. And they get pretty hard stuff right 99% of the time already. This will only improve. 2: Loading the frontpage of Reddit takes hundr…

LLMs have no soul, so I like content and curation from real people

All the web metadata I consume is organic and responsively farmed.

Re: The semantic web is now widely adopted

#25
post #18
post #8

The author gives two reasons why AI won't replace the need for metadata: 1: LLMs "routinely get stuff wrong" 2: "pricy GPU time" 1: I make a lot of tests on how well LLMs get categorization and data extraction right or wrong for my Product Chart ( https://www.productchart.com ) project. And they get pretty hard stuff right 99% of the time already. This will only improve. 2: Loading the frontpage of Reddit takes hundr…

How does Product Chart use LLMs?

We research all product data manually and then have AI cross-check the data and see how well it can replicate what the human has researched and whether it can find errors.

Actually, building the AI agent for data research takes up most of my time these days.

Re: The semantic web is now widely adopted

#26

In all honesty, llms are probably going to make all this entirely redundant. As such semantic web was not a natural follower to what we had before, and not web 3.0.

Have you read the article? It addresses this point towards the end.

And it fails to address why SemWeb failed in its heyday: that there's no business case for releasing open data of any kind "on the web" (unless you're wikidata or otherwise financed via public money) the only consequence being that 1. you get less clicks 2. you make it easier for your competitors (including Google) to aggregate your data. And that hasn't changed with LLMs, quite the opposite.

To think a turd such as JSON-LD can save the "SemWeb" (which doesn't really exist), and even add CSV as yet another RDF format to appease "JSON scientists" lol seems beyond absurd. Also, Facebook's Open Graph annotations in HTML meta-links are/were probably the most widespread (trivial) implementation of SemWeb. SemWeb isn't terrible but is entirely driven by TBL's long-standing enthusiasm for edge-labelled graph-like databases (predating even his WWW efforts eg [1]), plus academia's need for topics to produce papers on. It's a good thing to let it go in the last decade and re-focus on other/classic logic apps such as Prolog and SAT solvers.

[1]: https://en.wikipedia.org/wiki/ENQUIRE

Re: The semantic web is now widely adopted

#27
If even the semantic web people are declaring victory based on a post title and a picture for better integration with Facebook, then it's clear that Semantic Web as it was envisioned is fully 100% dead and buried.

The concept of OWL and the other standards was to annotate the content of pages, that's where the real values lie. Each paragraph the author wrote should have had some metadata about its topic. At the very least, the article metadata was supposed to have included information about the categories of information included in the article.

Having a bit of info on the author, title (redundant, as HTML already has a tag for that), picture, and publication date is almost completely irrelevant for the kinds of things Web 3.0 was supposed to be.

Re: The semantic web is now widely adopted

#28
post #22
post #8

The author gives two reasons why AI won't replace the need for metadata: 1: LLMs "routinely get stuff wrong" 2: "pricy GPU time" 1: I make a lot of tests on how well LLMs get categorization and data extraction right or wrong for my Product Chart ( https://www.productchart.com ) project. And they get pretty hard stuff right 99% of the time already. This will only improve. 2: Loading the frontpage of Reddit takes hundr…

> Reddit takes hundreds of http requests, parses megabytes of text, image and JavaScript code [...] to show some links to articles Yes, and I hate it. I closed Reddit many times because the wait time wasn't worth it.

https://old.reddit.com ?

Re: The semantic web is now widely adopted

#29

So much jumping to defend llms as the future. I'd like to point that llms hallucinate, could be injected, and often lack context which well structured metadata can provide. At least, I don't want for an llm to hollucinate the author's picture and bio based on hints in the article, thank you very much. I don't think that one is necessarily better than the other, but imagining that llms are a silver bullet when another…

Sure but do hallucinations matter then much just for categorisation? Hardly the end of the world if they make up a published date occasionally.

And prompt injection is irrelevant because the alternative we're considering is letting publishers directly choose the metadata.

Re: The semantic web is now widely adopted

#30
Are there any tools that employ LLMs to fill out the Semantic Web data? I can see that being a high-impact use case: people don’t generally like manually filling out all the fields in a schema (it is indeed “a bother”), but an LLM could fill it out for you – and then you could tweak for correctness / editorializing. Voila, bother reduced!

This would also address the two reasons why the author thinks AI is not suited to this task:

1. human stays in the loop by (ideally) checking the JSON-LD before publishing; so fewer hallucination errors

2. LLM compute is limited to one time per published content and it’s done by the publisher. The bots can continue to be low-GPU crawlers just as they are now, since they can traverse the neat and tidy JSON-LD.

——————

The author makes a good case for The Semantic Web and I’ll be keeping it in mind for the next time I publish something, and in general this will add some nice color to how I think about the web.

Post reply on HN