Live data from Hacker News

The semantic web is now widely adopted

csvbase.com

51–60 of 269 posts

Re: The semantic web is now widely adopted

#51
post #12
post #8

The author gives two reasons why AI won't replace the need for metadata: 1: LLMs "routinely get stuff wrong" 2: "pricy GPU time" 1: I make a lot of tests on how well LLMs get categorization and data extraction right or wrong for my Product Chart ( https://www.productchart.com ) project. And they get pretty hard stuff right 99% of the time already. This will only improve. 2: Loading the frontpage of Reddit takes hundr…

GPU compute price is dropping fast and will continue to do so.

But is it dropping faster than the needs of the next model that needs to be trained?

Re: The semantic web is now widely adopted

#52

In all honesty, llms are probably going to make all this entirely redundant. As such semantic web was not a natural follower to what we had before, and not web 3.0.

It's HN, most people don't read the article and jump into whatever conclusion they have at the moment despite not being an expert in the field.

As I already pointed out, none of the arguments the author brings up are really relevant. Resources and accuracy will not be a concern in 5 years.

What makes you think that I am not an expert btw?

It indeed seems like you appear to believ that what's written on the internet is true. So if someone writes that LLMs are not a contester to semantic web - then it might be true.

Could it be, that I merely challenge that author of the blog article and don't take his predictions for granted?

Re: The semantic web is now widely adopted

#53

In all honesty, llms are probably going to make all this entirely redundant. As such semantic web was not a natural follower to what we had before, and not web 3.0.

Have you read the article? It addresses this point towards the end.

yes

Re: The semantic web is now widely adopted

#54
post #9

> Before JSON-LD there was a nest of other, more XMLy, standards emitted by the various web steering groups. These actually have very, very deep support in many places (for example in library and archival systems) but on the open web they are not a goer. If archival systems and library's are using XML, wouldn't it be preferable to follow their lead and whatever standards they are using? Since they are the ones who ar…

The format really isn’t much of an issue. From an information point of view, the content of the different formats are identical, and translation among them is straightforward. Promoting JSON-LD potentially makes it more palatable to the modern web creators, perhaps increasing adoption. The bots have already adapted.

You're aware of straightforward translations to and from E-ARK SIP and CSIP? Between what formats?

As far as I can tell archivists don't care about "modern web creators", and they likely shouldn't, since archiving is a long term project. I know I don't, and I'm only building software for digital archiving.

Re: The semantic web is now widely adopted

#55
post #12

Earlier quoted context omitted.

GPU compute price is dropping fast and will continue to do so.

But is it dropping faster than the needs of the next model that needs to be trained?

Short answer is yes.

Also, GPU pricing is hardly relevant. From now on we will see dedicated co-processors on the GPU to handle these things.

They will keep on keeping up with the demand until we meet actual physical limits.

Re: The semantic web is now widely adopted

#57
post #8

The author gives two reasons why AI won't replace the need for metadata: 1: LLMs "routinely get stuff wrong" 2: "pricy GPU time" 1: I make a lot of tests on how well LLMs get categorization and data extraction right or wrong for my Product Chart ( https://www.productchart.com ) project. And they get pretty hard stuff right 99% of the time already. This will only improve. 2: Loading the frontpage of Reddit takes hundr…

Only slightly tongue in cheek, but if your measure of success is Reddit, perhaps a better example may serve your argument?

Re: The semantic web is now widely adopted

#58
post #9

> Before JSON-LD there was a nest of other, more XMLy, standards emitted by the various web steering groups. These actually have very, very deep support in many places (for example in library and archival systems) but on the open web they are not a goer. If archival systems and library's are using XML, wouldn't it be preferable to follow their lead and whatever standards they are using? Since they are the ones who ar…

If by that the author means JSON-LD has replaced MarcXML, BibTex records, and other bibliographic information systems, then that's very much not the case.

Re: The semantic web is now widely adopted

#59
post #11

The argument about LLMs is wrong, not because of reasons stated but because semantic meaning shouldn't solely be defined by the publisher. The real question is whether the average publisher is better than an LLM at accurately classifying their content. My guess is, when it comes to categorization and summarization, an LLM is going to handily win. An easy test is: are publishers experts on topics they talk about? The…

>My guess is, when it comes to categorization and summarization, an LLM is going to handily win. An easy test is: are publishers experts on topics they talk about? The truth of the internet is no, they're not usually.

LLMs are not experts either. Furthermore, from what I gather, LLMs are trained on:

>The entire world of SEO hacks, blogspam, etc

Re: The semantic web is now widely adopted

#60
post #46
post #8

The author gives two reasons why AI won't replace the need for metadata: 1: LLMs "routinely get stuff wrong" 2: "pricy GPU time" 1: I make a lot of tests on how well LLMs get categorization and data extraction right or wrong for my Product Chart ( https://www.productchart.com ) project. And they get pretty hard stuff right 99% of the time already. This will only improve. 2: Loading the frontpage of Reddit takes hundr…

Let's hope you never write articles about court cases then: https://www.heise.de/en/news/Copilot-turns-a-court-reporter-... The alleged low error rate of 1% can ruin your day/life/company, if it hits the wrong person, regards the wrong problem, etc. And that risk is not adequately addressed by hand-waving and pointing people to low error rates. In fact, if anything such claims would make me less confident in your pro…

This is the thing with errors and automation. A 1 % error rate in a human process is basically fine. A 1 % error rate in an automated process is hundreds of thousands of errors per day.

(See also why automated face recognition in public surveillance cameras might be a bad idea.)

Post reply on HN