The author gives two reasons why AI won't replace the need for metadata: 1: LLMs "routinely get stuff wrong" 2: "pricy GPU time" 1: I make a lot of tests on how well LLMs get categorization and data extraction right or wrong for my Product Chart ( https://www.productchart.com ) project. And they get pretty hard stuff right 99% of the time already. This will only improve. 2: Loading the frontpage of Reddit takes hundr…
GPU compute price is dropping fast and will continue to do so.
The semantic web is now widely adopted
51–60 of 269 posts
Re: The semantic web is now widely adopted
#52In all honesty, llms are probably going to make all this entirely redundant. As such semantic web was not a natural follower to what we had before, and not web 3.0.
It's HN, most people don't read the article and jump into whatever conclusion they have at the moment despite not being an expert in the field.
What makes you think that I am not an expert btw?
It indeed seems like you appear to believ that what's written on the internet is true. So if someone writes that LLMs are not a contester to semantic web - then it might be true.
Could it be, that I merely challenge that author of the blog article and don't take his predictions for granted?
Re: The semantic web is now widely adopted
#53Re: The semantic web is now widely adopted
#54> Before JSON-LD there was a nest of other, more XMLy, standards emitted by the various web steering groups. These actually have very, very deep support in many places (for example in library and archival systems) but on the open web they are not a goer. If archival systems and library's are using XML, wouldn't it be preferable to follow their lead and whatever standards they are using? Since they are the ones who ar…
The format really isn’t much of an issue. From an information point of view, the content of the different formats are identical, and translation among them is straightforward. Promoting JSON-LD potentially makes it more palatable to the modern web creators, perhaps increasing adoption. The bots have already adapted.
As far as I can tell archivists don't care about "modern web creators", and they likely shouldn't, since archiving is a long term project. I know I don't, and I'm only building software for digital archiving.
Re: The semantic web is now widely adopted
#55Earlier quoted context omitted.
GPU compute price is dropping fast and will continue to do so.
But is it dropping faster than the needs of the next model that needs to be trained?
Also, GPU pricing is hardly relevant. From now on we will see dedicated co-processors on the GPU to handle these things.
They will keep on keeping up with the demand until we meet actual physical limits.
Re: The semantic web is now widely adopted
#56Re: The semantic web is now widely adopted
#57The author gives two reasons why AI won't replace the need for metadata: 1: LLMs "routinely get stuff wrong" 2: "pricy GPU time" 1: I make a lot of tests on how well LLMs get categorization and data extraction right or wrong for my Product Chart ( https://www.productchart.com ) project. And they get pretty hard stuff right 99% of the time already. This will only improve. 2: Loading the frontpage of Reddit takes hundr…
Re: The semantic web is now widely adopted
#58> Before JSON-LD there was a nest of other, more XMLy, standards emitted by the various web steering groups. These actually have very, very deep support in many places (for example in library and archival systems) but on the open web they are not a goer. If archival systems and library's are using XML, wouldn't it be preferable to follow their lead and whatever standards they are using? Since they are the ones who ar…
Re: The semantic web is now widely adopted
#59The argument about LLMs is wrong, not because of reasons stated but because semantic meaning shouldn't solely be defined by the publisher. The real question is whether the average publisher is better than an LLM at accurately classifying their content. My guess is, when it comes to categorization and summarization, an LLM is going to handily win. An easy test is: are publishers experts on topics they talk about? The…
LLMs are not experts either. Furthermore, from what I gather, LLMs are trained on:
>The entire world of SEO hacks, blogspam, etc
Re: The semantic web is now widely adopted
#60The author gives two reasons why AI won't replace the need for metadata: 1: LLMs "routinely get stuff wrong" 2: "pricy GPU time" 1: I make a lot of tests on how well LLMs get categorization and data extraction right or wrong for my Product Chart ( https://www.productchart.com ) project. And they get pretty hard stuff right 99% of the time already. This will only improve. 2: Loading the frontpage of Reddit takes hundr…
Let's hope you never write articles about court cases then: https://www.heise.de/en/news/Copilot-turns-a-court-reporter-... The alleged low error rate of 1% can ruin your day/life/company, if it hits the wrong person, regards the wrong problem, etc. And that risk is not adequately addressed by hand-waving and pointing people to low error rates. In fact, if anything such claims would make me less confident in your pro…
(See also why automated face recognition in public surveillance cameras might be a bad idea.)