The whole field has been dominated by research, i.e. the wish to make simple things complicated (in order to publish papers) as opposed to engineering, i.e. making complicated things simple (in order to produce usable software efficiently). As a result the standards are horrendously - and needlessly - complicated. The few major practical outcomes like the schema.org, json-ld and the google annotation system, are resu…
I agree, the research is overly complicated. So it's a lot of extra work to sift through, but I've found a lot of gold in there. If you're looking for a simple, noise-free way to do the semantic web, I'm very confident that Tree Notation will enable it ( https://treenotation.org/ ). I've played around a bit with turning Schema.org into a Tree Language, and think that would be a fruitful exercise, but plenty more on t…
A Review of the Semantic Web Field
61–70 of 72 posts
Re: A Review of the Semantic Web Field
#62Maybe for its time it seemed like a good idea.. Like SOAP or manual features for image classification. Today, it's clear that languages and knowledge don't really work like that, and it's not practical to approach them this way. I've learned about the OWL and SPARQL 12 years ago, and it already felt like a very dated idea. But then who knows... Everybody have given up on NNs once too.
Re: A Review of the Semantic Web Field
#63Maybe for its time it seemed like a good idea.. Like SOAP or manual features for image classification. Today, it's clear that languages and knowledge don't really work like that, and it's not practical to approach them this way. I've learned about the OWL and SPARQL 12 years ago, and it already felt like a very dated idea. But then who knows... Everybody have given up on NNs once too.
The comparisons to NLP presents a good view on the problems. Its "easy" to write some logic rules to parse input text for a 50% demo. But then you want to improve & scale, and suddenly all the nuances, bites you. The rules get bigger, nested and complicated. Traditional NLP tried that avenue for a while, with decent success in small usecases, but for larger problems without success. (Compared to stuff like BERT & GPT…
Re: A Review of the Semantic Web Field
#64Earlier quoted context omitted.
I agree, the research is overly complicated. So it's a lot of extra work to sift through, but I've found a lot of gold in there. If you're looking for a simple, noise-free way to do the semantic web, I'm very confident that Tree Notation will enable it ( https://treenotation.org/ ). I've played around a bit with turning Schema.org into a Tree Language, and think that would be a fruitful exercise, but plenty more on t…
It is unclear to me what it would achieve compared to a spog (subject, predicate, object, graph) based representation like it exists in RDF based triplestores.
My take with ontologies is building consensus is hard.
Tree Notation offers a solution to the problem of: what should we agree on for the encoding? I assume that simpler is better, all else being equal. Then Tree Notation is the simplest, in terms of the thing with the fewest pieces(tokens).
To get to Tree Notation, nothing was added, only stripped. I started with an existing notation and stripped away each visible syntax token that wasn't needed. Surprisingly, not one is needed. Not one quote, parens, bracket, colon, etc.
So now if we can get consensus around going with the simplest thing, we have got a way to agree on an whether we should use XML, JSON-LD, turtle, etc. The simplest thing works (which would be Tree Notation, or a close relative—someone can rebrand the notation but the idea is largely the same). This does not suffer from the 927 problem, as there are a few classes of things where we do have 1 new language that is mathematically superior and of a different kind than others (binary notation, for example).
So after you have agreement on that encoding, versioning and forking and merging schemas is dead simple (just use Git—in Tree Notation all changes are semantic and noise free).
So now we've solved what encoding to use for our ontologies, and we have a very fast and efficient way to collaborate on them (it's just plain text and git).
That brings us to a third advantage which is more theoretical. Tree Notation maps words/nodes to a 3-D representation. This means that there would be an X-Y-Z isomorphism with an ontology and the real world. I don't really know where we go from there, but at least by this point we've moved the semantic web idea a lot further and can start looking at the next realm of possibilities.
Re: A Review of the Semantic Web Field
#65Earlier quoted context omitted.
I was "stuck" working with a bunch of leading academics and researchers on a SemWeb project using OWL/RDF, in collaboration with DARPA and the US Department of Defense, around 2008-2009. You are absolutely correct that they are hostile to anything outside of their "ideology". The awful, horrific performance of the RDF/OWL databases compared to the impure, heretical evil Neo4j that they despised for its practicality..…
Are you referring to the BFO crowd? I would be interested in your thoughts about BFO itself if thats the case.
My thoughts on ontologies in general are that they can certainly be powerful, and I've seen them used in the past in rules engines that powered fraud detection applications.
In the SemWeb community, in the late 2000s, they successfully convinced a bunch of CIOs of massive organizations, especially in the US Federal Government, that a key to being able to centralize and federate all of their data, and save money on duplicative systems, they could simply have semantic mappings on top of every IT system's databases, and query this semantic layer. Ideally, they could eliminate duplicated data, so that all systems would get data from the "Authoritative Data Source" system instead of duplicating it locally in the application's database.
I'm sure you can immediately see why this is wildly stupid and unrealistic. Imagine what it would look like if every single piece of data that I can technically obtain from another source has to remain in that source, and that storing that data locally with my application specific data is forbidden..... Suddenly, there is a massive increase in I/O, drop in performance, etc.
The whole project taught me a lesson about the politics of academia, and how there is a segment of the population that is highly educated, and has learned how to manufacture work for themselves outside of academia by pushing for high-level government officials to implement programs based on their theories..... MITRE was a big part of this particular project.
Re: A Review of the Semantic Web Field
#66The reason why tools like Protégé have not been sufficiently developed is because of infighting in the academic ontology community in addition to the reasons listed by the author. It has set the whole community back at least 5 years.
Re: A Review of the Semantic Web Field
#67Earlier quoted context omitted.
Many of us who have been in these battles over the decades have decided that the interchange format is almost irrelevant to the real challenge which is the modeling and semantic alignment. It's a useless parlour trick to merge graphs and call them integrated, approximately as it is to put several CSV files into an archive or simply loading unrelated tables into one RDBMS. Yes, you can run a processing engine on the a…
Relational model, XML/JSON etc. simply do not have a generic merge operation defined the same way as RDF does. This can be proved with pen and paper. And you still haven't addressed my second point about widespread industry use. It seems that SemWeb haters/sceptics always try to avoid this, why could that be?..
Who cares? This is not a problem anyone has, which is precisely why so few formats have a solution.
"widespread industry use"
It's not in "widespread" use. It's in niche use, and it's been in niche use for about two decades, and shows no sign of escaping that niche.
Human perception is a bit broken here. You show a list of 100 users and it looks like a tech is in "widespread use"... because you don't intuit that the market has hundreds of thousands of users, if not millions. (I'm being conservative. It's almost certainly millions.) RDF is niche. You can comfortably read an effectively-complete list of users over a coffee break. Try that trick with JSON.
Also, to be honest, referring to "haters" rather proves my point about just how quickly insults get trotted out. You almost literally just said "RDF!" with no further substantive conversation exactly the way I mentioned! I know about RDF. I used it ~2005 when working on some Mozilla stuff. It had every opportunity to overtake JSON, and was never in any danger of it.
In fact my current job for the last few weeks has been working on a massively cross-team data lake in the company I work for... and nobody is talking about RDF. Not me (and I do know it, actually), not any vendor that might provide useful technology, not any vendor that consumes data to provide reports on it (nobody consumes RDF in this space), nobody. Nominally a core use case for "semanticness", and it's a complete non-starter.
Re: A Review of the Semantic Web Field
#68Earlier quoted context omitted.
Relational model, XML/JSON etc. simply do not have a generic merge operation defined the same way as RDF does. This can be proved with pen and paper. And you still haven't addressed my second point about widespread industry use. It seems that SemWeb haters/sceptics always try to avoid this, why could that be?..
"simply do not have a generic merge operation defined the same way as RDF does." Who cares? This is not a problem anyone has, which is precisely why so few formats have a solution. "widespread industry use" It's not in "widespread" use. It's in niche use, and it's been in niche use for about two decades, and shows no sign of escaping that niche. Human perception is a bit broken here. You show a list of 100 users and…
I'm not talking about replacing JSON with RDF. Don't need data interchange -- don't use RDF. RDF is both at a different level of abstraction and solving problems of different scope.
Re: A Review of the Semantic Web Field
#69Earlier quoted context omitted.
Jena is everything but user friendly, it has a lot of weird edge cases, bugs, and a horrible API. RDF4J is okay for RDF, but completely ignores OWL. RDFLib is a bug ridden mess, have you ever used it, or checked their issue tracker, and commit history? With that amount of production breaking bugs that haven't been resolved for years, it might as well be unmaintained. Redland has last been updated years ago. Sure ther…
Research projects? Research is something coming out of academia, as you know. These are open-source projects with active developer communities. Jena has long been under Apache, RDF4J is now under Eclipse Foundation. Can you for once answer why large companies in the industry are using RDF/SPARQL as of today if it's so "dead"? Here's a list: http://sparql.club
I'm not saying that it's dead, worse it's hard to use, bad and unreliable.
Can you answer why large companies are still using Cobol, if it's so "dead"?
Legacy, lack of alternatives, managers that don't have technical expertise but that fall for the marketing.
The semantic web, MASSIVELY overpromises, and MASSIVELY under-delivers. Both truly in a "web-scale" way.
Re: A Review of the Semantic Web Field
#70Earlier quoted context omitted.
"simply do not have a generic merge operation defined the same way as RDF does." Who cares? This is not a problem anyone has, which is precisely why so few formats have a solution. "widespread industry use" It's not in "widespread" use. It's in niche use, and it's been in niche use for about two decades, and shows no sign of escaping that niche. Human perception is a bit broken here. You show a list of 100 users and…
Yes RDF is in its own niche -- data interchange. And that's where merge matters, when you for example need to merge protein data with genes and drugs etc. A bunch of pharma companies are using RDF Knowledge Graphs for that purpose. The need for data interchange comes with a certain company size, and that point RDF becomes the solution because there are no real alternatives. I'm not talking about replacing JSON with R…
Could you perhaps recommend some industry case studies or publications on that specific problem area of biopharmaceuticals?