Earlier quoted context omitted.
It comes from "the knowledge of being," and has been used to describe real-world knowledge representation, in particular hierarchical(-ish) semantic networks in AI since its early days.
When I see Palantir talk about it in a press release is that something real or just fluffy marketing?
CauseNet: Towards a causality graph extracted from the web
121–130 of 131 posts
Re: CauseNet: Towards a causality graph extracted from the web
#122This makes little sense to me. Ontologies and all that have been tried and have always been found to be too brittle. Take the examples from the front page (which I expect to be among the best in their set): human_activity => climate_change. Those are such a broad concepts that it's practically useless. Or disease => death. There's no nuance at all. There isn't even a definition of what "disease" is, let alone a way t…
I'm actively working with ontologies (disclaimer: as a researcher), and yours is the top comment, so I'll try to make some counterclaims here. No relation to this work tho. > Ontologies and all that have been tried and have always been found to be too brittle. I'd invite you to look at ontologies as nothing more than representations of things we know in some text-based format. If you've ever written an if statement,…
That's because we know how to interpret the concepts used in these representations, in relation to each other. It's just a syntactic change.
You might have a point if it's used as a kind of search engine: "show me wikipedia articles where X causes Y?" (although there is at least one source besides wikipedia, but you get my drift).
> Aside from my point above - haven't looked at the source data, but I doubt it stops at that level.
It does. It isn't even a triple, it's a pair: (cause, effect). There's no other relation than "causes". And if I skimmed the article correctly, they just take noun phrases and slap an underscore between the words and call it a concept. There's no meaning attached to the labels.
But the higher-order causations you mention are going to be pretty useless if there's no way on how to interpret them. It'll only work for highly specialized, unambiguous concepts, like myxomatosis (which is akin to encoding knowledge in the labels themselves), and the broad nature of many of the concepts will lead to quickly decaying usefulness when the length of the path increases. Here are some random examples (length 4 and 8, no posterior selection) from their "precision" set (197k pairs):
['mistake', 'deaths', 'riots', 'violence']
['higher_operating_income', 'increase_in_operating_income', 'increase_in_net_income', 'increase']
['mail_delivery', 'delays', 'decline_in_revenue', 'decrease']
['wastewater', 'environmental_problems', 'problems', 'treatment']
['sensor', 'alarm', 'alarm', 'alarm']
['thatch', 'problems', 'cost_overruns', 'project_delays']
['smoking_pot', 'lung_cancer', 'shortness_of_breath', 'conditions']
['older_medications', 'side_effects', 'physical_damage', 'loss']
['less_fat', 'weight_loss', 'death', 'uncertainties']
['diesel_particles', 'cancer', 'damages', 'injuries']
['malfunction_in_the_heating_unit', 'fire', 'fire_damage', 'claims']
['drug-resistant_malaria', 'deaths', 'violence', 'extreme_poverty']
['fairness_in_circumstances', 'stress', 'backache', 'aching_muscles']
['curved_spine', 'back_pain', 'difficulties', 'stress', 'difficulties', 'delay', 'problem', 'serious_complications']
['obama', 'high_gas_prices', 'recession', 'hardship', 'happiness', 'success', 'promotions', 'bonuses']
['financial_devastation', 'bankruptcy', 'stigma', 'homelessness', 'health_problems', 'deaths', 'pain', 'quality_of_life']
['methylmercury', 'neurological_damage', 'seizures', 'changes', 'crisis', 'growth', 'problems', 'birth_defects']
The latter is probably correct, but the chain of reasoning is false...This one is cherry-picked, but I found it to funny to omit:
['agnosticism', 'despair', 'feelings', 'aggression', 'action', 'riot', 'arrest', 'embarrassment', 'problems', 'black_holes']Re: CauseNet: Towards a causality graph extracted from the web
#123This makes little sense to me. Ontologies and all that have been tried and have always been found to be too brittle. Take the examples from the front page (which I expect to be among the best in their set): human_activity => climate_change. Those are such a broad concepts that it's practically useless. Or disease => death. There's no nuance at all. There isn't even a definition of what "disease" is, let alone a way t…
Re: CauseNet: Towards a causality graph extracted from the web
#124I read it as "casual" rather than "causal", got very dissapointed while reading the article! An inventory of casual knowledge would be really fun, although it's hard to think what it would consist of now that I think about it... There is this concept of "hidden knowledge" about all the things you know at work that no one really thinks about is knowledge so it's hard to let newcomers know about it. But that does sound…
Re: CauseNet: Towards a causality graph extracted from the web
#125Earlier quoted context omitted.
Ontology, not ontologies, have been tried. We have quite a good understanding that a system cannot be both sound a complete, regardless people went straight in to make a single model of the world.
> a system cannot be both sound a complete Huh, what do you mean by this? There are many sound and complete systems – propositional logic, first-order logic, Presburger arithmetic, the list goes on. These are the basic properties you want from a logical or typing system. (Though, of course, you may compromise if you have other priorities.)
first-order logic is sound, but not complete (Ie. I can express a set of strings you can not recognize in first-order logic).
Re: CauseNet: Towards a causality graph extracted from the web
#126The sample set contains: { "causal_relation": { "cause": { "concept": "boom" }, "effect": { "concept": "bust" } } } It's practically a hedge-fund-in-a-box.
Plus, regardless of what you might think of how valid that connection is, what they're actually collecting, absent any kind of mechanism, is a set of all apparent correlations...
Re: CauseNet: Towards a causality graph extracted from the web
#127Earlier quoted context omitted.
Democritus (b 460BCE) said, “I would rather discover one cause than gain the kingdom of Persia,” which suggests that finding true causes is rather difficult.
perhaps in a similar way that it is impossible to directly "observe" a wavefunction without collapsing it into an observale "effect".
Re: CauseNet: Towards a causality graph extracted from the web
#128Earlier quoted context omitted.
Democritus (b 460BCE) said, “I would rather discover one cause than gain the kingdom of Persia,” which suggests that finding true causes is rather difficult.
perhaps in a similar way that it is impossible to directly "observe" a wavefunction without collapsing it into an observale "effect".
Re: CauseNet: Towards a causality graph extracted from the web
#129Earlier quoted context omitted.
I'm actively working with ontologies (disclaimer: as a researcher), and yours is the top comment, so I'll try to make some counterclaims here. No relation to this work tho. > Ontologies and all that have been tried and have always been found to be too brittle. I'd invite you to look at ontologies as nothing more than representations of things we know in some text-based format. If you've ever written an if statement,…
I’m not sure trying to tease out high-integrity information from Wikipedia is a useful contribution at all. Our criteria of proof is whatever a private clique of wiki editors or worse their security-complex handlers say? I feel like LLMs have already achieved this and the results are about what you would expect.
But that's from the research POV. If your POV is immediately practical, then yes, it's about as useful as training a new employee using Wikipedia articles about your company.
Side-note: LLMs have actually proven quite useful given that they were trained on mostly what was available online.
Re: CauseNet: Towards a causality graph extracted from the web
#130Earlier quoted context omitted.
I'm actively working with ontologies (disclaimer: as a researcher), and yours is the top comment, so I'll try to make some counterclaims here. No relation to this work tho. > Ontologies and all that have been tried and have always been found to be too brittle. I'd invite you to look at ontologies as nothing more than representations of things we know in some text-based format. If you've ever written an if statement,…
> I'd invite you to look at ontologies as nothing more than representations of things we know in some text-based format. That's because we know how to interpret the concepts used in these representations, in relation to each other. It's just a syntactic change. You might have a point if it's used as a kind of search engine: "show me wikipedia articles where X causes Y?" (although there is at least one source besides…
> There's no other relation than "causes".
Looking at their Neo4j graph, they also retain the provenance of the causal relation in "claimedIn" relations between the reified triple of each cause-effect pair. So, that's at least marginally useful for fact-checking or quality evals. > they just take noun phrases and slap an underscore between the words and call it a concept.
Not to defend lazy approaches, but you could make this point about tokens also ("take any bunch of characters that happens often enough, and call it a token"). > It'll only work for highly specialized, unambiguous concepts
Fair point. Practically, there's not much use in this unless you really dedicate time to figure out what's meant by each concept, and prune junk. And by that time, you may as well make your own pairs.But from a research POV, hardly anyone will go through such effort, so I still find it quite useful. Some potential questions that I could derive from this (aside from ~110 works citing it since 2020, which is not bad for KRR work):
- What is the quality of causal relations (e.g., diagnostic decision trees) in medical articles?
- How can the original scientific provenance of cause-effect pairs best be represented?
- Which extracted causes are true variables in some effect, and to what extent/direction? (i.e., an ablation study at scale, provided you can find appropriate data)