Live data from Hacker News

The business of extracting knowledge from academic publications

theseedsofscience.pub

51–60 of 85 posts

Re: The business of extracting knowledge from academic publications

#51

Earlier quoted context omitted.

The parent is right, saying a database about proteins 'looks like one narrow niche knowledge base.' in this context is just nonsense.

in which context, and why your bold statements should be trusted blindly?

TL;DR: I explain what proteins are and why they're important, show some biological "flowcharts" , and end up with one "function definition" from the "source code" that makes you you, and has to do with you eating and breathing.

Long:

This is the bit that makes biology awesome to me, so excuse me for the small essay ;-)

Proteins are basically extremely advanced nanomachines which work together in larger systems to ultimately form a cell. Having a listing of all the proteins in a cell (the sum of the parts) is insufficient to grok the whole, but it's pretty darn important. The abilities and limitations determine and constrain what a cell can do, and ultimately influence what organisms and ecosystems are and are not capable of. Which is a big chunk of the science of biology.

Uniprot lists many/all of the genes/proteins that have been decoded so far. It's a bit odd to call that a "niche" in the context of the field of biology.

I'm not going to discourage you though: Proteins and protein systems are pretty darn awesome!

Example:

(using KEGG rather than uniprot, since it's got graphical maps, which is handy to get an intuition)

For instance, if you want to know why you need to eat and why you need to breathe (flowchart for what your cells do with starch and oxygen):

* https://www.genome.jp/pathway/map00500 start by finding starch on this map (gets split into glucose)

* https://www.genome.jp/pathway/map00010 which gets broken down into 2* pyruvate

* https://www.genome.jp/pathway/map00020 which gets processed

* https://www.genome.jp/pathway/map00190 and ultimately "burned" with oxygen.

Each step is a 'chemical reaction catalyzed by proteins'[1] (in the rectangles). You can dig in deeper to find your actual source code: Say we click on a random step (in this case near the top of glycolysis on map 00010)

* https://www.genome.jp/entry/K01810+K06859+K13810+K15916+5.3....

At the bottom you can find the gene listed for Homo Sapiens (HSA)

* https://www.genome.jp/entry/hsa:2821

And this lists the amino-acid (AA) sequence for the protein, and the nucleotide (NT) sequence found in humans. Since this is highly preserved functionality, that's probably (almost) exactly the source code that you have in each of your cells.

KEGG is nice to get an overview of some of the pathways that are fully understood with the maps.

[1] calling it a "chemical reaction" is sort of underselling many proteins. Proteins can have moving parts and can work together. I prefer to think of them as sophisticated nanomachines.

Re: The business of extracting knowledge from academic publications

#52
post #26

Earlier quoted context omitted.

Ingenuity's Pathway Analysis product has been the standard tool big pharma has used for 15+ years. Founded by folks from Stanford, they created their own knowledge-graph database a decade before they were popular, and have been paying PhD's to curate that database with genetic findings from papers over the last decades. Obviously their analysis also includes the publicly-available databases, and they have their own p…

I think the big case study everybody gives is Schrödinger who took 20 years to IPO. The problem is that the market is just not that big. Assume that every single pharma company buys your product - like Schrödinger - where does your revenue top out? You can beat that as a small pharma company with one or two assets that make it to market. So the question is: if the computational platform is so good then why not just b…

“If you’re so smart, then why aren’t you rich?”

Re: The business of extracting knowledge from academic publications

#53

Earlier quoted context omitted.

in which context, and why your bold statements should be trusted blindly?

TL;DR: I explain what proteins are and why they're important, show some biological "flowcharts" , and end up with one "function definition" from the "source code" that makes you you, and has to do with you eating and breathing. Long: This is the bit that makes biology awesome to me, so excuse me for the small essay ;-) Proteins are basically extremely advanced nanomachines which work together in larger systems to ult…

> Uniprot lists many/all of the genes/proteins that have been decoded so far. It's a bit odd to call that a "niche" in the context of the field of biology.

I actually checked uniprot, yes, it lists proteins (probably most of them), but ontology is raither narrow, it has few dozens properties, you can't for example query that DB with question: give me diseases which can be attributed to broken pathways synthesizing protein X, you would need to do a lot of manual work and check external databases of uncertain quality.

Another question is quality of that dataset, why it is so obvious that all those millions of pathways for hundreds thousands proteins are researched and described with 100% accuracy?

Re: The business of extracting knowledge from academic publications

#54

Earlier quoted context omitted.

TL;DR: I explain what proteins are and why they're important, show some biological "flowcharts" , and end up with one "function definition" from the "source code" that makes you you, and has to do with you eating and breathing. Long: This is the bit that makes biology awesome to me, so excuse me for the small essay ;-) Proteins are basically extremely advanced nanomachines which work together in larger systems to ult…

> Uniprot lists many/all of the genes/proteins that have been decoded so far. It's a bit odd to call that a "niche" in the context of the field of biology. I actually checked uniprot, yes, it lists proteins (probably most of them), but ontology is raither narrow, it has few dozens properties, you can't for example query that DB with question: give me diseases which can be attributed to broken pathways synthesizing pr…

> you can't for example query that DB with question: give me diseases which can be attributed to broken pathways synthesizing protein X, you would need to do a lot of manual work and check external databases of uncertain quality.

nod

Uniprot is more useful if you're looking for the actual "bare metal" NN and AA sequences. Which is rather important in its own right, obviously: Sooner or later you DO need the actual sequences if you're going to do something with them in real life.

But uniprot doesn't -itself- give you an understanding of what that code is then doing.

Re: The business of extracting knowledge from academic publications

#55

Earlier quoted context omitted.

Your sources are talking about the ratio of small molecule vs. large molecule drugs. Even if you're developing small molecule drugs you are likely targeting some aspect of protein signaling/gene expression. People are being dismissive of your comments because to say that proteins are niche in the context of pharma is like saying advertising is niche in the context of Meta and Google.

> People are being dismissive of your comments because to say that proteins are niche in the context of pharma is like saying advertising is niche in the context of Meta and Google. its all about how you define word "niche", for google, main revenue stream is supported by several pillars: search tech, infra tech, ads tech, ecosystem+network effect, human management. You remove one pillar, and everything is destroyed,…

I'd say it's less like pillars and more like (emergence) layers, with proteins being a pretty important layer .

I'm now not entirely sure what your experience is with bio sciences. You're definitely coming at it from an odd angle though!

Re: The business of extracting knowledge from academic publications

#56

Earlier quoted context omitted.

> People are being dismissive of your comments because to say that proteins are niche in the context of pharma is like saying advertising is niche in the context of Meta and Google. its all about how you define word "niche", for google, main revenue stream is supported by several pillars: search tech, infra tech, ads tech, ecosystem+network effect, human management. You remove one pillar, and everything is destroyed,…

I'd say it's less like pillars and more like (emergence) layers, with proteins being a pretty important layer . I'm now not entirely sure what your experience is with bio sciences. You're definitely coming at it from an odd angle though!

I didn't claim expertise, that's why I say "it looks like", "I suspect",

Re: The business of extracting knowledge from academic publications

#57

Earlier quoted context omitted.

I'd say it's less like pillars and more like (emergence) layers, with proteins being a pretty important layer . I'm now not entirely sure what your experience is with bio sciences. You're definitely coming at it from an odd angle though!

I didn't claim expertise, that's why I say "it looks like", "I suspect",

Well, maybe find some time and dive in a bit and see what can be found?

You never know, maybe you'll end up contributing to our understanding of life, maybe (indirectly) even save a few lives!

Re: The business of extracting knowledge from academic publications

#58

Earlier quoted context omitted.

I didn't claim expertise, that's why I say "it looks like", "I suspect",

Well, maybe find some time and dive in a bit and see what can be found? You never know, maybe you'll end up contributing to our understanding of life, maybe (indirectly) even save a few lives!

I am working on the service which potentially can answer questions like in this comment: https://news.ycombinator.com/item?id=38109294

life science is one of potential applications if there is an interest and money.

Re: The business of extracting knowledge from academic publications

#59
post #26

Earlier quoted context omitted.

Ingenuity's Pathway Analysis product has been the standard tool big pharma has used for 15+ years. Founded by folks from Stanford, they created their own knowledge-graph database a decade before they were popular, and have been paying PhD's to curate that database with genetic findings from papers over the last decades. Obviously their analysis also includes the publicly-available databases, and they have their own p…

I think the big case study everybody gives is Schrödinger who took 20 years to IPO. The problem is that the market is just not that big. Assume that every single pharma company buys your product - like Schrödinger - where does your revenue top out? You can beat that as a small pharma company with one or two assets that make it to market. So the question is: if the computational platform is so good then why not just b…

For a good computational platform, that can pinpoint unique and valuable candidates , isn't there an option to do license/consult in exchange for royalties ?

Re: The business of extracting knowledge from academic publications

#60
post #20

Earlier quoted context omitted.

The GP was talking about pharma. Proteins are not niche in pharma. Everything else may be niche in that domain, but proteins make up more than 95% of the targets of the pharmaceutical industry.

> The GP was talking about pharma. post is not about pharma specifically, but about general bio-medical literature, including genes, diseases, symptoms, etc. > proteins make up more than 95% of the targets of the pharmaceutical industry. do you have references to support this? Some google search says it is 280B market out of 1.5T total pharma market: https://www.alliedmarketresearch.com/protein-therapeutics-ma... htt…

The first Google search link you've provided is focused on proteins as an active ingredient (like an antibody), not the targets.

https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6314433/ Check out this article for example.

Post reply on HN