Live data from Hacker News

The business of extracting knowledge from academic publications

theseedsofscience.pub

31–40 of 85 posts

Re: The business of extracting knowledge from academic publications

#31
post #20

Earlier quoted context omitted.

The GP was talking about pharma. Proteins are not niche in pharma. Everything else may be niche in that domain, but proteins make up more than 95% of the targets of the pharmaceutical industry.

> The GP was talking about pharma. post is not about pharma specifically, but about general bio-medical literature, including genes, diseases, symptoms, etc. > proteins make up more than 95% of the targets of the pharmaceutical industry. do you have references to support this? Some google search says it is 280B market out of 1.5T total pharma market: https://www.alliedmarketresearch.com/protein-therapeutics-ma... htt…

This isn't a "do you have references" situation. You're wasting folks time. Nearly all drugs target proteins, with a few that target DNA or RNA.

Re: The business of extracting knowledge from academic publications

#32
post #31

Earlier quoted context omitted.

> The GP was talking about pharma. post is not about pharma specifically, but about general bio-medical literature, including genes, diseases, symptoms, etc. > proteins make up more than 95% of the targets of the pharmaceutical industry. do you have references to support this? Some google search says it is 280B market out of 1.5T total pharma market: https://www.alliedmarketresearch.com/protein-therapeutics-ma... htt…

This isn't a "do you have references" situation. You're wasting folks time. Nearly all drugs target proteins, with a few that target DNA or RNA.

You can just ignore my comments and walk away? To me you just another "internet expert" which I am not sure why should I blindly trust.

Re: The business of extracting knowledge from academic publications

#33

Earlier quoted context omitted.

Uniprot fills a pretty big niche - proteins. Your body is literally made out of proteins. They carry out most of the work and chemical reactions that amount to what we call life, and almost all drugs work by modulating some protein target in some way. UniprotIDs are the canonical identifiers that computational biologists use for proteins. But they are just one knowledge base. There are others. Each with their own foc…

> But they are just one knowledge base. There are others. So, is your claim that all topics/niches/verticals/steps in pharma development and productions are covered by all those databases with absolute quality and perfect UI/workflow? I follow some studies about trials meta-analysis, and my impression is that end result is mainly produced by manual work of some low paid postdocs, which makes them not trustworthy. > L…

The parent is right, saying a database about proteins 'looks like one narrow niche knowledge base.' in this context is just nonsense.

Re: The business of extracting knowledge from academic publications

#34

Earlier quoted context omitted.

> But they are just one knowledge base. There are others. So, is your claim that all topics/niches/verticals/steps in pharma development and productions are covered by all those databases with absolute quality and perfect UI/workflow? I follow some studies about trials meta-analysis, and my impression is that end result is mainly produced by manual work of some low paid postdocs, which makes them not trustworthy. > L…

The parent is right, saying a database about proteins 'looks like one narrow niche knowledge base.' in this context is just nonsense.

in which context, and why your bold statements should be trusted blindly?

Re: The business of extracting knowledge from academic publications

#35

Word of advice to all those who are chomping at the bit to disrupt pharma with AI. A pharmaceutical company that heavily leverages computation is called a pharmaceutical company. All modern pharmaceutical companies heavily leverage computational tools - including some powered by deep learning. Any company building a computational platform to accelerate drug discovery / development is not a pharmaceutical company. The…

I wouldn't be so quick to dismiss a niche for third party data analysis.

A very common feature across verticals is that in house data analysis is siloed.

While it's a difficult sell, getting multiple siloed data sources to agree to 3rd party analysis for shared gains can be extremely successful for everyone involved when pulled off.

So looking specifically at pharma, while analysis of published papers is meh, an often discussed component of research is the bias against publishing negative results.

So a hypothetical product like using ML across multiple firms' data of failed products and research in order to establish a model that could more quickly identify dud research avenues by leveraging industry data could only exist with the broadest dataset as a 3rd party product and would deliver gains that could be quite profitable for some of the largest companies out there.

Also, I think anyone who has worked with larger corporations on in house tech knows that even if the individual efforts are quite large and sophisticated, the sheer amount of bureaucracy that goes into every single thing at a sizable firm can mean significant advantages for startups vs in-house efforts, particularly when related to fast moving fields.

I'd agree that "moving into a niche without knowing it extremely well" can be fraught with issues and that attempting to get buy in from B2B firms for a startup is a nightmare, but I'd disagree with "large company does X in house so creating a startup to do X is a bad idea."

Re: The business of extracting knowledge from academic publications

#36

Earlier quoted context omitted.

Uniprot fills a pretty big niche - proteins. Your body is literally made out of proteins. They carry out most of the work and chemical reactions that amount to what we call life, and almost all drugs work by modulating some protein target in some way. UniprotIDs are the canonical identifiers that computational biologists use for proteins. But they are just one knowledge base. There are others. Each with their own foc…

> But they are just one knowledge base. There are others. So, is your claim that all topics/niches/verticals/steps in pharma development and productions are covered by all those databases with absolute quality and perfect UI/workflow? I follow some studies about trials meta-analysis, and my impression is that end result is mainly produced by manual work of some low paid postdocs, which makes them not trustworthy. > L…

I already gave you a real life example with the uniprot reference. Here is another flagship knowledge base that heavily leverages NLP extraction.[0] Here is another one that gets used in what seems like every network biology article.[1]

Meta analyses? Automating meta analyses is not a real need. They have their place, but it’s like a quaint cottage industry type thing - like custom haberdashery.

Also the most valuable knowledge is not in any publication. If you are reading about it in an article then you are already 2-3 years too late.

0. https://geneontology.org

1. https://string-db.org/

Re: The business of extracting knowledge from academic publications

#37

This was a fascinating read. The well-meaning and hard-working author went through several iterations of trying to make a profit on 'academic-knowledge-graph-adjacent' products, but things ultimately fell through. The article describes two separate things likely to appeal to HN readers. The first is that there is a lot of tacit knowledge not captured in scientific publications. The second is that the author and his t…

Academic here.

To be honest, while the author's depiction of academic publishing is mostly not wrong, they make it sound much worse than it actually is. Folk knowledge is a thing, but papers do contain most of the valuable knowledge if you know how to read them.

I think 95% of this person's failure to monetize their product comes from trying to sell it to an audience that is just quite broke, and the rest is probably mostly post hoc rationalization. Not only grad students and postdoc wages are low, in many countries (not the US) professors aren't well paid either (and buying software subscriptions from grant funds is often not allowed or difficult due to crazy bureaucracy).

As a full professor myself, I almost don't buy software for work. I suffer the torture of Microsoft Office, which my institution is subscribed to, I'm subscribed to Overleaf with grant money (for now, but I might be forced to cancel depending on how the funding goes) and I pay for ChatGPT out of pocket because trying to use grant money for that is bureaucratic hell. That's all. It would take a really transformative piece of software for me to subscribe to something else.

Re: The business of extracting knowledge from academic publications

#38

Earlier quoted context omitted.

> But they are just one knowledge base. There are others. So, is your claim that all topics/niches/verticals/steps in pharma development and productions are covered by all those databases with absolute quality and perfect UI/workflow? I follow some studies about trials meta-analysis, and my impression is that end result is mainly produced by manual work of some low paid postdocs, which makes them not trustworthy. > L…

I already gave you a real life example with the uniprot reference. Here is another flagship knowledge base that heavily leverages NLP extraction.[0] Here is another one that gets used in what seems like every network biology article.[1] Meta analyses? Automating meta analyses is not a real need. They have their place, but it’s like a quaint cottage industry type thing - like custom haberdashery. Also the most valuabl…

> that heavily leverages NLP extraction

I search https://www.google.com/search?q=site%3Ageneontology.org+nlp+... and don't see anything meaningful.

> Automating meta analyses is not a real need. They have their place, but it’s like a quaint cottage industry type thing - like custom haberdashery.

my opinion is that it has to be top level tool in modern science which could easily sort out lots of bs and contradictions in reported results.

Re: The business of extracting knowledge from academic publications

#39
post #35

Word of advice to all those who are chomping at the bit to disrupt pharma with AI. A pharmaceutical company that heavily leverages computation is called a pharmaceutical company. All modern pharmaceutical companies heavily leverage computational tools - including some powered by deep learning. Any company building a computational platform to accelerate drug discovery / development is not a pharmaceutical company. The…

I wouldn't be so quick to dismiss a niche for third party data analysis. A very common feature across verticals is that in house data analysis is siloed. While it's a difficult sell, getting multiple siloed data sources to agree to 3rd party analysis for shared gains can be extremely successful for everyone involved when pulled off. So looking specifically at pharma, while analysis of published papers is meh, an ofte…

I mean what you suggested already kind of exists in several forms.[0,1]

Third party data analyses is a thing, but this is often bundled with conducting experiments by CROs.

See my other comment about market size. The problem is that you can make A LOT more money selling drugs than you can selling software to pharma companies. So if your software is any good, then you should use it to make drugs and just be a pharma company.

0. https://www.opentargets.org/

1. https://www.citeline.com/en/products-services/clinical/pharm...

Re: The business of extracting knowledge from academic publications

#40
post #31

Earlier quoted context omitted.

This isn't a "do you have references" situation. You're wasting folks time. Nearly all drugs target proteins, with a few that target DNA or RNA.

You can just ignore my comments and walk away? To me you just another "internet expert" which I am not sure why should I blindly trust.

Your sources are talking about the ratio of small molecule vs. large molecule drugs. Even if you're developing small molecule drugs you are likely targeting some aspect of protein signaling/gene expression.

People are being dismissive of your comments because to say that proteins are niche in the context of pharma is like saying advertising is niche in the context of Meta and Google.

Post reply on HN