Live data from Hacker News

The business of extracting knowledge from academic publications

theseedsofscience.pub

21–30 of 85 posts

Re: The business of extracting knowledge from academic publications

#21

Word of advice to all those who are chomping at the bit to disrupt pharma with AI. A pharmaceutical company that heavily leverages computation is called a pharmaceutical company. All modern pharmaceutical companies heavily leverage computational tools - including some powered by deep learning. Any company building a computational platform to accelerate drug discovery / development is not a pharmaceutical company. The…

> I mean, just take a look at uniprot.[0] this looks like one narrow niche knowledge base. I am not expert in this domain, but there are probably some other use cases not covered by existing offerings. > literature extraction - because a team of 5-10 FTE ML / software engineers I think this problem is so hard and open ended, that your 5-10 non-star avg salary ml ftes likely produce very mediocre and likely not usable…

Uniprot fills a pretty big niche - proteins. Your body is literally made out of proteins. They carry out most of the work and chemical reactions that amount to what we call life, and almost all drugs work by modulating some protein target in some way. UniprotIDs are the canonical identifiers that computational biologists use for proteins.

But they are just one knowledge base. There are others. Each with their own focus area. And they are all associated with prominent bioNLP / biomedical AI research labs and employ human SME curators.

> 5-10 avg salary non-star ML FTEs likely produce very mediocre and likely non usable results

LOL!

Re: The business of extracting knowledge from academic publications

#22

Earlier quoted context omitted.

Corporate filings. Business intelligence. There is value in those areas. Science is too esoteric. And boy there are a lot of papers that don’t really signify anything at all but they fill up some pages and add to somebody’s paper count. It is a hoot that they sell access to individual scientific papers for $35 because if you think one will help you with some commercial problem you have the odds are the real value is…

> Corporate filings. Business intelligence. There is value in those areas also, lots of competition already

Competition is good, it means there is a market. There's plenty of empty niches with no competition exactly because there's no money there.

Re: The business of extracting knowledge from academic publications

#23

This was a fascinating read. The well-meaning and hard-working author went through several iterations of trying to make a profit on 'academic-knowledge-graph-adjacent' products, but things ultimately fell through. The article describes two separate things likely to appeal to HN readers. The first is that there is a lot of tacit knowledge not captured in scientific publications. The second is that the author and his t…

Law firms have about as much money as pharma, and the major legal research services (Westlaw et al) already have LLM based offerings.

Re: The business of extracting knowledge from academic publications

#24
post #20

Earlier quoted context omitted.

yes, proteins is a niche in grand scheme of things

The GP was talking about pharma. Proteins are not niche in pharma. Everything else may be niche in that domain, but proteins make up more than 95% of the targets of the pharmaceutical industry.

and an increasing number of drugs- in the "old days" it was almost entirely small molecules, but they are starting to peter out (both because the low-hanging fruit has already been plucked, and also, small molecules are a usually a terrible way to modulate biological activity in specific ways.

Re: The business of extracting knowledge from academic publications

#25
post #7

I agree with the OP but only partially. UniProt is a good counterexample of a database that has been built by extracting knowledge from publications and that is incredibly useful. But it took decades of expert hand-curators going through piles of articles to get to the current state. Also, proteomics articles report very simple outcomes that relatively easy to annotate. For instance, the subcellular localization of a…

Yes and UniProt is not a business, it's run by a foundation funded by US/UK/Swiss universities

Re: The business of extracting knowledge from academic publications

#26

Word of advice to all those who are chomping at the bit to disrupt pharma with AI. A pharmaceutical company that heavily leverages computation is called a pharmaceutical company. All modern pharmaceutical companies heavily leverage computational tools - including some powered by deep learning. Any company building a computational platform to accelerate drug discovery / development is not a pharmaceutical company. The…

Ingenuity's Pathway Analysis product has been the standard tool big pharma has used for 15+ years. Founded by folks from Stanford, they created their own knowledge-graph database a decade before they were popular, and have been paying PhD's to curate that database with genetic findings from papers over the last decades. Obviously their analysis also includes the publicly-available databases, and they have their own private genomic databases and expertise from helping many of the first-available genomic services go live.

Still, it's a tough business, and they were bought by Qiagen (to make an end-to-end solution, with moderate success, as the market is strangled by Illumina's control over NGS sequencing). And with exponential growth in information, their early relative advantage likely has waned.

Also note Veeva software, which provides infrastructure for pharma, is a public benefit company, legally devoted to its clients. That's how much power the customer has.

That said, there's something of a revolution in drug discovery as more is being done by small companies with ultra-focused expertise. If they get a viable candidate, a bigger company buys it to marshal through approvals, and then it can be sold again to a marketing company. Focusing on these pop-up drug-discovery companies that evolve out of someone's PhD could be a good niche; I would expect Pharma VC's to prefer funding companies using their favorite computing vendor, for reliability if not insight. Here computational discovery assistance would help a bit, but validating findings and science could be golden.

Re: The business of extracting knowledge from academic publications

#27

Earlier quoted context omitted.

> I mean, just take a look at uniprot.[0] this looks like one narrow niche knowledge base. I am not expert in this domain, but there are probably some other use cases not covered by existing offerings. > literature extraction - because a team of 5-10 FTE ML / software engineers I think this problem is so hard and open ended, that your 5-10 non-star avg salary ml ftes likely produce very mediocre and likely not usable…

Uniprot fills a pretty big niche - proteins. Your body is literally made out of proteins. They carry out most of the work and chemical reactions that amount to what we call life, and almost all drugs work by modulating some protein target in some way. UniprotIDs are the canonical identifiers that computational biologists use for proteins. But they are just one knowledge base. There are others. Each with their own foc…

> But they are just one knowledge base. There are others.

So, is your claim that all topics/niches/verticals/steps in pharma development and productions are covered by all those databases with absolute quality and perfect UI/workflow? I follow some studies about trials meta-analysis, and my impression is that end result is mainly produced by manual work of some low paid postdocs, which makes them not trustworthy.

> LOL!

I actually claim non-trivial expertise in this area (facts extraction from untrivial niche documents), and my observations is that even SOTA (aka results from top researchers) are hardly useable in real life, because this area is very hard. Could you support your "LOL!" with any references I could check?..

Re: The business of extracting knowledge from academic publications

#28
post #20

Earlier quoted context omitted.

yes, proteins is a niche in grand scheme of things

The GP was talking about pharma. Proteins are not niche in pharma. Everything else may be niche in that domain, but proteins make up more than 95% of the targets of the pharmaceutical industry.

> The GP was talking about pharma.

post is not about pharma specifically, but about general bio-medical literature, including genes, diseases, symptoms, etc.

> proteins make up more than 95% of the targets of the pharmaceutical industry.

do you have references to support this? Some google search says it is 280B market out of 1.5T total pharma market: https://www.alliedmarketresearch.com/protein-therapeutics-ma... https://www.statista.com/topics/1764/global-pharmaceutical-i...

Also, it is likely important steps how to research and produce protein drugs, but the end target is to cure diseases, so you need lots of additional data about diseases of all types, symptoms, pathways, trials, etc.

Re: The business of extracting knowledge from academic publications

#29
post #26

Word of advice to all those who are chomping at the bit to disrupt pharma with AI. A pharmaceutical company that heavily leverages computation is called a pharmaceutical company. All modern pharmaceutical companies heavily leverage computational tools - including some powered by deep learning. Any company building a computational platform to accelerate drug discovery / development is not a pharmaceutical company. The…

Ingenuity's Pathway Analysis product has been the standard tool big pharma has used for 15+ years. Founded by folks from Stanford, they created their own knowledge-graph database a decade before they were popular, and have been paying PhD's to curate that database with genetic findings from papers over the last decades. Obviously their analysis also includes the publicly-available databases, and they have their own p…

I think the big case study everybody gives is Schrödinger who took 20 years to IPO.

The problem is that the market is just not that big. Assume that every single pharma company buys your product - like Schrödinger - where does your revenue top out?

You can beat that as a small pharma company with one or two assets that make it to market. So the question is: if the computational platform is so good then why not just be pharma company?

Re: The business of extracting knowledge from academic publications

#30

Earlier quoted context omitted.

> Corporate filings. Business intelligence. There is value in those areas also, lots of competition already

Competition is good, it means there is a market. There's plenty of empty niches with no competition exactly because there's no money there.

As for corporate fillings, there are lots of strong products already, also data in those fillings is rather limited.

"Business intelligence" is very broad term, maybe it is possible to find market fit there, since area is moving very fast, but hard to judge without seeing specifics.

Post reply on HN