Live data from Hacker News

The business of extracting knowledge from academic publications

theseedsofscience.pub

11–20 of 85 posts

Re: The business of extracting knowledge from academic publications

#11

This was a fascinating read. The well-meaning and hard-working author went through several iterations of trying to make a profit on 'academic-knowledge-graph-adjacent' products, but things ultimately fell through. The article describes two separate things likely to appeal to HN readers. The first is that there is a lot of tacit knowledge not captured in scientific publications. The second is that the author and his t…

Corporate filings. Business intelligence. There is value in those areas. Science is too esoteric. And boy there are a lot of papers that don’t really signify anything at all but they fill up some pages and add to somebody’s paper count.

It is a hoot that they sell access to individual scientific papers for $35 because if you think one will help you with some commercial problem you have the odds are the real value is $0.00.

Re: The business of extracting knowledge from academic publications

#12
Big implications and ponderings here for the slew of RAG based applications that are about to hit the market.

I saw a blog post by a fellow from OpenAI earlier today which basically said "when you refer to a model like Bard, or Claude, or Llama -- you're moreso referring to the dataset, than you are the architecture."

His meaning was at training time, but perhaps a similar observation can be had in this context: The value of a retrieval / organization system is only as good as the average information quality in its target corpus

Re: The business of extracting knowledge from academic publications

#13
> Close to nothing of what makes science actually work is published as text on the web

I appreciate the author didn't find a good product-market-fit. But the above claim makes no sense and isn't supported in the article.

The author doubles down on the idea later:

> All that is to say: discovering relevant literature, compiling evidence, finding mechanisms turns out to be a tiny percentage of actual, real life R&D

The author admits to not having worked in research, and this is one place that lack of experience shows. This is the kind of thing that takes years, or decades to develop an appreciation for.

Discovering relevant literature, efficiently and well, is a thing that can make or break you as a research scientist.

Over time I've sensed the devaluation of literature search skills in science, but I've also noticed those that can't do good literature searches do bad research. They waste time, sometimes years, in re-discovery of old results. They commit resources to experiments that don't need to be done. They have objectively worse ideas because they don't actually understand the field they're working in. They can't see the holes in the literature that lead to opportunities.

Re: The business of extracting knowledge from academic publications

#14
Word of advice to all those who are chomping at the bit to disrupt pharma with AI.

A pharmaceutical company that heavily leverages computation is called a pharmaceutical company. All modern pharmaceutical companies heavily leverage computational tools - including some powered by deep learning.

Any company building a computational platform to accelerate drug discovery / development is not a pharmaceutical company. They are a software company who wants to sell software to pharmaceutical companies, which is a terrible business to be in.

The product the author is selling already exists for free. All of the big knowledge bases use automated ML tooling for curation and extraction. And the people who run them and QC them are world class experts in their domains. I mean, just take a look at uniprot.[0]

And the only types of pharma companies who would want a big knowledge graph would be the large ones with active programs across multiple therapeutic areas. Most of the small companies / start ups tend to be focused on getting one or two assets to market in a therapeutic domain. And the big companies possibly have their own ML teams doing literature extraction - because a team of 5-10 FTE ML / software engineers is like a rounding error in their R&D budget.

The other thing is that the most valuable knowledge is the stuff that is not in the literature. That’s why we do experiments.

[0] https://academic.oup.com/nar/article/51/D1/D523/6835362

Re: The business of extracting knowledge from academic publications

#15

Word of advice to all those who are chomping at the bit to disrupt pharma with AI. A pharmaceutical company that heavily leverages computation is called a pharmaceutical company. All modern pharmaceutical companies heavily leverage computational tools - including some powered by deep learning. Any company building a computational platform to accelerate drug discovery / development is not a pharmaceutical company. The…

> I mean, just take a look at uniprot.[0]

this looks like one narrow niche knowledge base. I am not expert in this domain, but there are probably some other use cases not covered by existing offerings.

> literature extraction - because a team of 5-10 FTE ML / software engineers

I think this problem is so hard and open ended, that your 5-10 non-star avg salary ml ftes likely produce very mediocre and likely not usable results.

Re: The business of extracting knowledge from academic publications

#16

This was a fascinating read. The well-meaning and hard-working author went through several iterations of trying to make a profit on 'academic-knowledge-graph-adjacent' products, but things ultimately fell through. The article describes two separate things likely to appeal to HN readers. The first is that there is a lot of tacit knowledge not captured in scientific publications. The second is that the author and his t…

Corporate filings. Business intelligence. There is value in those areas. Science is too esoteric. And boy there are a lot of papers that don’t really signify anything at all but they fill up some pages and add to somebody’s paper count. It is a hoot that they sell access to individual scientific papers for $35 because if you think one will help you with some commercial problem you have the odds are the real value is…

> Corporate filings. Business intelligence. There is value in those areas

also, lots of competition already

Re: The business of extracting knowledge from academic publications

#17

Word of advice to all those who are chomping at the bit to disrupt pharma with AI. A pharmaceutical company that heavily leverages computation is called a pharmaceutical company. All modern pharmaceutical companies heavily leverage computational tools - including some powered by deep learning. Any company building a computational platform to accelerate drug discovery / development is not a pharmaceutical company. The…

> I mean, just take a look at uniprot.[0] this looks like one narrow niche knowledge base. I am not expert in this domain, but there are probably some other use cases not covered by existing offerings. > literature extraction - because a team of 5-10 FTE ML / software engineers I think this problem is so hard and open ended, that your 5-10 non-star avg salary ml ftes likely produce very mediocre and likely not usable…

> this looks like one narrow niche knowledge base. I am not expert in this domain, but there are probably some other use cases not covered by existing offerings.

UniProt is not niche. It contains curated information about all proteins across numerous domains.

Re: The business of extracting knowledge from academic publications

#18
post #17

Earlier quoted context omitted.

> I mean, just take a look at uniprot.[0] this looks like one narrow niche knowledge base. I am not expert in this domain, but there are probably some other use cases not covered by existing offerings. > literature extraction - because a team of 5-10 FTE ML / software engineers I think this problem is so hard and open ended, that your 5-10 non-star avg salary ml ftes likely produce very mediocre and likely not usable…

> this looks like one narrow niche knowledge base. I am not expert in this domain, but there are probably some other use cases not covered by existing offerings. UniProt is not niche. It contains curated information about all proteins across numerous domains.

yes, proteins is a niche in grand scheme of things

Re: The business of extracting knowledge from academic publications

#19
post #12

Big implications and ponderings here for the slew of RAG based applications that are about to hit the market. I saw a blog post by a fellow from OpenAI earlier today which basically said "when you refer to a model like Bard, or Claude, or Llama -- you're moreso referring to the dataset , than you are the architecture." His meaning was at training time, but perhaps a similar observation can be had in this context: The…

"about" to hit the market? There's like 3000+ Co-Pilots already in the market.

But yes, more to come out next year obviously.

Re: The business of extracting knowledge from academic publications

#20
post #17

Earlier quoted context omitted.

> this looks like one narrow niche knowledge base. I am not expert in this domain, but there are probably some other use cases not covered by existing offerings. UniProt is not niche. It contains curated information about all proteins across numerous domains.

yes, proteins is a niche in grand scheme of things

The GP was talking about pharma. Proteins are not niche in pharma. Everything else may be niche in that domain, but proteins make up more than 95% of the targets of the pharmaceutical industry.
Post reply on HN