Live data from Hacker News

The business of extracting knowledge from academic publications

theseedsofscience.pub

71–80 of 85 posts

Re: The business of extracting knowledge from academic publications

#71
There is a wealth of Malaria data, painstakingly extracted and georeferenced from academic sources using Wellcome Trust funding [1].

I know this because I put it online in before I left them in 2012 [2] however I'm unable to use their new data explorer to extract any usable information.

In the past we provided direct access to the malaria endemicity surveys and georeferenced anopheles occurrence records but I'm now struggling to find that for some reason.

[1] https://malariaatlas.org/ [2] https://pubmed.ncbi.nlm.nih.gov/23680401/

Re: The business of extracting knowledge from academic publications

#72
Nice writeup, and a good explanation of why the academic literature is often less than helpful. There are several issues that aren't really mentioned however. For anyone looking into a specific field (such as the one mentioned in the article, metabolic engineering in yeast), you might want to consider:

1) Identification of all the major research groups working on the problem, including their physical locations and resources, lab standards and data standards, history of grants and proposals, etc. This will not be in the academic literature, and is fairly difficult and possibly expensive to acquire, and if you don't know what to look for, well, you need to hire experts with relevant experience. Think of it as due diligence (Theranos investors got burned because they didn't do this).

2) You have to aggregate a lot of papers to get a good picture of what each research group is up to. A single coherent project might generate a dozen papers scattered here and there, and importantly, they won't have published their failures in most cases, even though that information is just as valuable to an outsider attempting to replicate or advance their work as the successes are.

3) A common mistake is to neglect materials & methods and instead focus on results and discussion. A large fraction of the literature is based on poor methodology and so the results can't really be trusted, and fraud is remarkably widespread in academia, for various reasons from PhDs desperate to graduate to PIs who've made the practice their bread and butter for decades. Clearly written materials and methods sections that include all the information needed to replicate the work are an indication that it's fairly trustworthy. Deliberate obfuscation is a bad sign.

Re: The business of extracting knowledge from academic publications

#73
Skimming through the discussions here, I thought that perhaps we could change how we teach biology in high school and instead start by teaching stories of drug discovery using practical examples that motivate concepts that can then be discussed in greater detail at later courses.

Re: The business of extracting knowledge from academic publications

#74
post #59

Earlier quoted context omitted.

I think the big case study everybody gives is Schrödinger who took 20 years to IPO. The problem is that the market is just not that big. Assume that every single pharma company buys your product - like Schrödinger - where does your revenue top out? You can beat that as a small pharma company with one or two assets that make it to market. So the question is: if the computational platform is so good then why not just b…

For a good computational platform, that can pinpoint unique and valuable candidates , isn't there an option to do license/consult in exchange for royalties ?

No. Makes more sense to license the asset (molecule). And even then, big pharma companies won’t talk to you until your molecule is in phase 2.

If anybody with an AI pharma start up tells you they plan to make money by licensing their platform, then run the other way because it is a red flag.

If you want to see what a real “AI” pharma company looks like check out Vertex pharmaceuticals.

Re: The business of extracting knowledge from academic publications

#75
post #70

Earlier quoted context omitted.

So, that link says they discovered 1.5k FDA approved drugs which target proteins, while FDA has total 19k drugs approved: https://www.fda.gov/media/115824/download#:~:text=FDA%20regu... .

Also, and totally minor: there are nowhere close to 19k different approved small molecules. The same drug can be included in multiple products or formulations bringing that number you mentioned to 19k marketed products. Each generic formulation of iboprofen increases the latter count by 1. Counting all the marketed products of pure orange juice may add up to a large number but it is still one ingredient.

Yes, I compare drugs with drugs, not molecules with drugs.

Re: The business of extracting knowledge from academic publications

#76

Earlier quoted context omitted.

I happen to run a corporate filings product[1] so I'm curious to know in what way you find the data in filings limited. There are financial statements (ex. balance sheet) & disclosures (ex. litigation) in 100+ page annual reports so our tool makes it easier to find them. We also do AI (ex. sentiment analysis) and diffs (ex. redline / blackline) which yield their own insights. [1] https://last10k.com

> so I'm curious to know in what way you find the data in filings limited. to me every filling has maybe 20 essential numbers which are interesting: balance sheet, income statement and major sectors, everything else is some generic boilerplate, and there are dozens of services which will already sell it for cheap. Not sure what else you can sell to your clients..

I worked at a place where we developed information extraction systems that could be customized to the needs of particular customers. This was before transformers so the technology wasn't 100% ready.

Think of a global aircraft manufacturer turning maintenance documentation into a knowledge graph, a global clothing and shoes retailer building a model of what social media thinks about them, etc. I told other employees that our product could generate enough value for one customer that it would be worth it for one to buy us and... that's what happened.

Re: The business of extracting knowledge from academic publications

#77

Earlier quoted context omitted.

> Word of advice to all those who are chomping at the bit to disrupt pharma with AI. Literally the first line in the comment that started this thread.

Sure, now let's read the post?

I read the post before I made any criticism of your comments. We're talking on a thread within the larger context of comments on the post. But more importantly, if you read the post, you will see there is a theme of industrial biochemistry (IE, pharma and biotech) running through it, because pharma/biotech is the primary consumer of these products, and the vast majority of the revenue stream.

Re: The business of extracting knowledge from academic publications

#78

Earlier quoted context omitted.

Academic here. To be honest, while the author's depiction of academic publishing is mostly not wrong, they make it sound much worse than it actually is. Folk knowledge is a thing, but papers do contain most of the valuable knowledge if you know how to read them. I think 95% of this person's failure to monetize their product comes from trying to sell it to an audience that is just quite broke, and the rest is probably…

What do you use ChatGPT for? I know this has been discussed to death but never by someone outside of the mobile app writing business.

Quite a lot of things. A (probably non-exhaustive, off the top of my head) of things where it saves me the most time:

- Bureaucracy. Writing silly boilerplate, e.g. data management plans or gender perspective statements in grant proposals.

- Cutting or expanding text (we routinely have lots of forms and submissions where you need to write a text in a given word or character range).

- Polite emails in English to people I don't know much (e.g. "Write a polite professional email reminding this person that the deadline for reviewing paper Y expired yesterday...")

- Brainstorming. "Give me 10 ideas about research direction in topic X". It won't give great ideas, but it's good to set the mind rolling.

- Routine scripts/code used in experiments and papers: write a Python script to make a box plot with such and such data, or to take a file in this format and strip this unneeded content, etc. The typical kind of code that appears a lot in research, is trivial to code but consumes time and ChatGPT does it in seconds.

- Suggest titles (paper titles, grant proposal titles, etc.).

- Suggest ideas for exercises or exam questions (e.g. write an assignment that can be solved with the coin change algorithm but involves no coins or currency).

- How to do X in Excel (although the problem here is that my Excel is in Spanish - why, why did they decide to translate function names? - and it's not that good at that - but anyway, it's very useful).

The productivity boost is very noticeable, well worth the cost, even if it hurts to pay out of pocket for a tool used at work.

Re: The business of extracting knowledge from academic publications

#79
post #77

Earlier quoted context omitted.

Sure, now let's read the post?

I read the post before I made any criticism of your comments. We're talking on a thread within the larger context of comments on the post. But more importantly, if you read the post, you will see there is a theme of industrial biochemistry (IE, pharma and biotech) running through it, because pharma/biotech is the primary consumer of these products, and the vast majority of the revenue stream.

> there is a theme of industrial biochemistry

that's one of the themes (and you are already working hard to stretch drugs pharma to "biochemistry"), if you can't see other themes in his examples and screenshots, I think this discussion is not interesting to me.

Re: The business of extracting knowledge from academic publications

#80

I'm curious has one tried to do what the author did for a subject like physics or geology, mathematics or computer science? I for one would pay for a service that allowed me to discover interesting CS and math papers.

1. Define "interesting".

2. How much would you pay?

Post reply on HN