Live data from Hacker News

The business of extracting knowledge from academic publications

theseedsofscience.pub

61–70 of 85 posts

Re: The business of extracting knowledge from academic publications

#61

Earlier quoted context omitted.

Well, maybe find some time and dive in a bit and see what can be found? You never know, maybe you'll end up contributing to our understanding of life, maybe (indirectly) even save a few lives!

I am working on the service which potentially can answer questions like in this comment: https://news.ycombinator.com/item?id=38109294 life science is one of potential applications if there is an interest and money.

So the pathway to synthesize every protein is +/- the same: That's gene transcription[1] and translation[2]. If that's broken, you're in big trouble!

But if you mean in general if you're capable of looking at metabolic pathways where each protein catalyses a step in the pathway, that's definitely interesting. If a certain person has a flawed gene coding for protein X, that could indeed cause a problem.

To find valid answers, you might need to eg. track nodes and states in a graph, to figure all the consequences of a break. Not all types of storage systems/engines are equally good at that.

[1] https://en.wikipedia.org/wiki/Transcription_(biology)

[2] https://en.wikipedia.org/wiki/Translation_(biology)

edit: s/protein pathway/metabolic pathway/

Re: The business of extracting knowledge from academic publications

#62
post #60

Earlier quoted context omitted.

> The GP was talking about pharma. post is not about pharma specifically, but about general bio-medical literature, including genes, diseases, symptoms, etc. > proteins make up more than 95% of the targets of the pharmaceutical industry. do you have references to support this? Some google search says it is 280B market out of 1.5T total pharma market: https://www.alliedmarketresearch.com/protein-therapeutics-ma... htt…

The first Google search link you've provided is focused on proteins as an active ingredient (like an antibody), not the targets. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6314433/ Check out this article for example.

So, that link says they discovered 1.5k FDA approved drugs which target proteins, while FDA has total 19k drugs approved: https://www.fda.gov/media/115824/download#:~:text=FDA%20regu....

Re: The business of extracting knowledge from academic publications

#63

Earlier quoted context omitted.

I am working on the service which potentially can answer questions like in this comment: https://news.ycombinator.com/item?id=38109294 life science is one of potential applications if there is an interest and money.

So the pathway to synthesize every protein is +/- the same: That's gene transcription[1] and translation[2]. If that's broken, you're in big trouble! But if you mean in general if you're capable of looking at metabolic pathways where each protein catalyses a step in the pathway, that's definitely interesting. If a certain person has a flawed gene coding for protein X, that could indeed cause a problem. To find valid…

> Not all types of storage systems/engines are equally good at that.

Yes, I built system which traverses paths in graphs with 1B nodes and 10B links in 1h on affordable server. But that's only one part of the puzzle.

Re: The business of extracting knowledge from academic publications

#64

Earlier quoted context omitted.

So the pathway to synthesize every protein is +/- the same: That's gene transcription[1] and translation[2]. If that's broken, you're in big trouble! But if you mean in general if you're capable of looking at metabolic pathways where each protein catalyses a step in the pathway, that's definitely interesting. If a certain person has a flawed gene coding for protein X, that could indeed cause a problem. To find valid…

> Not all types of storage systems/engines are equally good at that. Yes, I built system which traverses paths in graphs with 1B nodes and 10B links in 1h on affordable server. But that's only one part of the puzzle.

Neat!

Re: The business of extracting knowledge from academic publications

#65

Word of advice to all those who are chomping at the bit to disrupt pharma with AI. A pharmaceutical company that heavily leverages computation is called a pharmaceutical company. All modern pharmaceutical companies heavily leverage computational tools - including some powered by deep learning. Any company building a computational platform to accelerate drug discovery / development is not a pharmaceutical company. The…

Thanks for commenting. I've been thinking a bit about taking some computational insights we have developed over years to bioinformatics.

During my Ph.D., my PI wanted to do some work on signalling pathway analysis. He was interested in knowing when and why sometimes body starts assuming that the sick state is the right state and start acting against drugs given to patients. This making treatment very hard.

I don't know what is status of that but one insight I got from him that most (all?) Pharma companies do not trust acdemic data. I didn't ask reasons but I'd assume because of replication issue (crisis). Knowing how data is gathered and published in signalling domain, I'd not blame them for having low trust in academic data.

Re: The business of extracting knowledge from academic publications

#66

This was a fascinating read. The well-meaning and hard-working author went through several iterations of trying to make a profit on 'academic-knowledge-graph-adjacent' products, but things ultimately fell through. The article describes two separate things likely to appeal to HN readers. The first is that there is a lot of tacit knowledge not captured in scientific publications. The second is that the author and his t…

Academic here. To be honest, while the author's depiction of academic publishing is mostly not wrong, they make it sound much worse than it actually is. Folk knowledge is a thing, but papers do contain most of the valuable knowledge if you know how to read them. I think 95% of this person's failure to monetize their product comes from trying to sell it to an audience that is just quite broke, and the rest is probably…

What do you use ChatGPT for? I know this has been discussed to death but never by someone outside of the mobile app writing business.

Re: The business of extracting knowledge from academic publications

#67
post #35

Earlier quoted context omitted.

I wouldn't be so quick to dismiss a niche for third party data analysis. A very common feature across verticals is that in house data analysis is siloed. While it's a difficult sell, getting multiple siloed data sources to agree to 3rd party analysis for shared gains can be extremely successful for everyone involved when pulled off. So looking specifically at pharma, while analysis of published papers is meh, an ofte…

I mean what you suggested already kind of exists in several forms.[0,1] Third party data analyses is a thing, but this is often bundled with conducting experiments by CROs. See my other comment about market size. The problem is that you can make A LOT more money selling drugs than you can selling software to pharma companies. So if your software is any good, then you should use it to make drugs and just be a pharma c…

What I suggested was simply an illustrative example off the top of my head. The fact there's multiple similar companies that already independently exist furthers my point.

As for your other point, I'd imagine there's a pretty big difference between the capabilities and infrastructure between bringing a drug to market and analyzing the data of people who bring drugs to market.

Re: The business of extracting knowledge from academic publications

#69
post #60

Earlier quoted context omitted.

The first Google search link you've provided is focused on proteins as an active ingredient (like an antibody), not the targets. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6314433/ Check out this article for example.

So, that link says they discovered 1.5k FDA approved drugs which target proteins, while FDA has total 19k drugs approved: https://www.fda.gov/media/115824/download#:~:text=FDA%20regu... .

Hi there. Apologies for any confusion. I was clarifying the intent of the first comment in this thread. I will add some impressionistic comments below in case it helps your future service.

Even when a drug target is unknown, or remains mysterious, very high chances are that the target is a protein. DNA or RNA as targets are niche (DNA is often avoided as a target on purpose and it’s hard to be specific to it without DNA-like material, and RNA is still hard to target effectively, though things are improving). There is not much else of use in the cells (lipids, sugars, cofactors, and metabolites, some examples of which have been targeted by a couple drugs each over the long history of trials, often unintentionally.

Small molecules are an excellent modality for an eventual approved therapy. They almost always (so far) targets a protein. They are hard to design but when they are done well they expose the target to something nature hasn’t seen before in order to get a desired effect. Sometimes people don’t care about the target itself (think recreational drugs, or phenotypic drug discovery), but the target typically remains a protein.

Proteins make up the machinery of the cell. You jam or modify them to achieve desired effects.

Re: The business of extracting knowledge from academic publications

#70
post #60

Earlier quoted context omitted.

The first Google search link you've provided is focused on proteins as an active ingredient (like an antibody), not the targets. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6314433/ Check out this article for example.

So, that link says they discovered 1.5k FDA approved drugs which target proteins, while FDA has total 19k drugs approved: https://www.fda.gov/media/115824/download#:~:text=FDA%20regu... .

Also, and totally minor: there are nowhere close to 19k different approved small molecules. The same drug can be included in multiple products or formulations bringing that number you mentioned to 19k marketed products. Each generic formulation of iboprofen increases the latter count by 1. Counting all the marketed products of pure orange juice may add up to a large number but it is still one ingredient.
Post reply on HN