Live data from Hacker News

The business of extracting knowledge from academic publications

markusstrasser.org

71–80 of 122 posts

Re: The business of extracting knowledge from academic publications

#71
post #23

Earlier quoted context omitted.

Unfortunately you'll run into the rapid fact that in the ML/AI space you get almost zero points for building something. You get a whole lot of points for discovering something, designing something, or a proof. But there's a very large amount of people focused entirely on aims that are very, very distant from actually making human lives genuinely better. Mostly because everyone quietly understands all the extraordinar…

>But there's a very large amount of people focused entirely on aims that are very, very distant from actually making human lives genuinely better. I can't speak to each and every person working on ML, but I thought I would share a fun use case I ran across the other day. There is a business in some foreign country that is similar to Uber Eats: customer goes to an app, browses for food from various restaurants, orders…

Guessing allergen content sounds like a disaster waiting to happen.

Re: The business of extracting knowledge from academic publications

#72

Interesting but I don't know how to make sense of it. How can it be that "close to nothing of what makes science actually work is published as text on the web"? - Is the information that makes science actually work mostly in images that the machines don't yet understand? - Was the information paywalled or in private databases and inaccessible to this researcher? - Are the papers mostly just advertisements for researc…

Science is a profession like others. When you are earning your Ph.D. you learn to think about the field by reading papers and discussing with peers and colleagues, yes. The intro of a well-structured research paper should follow this pattern: - This is a really important topic and here is why. - What is the current state of the art in this field? (this comes from reading 100-1000 publications on the topic and selecti…

Exactly! Scientific papers are not meant to stand on their own -- they are pieces of a much larger jigsaw puzzle. In order to make heads or tails out of a paper, one really needs to have a sense of where the paper fits into its larger picture. Building up necessary base of knowledge to develop that sense, both in terms of explicit knowledge and tacit knowledge, is part of what a PhD student is actually doing while they are working on their PhD, and is part of why the process takes as long as it does.

Also, the mechanical process of effectively reading a paper is highly non-linear, and is a skill in and of itself. In a lot of ways, it is more akin to high-level pattern matching than it is to more "normal" reading. At least at my institution, it is something that we actually teach our students to do in formal ways (the obligatory "How to read a scientific paper" lecture during the first term or two) and then make them practice over and over again for years (journal clubs, etc.). The original author eventually figured this out, which is to their credit.

Re: The business of extracting knowledge from academic publications

#73

Earlier quoted context omitted.

I may be a bit cynical, but at least in my former field (experimental physics), the main purpose of papers seems to be to "lock in" a finished achivement. You do the actual research, pass internal reviews and peer review, and then publishing the paper is just to make it "official". Unfortunately, many papers are never expected to be read. The crucial information exists, but you usually get it from personal communicat…

I did my phd in experimental physics and I have to say my realisation of this large point, that papers are little more than resume padding to lock in an achievement was a significant contributor towards destroying, and I use destroying seriously here, any faith or trust that peer review or publishing has anything at all to do with the scientific method at all. Your results replicate, or they don't. Your calculations,…

In academia there always is a difference between the way results are advertised and what conclusions are drawn internally. This is more true in some fields than others, I'm most familiar with it ML, Physics. Part of your skill as a researcher is to understand based on omissions, the datasets etc. the quiet part that isn't said out loud. Depending how you sell things you can get a Nature / Science paper with confusing inconsistent terminology, hand rolled C++ implementation, provided you are the first and a another method which might be 1000x times faster will only make it into PRL (yes I'm thinking of two specific papers, but won't say which).

Re: The business of extracting knowledge from academic publications

#74
post #55

It's interesting that OP did seemingy little research with respect to existing work in the field. https://www.nlm.nih.gov/medline/medline_overview.html Medline, a searchable online directory of medical research papers has existed for 50 years. The National Library of Medicine for many years was a leader in document search and retrieval before there was a web. In the 80's they were doing vector cosine document similar…

The best thing about the NLM's work in this space is how deeply it has been informed by the needs and workflows of the biomedical researchers, which is a perspective that has been sorely lacking in work coming from outsiders.

I did think that the author did a good job of outlining (some of) the basic structural issues that make this a tough field to monetize, but even setting those aside, there's no substitute for actually knowing your users and what they need, and that's something the NLM is amazing at.

(Disclosure, my PhD was funded by an NLM training grant, some of my research is funded extramurally by the NLM, and I have a lot of NLM colleagues, so I'm maybe a little bit biased)

Re: The business of extracting knowledge from academic publications

#76
post #61

> Why purchase access to a 3rd party AI reading engine or a knowledge graph when you can just hire hundreds of postdocs in Hyderabad to parse papers into JSON? (at a $6,000 yearly salary) I really like jobs people think AI can do in theory but can't really do them effectively irl. Where do I get a part-time gig like that if I think I am capable of reviewing and creating summary of non-STEM papers? Except for homework…

Yeah, you can't outsource that to Hyderabad. You'd need to know subject knowledge + very specific English and possibly other languages depending on the field (not saying Indians can't do this, but I've studied enough languages to know that doing high level/academic work in a non-native language is hell even when the language is pitched to students). And you'd have to know enough about the process and authors to know…

All good points. But you do have to recognize the tradeoff. Has AI come so far that it could perform better than industry specific human intelligence? You have to consider that maybe some Indian researchers could review the papers as they are doing that job as part time gig.

You have to test out both solution. And as these jobs are treated as contracts there is no significant commitment for choosing one over the other. We can't be certain if one method is better than the other without trying both of them out without prejudice.

I, for one am agnostic about either choice. Because AI is overhyped yet it has spillover benefits as a marketing-sales point but offshore human intelligence has a bad rep but could be effective if you have proper documentation, supervision and review framework.

Re: The business of extracting knowledge from academic publications

#77
post #54

The article's core claims are: > Extracting, structuring or synthesizing "insights" from academic publications (papers) or building knowledge bases from a domain corpus of literature has negligible value in industry. > Most knowledge necessary to make scientific progress is not online and not encoded. > Close to nothing of what makes science actually work is published as text on the web > The tech is not there to mak…

"follow the best institutions and ~50 top individuals" wasn't meant as a suggestion actually, just an observation of what most people do.

You're right they "could work if designed to solve a specific real world problem" but against what baseline? The baseline could be spending that time on actual deep tech projects and not NLP meta-science

Re: The business of extracting knowledge from academic publications

#78
post #54

The article's core claims are: > Extracting, structuring or synthesizing "insights" from academic publications (papers) or building knowledge bases from a domain corpus of literature has negligible value in industry. > Most knowledge necessary to make scientific progress is not online and not encoded. > Close to nothing of what makes science actually work is published as text on the web > The tech is not there to mak…

"follow the best institutions and ~50 top individuals" wasn't meant as a suggestion actually, just an observation of what most people do. You're right they "could work if designed to solve a specific real world problem" but against what baseline? The baseline could be spending that time on actual deep tech projects and not NLP meta-science

But you're right; open source projects for extracting infos (like PubTator) are valuable but ontologies/KGs need ongoing expert (ML, AI, SWEs, information architects, labelers) work (unlike most of Wikipedia or GH) so it's tough to make something that doesn't suck in a distributed open source fashion

Re: The business of extracting knowledge from academic publications

#79
post #61

Earlier quoted context omitted.

Yeah, you can't outsource that to Hyderabad. You'd need to know subject knowledge + very specific English and possibly other languages depending on the field (not saying Indians can't do this, but I've studied enough languages to know that doing high level/academic work in a non-native language is hell even when the language is pitched to students). And you'd have to know enough about the process and authors to know…

All good points. But you do have to recognize the tradeoff. Has AI come so far that it could perform better than industry specific human intelligence? You have to consider that maybe some Indian researchers could review the papers as they are doing that job as part time gig. You have to test out both solution. And as these jobs are treated as contracts there is no significant commitment for choosing one over the othe…

Oh yeah, I was just thinking currently. In five to ten years once AI/ML/etc. trickle out of tech/theory spaces and starts to be combined with subject expertise, I think we'll see really interesting things.

The other matter is that an Indian who could review papers that well would also cost more than 6k/year and would not be easily replaceable, which eliminates the main benefit of outsourcing for a company trying to operate in such a way in 2021.

In 2030? I'd say the odds are if somebody in Hyderabad can do that then they can start their OWN company rather than bother with us at all. Honestly, given India's role in pharmaceutical manufacture, I'd be shocked if things like that don't start popping up.

Re: The business of extracting knowledge from academic publications

#80
For anyone interested, my whole PhD was in biomedical hypothesis generation! I think the most "serious" attempts at building these systems have been focused around providing assistance to scientists, and not just coming up with new ideas on their own.

here's an actual medical paper that my first system, Moliere, was able to help discover:

https://link.springer.com/article/10.1007/s11481-019-09885-8

Post reply on HN