Live data from Hacker News

The business of extracting knowledge from academic publications

markusstrasser.org

81–90 of 122 posts

Re: The business of extracting knowledge from academic publications

#81
post #80

For anyone interested, my whole PhD was in biomedical hypothesis generation! I think the most "serious" attempts at building these systems have been focused around providing assistance to scientists, and not just coming up with new ideas on their own. here's an actual medical paper that my first system, Moliere, was able to help discover: https://link.springer.com/article/10.1007/s11481-019-09885-8

Is your PhD thesis online anywhere?

Re: The business of extracting knowledge from academic publications

#82

Interesting but I don't know how to make sense of it. How can it be that "close to nothing of what makes science actually work is published as text on the web"? - Is the information that makes science actually work mostly in images that the machines don't yet understand? - Was the information paywalled or in private databases and inaccessible to this researcher? - Are the papers mostly just advertisements for researc…

In my experience useful scientific knowledge is accumulated in people actively working. Documents (books, papers, guides, programs, talks, blog posts, etc) are communication tools, but are limited by the medium and the ability of the authors. People can consume documents and create analogies to their specific work, but from there it's the process of working that produces: experts, systems, tools. Sometimes those products are again documented.

Re: The business of extracting knowledge from academic publications

#83
> Close to nothing of what makes science actually work is published as text on the web

Unless there's some nuance I missed, I immensely disagree with this statement.

I'm currently in the biomedical literature review space, and I appreciate the detailed insights. I wonder if the author considered that literature review is used in a wide variety of domains outside pharma/drug discovery (where I perceived their efforts were focused). Regulatory monitoring/reporting, hospital guideline generation, etc.

This is a billion dollar industry, and I couldn't agree more that it's technologically underdeveloped. I do not agree that AI-based extraction is the solution, at least in the near-term. The formal methodologies used by reviewers/meta-analysts: search strategy generation, lit search, screening, extraction, critical appraisal, synthesis/statistical analysis, are IMO more nuanced than an AI can capture. They require human input or review. My business is betting on this premise :)

Re: The business of extracting knowledge from academic publications

#85
post #53
post #39

Earlier quoted context omitted.

I'm an academic (applied math) and want to respond to this: academic papers are the way they are for lots of reasons, many of which (not so good) have been mentioned on HN. There are a couple that I do not see very often however: (1) Many academics aren't aware non-academics read their papers at all: we work with other academics, go to conferences with other academics, and on the rare occasions we hear from readers,…

> What can help make the literature more accessible? Review articles, sometimes called surveys. I've always thought that new PhDs would be excellent authors for those, having digested lots literature for their dissertation.

> Review articles, sometimes called surveys.

Is this field specific? I have read survey articles in math and biology, and was told by some of my profs that they use these articles as an introduction to a new field.

A quick Google search seems to show these exist in CS (along with tutorial papers), physics, and chemistry but I'm having a little difficulty finding statistics survey papers (survey methods come up instead).

Is the problem that there aren't enough of them or they are behind paywalls?

Re: The business of extracting knowledge from academic publications

#86
post #80

For anyone interested, my whole PhD was in biomedical hypothesis generation! I think the most "serious" attempts at building these systems have been focused around providing assistance to scientists, and not just coming up with new ideas on their own. here's an actual medical paper that my first system, Moliere, was able to help discover: https://link.springer.com/article/10.1007/s11481-019-09885-8

Is your PhD thesis online anywhere?

https://sybrandt.com/documents/dissertation.pdf

Re: The business of extracting knowledge from academic publications

#87
post #5
post #3

One thing that strikes me about most academic knowledge tools is that they seem to focus on parsing the current set of academic literature and producing supposedly interesting insights out of them (which quickly tends to snowball into wanting some kind of generalized model for knowledge as a whole). What I think is much more interesting is creating tools that help people create better academic writing in the first pl…

I recall seeing a Show HN post a while back about a research focussed web browser that helps as a thinking tool: https://news.ycombinator.com/item?id=28446147

A sign-in/sign-up necessary just to see the browser in action? Hard pass.

Re: The business of extracting knowledge from academic publications

#88

I can confirm that in my current area of interest (how to synthesize a cello or saxophone sound), there are hundreds of academic papers published over decades, each of them says "our method sounds more realistic than others", but code and audio samples are never available, and verbal descriptions always skip crucial details. I have no doubt that academics have a ton of expertise, but their output in paper form is bas…

I may be a bit cynical, but at least in my former field (experimental physics), the main purpose of papers seems to be to "lock in" a finished achivement. You do the actual research, pass internal reviews and peer review, and then publishing the paper is just to make it "official". Unfortunately, many papers are never expected to be read. The crucial information exists, but you usually get it from personal communicat…

Counterpoint: citations are a valuable currency in science. Arguably one of the best ways to earn citations is to do good work and write clear papers.

Not saying incentives are perfectly aligned -- many citations are superficial ("this topic was studied before"), and papers count for a lot even if they're never cited, etc

Re: The business of extracting knowledge from academic publications

#89
post #53
post #39

Earlier quoted context omitted.

I'm an academic (applied math) and want to respond to this: academic papers are the way they are for lots of reasons, many of which (not so good) have been mentioned on HN. There are a couple that I do not see very often however: (1) Many academics aren't aware non-academics read their papers at all: we work with other academics, go to conferences with other academics, and on the rare occasions we hear from readers,…

> What can help make the literature more accessible? Review articles, sometimes called surveys. I've always thought that new PhDs would be excellent authors for those, having digested lots literature for their dissertation.

Also, I believe there is a hierarchy that goes something like: academic papers -> review articles -> specialized books -> text books.

The text changes to fit the audience, and the knowledge becomes more accepted (and or fundamental) further down the line.

Re: The business of extracting knowledge from academic publications

#90

Earlier quoted context omitted.

>And I got told in no uncertain terms what is the gist of this article: The value lies in the molecular and clinical data. In 2021 I would add digital pathology / imaging data. I feel like you are trying to tell me something REALLY valuable, but I don't quite understand it. Can you please elaborate?

My take: answering questions using clinical data > answering questions with papers

There is immense value in clinical data (all the information captured and siloed through EHR). Pharma companies pay for access to it to gather real-world evidence (RWE) how, for example, their drug performs. Molecular information is increasingly valuable too for research, biomarker development, patient cohort identification etc. The imaging data and pathology data are valuable because they are typically expertly annotated and can be used to train computer-vision algorithms etc. to solve medical problems - like diagnosis.
Post reply on HN