Live data from Hacker News

The business of extracting knowledge from academic publications

markusstrasser.org

1–10 of 122 posts

Re: The business of extracting knowledge from academic publications

#3
One thing that strikes me about most academic knowledge tools is that they seem to focus on parsing the current set of academic literature and producing supposedly interesting insights out of them (which quickly tends to snowball into wanting some kind of generalized model for knowledge as a whole). What I think is much more interesting is creating tools that help people create better academic writing in the first place (thinking tools if you will). This is however much more a UX problem rather than it being a pure engineering problem. That is why I think we see many more tools in the knowledge extraction space as most academics thinking about these kind of things probably have an engineering background. That combined with the the fact that we seemingly all want to throw machine learning at any problem we encounter.

Re: The business of extracting knowledge from academic publications

#4
Interesting but I don't know how to make sense of it. How can it be that "close to nothing of what makes science actually work is published as text on the web"?

- Is the information that makes science actually work mostly in images that the machines don't yet understand?

- Was the information paywalled or in private databases and inaccessible to this researcher?

- Are the papers mostly just advertisements for researchers to come gab with each other at conferences and doodle on cocktail napkins, and that's where all the "real science" happens?

- (From the comments) is the information needed to make sense of papers communicated privately or orally from PI's to postdocs and grad students, or within industrial research labs?

Something is missing from my mental picture here.

Don't real scientists mostly learn how to think about their fields by reading textbooks and papers? (This is a genuine question.) If so, isn't it likely that our tools just aren't advanced enough to learn like humans do? If not, what do humans use to learn how to think about science that's missing from textbooks and papers?

Re: The business of extracting knowledge from academic publications

#5
post #3

One thing that strikes me about most academic knowledge tools is that they seem to focus on parsing the current set of academic literature and producing supposedly interesting insights out of them (which quickly tends to snowball into wanting some kind of generalized model for knowledge as a whole). What I think is much more interesting is creating tools that help people create better academic writing in the first pl…

I recall seeing a Show HN post a while back about a research focussed web browser that helps as a thinking tool:

https://news.ycombinator.com/item?id=28446147

Re: The business of extracting knowledge from academic publications

#6
post #3

One thing that strikes me about most academic knowledge tools is that they seem to focus on parsing the current set of academic literature and producing supposedly interesting insights out of them (which quickly tends to snowball into wanting some kind of generalized model for knowledge as a whole). What I think is much more interesting is creating tools that help people create better academic writing in the first pl…

To your point but even more general: the ML/AI space is far too focused on replacing people rather than helping people. There is a suffocating cultural conceit that we are on the verge of general AI and oh my gosh what will the humans do, we better institute universal basic income right away, etc.

What a joke.

Try to help humans think better first. If you succeed at that, you might be on the right track towards developing cold fusion, er, general AI.

Re: The business of extracting knowledge from academic publications

#7
post #3

One thing that strikes me about most academic knowledge tools is that they seem to focus on parsing the current set of academic literature and producing supposedly interesting insights out of them (which quickly tends to snowball into wanting some kind of generalized model for knowledge as a whole). What I think is much more interesting is creating tools that help people create better academic writing in the first pl…

It likely wouldn’t take much to craft an ”arXiv Copilot” out of GitHub Copilot.

Re: The business of extracting knowledge from academic publications

#8

Interesting but I don't know how to make sense of it. How can it be that "close to nothing of what makes science actually work is published as text on the web"? - Is the information that makes science actually work mostly in images that the machines don't yet understand? - Was the information paywalled or in private databases and inaccessible to this researcher? - Are the papers mostly just advertisements for researc…

As the article states, papers are mostly career advancement tools and scientists are incentivized to put the least amount of useful information into them they can get away with. Real scientists mostly learn from their instructors who possess all the jealously guarded institutional knowledge.

Yes, it is very broken.

Re: The business of extracting knowledge from academic publications

#9
post #8

Interesting but I don't know how to make sense of it. How can it be that "close to nothing of what makes science actually work is published as text on the web"? - Is the information that makes science actually work mostly in images that the machines don't yet understand? - Was the information paywalled or in private databases and inaccessible to this researcher? - Are the papers mostly just advertisements for researc…

As the article states, papers are mostly career advancement tools and scientists are incentivized to put the least amount of useful information into them they can get away with. Real scientists mostly learn from their instructors who possess all the jealously guarded institutional knowledge. Yes, it is very broken.

It is always difficult to try to understand and implement the theory explained in papers which seem fine on the surface, but when you look more closely, you find a bunch of mistakes, there are giant holes in the details, and you end up trying to redo the whole paper.

There should be journals/websites/blogs dedicated to trying to reexplain / implement papers.

Re: The business of extracting knowledge from academic publications

#10
post #8

Interesting but I don't know how to make sense of it. How can it be that "close to nothing of what makes science actually work is published as text on the web"? - Is the information that makes science actually work mostly in images that the machines don't yet understand? - Was the information paywalled or in private databases and inaccessible to this researcher? - Are the papers mostly just advertisements for researc…

As the article states, papers are mostly career advancement tools and scientists are incentivized to put the least amount of useful information into them they can get away with. Real scientists mostly learn from their instructors who possess all the jealously guarded institutional knowledge. Yes, it is very broken.

It's hard for me to even comprehend how this could be true, but it does sound familiar enough from credible sources that maybe it's right regardless of what makes sense to me.
Post reply on HN