The business of extracting knowledge from academic publications
101–110 of 122 posts
Re: The business of extracting knowledge from academic publications
#102I can confirm that in my current area of interest (how to synthesize a cello or saxophone sound), there are hundreds of academic papers published over decades, each of them says "our method sounds more realistic than others", but code and audio samples are never available, and verbal descriptions always skip crucial details. I have no doubt that academics have a ton of expertise, but their output in paper form is bas…
Others have commented as well but I will reinforce: their output is basically unusable for you for the purpose you want to put it to.
Which is fair, but you should also recognize that you are not the audience of the papers and for good or for ill the system is not set up to help you with this.
Re: The business of extracting knowledge from academic publications
#103Earlier quoted context omitted.
I'm an academic (applied math) and want to respond to this: academic papers are the way they are for lots of reasons, many of which (not so good) have been mentioned on HN. There are a couple that I do not see very often however: (1) Many academics aren't aware non-academics read their papers at all: we work with other academics, go to conferences with other academics, and on the rare occasions we hear from readers,…
> What can help make the literature more accessible? Review articles, sometimes called surveys. I've always thought that new PhDs would be excellent authors for those, having digested lots literature for their dissertation.
A well written PhD or MSc thesis is often the best way into a new field, ime. If the committee is good on this aspect they'll insist you've put enough detail in for someone to follow along mostly self contained.
Re: The business of extracting knowledge from academic publications
#104One thing that strikes me about most academic knowledge tools is that they seem to focus on parsing the current set of academic literature and producing supposedly interesting insights out of them (which quickly tends to snowball into wanting some kind of generalized model for knowledge as a whole). What I think is much more interesting is creating tools that help people create better academic writing in the first pl…
Re: The business of extracting knowledge from academic publications
#105Earlier quoted context omitted.
As the article states, papers are mostly career advancement tools and scientists are incentivized to put the least amount of useful information into them they can get away with. Real scientists mostly learn from their instructors who possess all the jealously guarded institutional knowledge. Yes, it is very broken.
Hard disagree. With a caveat -- I do acknowledge that for an important number of professional academics your statement may be true, and I have heard a former post-doc at ETH Zurich describe their papers as career points (so also a grain of truth at elite institutes). But for most of the academics I have known and worked with, publications are taken quite seriously, and institutional knowledge is freely shared. There…
This feels a lot like the graduate students on the Academia StackExchange who are convinced the moment they present their idea it will get stolen, while every faculty member is like "I have a list of my own ideas that I don't have time to work on as long as my arm."
Re: The business of extracting knowledge from academic publications
#106Earlier quoted context omitted.
As the article states, papers are mostly career advancement tools and scientists are incentivized to put the least amount of useful information into them they can get away with. Real scientists mostly learn from their instructors who possess all the jealously guarded institutional knowledge. Yes, it is very broken.
I mean honestly this is just total bullshit. There is plenty of value in academic papers. It's just that there is very little money to be made in developing tools such as those mentioned by the OP as there is very little money in academia.
Re: The business of extracting knowledge from academic publications
#107Interesting but I don't know how to make sense of it. How can it be that "close to nothing of what makes science actually work is published as text on the web"? - Is the information that makes science actually work mostly in images that the machines don't yet understand? - Was the information paywalled or in private databases and inaccessible to this researcher? - Are the papers mostly just advertisements for researc…
"Interesting but I don't know how to make sense of it. How can it be that "close to nothing of what makes science actually work is published as text on the web"?"
I'm not convinced this assertion is true. Difficult to parse by a non-expert? Sure. Often stored in pictures rather than text? Absolutely - indeed, a number of journals directly ask reviewers if the text and graphs are duplicative and consider this a negative. Harder to "disrupt" and monetize than many companies have expected it to be? Certainly.
"Is the information that makes science actually work mostly in images that the machines don't yet understand?"
In my mind, this is the most credible bit of the author's complaints. A lot of science in biomedicine is done in information dense graphics. The author picks especially hard to approach ones, but this is definitely a thing.
"Was the information paywalled or in private databases and inaccessible to this researcher?"
For an outsider without institutional access to journals, this can be a problem. More acutely, there is some lag between "What I'm currently working on" and "What's in the literature" simply because the literature is slow (and has gotten way slower during the pandemic).
"Are the papers mostly just advertisements for researchers to come gab with each other at conferences and doodle on cocktail napkins, and that's where all the "real science" happens?"
In biomedicine, papers are the product. Conferences tend to have a couple uses:
1) Previews of coming attractions - things I'm working on that are close enough to done to talk about, but not so close as to be making the rounds yet. These talks will often have less detail than a paper would, because we've all sat through a presentation that does into a ton of implementation detail and they're agonizing. Also I only have fifteen minutes.
2) Looking for postdocs - either from the hiring or seeking end.
3) Building professional networks - this is mostly so when someone comes to me with a problem, I know who might be working on say...causal inference with time-varying exposures...and can reach out to them. Usually to ask what papers I should read to get caught up. Or to bring them in on a paper/grand proposal.
4) Looking for problems other people are having that I can solve, and then reaching out.
"(From the comments) is the information needed to make sense of papers communicated privately or orally from PI's to postdocs and grad students, or within industrial research labs?"
Only insofar as my graduate students and postdocs have access to my time and expertise, and a job that is expressly meant to encourage understanding things. "I don't know, why don't you spend a couple weeks figuring out how they did that" is a perfectly good use of a graduate student's time, but something I find is rarely encouraged elsewhere.
"Don't real scientists mostly learn how to think about their fields by reading textbooks and papers? (This is a genuine question.) If so, isn't it likely that our tools just aren't advanced enough to learn like humans do? If not, what do humans use to learn how to think about science that's missing from textbooks and papers?"
One of the things that's likely missing, because those are all finished products, is the process. For example, I spent an hour chatting with a graduate student about three or four different ways they can approach their problem - what assumptions come with each one, tradeoffs, etc. But only the branch that actually got used is going to be published.
Re: The business of extracting knowledge from academic publications
#108I can confirm that in my current area of interest (how to synthesize a cello or saxophone sound), there are hundreds of academic papers published over decades, each of them says "our method sounds more realistic than others", but code and audio samples are never available, and verbal descriptions always skip crucial details. I have no doubt that academics have a ton of expertise, but their output in paper form is bas…
I'm an academic (applied math) and want to respond to this: academic papers are the way they are for lots of reasons, many of which (not so good) have been mentioned on HN. There are a couple that I do not see very often however: (1) Many academics aren't aware non-academics read their papers at all: we work with other academics, go to conferences with other academics, and on the rare occasions we hear from readers,…
Full disclosure: your username is very easy to Google, and I should tell you we work on closely related topics. My opinion is shaped in part by several of the papers you've cited. I'm happy to disclose my identity and continue the conversation in private.
You wrote: "(1) Many academics aren't aware non-academics read their papers at all"
What, then, is the point of applied mathematics? Please know that I don't mean that in a dismissive way. I think there are a few reasonable answers. Chief among them is the belief that exploratory research is important in its own right and does not need to have an immediate non-academic use, as embodied by this quote of Hadamard:
“Practical application is found by not looking for it, and one can say that the whole progress of civilization rests on that principle.”
I believe the above quote as far as it concerns mathematics. But how do we get from there to the applied part? I have an internal dissonance about this that goes deeper than just semantics. Before I started my PhD, I had some vague belief that after writing up some research with an algorithm in it, you'd put it on the arxiv, and from there someone might one day need something like that, code it up, and use it. If I could put in a basic working implementation that was even better.
All the evidence I've seen so far tells me this is not so. The truth is no one is going to take the time to code up your algorithm, because no one has dozens of hours to spend understanding your paper, developing an algorithm suitable for an industrial problem, often just to get improvement on a niche subset of cases. I've been wondering how to estimate the number of algorithms described on the arxiv that are ever implemented and used in a non-academic setting -- my bet is (outside of ML), less than 1%.
I've heard many times that sophisticated higher-order methods for PDEs (finite element / volume, Galerkin, ...) are used in aeronautics, to determine the wind shape over an airplane wing. I've found out from talking to people in the industry at companies like Bombardier that for the most part they do second order finite difference like the rest of us. Why? Because you can code it up in an afternoon, whereas writing the more sophisticated methods can take weeks or months. As academics, we think that the theoretical work is the really hard part, and we neglect the human cost of writing and maintaining algorithms. We have it backwards: academics are (relatively) cheap; code (and changing code) is expensive. (Of course, I make these comments assuming a certain scale. We can come back to this.)
I think the fundamental issue is that I know few applied mathematicians who start with a problem and seek out a solution. Most often, you finish your (applied) math PhD armed with some machinery. If you want to get a professorship and you've done well, you typically turn the crank of your particular machine better and faster than most. In applied math we can say our model is motivated by some problem in the sciences/economics/whatever, but in my experience that just lets us erect a straw person (create a problem) and tear it down (solve the problem) using the machinery that only we have mastered. Just because a problem is hard doesn't make it important.
What to do, then? How do you work on "consequential" problems?
To be pithy about it, I've found it useful to think in terms of $ rather than h-index. In many cases, a consequential problem is one that, if you solve, you can monetize. You could frame this as asking what kind of mathematics could enable new technologies. In my experience it is very difficult to write down a mathematical question that, if answered, can lead to new technology. But if you manage to find such a question -- and it is possible -- it can be a goldmine.
I have more to say -- especially about how mathematicians need to get a reality check on the importance of hardware and its relevance in stochastic algorithms research -- but this is long enough as it is, and I don't want to just be a crazy person rambling in the corner. I'd be very curious to hear your thoughts.
Re: The business of extracting knowledge from academic publications
#109I can confirm that in my current area of interest (how to synthesize a cello or saxophone sound), there are hundreds of academic papers published over decades, each of them says "our method sounds more realistic than others", but code and audio samples are never available, and verbal descriptions always skip crucial details. I have no doubt that academics have a ton of expertise, but their output in paper form is bas…
It doesn't have to be this way. Here's the process I use in my lab:
1. Every paper that makes a claim of any kind based on code contains a link to a public Git repo.
2. The paper contains the Git hash identifying the exact commit used to justify the claim. Copy-paste it from the paper into your checkout.
3. The repo may have moved on with fixes and improvements, you can have those too.
If you are using version control already, it's not much work to do this. Of course, you have to be committed to making your code public.