That single change will speed the rate at which new results can be read and understood.
Retire This Idea: Scientific Knowledge Structured as “Literature”
61–70 of 88 posts
Re: Retire This Idea: Scientific Knowledge Structured as “Literature”
#62Author here. Delighted to see this near the top of HN this afternoon, and to see the healthy discussion it's provoked. Happy to answer any questions as best I can if you've got 'em.
Are you familiar with the field of bibliometrics?[0] It is a not-that smalish research area that works on quantitative research studies. Qualifying types of citations have been a pipe dream there for decades. There are a lot of scientific literature on the subject out there. [0] Or sciencemetrics, informetrics and other assorted subfields
For me one of the most salient things in that space -- albeit a different point than the one raised in the essay -- is the debate about what sorts of metrics are most appropriate (if any) for informing decisions about hiring, tenure, and the like. In some ways it's quite like the be-careful-what-you-wish-for problems any business has in choosing which metrics to care about. For instance, the h-index might shape a researcher's choice of whether to publish something as one larger paper or two smaller ones.
As Sam Altman put it, "It really is true that the company will build whatever the CEO decides to measure." I think a similar principle holds for research and scholarship.
Re: Retire This Idea: Scientific Knowledge Structured as “Literature”
#63Science works like that too, but on the blackboards and notebooks and today blogs and comments of the investigators, which lack real start and end dates and are regularly crossed out or replaced. The work is refined for publication because publications are essentially learning materials, they're written and edited to be as comprehensible as possible for any active researcher in the field to read for education or leisure; likewise in order for citations to be useful the cited works must include some background information which goes a long way in keeping the sciences connected, because it's not rare for analogous problems to pop up in unrelated fields, and it is crucial that researchers from outside each others' fields can read each others' work in order for interdisciplinary collaboration of this form to be possible. But all of this background information is just background noise to the people actively working on a problem and that's why it isn't on the blackboard or in the blog post, and it's why terse notations like dummy indices became popular in scratch work. The divide between the didactic and generative arrangements of scientific work is better explained thus: papers are binaries, not source code. Binaries don't go in source control.
Re: Retire This Idea: Scientific Knowledge Structured as “Literature”
#64Baby steps. Can we just ban using two-columns in PDFs? That single change will speed the rate at which new results can be read and understood.
Re: Retire This Idea: Scientific Knowledge Structured as “Literature”
#65I what is missing from the article is an understanding of the incentives of academics. Quite often putting a piece of research "to bed" is the goal. Why? Because any good academic has a long list of things they want to get to. Very few people want to get mired in the (inevitable) problems that arise in a paper for eternity. This is why there are review papers that summarize the state of the knowledge at a given point…
If papers were collaborative, one could simply propose his improvements directly where they fit in the reference paper, without having to write one from scratch, and readers would be immediately aware of these follow-up results, without having to search in dozens of papers.
Re: Retire This Idea: Scientific Knowledge Structured as “Literature”
#66I what is missing from the article is an understanding of the incentives of academics. Quite often putting a piece of research "to bed" is the goal. Why? Because any good academic has a long list of things they want to get to. Very few people want to get mired in the (inevitable) problems that arise in a paper for eternity. This is why there are review papers that summarize the state of the knowledge at a given point…
Most of the great scientists worked on a single idea their entire life and expanded upon it, improved it, corrected it and were never really done with it (einstein comes to mind). Scientific knowledge is never final I personally have a huge beef with the way life scientists publish their results in tiny tiny bits, that makes it extremely hard to cross-check them with other studies, to find out if they were later disp…
Re: Retire This Idea: Scientific Knowledge Structured as “Literature”
#67Re: Retire This Idea: Scientific Knowledge Structured as “Literature”
#68His approach may be a perfect fit for an experiment in meta-research. You could run periodic reviews of articles released in specific topical areas and create summaries of the findings. Any time those findings change, its a git commit, and you can track this change over time. It's something like Wikipedia, but designed for research. I would not be surprised if this already exists.
I see a challenge in figuring out the document structure which will be the most conducive to distributed version control (pull requests, etc), and can provide some insights on a historicalbasis. For example, you could run these summaries retrospectively, say looking at DNA research in the 1960s, and do a separate commit for every key finding through the years. It seems to me that you would pretty much have to specify a specialized coding language for scientific knowledge for this structure to work.
Re: Retire This Idea: Scientific Knowledge Structured as “Literature”
#69Earlier quoted context omitted.
Every paper has problems. Yes, even typos should be mentioned. Why not? They might trip up someone. Perhaps instead of structuring the whole thing around papers, it should be something like Wikipedia and editing articles in it.
I still don't see a concrete explanation of how this would work. Instead, I see you tossing off possible solutions that don't pass even a basic test of feasibility. Is bug tracker activity a sign of a good paper? Or a poor paper? The use-case I objected to, from the original article, is: > A dependency graph would tell us, at a click, which of the pillars of scientific theory are truly load-bearing. And it would tell…
What one should aim for is improving research quality and productivity.
Re: Retire This Idea: Scientific Knowledge Structured as “Literature”
#70I'm not sure why there's this holy grail of "the unified master version of" whatever.
Let me give an example. Say I write a paper on the shape of the non-dark-matter (stellar) density in the milky way by looking at y-type stars and I get an answer x. Now Bob comes along and looks at y2-type stars and gets answer x2. People have the idea that you just go to the first paper, add a footnote to y and x showing alternate values for y2 and x2...
But what that doesn't take into account is the fact that I used telescope a (a 10 meter hawaiian behemoth) and pointed it in one beam of the sky for 8 hours to get an ultra-deep pencil; but bob used telescope a2 (a modest 1.8 meter in la palma) that took an all sky survey and only goes very shallow. Now we add this in a footnote.
Next, there's a critical difference in the stars we studied. My y-type stars take 8 Gyr just to form, but Bob was using y2-type stars which live anywhere from 100 Myr to 15 Gyr. So I'm looking at the old stars and he's looking at all the stars. Since we know that different age stars live in different parts of the galaxy (old in the halo, young in the disk and bulge), our results are starting to look not as comparable as we thought... but it's minor, we'll add a footnote.
But then we realize that, since my old stars are giants and his all age stars are dwarfs, my stars are way brighter than his. Since my telescope is monstrous, and his is a small surveyor, my stars actually end up being observed to a distance 10 times that of his sample. In fact at these distances, the original model is a bad fit and we need to change from a power law to an Einasto profile. Bob can do that too, so our answers are easily comparable, but the Einasto law has more parameters so it would give a worse fit per parameter value than the power law he wanted to use originally... We add an appendix to the paper to explain this bit.
Then I notice that Bob's been using infrared data, and in the infrared there's a well known problem separating stars and galaxies in the data on telescope a2. In fact, Bob has to write a whole new section on some probabilistic tests and models he uses to adequately remove these galaxies from his y2 star sample. My telescope, observing in the optical at high resolution, has no such problem, so that section doesn't exist in my paper. Bob looks around awkwardly and stickies a hyperlink to his meta-analysis somewhere in my data section.
Then Jill comes along and says she doesn't agree at all with us; she got value x3 using the distribution of dwarf galaxies and if you believe in theory z, then _hers_ is the most accurate answer.
And we tell Jill to go write her own fucking paper.