Live data from Hacker News

Retire This Idea: Scientific Knowledge Structured as “Literature”

edge.org

61–70 of 88 posts

Re: Retire This Idea: Scientific Knowledge Structured as “Literature”

#62
post #54
post #43

Author here. Delighted to see this near the top of HN this afternoon, and to see the healthy discussion it's provoked. Happy to answer any questions as best I can if you've got 'em.

Are you familiar with the field of bibliometrics?[0] It is a not-that smalish research area that works on quantitative research studies. Qualifying types of citations have been a pipe dream there for decades. There are a lot of scientific literature on the subject out there. [0] Or sciencemetrics, informetrics and other assorted subfields

I'm aware of portions of that literature, but haven't studied it explicitly. What do you think are the most interesting things happening in that space at the moment?

For me one of the most salient things in that space -- albeit a different point than the one raised in the essay -- is the debate about what sorts of metrics are most appropriate (if any) for informing decisions about hiring, tenure, and the like. In some ways it's quite like the be-careful-what-you-wish-for problems any business has in choosing which metrics to care about. For instance, the h-index might shape a researcher's choice of whether to publish something as one larger paper or two smaller ones.

As Sam Altman put it, "It really is true that the company will build whatever the CEO decides to measure." I think a similar principle holds for research and scholarship.

Re: Retire This Idea: Scientific Knowledge Structured as “Literature”

#63
When all you have is a hammer, everything looks like a nail. The hubris on display here is: "it works for me, it must be right for everyone". What the author is missing is that the article-reference system is actually more flexible than a tree of dependencies. Consider the obvious example: citing Wikipedia. It's a pain, because Wikipedia changes, so you have to link to a certain revision and if you want to know the author you have to trawl the contributions, which nobody does. And this is true of Wikipedia because Wikipedia is a work in progress: it's not meant to be a presentation of knowledge but a repository of knowledge.

Science works like that too, but on the blackboards and notebooks and today blogs and comments of the investigators, which lack real start and end dates and are regularly crossed out or replaced. The work is refined for publication because publications are essentially learning materials, they're written and edited to be as comprehensible as possible for any active researcher in the field to read for education or leisure; likewise in order for citations to be useful the cited works must include some background information which goes a long way in keeping the sciences connected, because it's not rare for analogous problems to pop up in unrelated fields, and it is crucial that researchers from outside each others' fields can read each others' work in order for interdisciplinary collaboration of this form to be possible. But all of this background information is just background noise to the people actively working on a problem and that's why it isn't on the blackboard or in the blog post, and it's why terse notations like dummy indices became popular in scratch work. The divide between the didactic and generative arrangements of scientific work is better explained thus: papers are binaries, not source code. Binaries don't go in source control.

Re: Retire This Idea: Scientific Knowledge Structured as “Literature”

#64
post #61

Baby steps. Can we just ban using two-columns in PDFs? That single change will speed the rate at which new results can be read and understood.

Forget PDFs. JOVE and other video scientific journals should become the norm. Much less chance of fabrication and much better replication if you can demonstrate on video.

Re: Retire This Idea: Scientific Knowledge Structured as “Literature”

#65
post #5

I what is missing from the article is an understanding of the incentives of academics. Quite often putting a piece of research "to bed" is the goal. Why? Because any good academic has a long list of things they want to get to. Very few people want to get mired in the (inevitable) problems that arise in a paper for eternity. This is why there are review papers that summarize the state of the knowledge at a given point…

To me, the main problem with papers in their current shape is that they are required to be more or less self-contained. When one wants to state a result that improves a little bit the knowledge in a well established field, he has to waste time and space stating the definitions and preliminary results necessary to understand his result. This is counterproductive both for the author and the reader interested only in the small new bit of information.

If papers were collaborative, one could simply propose his improvements directly where they fit in the reference paper, without having to write one from scratch, and readers would be immediately aware of these follow-up results, without having to search in dozens of papers.

Re: Retire This Idea: Scientific Knowledge Structured as “Literature”

#66
post #11
post #5

I what is missing from the article is an understanding of the incentives of academics. Quite often putting a piece of research "to bed" is the goal. Why? Because any good academic has a long list of things they want to get to. Very few people want to get mired in the (inevitable) problems that arise in a paper for eternity. This is why there are review papers that summarize the state of the knowledge at a given point…

Most of the great scientists worked on a single idea their entire life and expanded upon it, improved it, corrected it and were never really done with it (einstein comes to mind). Scientific knowledge is never final I personally have a huge beef with the way life scientists publish their results in tiny tiny bits, that makes it extremely hard to cross-check them with other studies, to find out if they were later disp…

[deleted]

Re: Retire This Idea: Scientific Knowledge Structured as “Literature”

#68
It seems the problem is that the author is trying to create a representation of knowledge, whereas research papers are more of a logging of work done.

His approach may be a perfect fit for an experiment in meta-research. You could run periodic reviews of articles released in specific topical areas and create summaries of the findings. Any time those findings change, its a git commit, and you can track this change over time. It's something like Wikipedia, but designed for research. I would not be surprised if this already exists.

I see a challenge in figuring out the document structure which will be the most conducive to distributed version control (pull requests, etc), and can provide some insights on a historicalbasis. For example, you could run these summaries retrospectively, say looking at DNA research in the 1960s, and do a separate commit for every key finding through the years. It seems to me that you would pretty much have to specify a specialized coding language for scientific knowledge for this structure to work.

Re: Retire This Idea: Scientific Knowledge Structured as “Literature”

#69
post #48
post #38

Earlier quoted context omitted.

Every paper has problems. Yes, even typos should be mentioned. Why not? They might trip up someone. Perhaps instead of structuring the whole thing around papers, it should be something like Wikipedia and editing articles in it.

I still don't see a concrete explanation of how this would work. Instead, I see you tossing off possible solutions that don't pass even a basic test of feasibility. Is bug tracker activity a sign of a good paper? Or a poor paper? The use-case I objected to, from the original article, is: > A dependency graph would tell us, at a click, which of the pillars of scientific theory are truly load-bearing. And it would tell…

Why do we need all things to indicate paper quality? A lot of bug tracker activity might be both a good or bad sign, but more important that's not relevant.

What one should aim for is improving research quality and productivity.

Re: Retire This Idea: Scientific Knowledge Structured as “Literature”

#70
This solution is like trying to stir a pot of chili with an absinthe spoon.

I'm not sure why there's this holy grail of "the unified master version of" whatever.

Let me give an example. Say I write a paper on the shape of the non-dark-matter (stellar) density in the milky way by looking at y-type stars and I get an answer x. Now Bob comes along and looks at y2-type stars and gets answer x2. People have the idea that you just go to the first paper, add a footnote to y and x showing alternate values for y2 and x2...

But what that doesn't take into account is the fact that I used telescope a (a 10 meter hawaiian behemoth) and pointed it in one beam of the sky for 8 hours to get an ultra-deep pencil; but bob used telescope a2 (a modest 1.8 meter in la palma) that took an all sky survey and only goes very shallow. Now we add this in a footnote.

Next, there's a critical difference in the stars we studied. My y-type stars take 8 Gyr just to form, but Bob was using y2-type stars which live anywhere from 100 Myr to 15 Gyr. So I'm looking at the old stars and he's looking at all the stars. Since we know that different age stars live in different parts of the galaxy (old in the halo, young in the disk and bulge), our results are starting to look not as comparable as we thought... but it's minor, we'll add a footnote.

But then we realize that, since my old stars are giants and his all age stars are dwarfs, my stars are way brighter than his. Since my telescope is monstrous, and his is a small surveyor, my stars actually end up being observed to a distance 10 times that of his sample. In fact at these distances, the original model is a bad fit and we need to change from a power law to an Einasto profile. Bob can do that too, so our answers are easily comparable, but the Einasto law has more parameters so it would give a worse fit per parameter value than the power law he wanted to use originally... We add an appendix to the paper to explain this bit.

Then I notice that Bob's been using infrared data, and in the infrared there's a well known problem separating stars and galaxies in the data on telescope a2. In fact, Bob has to write a whole new section on some probabilistic tests and models he uses to adequately remove these galaxies from his y2 star sample. My telescope, observing in the optical at high resolution, has no such problem, so that section doesn't exist in my paper. Bob looks around awkwardly and stickies a hyperlink to his meta-analysis somewhere in my data section.

Then Jill comes along and says she doesn't agree at all with us; she got value x3 using the distribution of dwarf galaxies and if you believe in theory z, then _hers_ is the most accurate answer.

And we tell Jill to go write her own fucking paper.

Post reply on HN