Live data from Hacker News

Let’s Publish Everything

statmodeling.stat.columbia.edu

21–29 of 29 posts

Re: Let’s Publish Everything

#21
post #13
post #7

Earlier quoted context omitted.

(I agree with the parent poster, but this seems like the right place to add the following.) One thing that people who claim "the scarcity is gone, let's just publish everything" do miss is that there still is a scarcity: The time and attention of researchers. Let me expand a bit on that from my own field of studying reconnection in astrophysical plasma. As you can tell from the description it is not a well defined cl…

Great insight, thanks! It sounds like AI could play a part here in order to achieve the required efficiency, e.g. sorting out the right papers for you.

All I can say is that all the feeds that currently promise to solve that problem (Nasa ADS, Google Scholar, Researchgate and the feeds of several journals) are all pretty bad. I get several suggestions for medical/biophysical papers every day (probably due to "plasma" and "particle in _C_ell").

Re: Let’s Publish Everything

#22
post #14
post #7

Earlier quoted context omitted.

(I agree with the parent poster, but this seems like the right place to add the following.) One thing that people who claim "the scarcity is gone, let's just publish everything" do miss is that there still is a scarcity: The time and attention of researchers. Let me expand a bit on that from my own field of studying reconnection in astrophysical plasma. As you can tell from the description it is not a well defined cl…

The same is true within IT, you need to read a lot to make sure you're in the loop. I think it correlates to map structures, if every point in a map is connected to every other point there will be a large surface for each point in the beginning but as the amount of points grow the surface (of possible attention in this case) will shrink. How much area of attention is needed to perform the best in a creative research,…

Even if every paper is amazing, I still don't have the time to read about all topics. So even if every submitted paper is great and should be published, we still need sorting by topic.

Add to that the fact that "publish or perish" has lead to a decline in at least amount of "actually new information" per paper if not outright in quality of papers. And that is not going away easily.

Re: Let’s Publish Everything

#23
post #4

The same thing is trying to happen in scientific publication that's been happening in every other published medium: scarcity is gone. So why not embrace the abundance? Well if there's any contrary case to be made, it's that the explosion in quantity inevitably reduces the average quality. In the arts, low quality corresponds to that ineffable quality of shittiness . In news, it's fakeness. In science, it's irreproduc…

[deleted]

Re: Let’s Publish Everything

#24
post #4

The same thing is trying to happen in scientific publication that's been happening in every other published medium: scarcity is gone. So why not embrace the abundance? Well if there's any contrary case to be made, it's that the explosion in quantity inevitably reduces the average quality. In the arts, low quality corresponds to that ineffable quality of shittiness . In news, it's fakeness. In science, it's irreproduc…

This is a complex topic and I think Gelman (the linked poster) is either misinterpreting and/or confused by Kaufman and Glǎveanu (the article he's discussing). Just for some context, I agree with Gelman and KG both in their main arguments.

There's different issues colliding in the open science movement. One is what you're referring to, the fact that scarcity is gone. Combined with the overcrowding and hypercompetitiveness of science, you not only have what you are referring to, a decrease in verdicality, you also have a decrease in signal-to-noise in general. So the nonsense increases, but so do the traditional signals to quality. So nonsense appears in high-profile journals, and very quality work appears in low-profile journals or even as "unpublished" pieces.

The other problem though, what my colleagues refer to as the "science police", is an increasing tendency for certain groups to argue that a certain set of practices are not only good, but necessary for "good" science, and by implication, everything that does not is "bad" science, in a black-and-white kind of way.

For one thing, not all problems are with replication. Nonsense can replicate well, and very important legitimate phenomena can be difficult to replicate. If something is really not replicable at all, that's a problem, but replicability per se is only one part of scientific progress, and it comes in degrees with various causes.

It's also much more difficult to determine what is replicable sometimes than it might seem on the surface. Replicate what? What's important to replicate? How? Sometimes this is clear, but other times it is not.

Also, when you really delve into it, there's not really a good rationale for what, exactly, are the important ingredients for open science, or why. For example, is it really necessary to have preregistered studies? What's to keep someone from preregistering but then silently declining to publish null results? Or to "preregister" something they've already collected? If an important unanalyzed pre-existing dataset becomes available, is that "tainted" because it wasn't preregistered? Is it important to preregister, or just to make the data openly available? Is it better to use modeling to identify anomalies in studies, or to rely on preregistration? These issues aren't always clear.

I think there's a sense sometimes that the open science movement is not only trying to dismantle a broken system run by an established elite, but to replace it dogmatically with a new system run by a new elite, with its own imperfect rules. Already I've seen misuses of open science guidelines used to bully and discredit legitimate work (for example, by suggesting that someone is hiding something by not sharing data, when the data contains protected healthcare information and would be accessible to them anyway if they would just go through proper channels). This is tricky to discuss, as you might imagine, so it comes out in pieces like KG's piece. Gelman is asking "why not publish everything", which is responding (I think) to something different from what KG are responding to. Maybe I'm misreading KG, but I think they might also argue "why not publish everything"; they just have a different group they're addressing when they would say that.

Re: Let’s Publish Everything

#25
post #14

Earlier quoted context omitted.

The same is true within IT, you need to read a lot to make sure you're in the loop. I think it correlates to map structures, if every point in a map is connected to every other point there will be a large surface for each point in the beginning but as the amount of points grow the surface (of possible attention in this case) will shrink. How much area of attention is needed to perform the best in a creative research,…

Even if every paper is amazing, I still don't have the time to read about all topics. So even if every submitted paper is great and should be published, we still need sorting by topic. Add to that the fact that "publish or perish" has lead to a decline in at least amount of "actually new information" per paper if not outright in quality of papers. And that is not going away easily.

Yeah, I agree that force publishing something which is not ready is always bad.

Maybe publishing in itself is a way that does not scale well. Text has always been quite the slow way of making progress. And now with VR we have better ways to visualize and learn from others work.

How would you like to work, given free time and funding? Would you like to be able to keep track of every related research?

Re: Let’s Publish Everything

#26
post #7

Earlier quoted context omitted.

(I agree with the parent poster, but this seems like the right place to add the following.) One thing that people who claim "the scarcity is gone, let's just publish everything" do miss is that there still is a scarcity: The time and attention of researchers. Let me expand a bit on that from my own field of studying reconnection in astrophysical plasma. As you can tell from the description it is not a well defined cl…

Interesting. This seems to correlate almost exactly with discovery in music streaming too. Given the torrent of new music available, how do you discover the new music that you like? At the moment, AI is doing a pretty bad job of this - my Spotify Discover Weekly is an interesting listen, but I know it's not the best selection of new music out there suited to my tastes (to be fair, it's not really trying to be that, t…

Way back it was always through word of mouth that one discovered things. I still hold it true. You may have 10 friends how have their nieche and once in a while they recommend something they think you want to listen to.

I think that the music matter more based on whom recommended it. An algorithm won't have the same effect. It's missing the storyline on how you ended up with watching x movie or listening to y song.

Re: Let’s Publish Everything

#27
post #11

I am in sympathy with the author's goals but the major problem with publishing everything on arXiv like forums is that: (i) it becomes impossible to sort out the good stuff from the nonsense produced by cranks, and (ii) it unfairly advantages "high profile" groups. Today, if I have a grad student interested in security who wants some ideas for things to work on, I could ask them to go look up papers in Oakland, CCS a…

> I could ask them to go look up papers in Oakland, CCS and NDSS over the last couple of years and see if anything catches their fancy. Isn't this exactly the major thing that is wrong with research today? Limiting work/creativity to a few well known conferences done by elites for elites? I read blog posts, posted daily here on HN, that are way more informative, honest, and replicable than many papers published in th…

> Isn't this exactly the major thing that is wrong with research today? Limiting work/creativity to a few well known conferences done by elites for elites? I read blog posts, posted daily here on HN, that are way more informative, honest, and replicable than many papers published in the three conferences you named.

I totally disagree. The quality of papers at the "elite" conferences is way higher than most things I've read on HN. What is an example of an HN post that in your opinion is better than equivalent academic research in that area?

> Elite status happens exactly because there are conferences like the ones you mentioned. If you work with an advisor that publishes in Oakland, your chances of getting a paper in Oakland gets increased multiplicatively. And hint, that's not because your ideas (or papers) are better than anybody else's.

Sure, having an advisor on the Oakland PC helps a great deal. But it doesn't follow that your work is just the same as everyone else. Have you peer reviewed papers for these conferences? A majority of submissions, even at the "elite" conferences are just junk. That doesn't mean everything that gets published is not junk, but the stuff that does get published is significantly better than the average submission.

> Who cares? If the work is worth anything, people will cite it. If not, it will remain as is. Why does it matter? Why do you care if 10 people cited your work or 100 people if you are happy with the work?

Because the point of my research is not to sit in an ivory tower and produce academese that no one cares about. The goal is to have real impact on computer system design, and in my specific case, push practitioners towards methodologies that make systems more secure. That's not going to happen if no one reads our work.

Another way of looking at it is that a lot of our work is funded by taxpayer money. They aren't paying us to have fun proving lemmas that no one else cares about, the taxpayer would like us to produce research that results in tangible improvements in computer system design. In the system that we have today, the only way to have this tangible impact is to produce high quality papers that other people read, cite and build on top of.

Re: Let’s Publish Everything

#28
post #26

Earlier quoted context omitted.

Interesting. This seems to correlate almost exactly with discovery in music streaming too. Given the torrent of new music available, how do you discover the new music that you like? At the moment, AI is doing a pretty bad job of this - my Spotify Discover Weekly is an interesting listen, but I know it's not the best selection of new music out there suited to my tastes (to be fair, it's not really trying to be that, t…

Way back it was always through word of mouth that one discovered things. I still hold it true. You may have 10 friends how have their nieche and once in a while they recommend something they think you want to listen to. I think that the music matter more based on whom recommended it. An algorithm won't have the same effect. It's missing the storyline on how you ended up with watching x movie or listening to y song.

really good point. There are a number of bands that I listened to (and ended up liking) purely because of who told me about them.

Re: Let’s Publish Everything

#29
post #26

Earlier quoted context omitted.

Way back it was always through word of mouth that one discovered things. I still hold it true. You may have 10 friends how have their nieche and once in a while they recommend something they think you want to listen to. I think that the music matter more based on whom recommended it. An algorithm won't have the same effect. It's missing the storyline on how you ended up with watching x movie or listening to y song.

really good point. There are a number of bands that I listened to (and ended up liking) purely because of who told me about them.

I think the same applies to most stuff. Your idols influence alot. I see this as the best reason to not rely purely on algorithms.

Producing playlists is also a sort of art.

Post reply on HN