Live data from Hacker News

Publishers withdraw more than 120 gibberish papers

nature.com

21–30 of 54 posts

Re: Publishers withdraw more than 120 gibberish papers

#21
post #7

The papers in question were generated by the SCIGen paper generator. Jeremy Stribling, Dan Aguayo and I are the authors of that system, with Jeremy doing the lionshare of the work. The inspiration for SCIgen was TheSpark.com's high school paper generator, which ran from about 1999 to 2001. To call these autogenerated works gibberish is an affront to those of us who hand-assembled their context-free grammars!

This is great work. I'm curious about your motivations. Do you consider it a validation of your work that so many papers were published?

At the time, there was an arms race within the systems community to see who could flip a bigger bird to the organizers of the SCI spamference. David Mazieres and Eddie Kohler started it off when they submitted "Get Me Off Your Fucking Mailing List" [1] for publication. We wanted to one-up them and figured we could do so by submitting a random paper that might be accepted into the conference, allowing us to give an in-person talk. The plan worked up to a point, but when we asked the internet for $3k to send us to Orlando, the press and then the organizers of SCI caught wind and revoked the random paper they had previously accepted.

Everything since has been gravy. We're pumped that SCIgen is still a useful weapon against charlatans the world over (this now includes you, IEEE and Springer).

[1] http://www.scs.stanford.edu/~dm/home/papers/remove.pdf

Re: Publishers withdraw more than 120 gibberish papers

#23

It may not be long before passable high school and college essays can be undetectably generated by software. I wonder what we'll do then. At some point, I guess we'll have writer software that pretty much generates nicely articulated narrative for whatever it is we wish to say. "Write up something about global warming. Make it borderline skeptical, quote the Concerned Scientists and so forth. You know what to say."

Happily (?) high school and college essays (and probably potential scientific publications) will also be read chiefly by software in the near future, just as GRE essays are graded in part by software today. Whichever side has the best algorithm wins.

Re: Publishers withdraw more than 120 gibberish papers

#24

Buffalo buffalo Buffalo buffalo buffalo buffalo Buffalo buffalo? http://en.wikipedia.org/wiki/Buffalo_buffalo_Buffalo_buffalo...

Sometimes, but at others: Buffalo buffalo buffalo Buffalo buffalo Buffalo buffalo buffalo. [1]

Similarly, buffalo buffalo Buffalo buffalo buffalo buffalo Buffalo buffalo buffalo buffalo. [2]

________________________________

Notes:

1. C n v C n C n v

2. n v C n v v C n n v

Re: Publishers withdraw more than 120 gibberish papers

#26
post #19
post #16

Earlier quoted context omitted.

>Since conference proceedings are typically ignored researchers generally won't put their best effort into these papers. It's interesting, because this is completely the opposite in many CS fields (e.g. security). The top conferences carry much more weight than any journal. It's probably related to how fast the field changes.

I'm not in CS, but I've heard of this. It's definitely aberrant. Some other fields change just as rapidly so I'd just chalk this up as a cultural difference.

In many CS fields, workshops are the equivalent of conferences in other fields such as physics. It is important to judge each field on how the primary peer-review publication is done, not by some abstract rule that is supposed to be the only valid standard.

As a side note, I've heard people from the humanities scoff at physicists and their belief that journal papers are worth mentioning. The only real publications are books.

Re: Publishers withdraw more than 120 gibberish papers

#27
post #16
post #14

Conference proceedings (not to be confused with a handful of actual journals that just contain "proceedings" in their name) can be thought of like the paper programs you might be handed at a play. They're usually meant to be an accompaniment to the event rather than a stand-alone publication. Their intent is to collect contributions from conference participants into a handy reference that attendees can use if they wa…

>Since conference proceedings are typically ignored researchers generally won't put their best effort into these papers. It's interesting, because this is completely the opposite in many CS fields (e.g. security). The top conferences carry much more weight than any journal. It's probably related to how fast the field changes.

HCI and graphics are other examples: a SIGCHI proceedings paper carries way more weight than a paper in one the HCI journals. And a SIGGRAPH proceedings paper is better than than a paper in something like Journal of Computer Graphics Techniques, although they've recently converted the SIGGRAPH proceedings into a special issue of ACM Transactions on Graphics (which is now more prestigious than other issues of that journal, because it's really a proceedings volume), mostly to avoid the confusion from other fields that don't consider conferences to be publications. Those two conferences actually have nearly exclusive prestige in their field, to the extent that they're the only top publication venue: if you're working in HCI or graphics you more or less need SIGCHI or SIGGRAPH publications to be considered a legit active researcher. They also tend to be more selective in peer-review than the journals are, with acceptance rates in the 10-15% range, while journals are usually more like 20-35%.

In AI I think it's a bit mixed, and maybe a US—European difference. Europeans tend to value journal publications in places like JAIR, JMLR, various IEEE Transactions, etc. more. Americans also value the best of those journals, but I get the impression that conference publications at places like NIPS, AAAI, and ICML are at least as good, and definitely better than a lower-tier journal.

Re: Publishers withdraw more than 120 gibberish papers

#28
post #16
post #14

Conference proceedings (not to be confused with a handful of actual journals that just contain "proceedings" in their name) can be thought of like the paper programs you might be handed at a play. They're usually meant to be an accompaniment to the event rather than a stand-alone publication. Their intent is to collect contributions from conference participants into a handy reference that attendees can use if they wa…

>Since conference proceedings are typically ignored researchers generally won't put their best effort into these papers. It's interesting, because this is completely the opposite in many CS fields (e.g. security). The top conferences carry much more weight than any journal. It's probably related to how fast the field changes.

That's because it takes a long time to publish in a journal (like 1.5 years or longer), and by that time the topic is old. Publishing in a conference allows you to bring fresh-from-the-oven results to your colleagues, along with the freshly-baked bread smell ;)

Re: Publishers withdraw more than 120 gibberish papers

#29
post #19
post #16

Earlier quoted context omitted.

>Since conference proceedings are typically ignored researchers generally won't put their best effort into these papers. It's interesting, because this is completely the opposite in many CS fields (e.g. security). The top conferences carry much more weight than any journal. It's probably related to how fast the field changes.

I'm not in CS, but I've heard of this. It's definitely aberrant. Some other fields change just as rapidly so I'd just chalk this up as a cultural difference.

My read on it is that the traditional approach in CS was:

1. Non-peer-reviewed "workshops" or "symposia". Place for position papers, work-in-progress, etc. Typically screened only for basic suitability for the event. In some cases authors may only submit an abstract of their topic, not a full paper.

2. Peer-reviewed but non-archival "conferences". Place to disseminate properly written up recent work in relatively short form [6-12 pages double-column], get feedback from other researchers, etc. Typically peer-reviewed by 2-4 reviewers, on the basis of a full paper (not abstracts like in some fields). The top 10-40% of submitted papers are accepted for inclusion (depending on the conference's selectivity), and then accepted authors typically have a month or so to revise their papers on the basis of reviewer comments, before submitting the final version for publication in the proceedings.

3. Peer-reviewed archival "journals". Place to wrap up a project or line of work with a publication that takes 3 to 5 conference publications, and integrates them into a full writeup of the project, tied into a nice bow for posterity suitable for archiving in The Literature. Peer-reviewed in the same way as other fields. Typically longer, 20+ pages.

But what seems to have happened is nobody really reads #3, since all the new work comes out in #2, the peer-review standards of #2 are already high enough for practical filtering, and #3 is years behind. And then people (especially Americans, especially in certain fields) just stopped writing #3 as often altogether. The view seems to be that tying several conference papers into a bow might be nice from the perspective of having a sorted-out archival literature, but at the present time, your peers have already moved on, and there are diminishing returns from tying those 3 already-published SIGGRAPH papers into a summary journal article, compared to working on new research for the next SIGGRAPH paper instead. It's even seen a little as "milking" your previous work: if you take your conference publications and turn them into a journal article, some people will assume you're doing it because you need the journal line on your CV for some kind of tenure/hiring committee, since otherwise why would you be spending all this time "upgrading" old conference papers into a journal, rather than writing new conference papers on new research?

Re: Publishers withdraw more than 120 gibberish papers

#30
post #26
post #19

Earlier quoted context omitted.

I'm not in CS, but I've heard of this. It's definitely aberrant. Some other fields change just as rapidly so I'd just chalk this up as a cultural difference.

In many CS fields, workshops are the equivalent of conferences in other fields such as physics. It is important to judge each field on how the primary peer-review publication is done, not by some abstract rule that is supposed to be the only valid standard. As a side note, I've heard people from the humanities scoff at physicists and their belief that journal papers are worth mentioning. The only real publications ar…

Actually, there are workshops in many CS fields which are selective (not as much as big conferences, that's for sure) and as serious as conferences or journal on peer-reviewing.

The name "workshop" mainly means that the focus of the event is very specific and that the attendance is intended to be small so that real discussion with everyone can happen (contrary to big conferences where there are hundreds of people).

For example my last published paper was submitted to a workshop [1] and I had four different detailed reviews, which had nothing to envy to big conferences reviews I could have got for it. This is an example of a not very selective (because it has little submissions) workshop but which is very serious and would most certainly reject any computer generated paper.

Another example could be CHES [2] which is a workshop which grew big and is now the most prestigious conference (again, that means more prestigious than journals as in most CS fields) of its field (crypto hardware/side-channel security). In this case the name "workshop" sticks, but it is actually a full-fledged conference.

Yet another example of a small workshop [3] (still in the same field) which does serious peer reviewing (I had three looong and detailed reviews when I submitted my paper there). Last year they selected 8 papers for presentation the day of the workshop, and 4 of them were selected by a journal [4] for subsequent publication of an extended version (and after another small round of reviews since there were already the reviews of the workshop which were deemed serious enough by the journal's editors).

Maybe this is particular to the field of security (could it be due to the hacker subculture background it has in addition to its academic background?), but I guess my point is that it is not possible to generalize how the publication system works for all sciences, let alone for all academic research.

[1] http://pprew.org/

[2] http://www.chesworkshop.org/

[3] http://www.proofs-workshop.org/

[4] http://link.springer.com/journal/13389

Post reply on HN