Live data from Hacker News

Updated practice for review articles and position papers in ArXiv CS category

blog.arxiv.org

131–140 of 250 posts

Re: Updated practice for review articles and position papers in ArXiv CS category

#131
post #5

So what they no longer accept is preprints (or rejects…) It’s of course a pretty big deal given that arXiv is all about preprints. And an accepted journal paper presumably cannot be submitted to arXiv anyway unless it’s an open journal.

For position (opinion) or review (summarizing state of art and often laden with opinions on categories and future directions). LLMs would be happy to generate both these because they require zero technical contributions, working code, validated results, etc.

So what? People are experimenting with novel tools for review and publication. These restrictions are dumb, people can just ignore reviews and position papers if they start proving to be less useful, and the good ones will eventually spread through word of mouth, just like arxiv has always worked.

Re: Updated practice for review articles and position papers in ArXiv CS category

#132
post #8

Maybe it's time for a reputation system. E.g. every author publishes a public PGP key along with their work. Not sure about the details but this is about CS, so I'm sure they will figure something out.

I got that suggestion recently talking to a colleague from a prestigious university. Her suggestion was simple: Kick out all non-ivy league and most international researchers. Then you have a working reputation system. Make of that what you will ...

Keep in mind the fabulous mathematical research of people like Perelman [1], and one might even count Grothendieck [2].

[1] https://en.wikipedia.org/wiki/Grigori_Perelman [2] https://www.ams.org/notices/200808/tx080800930p.pdf

Re: Updated practice for review articles and position papers in ArXiv CS category

#133
post #89

There is a general problem with rewarding people for the volume of stuff they create, rather than the quality. If you incentivize researchers to publish papers, individuals will find ways to game the system, meeting the minimum quality bar, while taking the least effort to create the most papers and thereby receive the greatest reward. Similarly, if you reward content creators based on views, you will get view maximi…

Sure, just as long as we don't blame LLMs.

Blame people, bad actors, systems of incentives, the gods, the devils, but never broach the fault of LLMs and their wide spread abuse.

Re: Updated practice for review articles and position papers in ArXiv CS category

#134
I have a hunch that most of the slop is not just on CS but specifically about AI. For some reason, a lot of people's first idea when they encounter an LLM is "let's have this LLM write an opinion piece about LLMs", as if they want to test its self-awareness or hack it by self-recursion. And then they get a medley of the learning data, which if they are lucky contains some technical explanations sprinkled in.

That said, AI-generated papers have already been spotted in other disciplines besides cs, and some of them are really obvious (arXiv:2508.11634v1 starts with a review of a non-existing paper). I really hope arXiv won't react by narrowing its scope to "novel research only"; in fact there is already AI slop in that category and it is harder to spot for a moderator.

("Peer-reviewed papers only" is mostly equivalent to "go away". Authors post on the arXiv in order to get early feedback, not just to have their paper openly accessible. And most journals at least formally discourage authors from posting their papers on the arXiv.)

Re: Updated practice for review articles and position papers in ArXiv CS category

#135

it's clearly not sutainable to have the main website hosting CS articles not having any reviews or restrictions. (Except for the initial invite system) There were 26k submission in october: https://arxiv.org/stats/monthly_submissions Asking for a small amount of money would probably help. Issue with requiring peer reviewed journals or conferences is the severe lag, takes a long time and part of the advantage of arxiv…

I'll add the amount should be enough to cover at least a cursory review. A full review would be better. I just don't want to price out small players. The papers could also be categorized as unreviewed, quick check, fully reviewed, or fully reproduced. They could pay for this to be done or verified. Then, we have a reputational problem to deal with on the reviewer side.

I don't know about CS, but in mathematics the vast majority of researchers would not have enough funding to pay for a good quality full review of their articles. The peer review system mostly runs on good will.

Re: Updated practice for review articles and position papers in ArXiv CS category

#136

Earlier quoted context omitted.

I didn't agree with this idea, but then I looked at how much HN karma you have and now I think that maybe this is a good idea.

Ignoring the actual proposal or user, just looking at karma is probably a pretty terrible metric. High karma accounts tend to just interact more frequently, for long periods of time. Often with less nuanced takes, that just play into what is likely to be popular within a thread. Having a Userscript that just places the karma and comment count next to a username is pretty eye opening.

Yes, HN should probably publish karma divided by #comments. Or at least show both numbers.

Re: Updated practice for review articles and position papers in ArXiv CS category

#137
post #17

Earlier quoted context omitted.

Curious for the state on things here. Can we reliably tell if a text was LLM generated? I just heard of a prof screening assignments for this, but not sure how that would work.

Of course there are people who will sell you a tool to do this. I sincerely doubt it's any good. But then again they can apparently fingerprint human authors fairly well using statistics from their writing, so what do I know.

There are tools that claim accuracies in the 95%-99% range. This is useless for many actual applications, though. For example, in teaching, you really need to not have false positives at all. The alternative is failing some students because a machine unfairly marked their work as machine-generated.

And anyway, those accuracies tend to be measured on 100% human-generated vs. 100% machine-generated texts by a single LLM... good luck with texts that contain a mix of human and LLM contents, mix of contents by several LLMs, or an LLM asked to "mask" the output of another.

I think detection is a lost cause.

Re: Updated practice for review articles and position papers in ArXiv CS category

#138
post #25

Earlier quoted context omitted.

Their name, orcid, and email isn't enough?

You can’t get an arXiv account without a referral anyway. Edit: For clarification I’m agreeing with OP

You can create an arXiv.org account with basically any email address whatsoever[0], with no referral. What you can't necessarily do is upload papers to arXiv without an "endorsement"[1]. Some accounts are given automatic endorsements for some domains (eg, math, cs, physics, etc) depending on the email address and other factors.

Loosely speaking, the "received wisdom" has generally been that if you have a .edu address, you can probably publish fairly freely. But my understanding is that the rules are a little more nuanced than that. And I think there are other, non .edu domains, where you will also get auto-endorsed. But they don't publish a list of such things for obvious reasons.

[0]: Unless things have changed since I created my account, which was originally created with my personal email address. That was quite some time ago, so I guess it's possible changes have happened that I'm not aware of.

[1]: https://info.arxiv.org/help/endorsement.html

Re: Updated practice for review articles and position papers in ArXiv CS category

#139
post #5

So what they no longer accept is preprints (or rejects…) It’s of course a pretty big deal given that arXiv is all about preprints. And an accepted journal paper presumably cannot be submitted to arXiv anyway unless it’s an open journal.

Isnt arxiv also a likely LLM traing ground?

google internally started working on "indexing" patent applications, materials science publications, and new computer science applications, more than 10 years ago. You the consumer / casual are starting to see the services now in a rush to consumer product placement. You must know very well that major mil around the world are racing to "index" comms intel and field data; major finance are racing to "index" transactions and build deeper profiles of many kinds. You as an Internet user are being profiled by a dozen new smaller players. arxiv is one small part of a very large sea change right now

Re: Updated practice for review articles and position papers in ArXiv CS category

#140

The review paper is dead... so this is a good development. Like you can generate these things in a couple of iterations with AI and minor edits. Preprint servers could be dealing with 1000s of review/position papers over short periods of time and then this wastes precious screening work hours. It is a bit different in other fields where interpretations or know-how might be communicated in a review paper format that i…

What are review papers for anyway? I think they are either for 1) new grad students to end up with something nice to publish after reviewing the literature or, 2) older professors to write a big overview of everything that happened in their field as sort of a “bible” that can get you up to speed The former is useful as a social construct; I mean, hey, new grad students, don’t skimp on your literature review. Finding…

Ultimately, a key reason to write these papers in the first place is to guide practitioners in the field, right? Otherwise science itself is just a big (redacted term that can get people shadow-banned for simply using it).

As one of those practitioners, I've found good review/survey papers to be incredibly valuable. They call my attention to the important publications and provide at least a basic timeline that helps me understand how the field has evolved from the beginning and what aspects people are focusing on now.

At the same time, I'll confess that I don't really see why most such papers couldn't be written by LLMs. Ideally by better LLMs than we have now, of course, but that could go without saying.

Post reply on HN