Live data from Hacker News

Updated practice for review articles and position papers in ArXiv CS category

blog.arxiv.org

231–240 of 250 posts

Re: Updated practice for review articles and position papers in ArXiv CS category

#232

Earlier quoted context omitted.

Genuinely curious, did we ever manage to ban a piece of technology worldwide and effectively?

Do chlorofluorocarbons (CFCs) mostly banned by the Montreal Protocol count?

And lead in gasoline, and probably quite a few other things where we found a way to get similar end results with fewer annoying side effects.

Re: Updated practice for review articles and position papers in ArXiv CS category

#233

Earlier quoted context omitted.

LLMs are tools that make it easier to hack incentives, but you still need a person to decide that they'll use an LLM t do so. Blaming LLMs is unproductive. They are not going anywhere (especially since open source LLMs are so good.) If we want to achieve real change, we need to accept that they exist, understand how that changes the scientific landscape and our options to go from here.

everyone keeps claiming "they're here to stay" as if it's gospel. this constant drumbeat is rather tiresome and without much hard evidence.

[deleted]

Re: Updated practice for review articles and position papers in ArXiv CS category

#234

Earlier quoted context omitted.

LLMs are tools that make it easier to hack incentives, but you still need a person to decide that they'll use an LLM t do so. Blaming LLMs is unproductive. They are not going anywhere (especially since open source LLMs are so good.) If we want to achieve real change, we need to accept that they exist, understand how that changes the scientific landscape and our options to go from here.

everyone keeps claiming "they're here to stay" as if it's gospel. this constant drumbeat is rather tiresome and without much hard evidence.

[flagged]

Re: Updated practice for review articles and position papers in ArXiv CS category

#235

Earlier quoted context omitted.

everyone keeps claiming "they're here to stay" as if it's gospel. this constant drumbeat is rather tiresome and without much hard evidence.

Genuinely curious, did we ever manage to ban a piece of technology worldwide and effectively?

A large part of geopolitics is concerned with limiting the spread of weapons of mass destruction worldwide and to the greatest possible degree of efficacy. Moreover, the investment to train state-of-the-art models is greater than the Manhattan project and involves larger and more complex supply chains-- it cannot be done clandestinely. Because the scope of the project is large and resource-intensive there are not many bodies that would have to cooperate in order to place impassable obstacles on the path that is presently being taken. 'What if they won't cooperate toward this goal?' -- Worth considering, but the fact is that they can and are choosing not to. If the choice is there it is not an inevitability but a decision.

Re: Updated practice for review articles and position papers in ArXiv CS category

#236

Earlier quoted context omitted.

the problem is generally the same as with generative adversarial networks; the capability to meaningfully detect some set of hallmarks of LLMs automatically is equivalent to the capability to avoid producing those, and LLMs are trained to predict (ie. be indistinguishable from) their source corpus of human-written text. so the LLM detection problem is (theoretically) impossible for SOTA LLMs; in practice, it could be…

Sure, having a 100% reliable system is impossible as you have laid out. However, if I understand the announcement correctly, this is about volume, and I wonder if you could have a tool flag articles that show obvious signs of LLM usage.

The point is that this leads to an arms race. If Arxiv uses a top-of-line LLM for, say, 20 minutes per paper, cheating authors will use a top-of-line LLM for 21 minutes to beat that.

Re: Updated practice for review articles and position papers in ArXiv CS category

#237

Earlier quoted context omitted.

Define the metric as "people helped": then bussing them out to abandon them somewhere else isn't a solution, because the adjudicators can go "yes, you made the number go down, but you did so by decoupling the metric from what it was supposed to measure, so we're not rewarding you for it".

My spouse works in the homelessness field and the correct metric to follow is number of homeless given housing. It’s the “housing first” approach. Harder to game counting amount of people directly placed into homes - someone is paying rent and maintaining a trackable occupied space that you can verify that the client is actually utilizing - and this approach cannot be gamed by “bus them somewhere else” What many peop…

"Number of homeless given housing" is only the correct measure due to the nature of the domain-specific problem. I'm wary of this strategy in general, because the people responsible for deciding how things are accounted for are rarely experts enough to identify sensible domain-specific metrics, so they'll have to consult experts. But that creates a vulnerable point of significant interest to would-be grifters, and if they're not experts enough to assess expert consensus, you end up with metrics that don't work, baked in.

But yes, if we're only looking at homelessness, "how many formerly-homeless people have been given housing?" is a very good way to measure successful interventions.

Re: Updated practice for review articles and position papers in ArXiv CS category

#238
post #89

There is a general problem with rewarding people for the volume of stuff they create, rather than the quality. If you incentivize researchers to publish papers, individuals will find ways to game the system, meeting the minimum quality bar, while taking the least effort to create the most papers and thereby receive the greatest reward. Similarly, if you reward content creators based on views, you will get view maximi…

Who is getting rewarded for uploading tons of stuff to the arXiv?

Re: Updated practice for review articles and position papers in ArXiv CS category

#239

Earlier quoted context omitted.

> rewarding people for the volume ... rather than the quality. I suspect this is a major part of the appeal of LLMs themselves. They produce lines very fast so it appears as if work is being done fast. But that's very hard to know because number of lines is actually a zero signal in code quality or even a commit. Which it's a bit insane already that we use number of lines and commits as measures in the first place. T…

I think you will discover that few organizations use the size or number of edits as a metric of effort. Instead, you might be judged by some measure of productivity (such as resolving issues). Fortunately, language agents are actually useful at coding, when applied judiciously.

Yet it's common enough we see. You also bring up a 10x engineer joke. There's two types of 10x engineers: those that do 10x the work and those who solve 10x the jira tickets but are the cause of 100x of them.

The point is that people metric hack and very bureaucratic structures tend to incentivize metric hacking, not dissuade them. See Pournelle's Iron Law of Bureaucracy.

  > Fortunately, language agents are actually useful at coding, when applied judiciously.
I'm not sure this is in doubt by anyone. By definition it really must be true. The problem is that they're not being used judiciously but haphazardly. The problem is people in large organizations are more concerned with politics than the product they make.

If you cannot see how quality is decreasing then I'm not sure what to tell you. Yes, there are metrics where it's getting better but at the same time user frustration is increasing. AWS and Azure just had recent major outages. Cloudstrike took down lots of the world's network over an avoidable mistake. Microsoft is fumbling the windows upgrade. Apple intelligence was a disaster. YouTube search is beyond infuriating. Google search is so bad we turn to LLMs now. These are major issues and obvious. We don't even have the time to talk about the million minor issues like YouTube captions covering captions embedded in the video, which is not a majorly complicated problem to solve with AI and they're instead pushing AI upscale that is getting a lot of backlash.

So you can claim things are being used judiciously all you want, but I'm not convinced when looking at the results. I'm not happy that every device I use is buggy as shit and simultaneously getting harder to fix myself.

Re: Updated practice for review articles and position papers in ArXiv CS category

#240
In my experience, arXiv is not a preprint platform. It's a strange gatekeeper of science and should be avoided altogether. They have their favorites which they deem as "high quality" and everything else gets rejected. I am eagerly awaiting for people to dismiss arXiv altogether.
Post reply on HN