Earlier quoted context omitted.
We already got an LLM generated meta review that was very clearly just summarization of reviews. There were some pretty egregious cases of borderline hallucinated remarks. This was ACL Rolling Review, so basically the most prestigious NLP venue and the editors told us to suck it up. Very disappointing and I genuinely worry about the state of science and how this will affect people who rely on scientometric criteria.
With NeurIPS 2024 reviews going on right now, I'm sure that a whole lot of these kind of reviews are being generated daily.
GPT-fabricated scientific papers on Google Scholar
71–80 of 107 posts
Re: GPT-fabricated scientific papers on Google Scholar
#72Earlier quoted context omitted.
We already got an LLM generated meta review that was very clearly just summarization of reviews. There were some pretty egregious cases of borderline hallucinated remarks. This was ACL Rolling Review, so basically the most prestigious NLP venue and the editors told us to suck it up. Very disappointing and I genuinely worry about the state of science and how this will affect people who rely on scientometric criteria.
Most conferences have been flooded with submissions, and ACL is no exception. A consequence of that is that there are not sufficient numbers of reviewers available who are qualified to review these manuscripts. Conference organizers might be keen to accept many or most who offer to volunteer, but clearly there is now a large pool of people that have never done this before, and were never taught how to do this. Add so…
Pretty hard to combat. We just rebutted as if it were a real review - maybe it was - and hope that the chairs see it. Speaking to other folks, opinions are split over whether this sort of review should be flagged. I know some people who tried to query a review and it didn't help.
There were other small cues - the English was perfect, while other reviewers made small slips indicative of non-native speakers. One was simply the discrepancy between the tone of the review (generally very positive) and the middle-of-the-road rating and confidence. The structure of the review was very "The authors do X, Y, Z. This is important because A, B, C." and the reviewer didn't bother to fill out any of the other review sections (they just wrote single-word answeres to all of them).
The kicker was actually putting our paper in to 4o and asking it to write a review and seeing the same keywords pop up.
Re: GPT-fabricated scientific papers on Google Scholar
#73GPT might make fabricating scientific papers easier, but let's not forget how many humans fabricated scientific research in recent years - they did a great job without AI! For any who haven't seen/heard, this makes for some entertaining and eye-opening viewing! https://www.youtube.com/results?search_query=academic+fraud
Re: GPT-fabricated scientific papers on Google Scholar
#74Earlier quoted context omitted.
It might follow to say that current LLM;s arent trained to generate papers, BUT they also don't really need to reason. They just need to mimic the appearance of reason, follow the same pattern of progression. Ingesting enough of what amounts to executed templates will teach it to generate its own results as if output from the same template.
What is the difference between 'reasoning' and 'appearing to be reasoning' if the results are the same with the same input?
For example, I recently experimented with using ChatGPT to translate a Wikipedia article, on the grounds that it mighy maintain all the formatting and that Transformer models are also used by Google Translate.
As it was an experiment, I did actually check the results before submitting the translated article.
First roughly 3/4 were fine. Final quarter was completely invented but plausible, including references.
LLMs are very useful tools, I'll gladly use them to help with various tasks and they can (with low reliability but it has happened) even manage a whole project, but right now they should treated with caution and not left unsupervised — Peter principle, being promoted beyond their competence, still applies even though they're not human employees.
Re: GPT-fabricated scientific papers on Google Scholar
#75Earlier quoted context omitted.
With NeurIPS 2024 reviews going on right now, I'm sure that a whole lot of these kind of reviews are being generated daily.
With ICLR paper deadline coming up, I guess it's worth wargaming how GPT4 would review my submission.
One example was that GPT took an explicit geographic location from a figure caption and used that as a reference point when suggesting improvements (along the lines of "location X is under-represented on this map") I assume because it places some high degree of relevance to figures and the abstract when summarising papers. I think you might be able to combat this by writing defensively - in our case we might have avoided that by saying "more information about geographic diversity may be found in X and the supplementary information"
Re: GPT-fabricated scientific papers on Google Scholar
#76Re: GPT-fabricated scientific papers on Google Scholar
#77Earlier quoted context omitted.
With NeurIPS 2024 reviews going on right now, I'm sure that a whole lot of these kind of reviews are being generated daily.
With ICLR paper deadline coming up, I guess it's worth wargaming how GPT4 would review my submission.
Re: GPT-fabricated scientific papers on Google Scholar
#78Hope this research is better.
Re: GPT-fabricated scientific papers on Google Scholar
#79When I went to the APS March Meeting earlier this year, I talked with the editor of a scientific journal and asked them if they were worried about LLM generated papers. They said actually their main worry wasn't LLM-generated papers, it was LLM-generated reviews . LLMs are much better at plausibly summarizing content than they are at doing long sequences of reasoning, so they're much better at generating believable r…
Researchers get massive CVs, reviewers and editors get off easy, admins get to show great output numbers from their institutions, and of course the publishers continue making hand over fist.
It's a rather broken system.
Re: GPT-fabricated scientific papers on Google Scholar
#80For example, the authors’ statement that “[GPT’s] undeclared use—beyond proofreading—has potentially far-reaching implications for both science and society” suggests that, for them, using LLMs for “proofreading” is okay. But “proofreading” is understood in various ways. For some people, it would include only correcting spelling and grammatical mistakes. For others, especially for people who are not native speakers of English, it can also include changing the wording and even rewriting entire sentences and paragraphs to make the meaning clearer. To what extent can one use an LLM for such revision without declaring that one has done so?