Live data from Hacker News

Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

stefanzukin.com

31–40 of 107 posts

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#31

Even on hard, if you understand the terminology, the fake ones are mostly gibberish.

Right, I could pick out all the hard ones for fields I'm comfortable with, but struggled for some of the easy ones in fields I was less familiar with.

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#32

Seems like the model likes to repeat words in the title, particularly when hyphens are involved (I guess it considers them as different words?) e.g. "new dinosaur-like dinosaur" and "male-pattern traits in male rats" are a couple I saw.

That's mostly a GPT-2/Transformers quirk. Some approaches apply a repetition penalty to work around it.

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#33
I'm sure GPT2 abstracts would fly through many conferences screening processes. I've seen talks and posters that were utter non-sense but everybody was too polite to say anything to the person or advisors.

I've reviewed articles that were completely made up and the other reviewer didnt even detect that. Nor did the editor.

I've contacted editors about utterly wrong papers, criticized the article on pubpeer, and the article is still published... Because it would harm their notoriety. Thats one of the madenning ascpects of academic publishing.

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#37

Quite easy when you know one is fake. Flagging fake articles in a review queue, by abstract only, and when none may exist all the way up to all being fake .... Now that's a challenge. Also, if you train GPT on the whole corpus of Nature / Science / whatever articles up to, say, 2005, could you feed it leading text about discoveries after 2005 and see if it hypothesizes the justification for those discoveries in the s…

Considering the terrible quality of the writing in these examples, it's simple no matter how it's presented.

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#39
post #23

Even hard mode isn't that hard because GPT-2 tends to ramble on while saying nothing substantive. If I can't figure out what a paper is supposed to be talking about, it's fake. 4/4 on hard. Never read a Nature paper before.

(Easy/cliched joke, but) >If I can't figure out what a paper is supposed to be talking about, it's fake. Depends on the field...

The engineering/materials science/physics ones were fairly easy to identify for me. Usually it would be one or two sentences that were grammatically cohesive but would make a statement that didn't make any sense if you had even a basic understanding of the topic. One that stood out to me was an astrophysics paper that said a planet was orbiting solar wind. I don't have to be a PhD to know that's BS.

The medical and biotech ones are much harder.

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#40
The hard version usually requires me to understand why a number, measurement or chemical or other substance doesn't make sense in the context of what each paragraph is describing. This means I can't just skim it in order to spot the fake, I need to figure out that what it's saying is wrong.

That's close enough for this to be a success if the purpose was to persuade or fool laymen.

Post reply on HN