Quite easy when you know one is fake. Flagging fake articles in a review queue, by abstract only, and when none may exist all the way up to all being fake .... Now that's a challenge. Also, if you train GPT on the whole corpus of Nature / Science / whatever articles up to, say, 2005, could you feed it leading text about discoveries after 2005 and see if it hypothesizes the justification for those discoveries in the s…
Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
41–50 of 107 posts
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#42Quite easy when you know one is fake. Flagging fake articles in a review queue, by abstract only, and when none may exist all the way up to all being fake .... Now that's a challenge. Also, if you train GPT on the whole corpus of Nature / Science / whatever articles up to, say, 2005, could you feed it leading text about discoveries after 2005 and see if it hypothesizes the justification for those discoveries in the s…
A lot of the communication I have with folks is subtly flawed in logic or grammar, but that doesn’t make me think I’m working with a bunch of androids.
It’s natural and often even necessary to try to figure out an author’s intent when their writing doesn’t fully make sense.
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#43The hard version usually requires me to understand why a number, measurement or chemical or other substance doesn't make sense in the context of what each paragraph is describing. This means I can't just skim it in order to spot the fake, I need to figure out that what it's saying is wrong. That's close enough for this to be a success if the purpose was to persuade or fool laymen.
> A new era of hyperaridididididemia revealed by single-cell RNA-seq
So some are better than others. :-)
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#44Presumably when you're a native English speaker and have a broader interest the difficulty goes down a bit.
I like this project very much and would like to see some overall scores, and it might not hurt to allow for a verified result link to detect bragging rather than actual results (not that anybody on HN would ever brag about their score ;) ).
Overall: I'm not worried that generated papers will swamp the publications any day soon but for spam/click farms this must be a godsend and for sure it will cause trouble for search engines to classify real content from generated content.
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#45Earlier quoted context omitted.
(Easy/cliched joke, but) >If I can't figure out what a paper is supposed to be talking about, it's fake. Depends on the field...
The engineering/materials science/physics ones were fairly easy to identify for me. Usually it would be one or two sentences that were grammatically cohesive but would make a statement that didn't make any sense if you had even a basic understanding of the topic. One that stood out to me was an astrophysics paper that said a planet was orbiting solar wind. I don't have to be a PhD to know that's BS. The medical and b…
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#46Could we feed gpt-2 Turbo Encabulator? I want more Turbo Encabulator.
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#47Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#48Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#49Quite easy when you know one is fake. Flagging fake articles in a review queue, by abstract only, and when none may exist all the way up to all being fake .... Now that's a challenge. Also, if you train GPT on the whole corpus of Nature / Science / whatever articles up to, say, 2005, could you feed it leading text about discoveries after 2005 and see if it hypothesizes the justification for those discoveries in the s…
I find the whole thing ominous because there is no "there" there: there is no understanding in the GPT-2 system, but it's able to generate increasingly plausible text. This greatly increases the amount of plausible nonsense that can be used to drown out actual research. You could certainly replace a lot of pop-sci and start several political movements with GPT-2... all of which has no actual nutritional content.
So, basically it's achieved undergraduate level skills.
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#50Even on hard, if you understand the terminology, the fake ones are mostly gibberish.