Live data from Hacker News

Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

stefanzukin.com

21–30 of 107 posts

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#21
post #10

The sad thing is that often there's an equal mental effort to read GPT articles and the real ones. It's as if people are trying to make their papers as incomprehensible as possible.

At least in the samples I was presented, the more comprehensible articles were consistently the fake ones.

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#24

Quite easy when you know one is fake. Flagging fake articles in a review queue, by abstract only, and when none may exist all the way up to all being fake .... Now that's a challenge. Also, if you train GPT on the whole corpus of Nature / Science / whatever articles up to, say, 2005, could you feed it leading text about discoveries after 2005 and see if it hypothesizes the justification for those discoveries in the s…

Good point.

This challenge would be more interesting if there were "Neither is fake" and "Both are fake" buttons (and obviously, the test randomly showed two fake and two real articles in the mix)

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#25
post #23

Even hard mode isn't that hard because GPT-2 tends to ramble on while saying nothing substantive. If I can't figure out what a paper is supposed to be talking about, it's fake. 4/4 on hard. Never read a Nature paper before.

(Easy/cliched joke, but)

>If I can't figure out what a paper is supposed to be talking about, it's fake.

Depends on the field...

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#26
Cool demo.

With these GPT models, I don't get the appeal of creating fake text that at best can pass as real to someone who doesn't understand the topic and context. What's the use case? Generating more believable spam for social media? Anything else? Because there's no real knowledge representation or information extraction going on here.

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#27

Easy mode is cake. Hard mode is good enough that I'd like to see some sort of distance metric to the nearest real story, to be sure the model isn't accidentally copying truth.

Yes, I got a really short astronomical one about the discovery of a metallic core planet circling a G-Type star, and I only knew it was fake, because I would have heard about it!

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#28

Cool demo. With these GPT models, I don't get the appeal of creating fake text that at best can pass as real to someone who doesn't understand the topic and context. What's the use case? Generating more believable spam for social media? Anything else? Because there's no real knowledge representation or information extraction going on here.

For fun.

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#30
The generated abstracts may be gibberish but I wonder how often they contain little bits of brilliance, or make novel connections between ideas expressed in the training set. If we got a panel of domain experts to evaluate the snippets on this basis, thrir labels could be used to fine-tune the model in the direction of novel discovery. (This is almost certainly not a novel idea!)
Post reply on HN