Live data from Hacker News

Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

stefanzukin.com

11–20 of 107 posts

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#11
Quite easy when you know one is fake. Flagging fake articles in a review queue, by abstract only, and when none may exist all the way up to all being fake ....

Now that's a challenge.

Also, if you train GPT on the whole corpus of Nature / Science / whatever articles up to, say, 2005, could you feed it leading text about discoveries after 2005 and see if it hypothesizes the justification for those discoveries in the same way that the authors did?

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#12
post #10

The sad thing is that often there's an equal mental effort to read GPT articles and the real ones. It's as if people are trying to make their papers as incomprehensible as possible.

Incomprehensible language = look how smart I am now give me grant money

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#13

Quite easy when you know one is fake. Flagging fake articles in a review queue, by abstract only, and when none may exist all the way up to all being fake .... Now that's a challenge. Also, if you train GPT on the whole corpus of Nature / Science / whatever articles up to, say, 2005, could you feed it leading text about discoveries after 2005 and see if it hypothesizes the justification for those discoveries in the s…

Excellent point. Serializing them would make it more difficult.

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#14

STEM people love to bring up the sokal affair. the same STEM people also don't realize that many journals and conferences in STEM have been tricked by things like this (more specifically precursors using HMMs and etc). https://en.wikipedia.org/wiki/List_of_scholarly_publishing_s... edit: don't understand why i'm getting downvoted. is my comment not relevant to a post about the plausibility of abstracts generated by M…

The reason for the downvotes is probably the generalisation regarding what "STEM people love to bring up" and also "don't realize". It feels like an unprovoked strawman attack against an ambiguously defined group of people.

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#15
5/5 on hard node, but it’s tough sometimes, I don’t actually know much about biology. But if you’ve played around with GPT before, you get better at spotting the subtle logical errors it tends to make. I wonder whether the ability to identify machine generated texts will become a useful skill at some point.

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#17
post #15

5/5 on hard node, but it’s tough sometimes, I don’t actually know much about biology. But if you’ve played around with GPT before, you get better at spotting the subtle logical errors it tends to make. I wonder whether the ability to identify machine generated texts will become a useful skill at some point.

Or you just train a machine to do it and then generate a bunch and have this second machine sort out any it thinks are machine generated.

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#18
post #15

5/5 on hard node, but it’s tough sometimes, I don’t actually know much about biology. But if you’ve played around with GPT before, you get better at spotting the subtle logical errors it tends to make. I wonder whether the ability to identify machine generated texts will become a useful skill at some point.

Or you just train a machine to do it and then generate a bunch and have this second machine sort out any it thinks are machine generated.

You basically just described a GAN. Neat!

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#19
post #15

5/5 on hard node, but it’s tough sometimes, I don’t actually know much about biology. But if you’ve played around with GPT before, you get better at spotting the subtle logical errors it tends to make. I wonder whether the ability to identify machine generated texts will become a useful skill at some point.

Or you just train a machine to do it and then generate a bunch and have this second machine sort out any it thinks are machine generated.

The best fake-detecting model detecting fakes generated by the best generator model will always lag behind the latter model.

Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?

#20
post #18

Earlier quoted context omitted.

Or you just train a machine to do it and then generate a bunch and have this second machine sort out any it thinks are machine generated.

You basically just described a GAN. Neat!

GANs work by feeding back the mistakes and forcing the generator model to improve its cheating. In this case, filtering out titles that are ambiguous would act as an independent filter.
Post reply on HN