Priyanka Ranade, Aritran Piplai, Sudip Mittal, Anupam Joshi, and Tim Finin, Generating Fake Cyber Threat Intelligence Using Transformer-Based Models, Int. Joint Conf. on Neural Networks, IEEE, 2021. https://ebiq.org/p/969
Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
61–70 of 107 posts
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#62Even hard mode isn't that hard because GPT-2 tends to ramble on while saying nothing substantive. If I can't figure out what a paper is supposed to be talking about, it's fake. 4/4 on hard. Never read a Nature paper before.
Scores under 5 on what amounts to a coin flip doesn't strike me as so remarkable, especially when coupled with an incentivised reporting-bias as we see here. ("I got a high score! Proud to share!" Vs. "I got a low score, or an even score and look at all the people reporting high scores, think I might keep it to myself")
Being as it is, at this juncture, I think the AI may still have a chance to be strong with this one.
Also, were the AI to do well consistently, I'd think it might say more about the external unfamiliarity with, and the internal prevalence of, field-specific scientific jargon, than any AI's or human's innate intelligence.
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#63Cool demo. With these GPT models, I don't get the appeal of creating fake text that at best can pass as real to someone who doesn't understand the topic and context. What's the use case? Generating more believable spam for social media? Anything else? Because there's no real knowledge representation or information extraction going on here.
For fun .
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#64Earlier quoted context omitted.
Or you just train a machine to do it and then generate a bunch and have this second machine sort out any it thinks are machine generated.
The best fake-detecting model detecting fakes generated by the best generator model will always lag behind the latter model.
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#65The side-by-side display makes it pretty easy to distinguish the one from the other, simply compare them at a level where the one that makes the least sense is the one that is nonsense. Like that I score 10/11. But when looking at just the left side one suddenly the problem is much harder, and I'm happy to get better than even. Bits that don't help: not an English native writer. Seen too many real life papers with cr…
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#66Earlier quoted context omitted.
I find the whole thing ominous because there is no "there" there: there is no understanding in the GPT-2 system, but it's able to generate increasingly plausible text. This greatly increases the amount of plausible nonsense that can be used to drown out actual research. You could certainly replace a lot of pop-sci and start several political movements with GPT-2... all of which has no actual nutritional content.
I find it intriguing for exactly that reason. I agree there's no fundamental "there", but I suggest that you may find that ominous because it implies there's no fundamental understanding anywhere. Only stories that survive scrutiny. GPT can write "about" something from a prompt. This is not much different than me interpreting data that I'm analyzing. I'm constantly generating stories and checking them, until one stor…
Faking an entire 10 page paper with figures and citations is much harder. I'm sure it'll happen next week, but until then I can still say that's where real understanding is demonstrated.
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#67The sad thing is that often there's an equal mental effort to read GPT articles and the real ones. It's as if people are trying to make their papers as incomprehensible as possible.
Incomprehensible language = look how smart I am now give me grant money
The incomprehensibility comes from the fact that abstracts (and particularly NPG abstracts) are trying to do many things at once--and all in 200 words. In theory, the abstract should describe why your work is of broad general interest (so Nature's editors will publish it), while explaining the specific scientific question and answer(!) to a specialist audience of often-picky, sometimes-hostile peer reviewers, and conforming to a fairly specific style that doesn't reference the rest of the paper.
It's tough to do well, and even moreso for non-native English speakers.
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#68Earlier quoted context omitted.
The best fake-detecting model detecting fakes generated by the best generator model will always lag behind the latter model.
I think I see what you’re saying, but why is this so?
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#69Cool demo. With these GPT models, I don't get the appeal of creating fake text that at best can pass as real to someone who doesn't understand the topic and context. What's the use case? Generating more believable spam for social media? Anything else? Because there's no real knowledge representation or information extraction going on here.
I don't know if this translates to technical writing, but it's possible someone might complete a prompt on some specific topic(s) and then use that as a point to start from, especially if they're knowledgeable enough on the topic to correct the output. It's nice to be able to skip a lot of boilerplate words (how many words in this comment are actually the meat of this idea, and how many words are just there to tie all those morsels together?)
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#70I would be curious to try again with GPT-3.