Easy mode is cake. Hard mode is good enough that I'd like to see some sort of distance metric to the nearest real story, to be sure the model isn't accidentally copying truth.
Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
51–60 of 107 posts
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#52The side-by-side display makes it pretty easy to distinguish the one from the other, simply compare them at a level where the one that makes the least sense is the one that is nonsense. Like that I score 10/11. But when looking at just the left side one suddenly the problem is much harder, and I'm happy to get better than even. Bits that don't help: not an English native writer. Seen too many real life papers with cr…
I found that people would also refuse the test and would believe whatever the output of the model was due to my choice of subject.
Others that did a similar exercise and tried to verify their results using reddit had a great deal of people who would be able to spot fakes quite easily.
The biggest issue would be someone using a system to deliberately fool a targeted set of people which is easy given how ad networks are run.
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#53The hard version usually requires me to understand why a number, measurement or chemical or other substance doesn't make sense in the context of what each paragraph is describing. This means I can't just skim it in order to spot the fake, I need to figure out that what it's saying is wrong. That's close enough for this to be a success if the purpose was to persuade or fool laymen.
This was a hard-mode fake I just got: > A new era of hyperaridididididemia revealed by single-cell RNA-seq So some are better than others. :-)
Which is absolutely a real thing except that the exact quantum properties in fact didn't commute while they claimed they did commute and said for some reason simultaneous measurement required a third state anyway.
I don't know how I would have been able to distinguish that from completely reasonable methods for quantum error correction without knowing ahead of time which quantum states commute and which don't... pretty cool.
If I were skimming or half asleep I definitely wouldn't have caught a lot of these on hard, abstracts are always so poorly written and usually trying too hard to be complicated sounding by using big words when small ones would do just fine!
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#54Quite easy when you know one is fake. Flagging fake articles in a review queue, by abstract only, and when none may exist all the way up to all being fake .... Now that's a challenge. Also, if you train GPT on the whole corpus of Nature / Science / whatever articles up to, say, 2005, could you feed it leading text about discoveries after 2005 and see if it hypothesizes the justification for those discoveries in the s…
I find the whole thing ominous because there is no "there" there: there is no understanding in the GPT-2 system, but it's able to generate increasingly plausible text. This greatly increases the amount of plausible nonsense that can be used to drown out actual research. You could certainly replace a lot of pop-sci and start several political movements with GPT-2... all of which has no actual nutritional content.
GPT can write "about" something from a prompt. This is not much different than me interpreting data that I'm analyzing. I'm constantly generating stories and checking them, until one story survives it all. How do I generate stories!? Seriously. I'm sure I have a GPT module in my left frontal cortex. I use it all the time when I think about actions I take, and it's what I try to ignore when I meditate. Its ongoing narrative is what feeds back into how I feel about things, which affects how I interact with things and what things I interact with ... not necessarily as a goal-driven decision process, more as a feedback-driven randomized selection. Isn't this kind of the basis of Cognitive Behavioral Therapy, meditation, etc. See [1,2]. If you stick GPT and sentiment analysis into a room, will they produce a rumination feedback like a depressed person?
Anyway, if you can tell a coherent story to justify a result (once presented with a result), one that is convincing enough for people to believe and internalize the result in their future studies, how is that different from understanding that result and teaching it to others? The act of teaching is itself story generation. Mental models are just story-driven hacks that allow people to generalize results in an insanely complex system.
1. Happiness Hypothesis Jonathan Haidt 2. Buddhism and modern psychology, coursera
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#55Some of my favorites: "A new onset of primeval black magic in magic-ring crystals"
"The genetic network for moderate religiosity in one thousand bespectacled twins"
"Thermal vestige of the '70s and '00s disco ball trend"
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#56I'm sure GPT2 abstracts would fly through many conferences screening processes. I've seen talks and posters that were utter non-sense but everybody was too polite to say anything to the person or advisors. I've reviewed articles that were completely made up and the other reviewer didnt even detect that. Nor did the editor. I've contacted editors about utterly wrong papers, criticized the article on pubpeer, and the a…
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#57Quite easy when you know one is fake. Flagging fake articles in a review queue, by abstract only, and when none may exist all the way up to all being fake .... Now that's a challenge. Also, if you train GPT on the whole corpus of Nature / Science / whatever articles up to, say, 2005, could you feed it leading text about discoveries after 2005 and see if it hypothesizes the justification for those discoveries in the s…
I find the whole thing ominous because there is no "there" there: there is no understanding in the GPT-2 system, but it's able to generate increasingly plausible text. This greatly increases the amount of plausible nonsense that can be used to drown out actual research. You could certainly replace a lot of pop-sci and start several political movements with GPT-2... all of which has no actual nutritional content.
Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#58Re: Enigma: GPT-2 trained on 10K Nature Papers: Can you spot the difference?
#59Earlier quoted context omitted.
The engineering/materials science/physics ones were fairly easy to identify for me. Usually it would be one or two sentences that were grammatically cohesive but would make a statement that didn't make any sense if you had even a basic understanding of the topic. One that stood out to me was an astrophysics paper that said a planet was orbiting solar wind. I don't have to be a PhD to know that's BS. The medical and b…
Yes, this mirrors my experience. Fields that have my interest are pretty easy in isolation (just looking at one subject), but fields that are remote can be a challenge.
Three highly pathogenic β-coronaviruses have crossed the animal-to-human species barrier in the past two decades: SARS-CoV, MERS-CoV and SARS-CoV-2. To evaluate the possibility of identifying antibodies with broad neutralizing activity, we isolated a monoclonal antibody, termed B4, that cross-reacts with eight β-coronavirus spike glycoproteins, including all five human-infecting β-coronaviruses.