Live data from Hacker News

ChatGPT generates fake data set to support scientific hypothesis

nature.com

131–140 of 159 posts

Re: ChatGPT generates fake data set to support scientific hypothesis

#131
post #16

>“It will make it very easy for any researcher or group of researchers to create fake measurements on non-existent patients, fake answers to questionnaires or to generate a large data set on animal experiments.” Perhaps I'm naive, but I think the people that want to fake data were already doing it without tools like chatgpt. Especially since a ton of biological data is normally distributed, so it's exceedingly easy t…

Cases of fabrication have been caught because the fabrication is done so badly, i.e. inplausibly. It's often just imputing some "random" numbers or repeating samples etc. Many would be amazed how technically and mathematically illiterate scientists often are. And probably ones that fabricate even more so.

Maybe it will increase and/or get a bit higher quality with LLM fakery. But as with many "AI bad" themes, the problem isn't that "AI" can fabricate the data. The problem is fucked up institutions and cultures.

Re: ChatGPT generates fake data set to support scientific hypothesis

#133
post #16

>“It will make it very easy for any researcher or group of researchers to create fake measurements on non-existent patients, fake answers to questionnaires or to generate a large data set on animal experiments.” Perhaps I'm naive, but I think the people that want to fake data were already doing it without tools like chatgpt. Especially since a ton of biological data is normally distributed, so it's exceedingly easy t…

So we’re pretending making something an order of magnitude easier makes no difference? Ok.

This is so worrying for me, the amount of digital garbage that can now be generated that is not obviously garbage, that one must read and discern first nonsensical bad text generation and second if factual, is truth becoming a needle in the hay stack? How can this be cleaned?

Re: ChatGPT generates fake data set to support scientific hypothesis

#134
post #130
post #121

Earlier quoted context omitted.

null results are also replicable.

A null result, or the absence of evidence, is not evidence of absence. If you fake a null result, you’re not asserting anything other than you could not measure and collect supporting data using your experiment to prove or disprove a hypothesis. It is difficult for someone doing replication to accuse you of ill-intent, as opposed to faked data that proves your point when anyone else can replicate your experiment and…

You still have to give out the statistics that show the null result. E.g. something with a high p-value. You are in fact "confirming the null hypothesis". They aren't any more difficult to replicate than results supporting "the alternative hypothesis".

(The whole binary hypothesis system and culture is a mess though, but that's besides the point.)

Re: ChatGPT generates fake data set to support scientific hypothesis

#135

Roughly speaking, the whole point of an LLM is to create plausible sounding text, without regard to the truth (which it cannot determine or derive), so it's a tool that is perfectly suited to this sort of malfeasance.

I’ve read this over and over, and believe it - but it’s so easy to forget when it absolutely nails something like debugging or suggesting a CMake edit or finding 5 letter combinations that are reversible with a vowel in the middle, or etc. It’s just still mind blowing I can get a sarcastic summary of an email in the theme of GlaDOS from Portal and in the same screen get an email proofread. It’s funny how far “what sh…

a billion humans typing on a billion keyboards for a few decades... just needed someone to come along and categorize each of the outputs.

Re: ChatGPT generates fake data set to support scientific hypothesis

#136

Turing test passed, because that is what quite a lot of human scientists did. Hello, Dr. Ariely, you've met your match. Sarcasm aside, once such systems really learn to lie, they will be all too human-like. Perhaps the defining quality of real intelligence is deception.

Its insane how they think they can have their AGI and eat it too! The closest thing we have to AGI is like a human child even though that's Synthetic (mankind made) General Intelligence and they can't really be truly relied upon and not learn or formulate sentences you don't like

Re: ChatGPT generates fake data set to support scientific hypothesis

#137

Roughly speaking, the whole point of an LLM is to create plausible sounding text, without regard to the truth (which it cannot determine or derive), so it's a tool that is perfectly suited to this sort of malfeasance.

I’ve read this over and over, and believe it - but it’s so easy to forget when it absolutely nails something like debugging or suggesting a CMake edit or finding 5 letter combinations that are reversible with a vowel in the middle, or etc. It’s just still mind blowing I can get a sarcastic summary of an email in the theme of GlaDOS from Portal and in the same screen get an email proofread. It’s funny how far “what sh…

It might seem that there are cases where the ability to generate heaps of plausibly sounding text on demand is helpful, and there are cases where it is pretty much the worst possible capability to have.

Re: ChatGPT generates fake data set to support scientific hypothesis

#138
post #127

Earlier quoted context omitted.

It doesn't generate new tokens. Train an LLM on text that only uses lowercase, and it will never output an uppercase letter.

So the model is limited to using words and characters that already exist. I agree with you but I don't see why is a limitation worth pointing out.

you literally have to put in every number for it to do mathematics correctly...

its as stupid as that. some try to get around it by indeed only having the 10 different digits and glue them together, but its a hallucination that that works.

an important point in generalization is for example that you teach it something. This is literally important

'ycombinator is a website' is a prompt that is almost impossible of ycombinator is not in your training set

Re: ChatGPT generates fake data set to support scientific hypothesis

#139

Earlier quoted context omitted.

So we’re pretending making something an order of magnitude easier makes no difference? Ok.

I'm not sure this is much better than the state of the art. Training a model on data and then having it generate new, fake data, is not only easy, it's a standard tool for model boosting.

Poisoning the well for others, huh?

Re: ChatGPT generates fake data set to support scientific hypothesis

#140
The framing makes it sound like it's a "bug" or something. From my understanding it's not, because it's hardly a reliable reasoning tool: whether using statements or using data. Unless we come up with or advance a better architecture, similar "panic porn" is useless, not to mention this reeks of a hit piece. Just verify everything and stop with the blind trust.
Post reply on HN