Live data from Hacker News

ChatGPT generates fake data set to support scientific hypothesis

nature.com

121–130 of 159 posts

Re: ChatGPT generates fake data set to support scientific hypothesis

#121
post #16

>“It will make it very easy for any researcher or group of researchers to create fake measurements on non-existent patients, fake answers to questionnaires or to generate a large data set on animal experiments.” Perhaps I'm naive, but I think the people that want to fake data were already doing it without tools like chatgpt. Especially since a ton of biological data is normally distributed, so it's exceedingly easy t…

The easiest data to fake is null results because anything else is replicable, hence the importance of replication.

null results are also replicable.

Re: ChatGPT generates fake data set to support scientific hypothesis

#122
"Gordon’s great insight was to design a program which allowed you to specify in advance what decision you wished it to reach, and only then to give it all the facts. The program’s task, which it was able to accomplish with consummate ease, was simply to construct a plausible series of logical-sounding steps to connect the premises with the conclusion."

From the thumping good detective-ghost-horror-who dunnit-time travel-romantic-musical-comedy-epic: Dirk Gently's Holistic Detective Agency

Re: ChatGPT generates fake data set to support scientific hypothesis

#123

It's only a matter of time until someone comes up with a GPT that takes whatever off-axis theories a research paper writer wishes to promulgate, and searches the entire corpus of academic literature for references that can be strung together in such a way as to support any argument one likes. A quack's dream come true, substantiating an argument by backsolving from its feeble or malevolent conclusion to a set of well…

There is in fact a treasure throve of high level research that is to controversial for any expert to go near.

Wild things will happen if screaming hoax from ignorance can no longer shut down constructive efforts.

It will simply combine what is written about germ theory or heavier than air flying machines and produce sensible responses.

The patent db's are full of treasures if you have oh 1000 years? to study it. Maybe 10 000?

It should also be possible to take a seemingly unworkable idea that makes no sense and gather just what is needed to bring it into reality.

For stuff you can build or otherwise test properly it makes no difference what people think is possible.

People think very little is possible, we always did! Everything that can be discovered has been discovered has been the mantra for thousands of years. This while the things people actually accomplish seem to get more and more astonishing.

Re: ChatGPT generates fake data set to support scientific hypothesis

#124
post #109
post #96

Earlier quoted context omitted.

Being able to do it at scale is an issue because of paper mills.

`np.random.normal(0, 1)`. You too can generate infinite amount of normal distribution data to use for fake datasets, free of AI.

People are probably thinking data more complex than normal distributions (though I'm also not sure if GPT-4 is the best method for that)

Re: ChatGPT generates fake data set to support scientific hypothesis

#125

Roughly speaking, the whole point of an LLM is to create plausible sounding text, without regard to the truth (which it cannot determine or derive), so it's a tool that is perfectly suited to this sort of malfeasance.

I’ve read this over and over, and believe it - but it’s so easy to forget when it absolutely nails something like debugging or suggesting a CMake edit or finding 5 letter combinations that are reversible with a vowel in the middle, or etc.

It’s just still mind blowing I can get a sarcastic summary of an email in the theme of GlaDOS from Portal and in the same screen get an email proofread.

It’s funny how far “what should the next word be” can go.

Re: ChatGPT generates fake data set to support scientific hypothesis

#126
post #96
post #16

>“It will make it very easy for any researcher or group of researchers to create fake measurements on non-existent patients, fake answers to questionnaires or to generate a large data set on animal experiments.” Perhaps I'm naive, but I think the people that want to fake data were already doing it without tools like chatgpt. Especially since a ton of biological data is normally distributed, so it's exceedingly easy t…

Being able to do it at scale is an issue because of paper mills.

The issue is that we consider passing peer review to be enough for something to be taken as truth when it should really be reproducibility/replicability. That, and the incentives currently driving academia are absolutely ridiculous. Paper mills are merely a symptom of academia's ills, not the cause of them.

Re: ChatGPT generates fake data set to support scientific hypothesis

#127
post #49

Earlier quoted context omitted.

I don't think that's accurate, it generates novel outputs that were not observed in the training data.

It doesn't generate new tokens. Train an LLM on text that only uses lowercase, and it will never output an uppercase letter.

So the model is limited to using words and characters that already exist. I agree with you but I don't see why is a limitation worth pointing out.

Re: ChatGPT generates fake data set to support scientific hypothesis

#128
The paper[1] itself reads like marketing material that was itself written by GPT-4. Why was this published? And why is Nature reporting on it? As someone who generates fake/simulated datasets for a living, the scientific value of this paper is totally lost on me.

[1] https://doi.org/10.1001/jamaophthalmol.2023.5162

Re: ChatGPT generates fake data set to support scientific hypothesis

#130
post #121

Earlier quoted context omitted.

The easiest data to fake is null results because anything else is replicable, hence the importance of replication.

null results are also replicable.

A null result, or the absence of evidence, is not evidence of absence. If you fake a null result, you’re not asserting anything other than you could not measure and collect supporting data using your experiment to prove or disprove a hypothesis. It is difficult for someone doing replication to accuse you of ill-intent, as opposed to faked data that proves your point when anyone else can replicate your experiment and get totally different or even contradictory results.
Post reply on HN