Live data from Hacker News

ChatGPT generates fake data set to support scientific hypothesis

nature.com

51–60 of 159 posts

Re: ChatGPT generates fake data set to support scientific hypothesis

#51
It's only a matter of time until someone comes up with a GPT that takes whatever off-axis theories a research paper writer wishes to promulgate, and searches the entire corpus of academic literature for references that can be strung together in such a way as to support any argument one likes.

A quack's dream come true, substantiating an argument by backsolving from its feeble or malevolent conclusion to a set of well-known premises but-with-citations. converting untenable speculation into something that passes many superficial tests of legitimacy, which is more than enough to boost it into broader and less critical visibility.

"thick with citations, therefore truthy" is a big blind spot in the casual heuristic used ro gauge the quality of a given piece of research writing, especially at the undergrad level where this tool, lets call it CheatGPT, would be stupendously popular.

Re: ChatGPT generates fake data set to support scientific hypothesis

#52
post #16

>“It will make it very easy for any researcher or group of researchers to create fake measurements on non-existent patients, fake answers to questionnaires or to generate a large data set on animal experiments.” Perhaps I'm naive, but I think the people that want to fake data were already doing it without tools like chatgpt. Especially since a ton of biological data is normally distributed, so it's exceedingly easy t…

So we’re pretending making something an order of magnitude easier makes no difference? Ok.

It's not an order of magnitude easier. It won't make a difference

Re: ChatGPT generates fake data set to support scientific hypothesis

#53

I think it would be more clear to say, "ChatGPT assembled a fake data set as a continuation to scientific hypothesis". ChatGPT does not generate data. It reassembles the data (text) it was given, including the text in its training corpus.

Would you say the same about all systems that are considered "generative"?

Re: ChatGPT generates fake data set to support scientific hypothesis

#55
post #39

Earlier quoted context omitted.

The world of Star Trek is one where humanity learns from the devastation of world war three and over a century, creates a society where almost everything is run by computers, money doesn't exist, and people work primarily for their personal satisfaction. But I think true AGI is frowned upon there.

"Humanity learns" is impossible. The unit of analysis is the individual. Some state may cross between individuals via education, but the individuals still must learn. History shows that knowledge transmission remains a sticky wicket.

humanity learns just mean sufficient individuals learn to take power and enact such a society

Re: ChatGPT generates fake data set to support scientific hypothesis

#56
post #22

Earlier quoted context omitted.

Exactly. How is this news? What would be surprising is if ChatGPT _could_ generate fake data that passed analysis.

This could be the key to gain ultimate understanding

AGI Silicon Valley style: Fake it till you make it, just like everything else.

Re: ChatGPT generates fake data set to support scientific hypothesis

#57
post #16

>“It will make it very easy for any researcher or group of researchers to create fake measurements on non-existent patients, fake answers to questionnaires or to generate a large data set on animal experiments.” Perhaps I'm naive, but I think the people that want to fake data were already doing it without tools like chatgpt. Especially since a ton of biological data is normally distributed, so it's exceedingly easy t…

So we’re pretending making something an order of magnitude easier makes no difference? Ok.

Are we pretending Faker hasn't been a staple library in software testing for years?

Re: ChatGPT generates fake data set to support scientific hypothesis

#58
As an academic researcher, I find analysis of GPT-4 itself — as it pertains to other fields — to be essentially meaningless. There are no guarantees that the version of the model that was used will be available in the future (the API endpoints seem to have a ~12 month future-looking guarantee at most).

Don't get me wrong:

1. GPT-4 is incredibly interesting

2. Studying GPT-4 is interesting for people working in that field

But when I see people writing about how GPT-4 can pass the USMLE (etc), it has no lasting meaning. It might as well be marketing for OpenAI, and to me it has roughly that amount of academic importance.

Re: ChatGPT generates fake data set to support scientific hypothesis

#59
post #16

>“It will make it very easy for any researcher or group of researchers to create fake measurements on non-existent patients, fake answers to questionnaires or to generate a large data set on animal experiments.” Perhaps I'm naive, but I think the people that want to fake data were already doing it without tools like chatgpt. Especially since a ton of biological data is normally distributed, so it's exceedingly easy t…

So we’re pretending making something an order of magnitude easier makes no difference? Ok.

ChatGPT will unlock new levels of both good and bad. The question is what the ratio between the two is going to be.

Re: ChatGPT generates fake data set to support scientific hypothesis

#60
If you haven't read the breakdown of the epic Jan Hendrik Schön scandal in which he published about a half-dozen fraudulent papers in Science and Nature based on fabricated results regarding organic (chemically speaking) semiconductor devices cooked up out of thin air, start here. Required reading for any young graduate student IMO. The shame of Bell Labs, Science and Nature, all taken for a ride:

https://en.wikipedia.org/wiki/Plastic_Fantastic

If that fraudster had started out with ChatGPT4, the fraud might have persisted for another decade (because organic semiconductors don't seem to have the capabilities he believed they had), because he was only detected via replicated datasets. If he'd had ChatGPT4 to generate new plausible datasets, well...?

I guarantee you that a significant fraction of the people in academia who 'got there first' on significant discoveries in science did so by fabricating data along the lines of Schön. They just guessed right, and fabricated data, and then more serious careful scientists were able to replicate their bogus work later.

Schön guessed wrong, and every effort to replicate his work failed, and Bell Labs, Science and Nature were left with egg on their face, which they're still trying to wipe off. ChatGPT4 and its shady parents and affiliates will only make this problem worse, not better.

"Benefit to humanity" my ass.

[edit: if you wonder why I sound so salty I read all those Schön papers with interest and fascination when I was a young graduate student myself, now I'm older and seriously jaded.]

Post reply on HN