Live data from Hacker News

ChatGPT generates fake data set to support scientific hypothesis

nature.com

61–70 of 159 posts

Re: ChatGPT generates fake data set to support scientific hypothesis

#61
post #50

Earlier quoted context omitted.

"Humanity learns" is impossible. The unit of analysis is the individual. Some state may cross between individuals via education, but the individuals still must learn. History shows that knowledge transmission remains a sticky wicket.

Firm disagree! Example: take a tour of the tower of London and learn about all the nasty medieval tortures we used to inflict on people. Now we don't do that anymore. Humans learn things collectively via culture and cultural transmission has been an extremely effective tool of knowledge preservation over the generations.

Arguably culture has been degrading somewhat lately

Re: ChatGPT generates fake data set to support scientific hypothesis

#62

It's only a matter of time until someone comes up with a GPT that takes whatever off-axis theories a research paper writer wishes to promulgate, and searches the entire corpus of academic literature for references that can be strung together in such a way as to support any argument one likes. A quack's dream come true, substantiating an argument by backsolving from its feeble or malevolent conclusion to a set of well…

This is already possible with search engines, there is enough information on the internet that you can substantiate just about any claim regardless of how much evidence there is to the contrary. (see flat-earth, plenty of plausible sounding claims with real, albeit, cherry picked evidence).

Re: ChatGPT generates fake data set to support scientific hypothesis

#63
post #16

>“It will make it very easy for any researcher or group of researchers to create fake measurements on non-existent patients, fake answers to questionnaires or to generate a large data set on animal experiments.” Perhaps I'm naive, but I think the people that want to fake data were already doing it without tools like chatgpt. Especially since a ton of biological data is normally distributed, so it's exceedingly easy t…

So we’re pretending making something an order of magnitude easier makes no difference? Ok.

I'm not sure this is much better than the state of the art. Training a model on data and then having it generate new, fake data, is not only easy, it's a standard tool for model boosting.

Re: ChatGPT generates fake data set to support scientific hypothesis

#64
post #22

Earlier quoted context omitted.

Exactly. How is this news? What would be surprising is if ChatGPT _could_ generate fake data that passed analysis.

This could be the key to gain ultimate understanding

This.

Re: ChatGPT generates fake data set to support scientific hypothesis

#65
post #49

I think it would be more clear to say, "ChatGPT assembled a fake data set as a continuation to scientific hypothesis". ChatGPT does not generate data. It reassembles the data (text) it was given, including the text in its training corpus.

I don't think that's accurate, it generates novel outputs that were not observed in the training data.

It doesn't generate new tokens.

Train an LLM on text that only uses lowercase, and it will never output an uppercase letter.

Re: ChatGPT generates fake data set to support scientific hypothesis

#66
post #62

It's only a matter of time until someone comes up with a GPT that takes whatever off-axis theories a research paper writer wishes to promulgate, and searches the entire corpus of academic literature for references that can be strung together in such a way as to support any argument one likes. A quack's dream come true, substantiating an argument by backsolving from its feeble or malevolent conclusion to a set of well…

This is already possible with search engines, there is enough information on the internet that you can substantiate just about any claim regardless of how much evidence there is to the contrary. (see flat-earth, plenty of plausible sounding claims with real, albeit, cherry picked evidence).

Yes of course, this is already possible with AI writing assistance as well, if you're willing to plug in some of the phrases they come up with into a search engine to figure out where they may have come from. But you still have to do the work of stringing the arguments together into a cohesive structure and figuring out how to find research that may be well outside the domains you're familiar with.

But I'm talking about writing a thesis statement, "eating cat boogers makes you live 10 years longer for Science Reasons" and have it string together a completely passable and formally structured argument along with any necessary data to convince enough people to give your cat booger startup revenue to secure next round, because that seems to be where all these games are headed. The winner is the one who can outrun the truth by hashing together a lighter weight version of it, and though it won't stand up to a collision with real thing, you'll be very far from the explosion by the time it happens.

Re: ChatGPT generates fake data set to support scientific hypothesis

#67

Earlier quoted context omitted.

This could be the key to gain ultimate understanding

AGI Silicon Valley style: Fake it till you make it, just like everything else.

I mean, that's how adversarial networks work, isn't it?

Re: ChatGPT generates fake data set to support scientific hypothesis

#68

As an academic researcher, I find analysis of GPT-4 itself — as it pertains to other fields — to be essentially meaningless. There are no guarantees that the version of the model that was used will be available in the future (the API endpoints seem to have a ~12 month future-looking guarantee at most). Don't get me wrong: 1. GPT-4 is incredibly interesting 2. Studying GPT-4 is interesting for people working in that f…

It shows current state and progress. No strong reason to believe future models will preform worse.

Re: ChatGPT generates fake data set to support scientific hypothesis

#69
It's awesome to know that the scientific process that's been in use for centuries, based on peer review and the importance of replication, is still so powerful that AI generated fake data can't do anything to undermine it. Reality doesn't care how the data was faked when you replicate an experiment!

Now, if only scientists and institutions would actually bother using those tools we developed centuries ago. Unfortunately if they don't - you don't exactly need chatgpt to fake data you know? Replication crisis etc etc.

Long story short: Man is it awesome that the scientific method is resilient to this! Too bad nobody uses it.

Re: ChatGPT generates fake data set to support scientific hypothesis

#70

I think it would be more clear to say, "ChatGPT assembled a fake data set as a continuation to scientific hypothesis". ChatGPT does not generate data. It reassembles the data (text) it was given, including the text in its training corpus.

Would you say the same about all systems that are considered "generative"?

Probably.

It's easy to mistake entropy for novelty. Computers don't create: they compute. Calling an LLM "Artificial Intelligence" is a bit like mistaking a pseudorandom number generator for true noise.

Post reply on HN