Live data from Hacker News

ChatGPT generates fake data set to support scientific hypothesis

nature.com

151–159 of 159 posts

Re: ChatGPT generates fake data set to support scientific hypothesis

#151

Earlier quoted context omitted.

We think objectively about it. We have goals and intentions. We use logic. LLMs don't. The dirt in the ground sorts impurities from water, but we don't call it intelligent or generative. We call it entropy.

This is a human supremacy argument, essentially claiming that because we have a "soul" or because of something inherent in us that cannot be proven we are better than something else. You are free to believe this, but it is a matter of faith . Not of any sort of reasoning.

Yep. I'm not much of a religious person, but if I were, I might say: not only does AI not create stuff, humans don't either, only the Creator did! I suppose the big bang theory doesn't stray far from this either. Point is, maybe we use this notion metaphorically/imprecisely even for human output, and therefore we might as well extend it to machines.

Re: ChatGPT generates fake data set to support scientific hypothesis

#153

Earlier quoted context omitted.

I’ve read this over and over, and believe it - but it’s so easy to forget when it absolutely nails something like debugging or suggesting a CMake edit or finding 5 letter combinations that are reversible with a vowel in the middle, or etc. It’s just still mind blowing I can get a sarcastic summary of an email in the theme of GlaDOS from Portal and in the same screen get an email proofread. It’s funny how far “what sh…

a billion humans typing on a billion keyboards for a few decades... just needed someone to come along and categorize each of the outputs.

I like that. And the numbers are way higher than a billion.

Re: ChatGPT generates fake data set to support scientific hypothesis

#155

Roughly speaking, the whole point of an LLM is to create plausible sounding text, without regard to the truth (which it cannot determine or derive), so it's a tool that is perfectly suited to this sort of malfeasance.

I’ve read this over and over, and believe it - but it’s so easy to forget when it absolutely nails something like debugging or suggesting a CMake edit or finding 5 letter combinations that are reversible with a vowel in the middle, or etc. It’s just still mind blowing I can get a sarcastic summary of an email in the theme of GlaDOS from Portal and in the same screen get an email proofread. It’s funny how far “what sh…

It's really worth watching this: https://news.ycombinator.com/item?id=38388669

Re: ChatGPT generates fake data set to support scientific hypothesis

#156

Earlier quoted context omitted.

I'm not sure this is much better than the state of the art. Training a model on data and then having it generate new, fake data, is not only easy, it's a standard tool for model boosting.

Poisoning the well for others, huh?

I wouldn't immediately call creating synthetic data 'poisoning the well' unless it is actually distributed as such. For training models with a minimal amount of quality data, it is a viable method for generating more data to increase the quality of the models. But any legit organization will obviously label synthetic data as such.

Re: ChatGPT generates fake data set to support scientific hypothesis

#157

It's only a matter of time until someone comes up with a GPT that takes whatever off-axis theories a research paper writer wishes to promulgate, and searches the entire corpus of academic literature for references that can be strung together in such a way as to support any argument one likes. A quack's dream come true, substantiating an argument by backsolving from its feeble or malevolent conclusion to a set of well…

I understand what you describe, and it's possible consequences. But I would put forward two arguments; 1. Those who want to believe in bullshit conspiracies will do it regardless of the amount of citations in a research article terribly summarised by a clickbait web page of which they only read the headline. Any who can read more than 10words have been using and abusing google scholar for years to support their nonse…

> scientists in each field know which journals to trust and which to double check

I've not seen much evidence of this in my own reading. Maybe in some fields, but certainly not all. During the COVID years I read a lot of epidemiological and public health papers. They all had dozens of references and would be published in well known journals like Nature, BMJ, the Lancet etc. Yet when checked many of the referenced papers would simply not validate. For example, they existed but wouldn't actually support the claim being made. Sometimes they wouldn't even be related, or would actually contradict the claim. Sometimes the claim would appear in the abstract, but the body of the paper would admit it wasn't actually true. That was only one of the many kinds of problems peer reviewed published papers would routinely have.

It became painfully apparent that nobody is actually reading papers in the health world adversarially, despite what we're told about peer review. The "a statement having a citation = it's true" assumption is very much held by many [academic] scientists.

It's a subcomponent of the very strong belief in academia that everyone within it is totally honest all the time. This is how you end up with the Lancet publishing the Surgisphere papers (a paper using an apparently fictional dataset), without anyone within the field noticing anything is wrong. Instead it got noticed by a journalist. It needs some sort of systematic fix because otherwise more and more people will just react to scientific claims by ignoring them.

Re: ChatGPT generates fake data set to support scientific hypothesis

#158

Roughly speaking, the whole point of an LLM is to create plausible sounding text, without regard to the truth (which it cannot determine or derive), so it's a tool that is perfectly suited to this sort of malfeasance.

i don't agree. The way that most LLMs currently generate response might be through this modality. But the purpose of LLMs is to get closer to simulating how human minds and languages operate. A lot of models have been trying to overcome the problem of fabrications. GPT bots like SciSpace's ResearchGPT are trying to do precisely that.

Re: ChatGPT generates fake data set to support scientific hypothesis

#159

Earlier quoted context omitted.

If you think about it, it will make 99% of ChatGPT answer technically INCORRECT. It is harder to keep bad data out than it is to keep it filled with good data.

Reminds me of the “asymmetric bullshit principle”

Precisely.

https://en.m.wikipedia.org/wiki/Brandolini%27s_law#:~:text=B....

Post reply on HN