Live data from Hacker News

ChatGPT generates fake data set to support scientific hypothesis

nature.com

141–150 of 159 posts

Re: ChatGPT generates fake data set to support scientific hypothesis

#141
post #130

Earlier quoted context omitted.

A null result, or the absence of evidence, is not evidence of absence. If you fake a null result, you’re not asserting anything other than you could not measure and collect supporting data using your experiment to prove or disprove a hypothesis. It is difficult for someone doing replication to accuse you of ill-intent, as opposed to faked data that proves your point when anyone else can replicate your experiment and…

You still have to give out the statistics that show the null result. E.g. something with a high p-value. You are in fact "confirming the null hypothesis". They aren't any more difficult to replicate than results supporting "the alternative hypothesis". (The whole binary hypothesis system and culture is a mess though, but that's besides the point.)

This is true.

However, I think that no one will do the scut work necessary to find that a null result was faked, and even if they do since you the researcher got very little status out of it then it’s believable that you made a mistake, and didn’t falsify data.

Re: ChatGPT generates fake data set to support scientific hypothesis

#142
post #55

Earlier quoted context omitted.

"Humanity learns" is impossible. The unit of analysis is the individual. Some state may cross between individuals via education, but the individuals still must learn. History shows that knowledge transmission remains a sticky wicket.

humanity learns just mean sufficient individuals learn to take power and enact such a society

Yes, but my point us that repeating "humanity learns" can lead us down garden paths into thinking there is some species-level recollection, when history reveals a mixed bag at best.

Re: ChatGPT generates fake data set to support scientific hypothesis

#143
post #113

How long til I can produce a paper that says smoking is good for you, and will extend life or increase quality of life? ;-)

There are plenty such papers out there - now obviously disproven. At some point smoking was prescribed as a cure or various heaelty issues. Letting you know so you don't feel impressed by what was "created" when a chat bot generates such a paper.

Re: ChatGPT generates fake data set to support scientific hypothesis

#144

Earlier quoted context omitted.

You still have to give out the statistics that show the null result. E.g. something with a high p-value. You are in fact "confirming the null hypothesis". They aren't any more difficult to replicate than results supporting "the alternative hypothesis". (The whole binary hypothesis system and culture is a mess though, but that's besides the point.)

This is true. However, I think that no one will do the scut work necessary to find that a null result was faked, and even if they do since you the researcher got very little status out of it then it’s believable that you made a mistake, and didn’t falsify data.

Probably not, especially as null results are almost impossible to publish (which is mad).

Re: ChatGPT generates fake data set to support scientific hypothesis

#145
post #62

It's only a matter of time until someone comes up with a GPT that takes whatever off-axis theories a research paper writer wishes to promulgate, and searches the entire corpus of academic literature for references that can be strung together in such a way as to support any argument one likes. A quack's dream come true, substantiating an argument by backsolving from its feeble or malevolent conclusion to a set of well…

This is already possible with search engines, there is enough information on the internet that you can substantiate just about any claim regardless of how much evidence there is to the contrary. (see flat-earth, plenty of plausible sounding claims with real, albeit, cherry picked evidence).

AI criticism is essentially people claiming that having access to something they don't like will end the world. As you say, we already have a good example of this and while it is mostly bad and getting worse it's not world-ending.

Re: ChatGPT generates fake data set to support scientific hypothesis

#146

It's only a matter of time until someone comes up with a GPT that takes whatever off-axis theories a research paper writer wishes to promulgate, and searches the entire corpus of academic literature for references that can be strung together in such a way as to support any argument one likes. A quack's dream come true, substantiating an argument by backsolving from its feeble or malevolent conclusion to a set of well…

I understand what you describe, and it's possible consequences. But I would put forward two arguments; 1. Those who want to believe in bullshit conspiracies will do it regardless of the amount of citations in a research article terribly summarised by a clickbait web page of which they only read the headline. Any who can read more than 10words have been using and abusing google scholar for years to support their nonsense, the ability to find a reference does not equate to the ability to critically appraise the contents. 2. Science is a small world. In each specific field you get to know the big names and institutes, and those are used as a better gauge of the quality of a paper. The peer-review, for all it's pitfalls including its dependence on volunteers, does a good job at stemming a lot of bullshit. I'd proffer it's not the articles but rather the multiple for-profit publishing houses setting up multiple journals through which they funnel pay-to-publish articles that are contributing to the dilution of trust in published science. Again on that note, scientists in each field know which journals to trust and which to double check.

Re: ChatGPT generates fake data set to support scientific hypothesis

#147

Earlier quoted context omitted.

I think that by that same logic, you could say human artists, writers, etc. don't create either, they just move existing matter (which typically isn't created nor destroyed) from one place to another. You could also say an electric company doesn't generate electricity but merely converts other energy into it -- one of the most popular uses of the word "generator" yet the atoms/energy are already here; we just arrange…

We think objectively about it. We have goals and intentions. We use logic. LLMs don't. The dirt in the ground sorts impurities from water, but we don't call it intelligent or generative. We call it entropy.

This is a human supremacy argument, essentially claiming that because we have a "soul" or because of something inherent in us that cannot be proven we are better than something else. You are free to believe this, but it is a matter of faith. Not of any sort of reasoning.

Re: ChatGPT generates fake data set to support scientific hypothesis

#148
post #32

Earlier quoted context omitted.

Exactly. It's no different from the fake data sets created by hand in scientific misconduct cases for years, just I guess easier. I guess that's not a good thing, but given even making a fake data set by hand is far easier than generating real data, I'm not sure if this will suddenly make more people fake data.

The fake data that's been caught so far in just about every one of the cases that I've read about has been ridiculously poorly constructed fake data. I think the people that are good at faking data simply don't get caught.

Having spent significant time implementing machine learning papers before the LLM age, I can promise you over 90% of papers you'll find are full of shit. The claims they make are true in only the most contrived of circumstances and don't hold up under any kind of scrutiny. How exactly they came to these lies (data lies, result lies, omitting lies) is really immaterial. The concept everyone is apparently struggling with is that producing a paper that is entirely lies is not doing the scientific world a disservice: It is not unusual and already happens at scale. Making it even easier might motivate someone to actually figure out a way to ensure papers are reproducible and not full of shit. In essence this is a good thing.

Re: ChatGPT generates fake data set to support scientific hypothesis

#149
post #121

Earlier quoted context omitted.

The easiest data to fake is null results because anything else is replicable, hence the importance of replication.

null results are also replicable.

But you won't find null results in a journal, so problem solved.

Re: ChatGPT generates fake data set to support scientific hypothesis

#150
post #55

Earlier quoted context omitted.

humanity learns just mean sufficient individuals learn to take power and enact such a society

Yes, but my point us that repeating "humanity learns" can lead us down garden paths into thinking there is some species-level recollection, when history reveals a mixed bag at best.

Completely agree transmission of knowledge is a sticky wicket. Good reminder to stay logistically grounded. I have found many of the ambitious thinkers close to me tend to aspire to a "humanity learns" moment but if they're anywhere near politics they tend to be tempered fairly well by the logistical realities of bringing ideas to pass
Post reply on HN