Live data from Hacker News

GPT Unicorn has drawn a unicorn

gpt-unicorn.adamkdean.co.uk

111–120 of 207 posts

Re: GPT Unicorn has drawn a unicorn

#111

It has become common knowledge that GPT4 (and also 3.5) have problems with deterministic outputs (even at T=0). So what we're seeing here is just the effect of random sampling, not any actual change to the model itself. If you scroll down, you'll see other close attempts by the exact same model that could already be counted as a win depending on who you ask. Edit: This comment section is a super fascinating case stud…

100%. This person is trying to find patterns in random noise and believes they are meaningful. The original post hurts my head with its bad logic.

>The original post hurts my head with its bad logic.

Huh? What "original post"? This is an experiment, today the model drew something resembling a unicorn. Tomorrow we will see how the experiment goes again. I see no associated analysis, so what makes your "head hurt".

Re: GPT Unicorn has drawn a unicorn

#112
post #71

Earlier quoted context omitted.

I believe the logic is fine. You seem to think you need multiple data points from the same version of the model (i.e. multiple samples per day at least) would be necessary to judge the actual performance on each particular day. That's worse logic. How would you visualize the very large sample you would get? Even with the current 118 samples (one per day) it's already difficult to find a pattern. Would you "average" t…

I agree that trying to determine the distribution of these drawings is hard because it isn't a simple floating point number in its current form. But maybe you could covert it to a linear monotonic measure? You could pass it to an image recognition model and see record the degree to which it thinks it is an: animal horse unicorn Basically if it fails to be a unicorn, see if it is a horse and if it fails to be a horse…

>You could pass it to an image recognition model and see record the degree to which it thinks it is...

We don't care what another algorithm "thinks". We want to see if what it draws is humanly interpretable as a unicorn.

Re: GPT Unicorn has drawn a unicorn

#113

It has become common knowledge that GPT4 (and also 3.5) have problems with deterministic outputs (even at T=0). So what we're seeing here is just the effect of random sampling, not any actual change to the model itself. If you scroll down, you'll see other close attempts by the exact same model that could already be counted as a win depending on who you ask. Edit: This comment section is a super fascinating case stud…

> This comment section is a super fascinating case study on the inherent flaws in human cognition. Especially when it comes to seeing patterns in random noise. The fact that some people believe that the model really has to have changed in the past few days is amazing You need only to look at the discourse around the Tesla FSD superusers to see this: they report a glitch at an intersection one day, then believe the ne…

Thats even worse for autonomous cars, there is so such data and noise there is no way to reproduce the issue, it's complete chaos. Whereas with a LLM if we control the seed we can 100% reproduce the same result

Re: GPT Unicorn has drawn a unicorn

#114

It has become common knowledge that GPT4 (and also 3.5) have problems with deterministic outputs (even at T=0). So what we're seeing here is just the effect of random sampling, not any actual change to the model itself. If you scroll down, you'll see other close attempts by the exact same model that could already be counted as a win depending on who you ask. Edit: This comment section is a super fascinating case stud…

I have not kept up with GPT architecture, other than noticing that other people have noticed that T=0 is clearly not deterministic for these things (and that that results from a bug, not an intentional feature). This much was obvious when the supposedly genius idiots rolled out GPT-3. It's wonderful to see the whole world bend over and just take it up the ass from a bunch of people who can't figure out why their code…

We’re pretty sure the nondeterminism is batching + mixture of experts + contention for specific experts

Re: GPT Unicorn has drawn a unicorn

#115

It has become common knowledge that GPT4 (and also 3.5) have problems with deterministic outputs (even at T=0). So what we're seeing here is just the effect of random sampling, not any actual change to the model itself. If you scroll down, you'll see other close attempts by the exact same model that could already be counted as a win depending on who you ask. Edit: This comment section is a super fascinating case stud…

yeah, see also image-2023-04-25 which is way earlier and comes really close, surrounded by garbage

Re: GPT Unicorn has drawn a unicorn

#116

Earlier quoted context omitted.

I have not kept up with GPT architecture, other than noticing that other people have noticed that T=0 is clearly not deterministic for these things (and that that results from a bug, not an intentional feature). This much was obvious when the supposedly genius idiots rolled out GPT-3. It's wonderful to see the whole world bend over and just take it up the ass from a bunch of people who can't figure out why their code…

We’re pretty sure the nondeterminism is batching + mixture of experts + contention for specific experts

If by batching you mean bad code that fails to sort or relies on hardware to best-guess how things sort, then sure, that's called a bug. Also, "we're pretty sure" is rather self-important while also admitting total, abject failure to produce a deterministic result. You shouldn't blame yourself. A lot of people had the same feeling after staking their life on the revolutionary properties of NFTs.

Re: GPT Unicorn has drawn a unicorn

#117

Earlier quoted context omitted.

100%. This person is trying to find patterns in random noise and believes they are meaningful. The original post hurts my head with its bad logic.

> The original post hurts my head with its bad logic. Huh? What "original post"? This is an experiment, today the model drew something resembling a unicorn. Tomorrow we will see how the experiment goes again. I see no associated analysis, so what makes your "head hurt".

https://adamkdean.co.uk/posts/gpt-unicorn-a-daily-exploratio...

> The idea behind GPT Unicorn is quite simple: every day, GPT-4 will be asked to draw a unicorn in SVG format. This daily interaction with the model will allow us to observe changes in the model over time, as reflected in the output.

Re: GPT Unicorn has drawn a unicorn

#118

Earlier quoted context omitted.

> This comment section is a super fascinating case study on the inherent flaws in human cognition. Especially when it comes to seeing patterns in random noise. The fact that some people believe that the model really has to have changed in the past few days is amazing You need only to look at the discourse around the Tesla FSD superusers to see this: they report a glitch at an intersection one day, then believe the ne…

Thats even worse for autonomous cars, there is so such data and noise there is no way to reproduce the issue, it's complete chaos. Whereas with a LLM if we control the seed we can 100% reproduce the same result

>> if we control the seed we can 100% reproduce the same result

No, that's the problem. You can't. You should be able to, but you can't. If you could, they wouldn't be scary. But we have Temperature Zero, different results. Because no one gave enough of a shit when coding them, and no one gives enough of a shit to try to fix the issue.

This is what in any other industry would be called gross negligence.

Re: GPT Unicorn has drawn a unicorn

#119
post #90

This is why, as a product manager, you should always test 20 hypotheses per month. At p-value of 0.05 this basically guarantees a successful product feature test every month!

Just make sure to stop immediately after the trial that validates the hypothesis.

Re: GPT Unicorn has drawn a unicorn

#120

Earlier quoted context omitted.

I agree that trying to determine the distribution of these drawings is hard because it isn't a simple floating point number in its current form. But maybe you could covert it to a linear monotonic measure? You could pass it to an image recognition model and see record the degree to which it thinks it is an: animal horse unicorn Basically if it fails to be a unicorn, see if it is a horse and if it fails to be a horse…

> You could pass it to an image recognition model and see record the degree to which it thinks it is... We don't care what another algorithm "thinks". We want to see if what it draws is humanly interpretable as a unicorn.

Then you could use mechanical turk to have them rate each image to figure out how close it is to a Unicorn...
Post reply on HN