Live data from Hacker News

GPT-4o

openai.com

451–460 of 1001 posts

Re: GPT-4o

#451
A test I've been using for each new version still fails.

Given the lyrics for Three Blind Mice, I try to get ChatGPT to create an image of three blind mice, one of which has had its tail cut off.

It's pretty much impossible for it to get this image straight. Even this new 4o version.

Its ability to spell in images has greatly improved, though.

Re: GPT-4o

#452

I found these videos quite hard to watch. There is a level of cringe that I found a bit unpleasant. It’s like some kind of uncanny valley of human interaction that I don’t get on nearly the same level with the text version.

While it is probably pretty normal for California, the insincere flattery and patronizing eagerness are definitely grating But then you have to stack that up against the fact that we are examining a technology and nitpicking over its tone of voice .

I feel like it's largely an effect of tuning it to default as "a ultra helpful assistant which is happy to help with any request via detailed responses in candid and polite manner..." kind of thing as you basically lose free points any time it doesn't jump on helping with something, tries to use short output and generates a more incorrect answer as a result, or just plain has to be initialized with any of this info.

It seems like both the voice and responses can be tuned pretty easily though so hopefully that kind of thing can just be loaded in your custom instructions.

Re: GPT-4o

#453
post #451

A test I've been using for each new version still fails. Given the lyrics for Three Blind Mice, I try to get ChatGPT to create an image of three blind mice, one of which has had its tail cut off. It's pretty much impossible for it to get this image straight. Even this new 4o version. Its ability to spell in images has greatly improved, though.

GPT-4o with image output is not yet available. So what did you even test? Dall-E 3?

Re: GPT-4o

#454
post #142

Too bad they consume 25x the electricity Google does. https://www.brusselstimes.com/world-all-news/1042696/chatgpt...

And in this 25x you get your answer.

What if we actually counted the electricity that the websites use instead of just the search engine page ?

Re: GPT-4o

#455
post #43
post #24

The most impressive part is that the voice uses the right feelings and tonal language during the presentation. I'm not sure how much of that was that they had tested this over and over, but it is really hard to get that right so if they didn't fake it in some way I'd say that is revolutionary.

(I work at OpenAI.) It's really how it works.

Will the new voice mode allow mixing languages in sentences?

As a language learner, this would be tremendously useful.

Re: GPT-4o

#456

Impressed by the model so far. As far as independent testing goes, it is topping our leaderboard for chess puzzle solving by a wide margin now: https://github.com/kagisearch/llm-chess-puzzles?tab=readme-o...

would love if you could do multiple samples or even just resampling and get a boostrapped CI estimate

Re: GPT-4o

#458

OAI just made an embarrassment of Google's fake demo earlier this year. Given how this was recorded, I am pretty certain it's authentic.

I don't doubt this is authentic, but if they really wanted to fake those demos, it would be pretty easy to do using pre-recorded lines and staged interactions.

For what it's worth, OpenAI also shared videos of failed demos:

https://vimeo.com/945591584

I really value how open they are being about its limitations.

Re: GPT-4o

#460

I admit I drink the koolaid and love LLMs and their applications. But damn, the way it’s responds in the demo gave me goosebumps in a bad way. Like an uncanny valley instincts kicks in.

I also thought the screwups, although minor, were interesting. Like when it thought his face was a desk because it did not update the image it was "viewing". It is still not perfect, which made the whole thing more believable.

I was shocked at how quickly and naturally they were able to correct the situation.
Post reply on HN