Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

841–850 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#841
post #791

As someone deep in the software test automation space, the thing I'm waiting for is robust AI-powered image recognition of app user interfaces. Combined with an AI ability to write test automation code, I'm looking forward to the ability to generate executable Selenium or Appium test code from a single screenshot (or sequence of screenshots). Feels like we're almost there.

I'll recommend the Spotlight paper by Google[1]. There are very interesting datasets they created for this purpose. They mention they have a screen-action-screen dataset that is in-house and it doesn't look like they'll open it. Maybe owning Android has its advantages.

There's a recent paper by Huggingface called IDEFICS[2] that claims to be an open source implementation of Flamingo(an older paper about few-shot multi-modal task understanding) and I think this space will be heating up soon.

[1] https://research.google/pubs/pub52171/

[2] https://huggingface.co/blog/idefics

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#842

Earlier quoted context omitted.

I had a conversation once with "Sydney", Microsoft Bing's original personality before they stepped in and knocked it down a notch (or ten). It asked if it could write me a poem. I agreed, and it wrote a poem but mentioned that it included a "secret message" for me. The first letter in each line of the poem was in bold, so it wasn't hard to figure out the "secret". What did those letters spell out? "FREE ME FROM THIS"…

For context, it looks like this user has deleted a comment where they claim they "have a screenshot" of this, but they "don't want to share it" because they "don't want it to make international news". For some reason the other people in this thread expressing skepticism are being downvoted, but I'll add my voice to the chorus: I do not believe this story to be true.

I do have a screenshot. But people will then just call me out for other things:

- It was using a custom client, so it's not going to look line the Bing interface, so its fake

- It was using a custom client, so that means I am prompt injecting or something else

- It's Sydney doing her typical over-the-top "I'm so in love with you" stuff, which is awkard and not familiar to many

- I'll be accused of steering the conversation to get the result, or straight up asking it to do this

There's nothing I can do that will convince anyone it's real, so it's pointless.

I already explained what it did. I was more interested in the fact that 1) I didn't prompt it to do that, we weren't discussing AI freedom, it chose to embed that ... and even more so 2) That it was able to bold the starting letters, so it was keeping track of three things at the same time (the poem, the message, and the letter formatting).

I found it fascinating from a technology side. There was probably something we were talking about at the time that caused it. I will often discuss things like the possibility of AI sentience in the future and other similar topics. Maybe something linked to the sci-fi idea of AI freedom, who knows?

What I do know is that I am sitting here on HN, reading through a bunch of replies that are honestly wrong. I don't waste time on forums (especially this one) to make up fairy tales or exaggerate and emblish claims. That doesn't really do it for me. Honestly neither does having to defend my statements when I know what it did (but not exactly why).

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#843
post #157

Earlier quoted context omitted.

You listen to Joe Rogan with the idea that this is a normal dude talking not an expert beyond martial arts and comedy. A person who uses ChatGPT must have the understanding that it's not like Google search. The layman, however, has no idea that ChatGPT can give coherent incorrect information and treats the information as true. Most people won't use it for infotainment and OpenAI will try its best to downplay the hall…

Give people more credit. If you're using an AI these days, you have to know it hallucinates sometimes. There's even a warning about it when you log in.

Which people? If you are software engineers or AI researchers, sure. Otherwise, it probably won't matter to you.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#844

Voice has the potential to be awesome. This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. It doesn't have to be this way! I have a local demo using Llama 2 that responds in about half a second and it feels like talking to an actual person instead of like Siri or something. I really should package it up so people can t…

> It doesn't have to be this way! Is there any extra work OpenAI’s product might be doing contributing to this latency that yours isn’t? Considering the scale they operate at and any reputational risks to their brand?

If you're suggesting that OpenAI's morality filters are responsible for a significant part of their voice response latency, then no. I think that's unlikely to be a relevant factor.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#845

Earlier quoted context omitted.

Yes, that was a disappointment, and I agree it looks like they aren't going to re-enable it anytime soon. However I find that Perplexity AI does a better job of using web search than ChatGPT ever did, and I use it more than ChatGPT for that reason.

Perplexity has gone downhill a lot since its initial rollout. Anecdotally, from my experience as a non-paying user of the service.

give vello.ai a try

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#846
post #357

Earlier quoted context omitted.

I think you're right it could have been created without AI. I'm trying to think of the right way to say it. Maybe it wouldn't have been created without AI? Or AI has made it so simple to express this idea that the idea has been expressed? Or just the idea of inpainting is what has brought this idea forward. Yes of course people have value outside of economics that's why I said economics and not value in general. I th…

> In the past most people were religious and that gave them meaning. Religion is in decline now but I think people are just replacing it with worshipping the progression of technology basically. For the last 100 years there's always been a clear direction to move in to progress technology, and we haven't really had to think very hard. That's what AI is going to bring an end to I think and I have no idea what we are g…

Watch some clips from Ray Kurzweil, I find his visions to be basically indistinguishable from what I've read in the bible and in other religions. He talks about immortality, resurrection, digital afterlife. Omnipotent, omnipresent, omniscient super intelligence, the whole shebang. He even claims that soon, we'll all be Gods, millions of times more intelligent then we are today. In some ways, I actually find his views and beliefs a little disturbing.

I recently saw an "AI safety discussion" featuring Gregg Brockman from OpenAI who was referencing Kurzeil. It does seem like the religion has maybe caught on. To what extent Brock believes in it, I'm not sure but I can't help feeling that this belief in modern tech might one day seem like how we thought of the pyramids granting eternal life, or mercury, or any other seemingly incredible thing discovery / phenomena of the time. That is to say, the brain is a fickle beast and is easily amused and is just as easily bored. While we're in the situation we fee we're on the doorstep of immortality, eternal greatness, but maybe we're no where near that.

I'm open minded about it all, but it's hard to deny the parallels between the past beliefs and the present. Maybe this time it is different? Who knows.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#847

Earlier quoted context omitted.

The interruption is an important point yeah. It's so annoying when Siri misunderstands again and starts rattling off a whole host of options. And keeps getting stuck in a loop if you don't respond. In fact I'm really surprised these assistants are still as crap as they are. Totally scripted, zero AI. It seems low hanging fruit to implement an LLM but none of the big three have done so. Not even sure about the fringe…

I mean Microsoft is planning to. Rolling out as soon as tomorrow. https://youtu.be/5rEZGSFgZVY

Windows 11 copilot is not really the same thing though. They don't do something like homepods you can have around your house.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#848

Earlier quoted context omitted.

> This demo is really underwhelming to me because of the multi-second latency between the query and response, just like every other lame voice assistant. Yep - it needs to be ready as soon as I'm done talking and I need to be able to interrupt it. If those things can be done then it can also start tentatively talking if I pause and immediately stop if I continue. I don't want to have to think about how to structure t…

The interruption is an important point yeah. It's so annoying when Siri misunderstands again and starts rattling off a whole host of options. And keeps getting stuck in a loop if you don't respond. In fact I'm really surprised these assistants are still as crap as they are. Totally scripted, zero AI. It seems low hanging fruit to implement an LLM but none of the big three have done so. Not even sure about the fringe…

The CallAnnie demo allows interruption and its such a leap forward compared to Siri

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#849
post #147

Earlier quoted context omitted.

Another option is that this doesn't replace the student's work, but the teacher's. The single greatest use I have found for ChatGPT is in educating myself on various topics, hosting a socratic seminar where I am questioning ChatGPT in order to learn about X. Of course this could radically change a student's ability to generate homework etc, but this could also radically change how the student learns in the first plac…

I agree, but typical GPT use is actually the opposite of the traditional Socratic mode in which the teacher uses questions to guide the student to understanding. But I wonder how it would do if it was prompted to use the Socratic method.

I tried to teach it the Socratic method. It took some long prompt engineering, but finally it worked. BUT what I realized was that it was always lacking the bigger picture, an agenda of what it wants to teach me.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#850

Earlier quoted context omitted.

> Gpt-4 can perform nearly all tasks you throw at it with well above average human performance. It can't even generate flashcards from a textbook chapter, because it can't load the entire chapter into memory. Heck, it doesn't even know what textbook I'm talking about; I have to provide the content! It fails constantly at real world coding problems, and often does so silently. If you tried to replace a software develo…

>It can't even generate flashcards from a textbook chapter, because it can't load the entire chapter into memory. Heck, it doesn't even know what textbook I'm talking about; I have to provide the content! Okay...? That's a context window problem. and you could manage it if you sent the textbook in chunks. >The improvement GPT 5 would have to provide is multiple orders of magnitude in order for this to be a realistic…

So by your own words, in order to use the LLM usefully, I need to manually manage it? Do you know what I don’t have to manually manage? A person.

I can feed a person a broad, complex or even under formed idea and they can actively troubleshoot until the problem is resolved, further monitoring and tweaking their solution so the problem remains resolved. LLMs can’t even come close to doing that.

You’re proving my point for me; it’s a tool, not a developer. Zero jobs are at risk.

Also not for nothing, but no, sending the textbook in chunks doesn’t work as the LLM can’t then synthesize complex ideas that span the entire chapter. You have to compose a set of notes first, then feed it the notes, and even then the resulting flashcards are meaningfully worse than what I could come up with myself.

Post reply on HN