Live data from Hacker News

We are beginning to roll out new voice and image capabilities in ChatGPT

openai.com

821–830 of 914 posts

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#821
post #600

Earlier quoted context omitted.

That bike example seemed a mix of underwhelming (for being the demo video) and even confusing. 1. It's not smart enough to recognize from the initial image this is a bolt style seat lock (which a human can). 2. The manual is not shown to the viewer, so I can't infer how the model knows this is a 4mm bolt (or if it is just guessing given that's the most likely one). 3. I don't understand how it can know the toolbox is…

The bike shown in the first image is Specialized Sirrus X. You can make out from the image of the manual that it says "spacer/axle/bolt specifications". Searching for this yields the following Specialized bike manual which is similar: https://www.manualslib.com/manual/1974494/Specialized-Epic-E... -- there are some notable differences, but the Specialized Sirrus X manuals that are online aren't in the same style. The…

It bugged me that they made no mention of torque. The manual is really clear on that part with a big warning:

> WARNING! Correct tightening force on fasteners (nuts, bolts, screws) on your bicycle is important for your safety. If too little force is applied, the fastener may not hold securely. If too much force is applied, the fastener can strip threads, stretch, deform or break. Either way, incorrect tightening force can result in component failure, which can cause you to lose control and fall. Where indicated, ensure that each bolt is torqued to specification. The following is a summary of torque specifications in this manual...

The seat collar also probably has the max torque printed on it.

When they asked if they had the right tool, I would have preferred to see an answer along the lines of "ideally you should be using a torque wrench. You can use the wrench you have currently, but be careful not to over tighten."

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#822
post #815

I keep hoping to be able to give it a jpg of handwritten text and it'll give me back ASCII text.

This... would be amazing. Handwritten OCR has been hit or miss, requiring a collection of penstroke data for most recognizers to work, and they work poorly at that.

It strikes me as an ideal task for AI.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#823

Earlier quoted context omitted.

I always thought a better future would be full of more and more distilled, accurate, useful knowledge and truthful people to promote that. Comments like yours make me think that no one cares about this...and judging by a lot of the other comments, I guess they don't. Probably going to be people, wading through a sea of AI generated shit, and the individual is supposed to just forever "apply critical thinking" to it a…

There aren't any real world sources of truth you can avoid applying critical thinking to. Much published research is false, and when it isn't, you need to know when it's expired or what context it's valid in.

But do we need 9999999x the amount of information to critically be thinking about, is this going to be helpful ?

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#824

Earlier quoted context omitted.

Since this is getting a bit of interest, here's one more demo of this https://youtu.be/cvKUa5JpRp4 This demo shows even lower latency, plus the ability to handle very large menus with lots of complicated sub-options (this restaurant has over a billion option combinations to order a coffee). The latency is negative in some places, meaning the system finishes predicting before I finish speaking.

This is cool. But I want to see how it handles you going back one and tweaking it.

We've built something similar that allows you to tweak/update notes & reminders https://qwerki.com/ (private beta) here's the video demo https://www.youtube.com/shorts/2hpBTxjplIE we've since moved to training our own LLAMA as it's more responsive & we have better reliability.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#825

Earlier quoted context omitted.

you're absolutely right. also, I did not know until today's thread that OpenAI's stated goal is building AGI. which is probably never going to happen, ever, no matter how good technology gets. which means yes, we are absolutely looking at AltaVista here, not Google, because if you subtract a cult from an innovative business, you might be able to produce a profitable business.

Why isn’t AGI ever going to happen? Ever?

Because the goalposts are currently somewhere near Neptune, and expected to catch up to Voyager sometime in the next couple years.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#826
post #777

Earlier quoted context omitted.

Don’t they do it already? There are a lot of languages where intonation is absolutely necessary to distinguish between some words, so I would be surprised that this not already taken into account by the major voice assistants.

In English, intonation changes the meaning of the word but not the word itself. From what I understand, in tonal languages tone changes the whole word. I don't think ML understands that difference yet.

Yeah they do. I was able to get ChatGPT-4 to transcribe 我哥哥高過他的哥哥, which says that they can. I did have to set the app to Chinese, and the original didn't work so I had to modify what I said slightly.

https://www.tiktok.com/t/ZT86psPxY/

Roughly translated, my older brother is taller than that other guy's older brother.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#827

Earlier quoted context omitted.

GPT-4 is not the same product. I know it seems like it due to the way they position 3.5 and 4 on the same page, but they are really quite separate things. When I signed up for ChatGPT plus I didn't even bother using 3.5 because I knew it would be inferior. I still have only used it a handful of times. GPT-4 is just so much farther ahead that using 3.5 is just a waste of time.

Would you mind sharing some threads where you thought ChatGPT was useful? These discussions always feel like I’m living on a different planet with a different implementation of large language models than others who claim they’re great. The problems I run into seem to stem from the fundamental nature of this class of products.

Here's a convo I had yesterday when thinking about how to print a Binary Search Tree.

https://chat.openai.com/share/338e7397-0201-44f4-a2c3-75b733...

I use ChatGPT for all sorts of things - looking into visas for countries, coding, reverse engineering companies from job descriptions, brainstorming etc etc.

It saves a lot of time and gives way more value than what you pay for it.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#828
post #731
post #327

Earlier quoted context omitted.

I don't disagree but having to have a learning phase for patterns sounds a bit like people clinging to an old way of things.

You mean like prompt engineering? What you’re describing as “clinging to an old way of things” is how every single thing has been, ever, new or old.

I don't know why you come here and say something so obviously untrue.

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#829
post #603
post #327

Earlier quoted context omitted.

I don't disagree but having to have a learning phase for patterns sounds a bit like people clinging to an old way of things.

It’s better to have a pattern than having no pattern with ChatGPT to tell when it’s hallucinating or not. I wish MLs were more useful than search engines, but they have still a long way to go to replace them (if they ever do).

Google still thinks I want to click on the sites I haven't clicked on in a decade even though they are first results. Search engines have a long way to go to catch up to GPT

Re: We are beginning to roll out new voice and image capabilities in ChatGPT

#830

Earlier quoted context omitted.

In English, intonation changes the meaning of the word but not the word itself. From what I understand, in tonal languages tone changes the whole word. I don't think ML understands that difference yet.

Yeah they do. I was able to get ChatGPT-4 to transcribe 我哥哥高過他的哥哥, which says that they can. I did have to set the app to Chinese, and the original didn't work so I had to modify what I said slightly. https://www.tiktok.com/t/ZT86psPxY/ Roughly translated, my older brother is taller than that other guy's older brother.

Of course speech recognition works for Chinese. What it doesn't do is transcribe intonation and prosody in non-tonal languages. It's not even clear how one would transcribe such a thing as I'm not aware of a standard notation.
Post reply on HN