Live data from Hacker News

Introducing ChatGPT and Whisper APIs

openai.com

271–280 of 696 posts

Re: Introducing ChatGPT and Whisper APIs

#271
post #33

Earlier quoted context omitted.

I thought it was Midjourney who stole their thunder. Stable Diffusion is free but it's much harder to get good results with it. Midjourney on the other hand spits out art with a very satisfying style.

You are like 2 months out of date. Stable diffusion now has a massive ecosystem around it (civitai/automatic1111), that when used well, completely crushes any competitors in terms of the images it produces. Midjourney is still competitive, but mostly because its easier to use. Dalle2 will get you laughed out of the room in any ai art discussion.

>Dalle2 will get you laughed out of the room in any ai art discussion.

and claiming AI art is art would get you laughed out of any art discussion.

personally I think AI art is really cool, but to discount what Dalle 2 did for AI art is unfair.

Re: Introducing ChatGPT and Whisper APIs

#272

I can't seem to find it mentioned but is there an unrestricted/uncensored mode for this? I'd love to have some fun with a few friends and hook it up in a matrix room for us

https://news.ycombinator.com/item?id=34972791

That's just a list of exploits that will be fixed as soon as they come to OpenAI's attention, if they haven't already been fixed. Is anyone actually committed to providing uncensored models as either paid services or open distributions?

Re: Introducing ChatGPT and Whisper APIs

#273
post #88

Whisper as an API is great, but having to send the whole payload upfront is a bummer. Most use cases I can build for would want streaming support. Like establish a WebRTC connection and stream audio to OpenAI and get back a live transcription until the audio channel closes.

FWIW, AssemblyAI has great trasncript quality in my experience, and they support streaming: https://www.assemblyai.com/docs/walkthroughs#realtime-stream...

We're using AssemblyAI too, and I agree that their transcription quality is good. But as soon as Whisper supports world-level timestamps, I think we'll seriously consider switching as the price difference is large ($0.36 per hour vs $0.9 per hour).

Re: Introducing ChatGPT and Whisper APIs

#274

Earlier quoted context omitted.

I can only see it affecting 'art', where you might want to have characters that are despicable say despicable things. But really we shouldn't be using AI to make our art for us anyway. Help, sure, but it shouldn't be literally writing our stories.

So you feel that when progress enables us to provide more abundance for humanity, we should artificially limit that abundance for everyone so that a few people aren't inconvenienced?

[flagged]

Re: Introducing ChatGPT and Whisper APIs

#275

Feels like people running small websites monetised with ads will get killed. Why go to a recipe website to search for healthy meals for your kids if you can ask chat gpt in Instacart? :(

No more scrolling through a thousand lines of SEO optimization disguised as some deep heartfelt backstory before getting to the actual recipe? That sounds like a huge win to me.

My love for the perfect Tuna Salad Sandwich began when I was but a young IPython notebook. My lead developer would occasionally eat at the desk and crumbs would fall into the keyboard....

Re: Introducing ChatGPT and Whisper APIs

#276

Earlier quoted context omitted.

Yeah, it used to be I'd set Google results to just one year back, now I'm having to set it to one month.

Can you explain what you do that for?

Avoiding out of date advice, also filtering just for newest trends and techniques

Re: Introducing ChatGPT and Whisper APIs

#277
post #33

Earlier quoted context omitted.

I thought it was Midjourney who stole their thunder. Stable Diffusion is free but it's much harder to get good results with it. Midjourney on the other hand spits out art with a very satisfying style.

You are like 2 months out of date. Stable diffusion now has a massive ecosystem around it (civitai/automatic1111), that when used well, completely crushes any competitors in terms of the images it produces. Midjourney is still competitive, but mostly because its easier to use. Dalle2 will get you laughed out of the room in any ai art discussion.

I must be horribly out of date then - I thought Midjourney was the cut down DALL-E approximation, created to givr something to play with to people who couldn't get on the various waiting lists, or can't afford to run SD on their own.

Re: Introducing ChatGPT and Whisper APIs

#278
Right on time.

Google's Speech-to-Text is $0.024 per minute ($0.016 per minute with logging) with 60 free minutes per month. Files below 1 minute can be posted to the server, anything longer needs to be uploaded into a bucket, which complicates things, but at least they're GDPR compliant.

Whisper is $0.006 per minute with the following data usage policies

- OpenAI will not use data submitted by customers via our API to train or improve our models, unless you explicitly decide to share your data with us for this purpose. You can opt-in to share data.

- Any data sent through the API will be retained for abuse and misuse monitoring purposes for a maximum of 30 days, after which it will be deleted (unless otherwise required by law).

I've been using Whisper on a server (CPU only) to transcribe recordings made during a bike ride with a lavalier microphone, so it's pretty noisy due to the wind and the tires and Whisper was better than Google.

Plus, Whisper, when used with `response_format="verbose_json"`, outputs the variables `temperature`, `avg_logprob`, `compression_ratio`, `no_speech_prob` which can be used very effectively to filter out most of the hallucinations.

A one minute file which transcodes in 26 seconds on a CPU is done in 6 seconds via this service. Another one minute file with a lot of "silence" needs around 56 seconds on a CPU and was ready in 4.3 seconds via the service. "Silence" means that maybe 5 seconds of the file contain speech while the rest is wind and other environmental noises. Another relatively silent one went from 90 seconds down to 5.4. On the CPU I was using the medium model while the service is using large-v2

A couple of days ago I posted an example to a thread [0], where I was getting the following with Whisper

---

00:00.000 --> 00:05.000 Also temperaturmäßig ist es recht gut. [So temperature wise, it's pretty good.]

00:05.000 --> 00:09.000 Der eine hat 12 Grad, der andere 10. [One has 12 degrees, the other 10. (I have two temperature sensors mounted on the bike, ESP32 streaming the data to the phone via BLE)]

00:09.000 --> 00:12.000 Also sagen wir mal, 10 Grad. [So let's say 10 degrees.]

00:14.000 --> 00:19.000 Es ist bewölkt und windig. [It's cloudy and windy.]

00:20.000 --> 00:24.000 Aber irgendwie vom Wetter her gut. [But somehow from the weather it's good.]

00:24.000 --> 00:31.000 Ich habe heute überhaupt nichts gegessen und sehr wenig getrunken. [I ate nothing at all today and drank very little.]

00:54.000 --> 00:59.000 Vielen Dank für's Zuschauen! [Thanks for watching!] ---

While Google was outputting

"Also temperaturmäßig es ist recht gut, der eine hat 12° andere 10. Es ist angemalte 10 Grad. Es ist bewölkt und windig, aber er hat sie vom Wetter her gut, ich wollte überhaupt nichts gegessen und sehr wenig getrunken."

["So temperature-wise it's pretty good, one has 12° other 10. It's painted 10 degrees. It's cloudy and windy, but he has it good from the weather, I did not want to eat anything at all and drank very little."]

---

Apart from the hallucinated line, Whisper got everything correct, and the hallucinated line was able to be discarded due to the variables like `avg_logprob`.

[0] https://news.ycombinator.com/item?id=34877020#34880531

Re: Introducing ChatGPT and Whisper APIs

#279
post #208

Earlier quoted context omitted.

If you think you're overpaying just hit the API yourself.

Any idea how to encode the previous messages when sending a followup question? E.g.: 1. I ask Q1 2. API responds with A1 3. I ask Q2, but want it to preserve Q1 and A1 as context Does Q2 just prefix the conversation like this? „I previously asked {Q1}, to which you answered {A1}. {Q2}“

https://platform.openai.com/docs/guides/chat/introduction

"The main input is the messages parameter. Messages must be an array of message objects, where each object has a role (either “system”, “user”, or “assistant”) and content (the content of the message). Conversations can be as short as 1 message or fill many pages."

"Including the conversation history helps when user instructions refer to prior messages. In the example above, the user’s final question of “Where was it played?” only makes sense in the context of the prior messages about the World Series of 2020. Because the models have no memory of past requests, all relevant information must be supplied via the conversation. If a conversation cannot fit within the model’s token limit, it will need to be shortened in some way."

So it looks like you pass in the history with each request.

Re: Introducing ChatGPT and Whisper APIs

#280

Big news! Many apps will be integrating ChatGPT. Worried about AI-generated content flood the search engine, making it harder to do in-depth research.

This is a good thing.

The future is curation and cultivation. We've been living in an age of information abundance and markets haven't adapted. The age of "crawl every website, index everything, and let people search it" is coming to an end. There is just too much content and too much of it is low quality. With or without AI.

This abundance problem isn't just a WWW problem. Movies, TV, music, podcasts, short form content, food, widgets, wibbles and wobbles all suffer from abundance these days. We are quickly exiting the age of supply chain driven scarcity and getting a marketplace flooded with options. Capitalism has delivered on basically everything it's promised with some asterisks and, if we don't give into consumerism, we want for little and have everything we need at our fingertips.

I've personally opened up my pocket book to curation services. I know brands that I trust. I know services that reliably surface quality content. I suspect the next few decades are going to trend towards services that separate noise from signal - and I suspect AI is going to be a big part of that.

Post reply on HN