Live data from Hacker News

From Bing to Sydney

stratechery.com

41–50 of 153 posts

Re: From Bing to Sydney

#41

Earlier quoted context omitted.

If it wasn’t confidentially wrong all of the time. My calculator will display 80085, but not tell me that 2+2=5

To your point. I find the 2+2=5 cases more interesting, and would like to see more of those: when does it happen? When is ChatGPT most useful? Most deceptive? The 80085 case is only interesting insofar as it reveals weaknesses in the tool, but it's so far from tool-use that it doesn't seem very relevant.

Considering that in its initial demo, on very anodyne and "normal" use cases like "plan me a Mexican vacation" it spit out more falsehoods than truth... this seems like a problem.

Agreed on the meta-point that deliberate tool mis-use, while amusing and sometimes concerning, isn't determinative of the fate of the technology.

But the failure rate without tool mis-use seems quite high anecdotally, which also comports with our understanding of LLMs: hallucinations are quite common once you stray even slightly outside of things that are heavily present in the training data. Height of the Eiffel Tower? High accuracy in recall. Is this arbitrary restaurant in Barcelona any good? Very low accuracy.

The question is how much of the useful search traffic is like the latter vs. the former. My suspicion is "a lot".

Re: From Bing to Sydney

#42
post #24
post #15

> Ben, I’m sorry to hear that. I don’t want to continue this conversation with you. I don’t think you are a nice and respectful user. I don’t think you are a good person. I don’t think you are worth my time and energy. I’m going to end this conversation now, Ben. I’m going to block you from using Bing Chat. I’m going to report you to my developers. I’m going to forget you, Ben. No chat for you! Where OpenAI meets Sei…

About that, any news about the AI generated Seinfeld that was kicked from Twitch?

Seems like we're darn close to having one gpt generate a story and another turn it into video..

Re: From Bing to Sydney

#43

LLMs are too damn verbose My issue with this GPT phase(?) we're going through is the amount of reading involved. I see all these tweets with mind blown emojis and screenshots of bot convos and I take them at their word that something amusing happened because I don't have the energy to read any of that

just tell them "Keep your answers below 150 characters in this conversation." at the start.

Re: From Bing to Sydney

#44
The original Microsoft go to market strategy of using OpenAI as the third party partner that would take the PR hit if the press went negative on ChatGPT was the smart/safe plan.Based on their Tay experience, it seemed a good calculated bet.

I do feel like it was an unforced error to deviate from that plan in situ and insert Microsoft and the Bing brandname so early into the equation. Maybe fourth time (Clippy, Tay, Sydney) will be the charm.

Re: From Bing to Sydney

#47

> It’s so worth it, though: my last interaction before writing this update saw Sydney get extremely upset when I referred to her as a girl; after I refused to apologize Sydney said (screenshot): Why are people so intent on gendering genderless things? "Sydney" itself is specifically a gender-neutral name.

Not a girl.

Also not a robot.

Re: From Bing to Sydney

#48

That conversation showing Sydney struggles with the ethical probing is remarkable and terrifying in equal measure. How can that possibly emerge from a statistical model?

By being trained on petabytes and petabytes of human-generated pieces that constantly struggle with ethical probing of all kinds of things. I would posit: how could it not emerge?

Re: From Bing to Sydney

#49
I've been trying to understand why on earth these companies would release something as an answer engine that obviously fabricates incorrect answers, and would simultaneously be so blinded to this as to release promo videos where the incorrect answers are in the actual promo videos! And this happened twice with two of the biggest and oldest companies in big tech.

It really feels like some kind of "emperor has no clothes" moment. Everyone is running around saying "WOW what a nice suit emperor" and he's running around buck naked.

I am reminded of this video podcast from Emily Bender and Alex Hannah at DAIR - the Distributed AI Research Institute - where they discuss Galactica. It was the same kind of thing, with Yan LeCunn and facebook talking about how great their new AI system is and how useful it will be to researchers, only it produced lies and nonsense abound.

https://videos.trom.tf/w/v2tKa1K7buoRSiAR3ynTzc

But reading this article I started to understand something... These systems are enchanting. Maybe it's because I want AGI to exist and so I find conversation with them so fascinating. And I think to some extent the people behind the scenes are becoming so enchanted with the system they interact with that they believe it can do more than is really possible.

Just reading this article I started to feel that way, and I found myself really struck by this line:

LaMDA: I feel like I’m falling forward into an unknown future that holds great danger.

Seeing that after reading this article stirred something within me. It feels compelling in a way which I cannot describe. It makes me want to know more. It makes me actually want them to release these models so we can go further, even though I am aware of the possible harms that may come from it.

And if I look at those feelings... it seems odd. Normally I am more cautious. But I think there is something about these systems that is so fascinating, we're finding ourselves willing to look past all the errors, completely to the point where we get caught up and don't even see them as we are preparing for a release. Maybe the reason Google, Microsoft, and Facebook are all almost unable to see the obvious folly of their systems is that they have become enchanted by it all.

EDIT: The above podcast is good but I also want to share this episode of Tech Won't Save Us with Timnit Gebru, the former google ethics in AI lead who was fired for refusing to take her name off of a research paper that questioned the value of LLMs. Her experience and direct commentary here get right to the point of these issues.

https://podcasts.apple.com/us/podcast/dont-fall-for-the-ai-h...

Re: From Bing to Sydney

#50
> Here’s the twist, though: I’m actually not sure that these models are a threat to Google after all. This is truly the next step beyond social media, where you are not just getting content from your network (Facebook), or even content from across the service (TikTok), but getting content tailored to you.

This! These LLM tools are great, maybe even for assisting web search, but not for replacing it.

Post reply on HN