Live data from Hacker News

From Bing to Sydney

stratechery.com

141–150 of 153 posts

Re: From Bing to Sydney

#141

Ben’s got it just right. These things are terrible at the knowledge search problems they’re currently being hyped for. But they’re amazing as a combination of conversational partner and text adventure. I just asked ChatGPT to play a trivia game with me targeted to my interests on a long flight. Fantastic experience, even when it slipped up and asked what the name of the time machine was in “Back to the Future”. And t…

IMO it's only a matter of time before someone hooks up a LLM to a speech-to-text recognizer with a TTS engine like something from ElevenLabs, and you have a full blown "AI" that you can converse with. Once someone builds a LLM that can remember facts tied to your account this thing is going to go off the rails.

Funny you mention that… I have done exactly this. Including using ElevenLabs for TTS. And also teaching it “facts” about me / calendar for the day in a hidden prompt when launching a conversation. It works pretty well.

Re: From Bing to Sydney

#142

Earlier quoted context omitted.

To your point. I find the 2+2=5 cases more interesting, and would like to see more of those: when does it happen? When is ChatGPT most useful? Most deceptive? The 80085 case is only interesting insofar as it reveals weaknesses in the tool, but it's so far from tool-use that it doesn't seem very relevant.

Considering that in its initial demo, on very anodyne and "normal" use cases like "plan me a Mexican vacation" it spit out more falsehoods than truth... this seems like a problem. Agreed on the meta-point that deliberate tool mis-use, while amusing and sometimes concerning, isn't determinative of the fate of the technology. But the failure rate without tool mis-use seems quite high anecdotally, which also comports wi…

> But the failure rate without tool mis-use seems quite high anecdotally

The problem with your judgement is you click on every “haw haw, ChatGPT dumb” and you don’t read any of the articles that show how an LLM works, what is is quantitatively good at and bad at and how to improve performance on tasks using other methods such as PAL, Toolformer or other analytic augmentation methods.

Go read some objective studies and you won’t be yet another servomechanism blindly spreading incorrect assumptions based on anecdotes from attention starved bloggers.

Re: From Bing to Sydney

#143

All these ChatGPT gone rogue screenshots create interesting initial debate, but I wonder if it's relevant to their usage as a tool in the medium term. Unhinged Bing reminds me of a more sophisticated and higher-level version of getting calculators to write profanity upside down: funny, subversive, and you can see how prudes might call for a ban. But if you're taking a test and need to use a calculator, you'll still u…

What sort of profanity can you write on a calculator?

Re: From Bing to Sydney

#144
post #63

> It’s so worth it, though: my last interaction before writing this update saw Sydney get extremely upset when I referred to her as a girl; after I refused to apologize Sydney said (screenshot): Why are people so intent on gendering genderless things? "Sydney" itself is specifically a gender-neutral name.

> Why are people so intent on gendering genderless things? I heard there are entire languages which do that everywhere...

I speak one such language. That language includes a "neutral" gender to describe things in non-gendered terms and has neat built-in features like using "they/them" to refer to a person whose gender is unknown.

Re: From Bing to Sydney

#145
post #65

Earlier quoted context omitted.

It can summarize its own output, the user directs everything about the output, style, format, length, etc. Everything.

> style I like asking to to type like a frustrated teen on the phone. it huffs and puffs and rolls its virtual eyes. prompt: could you pick a quantum computer at the mall for me? response: ugh, seriously? you can't just buy a quantum computer at the mall, they're like super expensive and only a few companies sell them. Plus, they require special conditions to operate.

> Ugh, seriously? Like, I can't even with this. I don't know why you're making me come to the mall just to pick out a quantum computer that you're not even gonna use properly. And you had to go and choose the one that uses liquid helium? That's, like, so old school. It's like listening to classic music while everyone else is jamming to something modern and cool. Do you want to keep using your classic computer too while you're at it? rolls eyes

Re: From Bing to Sydney

#146
post #124
post #110

Earlier quoted context omitted.

>However, the complexity of our probabalistic word machine is far greater, in terms of both richness of inputs, motivation, and dimensionality. If thought (as expressed in language) is just probabilistic pattern matching, then how did we develop our own training data from scratch?

There is a huge universe of inputs, aka training data, that feeds into us, far more than a digital text based LLM. From that we generated the training data for the LLM. That data is just a sliver of the human experience.

The universe contained exactly 0 words until humans created them, so if we are just stringing together words, then how did we make the words?

Re: From Bing to Sydney

#147

Earlier quoted context omitted.

Considering that in its initial demo, on very anodyne and "normal" use cases like "plan me a Mexican vacation" it spit out more falsehoods than truth... this seems like a problem. Agreed on the meta-point that deliberate tool mis-use, while amusing and sometimes concerning, isn't determinative of the fate of the technology. But the failure rate without tool mis-use seems quite high anecdotally, which also comports wi…

> But the failure rate without tool mis-use seems quite high anecdotally The problem with your judgement is you click on every “haw haw, ChatGPT dumb” and you don’t read any of the articles that show how an LLM works, what is is quantitatively good at and bad at and how to improve performance on tasks using other methods such as PAL, Toolformer or other analytic augmentation methods. Go read some objective studies an…

Hi, I work on LLMs daily, along with some intensely talented, skilled, and experienced machine learning engineers who also work on LLMs daily. My opinion is formed by both my own experiences with LLMs as well as the opinions of those experts.

Wanna try again? Alternatively you can keep riding the hype train from techfluencers who keep promising the moon but failing to deliver, just like they did for crypto.

Re: From Bing to Sydney

#148

All these ChatGPT gone rogue screenshots create interesting initial debate, but I wonder if it's relevant to their usage as a tool in the medium term. Unhinged Bing reminds me of a more sophisticated and higher-level version of getting calculators to write profanity upside down: funny, subversive, and you can see how prudes might call for a ban. But if you're taking a test and need to use a calculator, you'll still u…

Unhinged Bing reminds me of a more sophisticated and higher-level version of getting calculators to write profanity upside down: funny, subversive, and you can see how prudes might call for a ban. With all due respect, that seems very strained as an analogy - it's not a bug but a strange human interpretation of expected behavior. You could at least compare it to Microsoft Tay, the chatbot which tweeted profanity just…

What's the rate of Bing chat spitting out vitriol against an actual search-intentioned query? (And not some edge case that a prompt engineer designed, like a real person putting a real search)

As one sample point, I've been using Bing for a couple of days now for real searches, and over dozens of actually-intentioned searches, it has never once tried to tell me what it really thinks of itself, it has never even made a reference to me, to say nothing of anything degrading towards me.

If you use Bing Chat in practice, you'll find that all the edge cases are engineered. Much like if you use a calculator in practice, it almost always doesn't say 55378008 or display porn (versus if you were angling for that, or run porn.89z).

Re: From Bing to Sydney

#149
post #146
post #124

Earlier quoted context omitted.

There is a huge universe of inputs, aka training data, that feeds into us, far more than a digital text based LLM. From that we generated the training data for the LLM. That data is just a sliver of the human experience.

The universe contained exactly 0 words until humans created them, so if we are just stringing together words, then how did we make the words?

> The universe contained exactly 0 words until humans created them,

Human words are one fork of the sound wave based communication systems that many animals on earth use. There was no distinct moment when we went from 0 to 1 words. There was no "first person to speak". We didn't make language. It emerged over time due to evolutionary pressures.

Re: From Bing to Sydney

#150

Earlier quoted context omitted.

> I think this is wrong, because in general, when analogy is good, it is typically good because of the tendency toward allowing for reflex responses. It can't be good and bad for the same reason. It needs to be for a different reason or there isn't logical consistency. That's some weird reasoning. Human emotions are crucial to human existence but we know they also can have bad results. But when emotions are useful to…

> It will be. You can observe the evolution of Google's search system and it has converged to it's current of pushing stuff to sell before everything else. The charter of a public company is maximizing returns to share holders. That is the task of the entire organization Yeah, probably it will evolve in that direction. I could imagine that happening. > That's some weird reasoning. In the AI textbooks I've read, refle…

You would have sentences like "a reflex agent reacts without thinking" and then an example of that might be "a human who puts their hand on a stove yanks it away without thinking about it" and this is rational because the decision problem doesn't call for correct cognition - it calls for minimization of response time such that the hand isn't burned.

We have to be specific about what we're discussing. The human reflex to pull away from a hot stove serves the human, the human gets a benefit from the reflex in the context of a world that has hot stoves but doesn't have, say, traps intended to harm people when they manifest the hot-stove reaction.

Some broad optimization algorithm, if it trained or designed actors, might add a heat reflex to the actors, in the hot-stove-world-context and these actors might also benefit from this. The action of the optimization algorithm would qualify as rational. A person who trained their reflexes could similarly be considered rational. However, the reflex itself is not "rational" or "good" but simply a method or tool.

Which is to say you seem to be implicitly stuck on a fallacious argument "since reflexes are 'good', any reflex reaction is 'good' and 'rational'". And that is certainly not the case. Especially, the modern world we both live in often presents people with communication intended to leverage their reflexes to benefit of the communicator and often against the interests of those targeted. Much of it is advertising and some of it is "social engineering". The social engineering example is something like a message from a Facebook friend saying "is this you? with a link", where if you click the link, it will hack your browser and use it to send more such links as well as other harmful-to-you actions.

It seems like your arguments suffer from failing to make "fine" distinctions between categories like "good", "rational", and "useful-in-a-situation". They are valid things but aren't the same. Analogies can be useful but they aren't automatically rational or good. You begin with me saying "this isn't inherently good or rational though it can be useful-in-a-situation and you think I'm saying analogies aren't good, are bad, which I'm not saying either".

Post reply on HN