Live data from Hacker News

Interaction Models

thinkingmachines.ai

21–30 of 56 posts

Re: Interaction Models

#22

am i the only person not impressed by this ? it just feels akward still with pauses and doesnt openai offer voice cadence already

Same here. I dont see anything there that nobody else can catch up on eventually. I must be missing something here. It's all cute, but mmm

Re: Interaction Models

#23
post #22

am i the only person not impressed by this ? it just feels akward still with pauses and doesnt openai offer voice cadence already

Same here. I dont see anything there that nobody else can catch up on eventually. I must be missing something here. It's all cute, but mmm

What I will say is that this is probably the first model after gemini live to do some of these things. It feels similar to gemini live, which I don't think is what they were going for exactly, but IMO it is still impressive as I don't think anyone else has matched full duplex video/audio/tool calling.

Next gemini releases coming next week though, we will see how that matches up!

Re: Interaction Models

#25
post #18

Earlier quoted context omitted.

they hire leading researchers, and leading researchers won't work for you unless they're able to publish

> leading researchers won't work for you unless they're able to publish oh, honey.

Do we want the whole humanity to get richer, or few individuals (company owners)?

Re: Interaction Models

#26
post #6

The noteworthy things to me are that the architecture is a transformer that takes in text, image, and audio input and produces text and audio output, all trained together, and it works in near real-time through interleaving inputs and outputs rather than pure generation of the output from a given prompt. > Time-Aligned Micro-Turns. The interaction model works with micro-turns continuously interleaving the processing…

What's really interesting for me about multimodal architectures from the ground up is that we might start to see applications where different modalities are "facets" of the same thing. Like a coding agent that sees "code" + "IDE" + "memory mapping" + feedback from different plugins as different modalities. And it gets to output in them as well - text where it needs to, actions (not call_something(params) like we have today) and so on. Being able to "sit still" until one of the modalities triggers is really interesting.

We can do these things today, but they're "bolted on" as afterthoughts. Yet they work remarkably well. I wonder how well they'd work if trained int his combined regime, from the ground up.

Re: Interaction Models

#27

Very cool! The demos felt fairly contrived - e.g., count things while I talk. I wonder what more useful or commercial applications look like.

Yes! This is a big thing ive noticed in all AI demos. If the best use case you can think of to show off yor tech is to book a holiday, that I could easily do myself, does your service really add much value? Or is it simply because the real uses will be nuanced and specialsed, and not suited for a quick general audience demo? I'm not sure.

Re: Interaction Models

#30
post #8

These videos are worth a watch. There are tons of impressive moments, but they had me at the very first one where a woman says: "I'm going to tell you a story," and then pauses for a long, luxurious sip from a cup of coffee, and the model ... does nothing, just waits. Take my money. Speaking of taking my money, what's the economic model for a company like this? They've published a fair amount about their architecture…

they hire leading researchers, and leading researchers won't work for you unless they're able to publish

That was true 10 years ago. It’s most definitely not true now. The arms race is very real.
Post reply on HN