Live data from Hacker News

Interaction Models

thinkingmachines.ai

31–40 of 56 posts

Re: Interaction Models

#31
post #10
post #8

These videos are worth a watch. There are tons of impressive moments, but they had me at the very first one where a woman says: "I'm going to tell you a story," and then pauses for a long, luxurious sip from a cup of coffee, and the model ... does nothing, just waits. Take my money. Speaking of taking my money, what's the economic model for a company like this? They've published a fair amount about their architecture…

> They've published a fair amount about their architecture - enough that I imagine frontier labs could implement. i think the real ones know this is the tip of the iceberg? hparam tuning, data recipes, data collection, custom kernels, rl/eval infra, all immensely deep topics that would condense multiple decades of phd lifetimes to produce SOTA performance (in both senses of the word) like this. i would also calibrate…

I agree that full duplex is the amazing bit. For instance, the three engineers shouting trivia questions while a timer is running — that’s extremely novel as far as I can tell.

I’d like to believe from the demos that this ability to wait kind of falls out of the model as an emergent property — perhaps coming out of a small RL loop - rather than a specific behavior trained, a-la a VAD component in a stack. Either way, I would guess that VAD absolutely cannot do this right now — interruptions are highly annoying on all voice interaction experiences, and if it were a simple matter of better post training, SOMEONE would have done this, e.g. elevenlabs.

But, I disagree on your idea that this is too expensive/too hard to replicate. For me, yes. But, there’s an existence proof — a small team at a new company just did this without a real roadmap, certainly for less than $1b dollars and probably in less than two years. They are almost certainly less skilled at your list of needs to replicate than teams at the frontier labs, who have been given a roadmap.. So I don’t think it’s as difficult as you propose, from an organizational skills perspective.

Re: Interaction Models

#32
post #20
post #8

These videos are worth a watch. There are tons of impressive moments, but they had me at the very first one where a woman says: "I'm going to tell you a story," and then pauses for a long, luxurious sip from a cup of coffee, and the model ... does nothing, just waits. Take my money. Speaking of taking my money, what's the economic model for a company like this? They've published a fair amount about their architecture…

In China it's become well known that promising new companies will get an offer from either Alibaba or Tencent. In the US, it's probably simmilar. Everything that's out in the open can get acquired or simply copied. Maybe that is what Thinking Machines is hoping as well?

Publish a Demo -> acquihire for anthropic/oAI/GOOG/META stock and cash is an understandable economic model. In this case, I feel like they built more than would be needed though — and I hope they deploy something useful, I’d love to play with it.

Re: Interaction Models

#34
post #6

The noteworthy things to me are that the architecture is a transformer that takes in text, image, and audio input and produces text and audio output, all trained together, and it works in near real-time through interleaving inputs and outputs rather than pure generation of the output from a given prompt. > Time-Aligned Micro-Turns. The interaction model works with micro-turns continuously interleaving the processing…

> interleaving the processing of 200ms worth of input and generation of 200ms worth of output.

How does this work? Don't LLMs/transformers need whole context to output next chunk of tokens?

Re: Interaction Models

#35
post #22

am i the only person not impressed by this ? it just feels akward still with pauses and doesnt openai offer voice cadence already

Same here. I dont see anything there that nobody else can catch up on eventually. I must be missing something here. It's all cute, but mmm

A bunch of companies made light bulbs after Edison, that doesn't mean that light bulbs weren't an interesting invention.

Re: Interaction Models

#36

am i the only person not impressed by this ? it just feels akward still with pauses and doesnt openai offer voice cadence already

hard agree, there's already "voice ai" companies that use the normal models and have this "interaction" engine on top of them to produce better results than I've seen in these demos. idk why people are impressed

Re: Interaction Models

#37

Aside from how impressive the model is, the demos here are very well done! Quirky and short, unlike what we're used to from Anthropic and OpenAI.

Agree that this is interesting/impressive, and the demos are nice.

But I completely cracked up at the unexpected physical comedy of the woman in the "slouching" demo, haha omg that was comedy gold, no notes...

I do appreciate less of that flavor of demo that we get from OpenAI/Anthropic, and more of this "human"-feeling vibe. Dare I go as far as calling this an example of "human-centered design" even (https://en.wikipedia.org/wiki/Human-centered_design)?

Re: Interaction Models

#39
post #8

These videos are worth a watch. There are tons of impressive moments, but they had me at the very first one where a woman says: "I'm going to tell you a story," and then pauses for a long, luxurious sip from a cup of coffee, and the model ... does nothing, just waits. Take my money. Speaking of taking my money, what's the economic model for a company like this? They've published a fair amount about their architecture…

hasn't the economic model always been enterprise llms?

tinker - for fine tuning a custom enterprise model,

interaction models - for working as a digital paired employee (as opposed to a company having to reinvent their entire process around ai agents)

Re: Interaction Models

#40
post #20

Earlier quoted context omitted.

In China it's become well known that promising new companies will get an offer from either Alibaba or Tencent. In the US, it's probably simmilar. Everything that's out in the open can get acquired or simply copied. Maybe that is what Thinking Machines is hoping as well?

Publish a Demo -> acquihire for anthropic/oAI/GOOG/META stock and cash is an understandable economic model. In this case, I feel like they built more than would be needed though — and I hope they deploy something useful, I’d love to play with it.

Purely out of curiousity, I see you are using an em dash. Did you use voice transcription or something? It looks hand-typed though. I'm confused.
Post reply on HN