I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
AMD acquires Taalas to boost inference performance by etching models in silicon
431–440 of 712 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#432Earlier quoted context omitted.
People already buy new phones every year, this just creates even more reason to do so
Your location/income bias is showing. Most people do not buy new phones every year.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#433Earlier quoted context omitted.
Just to see how fast it is try chatjimmy.ai
It is really fast and ... really hallucinates. I asked "Does the Wang corporation still exist? If not, what happened to it?" and it replied (in part): "Yes, the Wang Corporation, the company that originally developed and marketed the Wang 2200 computer, still exists as a rebranded company under the name PPL (Precision Pencil and Label), but it has undergone significant changes and challenges over the years. Here's a…
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#434I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
They are: https://openai.com/index/cerebras-partnership/ My guess is they only consider Luna "good enough" to justify the immense up-front investment to put it onto silicon, but Luna at 10x the current speed would be killer. If they're really pursuing live voice conversations with a hardware assistant, latency is more important than accuracy (for complex questions the assistant could always say something like "wait a…
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#435Earlier quoted context omitted.
Did you use chatjimmy? It's somewhat terrifying to use when you think of the potential results with a better model. Ok, real life example: I now spend most of my time, as a developer, waiting for the agent to do its thing (after careful prompting, I'm also thinking about work stuff, don't worry I'm not useless). What if it gave back the same excellent results, but instantaneously? Why, then, I certainly would become…
I tried it. I asked where Bruce Lee was born. It stated he was born in Hong Kong. I challenged it and it went further naming a hospital there. I stated he was born in San Francisco and it apologized and then said his father was a missionary traveling in America, which was also wrong. Bruce’s father was a famous Cantonese Opera singer and actor. This model had zero information right, while being fast in responding. Un…
> Bruce Lee was born in San Francisco, California, USA on November 27, 1940.
> Bruce Lee's father was a Chinese opera singer
That being said, this is not a good test. It is a language model (a very small one), not an encyclopedia.
ChatJimmy interface is just a tech demo. Without tool calling functionality we can't expect it to be factually correct.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#436Earlier quoted context omitted.
They are: https://openai.com/index/cerebras-partnership/ My guess is they only consider Luna "good enough" to justify the immense up-front investment to put it onto silicon, but Luna at 10x the current speed would be killer. If they're really pursuing live voice conversations with a hardware assistant, latency is more important than accuracy (for complex questions the assistant could always say something like "wait a…
How ? Do LLMs actually "know' when they don't "know" ?
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#437Earlier quoted context omitted.
Baking the base models on to ROM makes a lot of economic sense. Less so for consumers though, because it'd mean the phone is out of date in 3 months when a better model comes along.
I don’t think average user _needs_ to solve frontier challenges. ”Call to Jane”, ”turn on the lights” and ”what’s the weather this afternoon” is more like it I would guess. Ofc if the model has some critical bugs that’s another matter.
Maybe baking in a model that is "certified" to have some unconditioned truths + rest is pulled from external models/store could make sense. But AFAIK that doesn't exist and I'm not sure it can possibly be made. Perhaps society as a whole at least can work on an open corpus of training data, but I'm not holding my breath on this.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#438Earlier quoted context omitted.
They are: https://openai.com/index/cerebras-partnership/ My guess is they only consider Luna "good enough" to justify the immense up-front investment to put it onto silicon, but Luna at 10x the current speed would be killer. If they're really pursuing live voice conversations with a hardware assistant, latency is more important than accuracy (for complex questions the assistant could always say something like "wait a…
How ? Do LLMs actually "know' when they don't "know" ?
To answer your question: A large language model itself does not know this (afaik). But chatbots are not "just LLMs" but a whole bunch of systems (and models) around them.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#439Earlier quoted context omitted.
Did you use chatjimmy? It's somewhat terrifying to use when you think of the potential results with a better model. Ok, real life example: I now spend most of my time, as a developer, waiting for the agent to do its thing (after careful prompting, I'm also thinking about work stuff, don't worry I'm not useless). What if it gave back the same excellent results, but instantaneously? Why, then, I certainly would become…
This is the same pitch that people make about AI today. Speed isn’t the differentiator, quality is
AI previously provided speed but not quality. As soon as quality reached an acceptable threshold, the speed became the reigning factor.
In my opinion the quality is still much lower, but speed means the cost is significantly lower also.