Live data from Hacker News

Microsoft and OpenAI end their exclusive and revenue-sharing deal

bloomberg.com

751–760 of 915 posts

Re: Microsoft and OpenAI end their exclusive and revenue-sharing deal

#751
post #685
post #612

Earlier quoted context omitted.

If you do the math (I did), in 2 years, open source models that you can run on a future MacBook Pro will be as capable as the frontier cloud models are today. Memory bandwidth is growing rapidly, as is the die area dedicated to the neural cores. And all the while, we have the silicon getting more power efficient and increasingly dense (as it always does). These hardware improvements are coming along as the open sourc…

A Opus 4.7/Gpt5.5 class model is 5 trillion parameters[1]. To run a 8 bit quantized version of that you need roughly 5TB of RAM. Today that is around 18 NVidia B300. That's around $900,000, without including the computers to run them in. It's true that the capability of open source models is improving, but running actual frontier models on your MPB seems a way off. [1] https://x.com/elonmusk/status/204212356166685523…

Opus and Gpt are generic LLMs with knowledge on all sort of topics. For specific use cases you probably don't need all the parameters? Suppose you want to generate code with opencode, what part of the generic LLM is needed and what parts can be removed?

Re: Microsoft and OpenAI end their exclusive and revenue-sharing deal

#752

Earlier quoted context omitted.

people are trying, especially for inference. For training, it’s just too high risk to tank your training I think. TPUs are at least dogfooded by Google deepmind, no team AFAIK has gotten the AMD stack to train well.

Interesting. Why? My current mental model is that AMD chips are just a bit behind, so, less efficient, but no biggie. Do labs even use CUDA?

i'm doing inference on a free mi300x instance from AMD right now. not sure if the software stack is just old or what, but here's what i've observed: stuck on an old version of vllm pre-Transformers 5 support. it lacks MoE support for qwen3 models. oss-120b is faaaar slower than it should be.

int8 quantization seems like it's almost supported, but not quite. speeds drop to a fraction of full precision speed and the server seems like it intermittently hangs. int4 quantization not supported. fp8 quantization not supported.

again, maybe AMD is just being lazy with what they've provided, but it's not a great look.

right now the fastest smart model i can run is full precision qwen3-32b. with 120 parallel requests (short context) i'm getting PP @ 4500 tokens/sec and TG @ 1300 tokens/sec

Re: Microsoft and OpenAI end their exclusive and revenue-sharing deal

#753

Earlier quoted context omitted.

> emergence of a new kind of intelligence Curious about your definition of these terms. Just because you are impressed by the capabilities of some tech (and rightfully so), doesn't mean it's intelligent. First time I realized what recursion can do (like solving towers of hanoi in a few lines of code), I thought it was magic. But that doesn't make it "emergence of a new kind of intelligence".

> Curious about your definition of these terms. Likewise - I think sometimes we ascribe a mythical aura to the concept of “intelligence” because we don’t fully understand it. We should limit that aura to the concept of sentience, because if you can’t call something that can solve complex mathematical and programming problems (amongst many other things) intelligent, the word feels a bit useless.

> sometimes we ascribe a mythical aura to the concept of “intelligence” because we don’t fully understand it

Agreed! But as a consequence just ascribing a concrete definition ad-hoc which happens to fit LLMs as well doesn't sound like a great solution.

Re: Microsoft and OpenAI end their exclusive and revenue-sharing deal

#754

Earlier quoted context omitted.

You're free to hold an LLM accountable in the exact same way: fire it if you don't like its work.

Giving something that has no internal concept of time (or identity for that matter) a prison sentence of n years seems kinda ineffectual.

Prison sentence? For writing sloppy code? Now that's an interesting idea...

Re: Microsoft and OpenAI end their exclusive and revenue-sharing deal

#755

Earlier quoted context omitted.

That's not "math". That's a "wild guess", or baseless extrapolation at best.

My son doubled in size in the first 8 months of his life. At age 12, he will be larger than the Moon.

One of my favorite xkcd

https://xkcd.com/605/

Re: Microsoft and OpenAI end their exclusive and revenue-sharing deal

#757

Earlier quoted context omitted.

Do we have proof that it's cheaper in terms of $/token/intelligence?

I think the public pricing usually has it cheaper (relatively). Obviously since AI is constantly evolving it's not going to compare as favourably farther to a major Gemini release I was mainly referring to the TPU hardware advantage + GCP running and designing their own datacenter stack.

Does TpU actually have an advantage over Nvidia GPUs?

Re: Microsoft and OpenAI end their exclusive and revenue-sharing deal

#758
post #750

Earlier quoted context omitted.

That’s insane. There should be a big team of people at AMD whose whole job is just to dogfood their stuff for training like this. Speaking of which, Amazon is in the same boat, I’m constantly surprised that Amazon is not treating improving Inferentia/Trainium software as an uber-priority. (I work at Amazon)

> There should be a big team of people at AMD whose whole job is just to dogfood their stuff if they had this management attitude, they wouldn't have been so far behind so as to need this action in the first place!

I'll just leave this here from 10 years ago:

> “Are we afraid of our competitors? No, we’re completely unafraid of our competitors,” said Taylor. “For the most part, because—in the case of Nvidia—they don’t appear to care that much about VR. And in the case of the dollars spent on R&D, they seem to be very happy doing stuff in the car industry, and long may that continue—good luck to them.

https://arstechnica.com/gadgets/2016/04/amd-focusing-on-vr-m...

"car industry" is linked to the GPU-accelerated self-driving car work, ie, making neural networks run fast on GPUs: https://arstechnica.com/gadgets/2016/01/nvidia-outs-pascal-g...

Re: Microsoft and OpenAI end their exclusive and revenue-sharing deal

#759
post #519
post #452

Earlier quoted context omitted.

What are you using LLMs for? To learn about world’s politics? Oh boy I have a news for you…

One of the first things I did when openAI came out was asking it "which active politican is a spy?" - and it was blocked from the start. I asked early, at the time people were posting various jailbreaks, never worked. On a side note, any self hosted model I can get for my PC? I have 96 GB of RAM.

> and it was blocked from the start.

Wait - what?

Re: Microsoft and OpenAI end their exclusive and revenue-sharing deal

#760
post #685
post #612

Earlier quoted context omitted.

If you do the math (I did), in 2 years, open source models that you can run on a future MacBook Pro will be as capable as the frontier cloud models are today. Memory bandwidth is growing rapidly, as is the die area dedicated to the neural cores. And all the while, we have the silicon getting more power efficient and increasingly dense (as it always does). These hardware improvements are coming along as the open sourc…

A Opus 4.7/Gpt5.5 class model is 5 trillion parameters[1]. To run a 8 bit quantized version of that you need roughly 5TB of RAM. Today that is around 18 NVidia B300. That's around $900,000, without including the computers to run them in. It's true that the capability of open source models is improving, but running actual frontier models on your MPB seems a way off. [1] https://x.com/elonmusk/status/204212356166685523…

Do that will only be possible with something like better 3D NAND flash memory, needs a new hardware. People are already trying to bring that the market. Contemplated taking a compiler position in such a company.
Post reply on HN