Live data from Hacker News

OpenAI unveils its first custom chip, built by Broadcom

techcrunch.com

321–330 of 496 posts

Re: OpenAI unveils its first custom chip, built by Broadcom

#321

Earlier quoted context omitted.

A lot of companies that serve a single family of LLMs seem to prefer nvidia though. Why is that? It's not just good drivers, which is what moats them for games and ML. It's a multi-decade work of making chips that are nice to program for and software infrastructure around them. Apple and Google have excelent chips, yet they needed to invest a lot in long-tail software projects to make those chips do actual premium wo…

> A lot of companies that serve a single family of LLMs seem to prefer nvidia though. Why is that? If you write your tools for CUDA, you’re going to prefer hardware the runs CUSA. How is there anything more to it than this?

Cool. That's it.

What will people use to write for Jalapeño matters.

Nvidia has multi-decade heritage. Apple spent almost a decade in MLX. Snapdragon failed partly here. OpenAI announced nothing regarding to that, so this big moat that multiple companies have (nvidia the most prominent) is nil for them.

Re: OpenAI unveils its first custom chip, built by Broadcom

#322

So I’ve been wondering about “one or two levels back” chip design. If I understand it, 28nm chips (pre EUV) is just about suitable to run (not train just inference) frontier models. And so if I was a mid-level State would it be worth while to take my nascent chip industry and push it out to build a 28nm foundry and supporting eco-system. The models will come but the real challenge of the future is having enough compu…

You are underestimating how difficult it is even for a large nation state to attract the kind of talent and investment it would take to set up a chip industry. It is out of reach for anyone outside of the 3-5 largest national economies and a few big American/Chinese multinational corporations.

Re: OpenAI unveils its first custom chip, built by Broadcom

#323

I wanna see an inference chip where the weights are part of the rom of the chip. There would be 1 multiplier per weight (and since they're constant, the whole thing turns into a bunch of simple adders), and the total pipelined system throughput would be one token per clock cycle. That means you can probably have millions of users simultaneously using a single bit of silicon, with perhaps 500 million tokens per second…

> weights [as] part of the rom of the chip

Not really that: you are pointing to Compute-In-Memory (CIM) - techniques where the data (here, a multiplier value) is part of the processor (here, the multiplying circuit).

The problem of "fetch and process" is bypassed completely architecturally: the data is there where the processing happens - it's not moved, there is no latency.

Re: OpenAI unveils its first custom chip, built by Broadcom

#324
post #228

Earlier quoted context omitted.

>> I wanna see an inference chip where the weights are part of the rom of the chip. I've been wondering about that for a while now. For a lot of tasks putting weights in ROM is probably OK. OTOH: >> There would be 1 multiplier per weight... I'm not sure that is a good idea. Maybe if its quantized down to 2 bits... Otherwise maybe a small ROM near each multiplier (or row of them or whatever) so the multipliers could h…

analog chips could also be very interessting instead of using digital signals and processing them against the weights in the ROM. I have no idea if that scales with such big models though.

The drawback is in keeping signal fidelity (e.g. dissipation, temperature etc.) and in the conversion between analogue and digital.

Nonetheless, yes, there are already implemented solutions for small NNs (I understand mostly acting as triggers).

Re: OpenAI unveils its first custom chip, built by Broadcom

#325

I wanna see an inference chip where the weights are part of the rom of the chip. There would be 1 multiplier per weight (and since they're constant, the whole thing turns into a bunch of simple adders), and the total pipelined system throughput would be one token per clock cycle. That means you can probably have millions of users simultaneously using a single bit of silicon, with perhaps 500 million tokens per second…

“ Wafer level faults probably won't matter though - neural nets are resistant to a few missing or wrong weights.” Brain science people “love” traumatic brain injury cases because it can help explore what happens when bits of the “brain wafer” get damaged. We’ve learned a lot from such things. I wonder if people are intentionally “destroying” parts of the model weights to learn more about what happens? Like could you…

Of course tampering with chunks or nodes in the NNs is a way to study the "spawned" (through gradient descent etc.) configuration and "reverse-engineer the black box" to get "AI transparency".

Anthropic published an important work around one year and a half ago.

Re: OpenAI unveils its first custom chip, built by Broadcom

#326

This is very cool to see - seems like soooo much efficiency waiting to be unlocked at the chip level. What's everyone think of Taalas? They're actually burning the LLM model into the silicon, with some onboard memory for fine-tuning. They claim huge cost / latency wins. Super fast demo live at: https://chatjimmy.ai/ https://taalas.com/ https://www.reddit.com/r/singularity/comments/1r9frzk/taalas...

It seems technically interesting, but they seem very sparse on details. I don't know if I like the idea of a single unchanging model forever on a chip. How much more expensive would the silicon be if they used rewritable ROM for the weights? Such an arrangement would permit fine-tunes of the model it was designed for, which might minimize concerns about the model becoming outdated.

There is no memory storage of weights in the Taalas cards but translation of the weight multiplier into a circuit.

Re: OpenAI unveils its first custom chip, built by Broadcom

#327

This is starting to sound like startup scope creep. Instead of making the AI model it’s now custom silicon, web browsers, and consumer electronics?

Maybe, but they’re also a massive company. At some point Google stopped being a startup and become a massive company with margins to look after

Re: OpenAI unveils its first custom chip, built by Broadcom

#328

Earlier quoted context omitted.

I tried making a button using Claude entirely (including the 3D printed enclosure) and it effed up pretty hard with the traces and the header spacing. The project was a big red arcade button that plays the "ah-my-groin.mp3" when pushed (from Simpsons). It did cool work on saving battery life, and the 3d enclosure was awesome, but yeah, I'm convinced I'd have to do another version or two of the custom chip until it ca…

PCB layout is an art, and doesn't seem to map well to LLMs (I tried for shits and giggles recently). Claude in general, kind of like code, does a lot of redundant belt and suspenders stuff in the schematics it generates (if it can generate them at all). It's one of those things that's really not there yet outside of the simplest designs.

DeepPCB has an AI autorouter [1] that uses reinforcment learning and works really well. Recently they also released an AI agent that analyzes your board, proposes plans and can route your board for you [2]. They have a KiCad plugin [3] and you can try it for free.

[1] https://deeppcb.ai/reinforcement-learning-pcb-routing-explai... [2] https://deeppcb.ai/cooper/ [3] https://deeppcb.ai/deeppcb-kicad-plugin-ai-pcb-routing/

Disclaimer: I work at InstaDeep, the company behind DeepPCB, but I don't work on this product.

Re: OpenAI unveils its first custom chip, built by Broadcom

#329
post #306

Earlier quoted context omitted.

Nah, you just need to get the CEO behind it. Most coordination issues get solved when the CEO is breathing down your neck to get something done. Trouble is that they don't do this enough.

Eh, zero guarantees on that one. The Fire Phone was Jeff Bezos' personal baby, and we know how that went. Then there was the Apple G4 Cube with Steve Jobs, the Model X' Falcon Wing doors and Elon, and lets not even talk about the Metaverse and Zuck.

Actually, you've provided examples that prove the point. None of those were especially good (though everyone wanted the G4 Cube), and yet they made it to market anyway. Why?

Because the CEO was behind it, breathing down their necks.

Re: OpenAI unveils its first custom chip, built by Broadcom

#330
post #236

Earlier quoted context omitted.

> seems like soooo much efficiency waiting to be unlocked at the chip level Well if you are exclusively using GPUs that are general purpose, of course you leave so much efficiency on the table. That’s why Google started making TPUs more than a decade ago. I remember that kerfuffle when Google fired Timnit Gebru when Gebru’s paper used GPUs to calculate the environment impact of LLMs while ignoring the efficiency of T…

That ... wasn't the kerfuffle

She wrote the stochastic parrots paper.

Google’s internal review blocked it from publication. Stated reasons were about paper quality. You can speculate whether that was the real reason.

Gebru issued an ultimatum email and said she would resign if some list of conditions weren’t met.

Google said “thanks, we accept your resignation”.

She claims it is retaliation, but it seems more like an own-goal if you ask me. She basically handed Google the solution to their problem.

Practical lesson: don’t tell your employer you might quit before you’re ok with leaving.

Post reply on HN