Live data from Hacker News

Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

cerebras.ai

41–50 of 100 posts

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#41
post #4

Earlier quoted context omitted.

Their CEO is a felon who plead guilty to accounting fraud: https://milled.com/theinformation/cerebras-ceos-past-felony-... Experienced investors will not touch them: https://www.nbclosangeles.com/news/business/money-report/cer... I estimated last year that they can only produce about 300 chips per year and that is unlikely to change because there are far bigger customers for TSMC that are ahead of them in priority fo…

Openai wanted to buy them. G42 the largest player in middle east owne a big chunk. You are simply wrong about big investors not touching them but my guess is they will be bought soon by Meta or Apple.

> Apple

I can't imagine Apple being interested.

Their priority is figuring out how to optimise Apple Silicon for LLM inference so it can be used in laptops, phones and data centres.

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#42
post #40

Earlier quoted context omitted.

Only an insignificant minority of companies are running their own AI LLM models. Everyone else is perfectly fine using whatever Azure, GCP etc provide. Enterprise companies don't need to be the fastest or have the best user experience. They need to be secure, trusted and reliable. And you get that by using cloud offerings by default and only going third party when there is a serious need.

If you think that cloud offerings are secure and trustworthy by default you truly must be living under a rock.

I have worked for a dozen companies all earnt more than $20b a year in revenue. That includes two banks and a hedge fund. All use the cloud.

You must be living under a rock if you think the cloud isn't secure enough for the enterprise.

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#43
post #40

Earlier quoted context omitted.

If you think that cloud offerings are secure and trustworthy by default you truly must be living under a rock.

I have worked for a dozen companies all earnt more than $20b a year in revenue. That includes two banks and a hedge fund. All use the cloud. You must be living under a rock if you think the cloud isn't secure enough for the enterprise.

Some quant-heads endorsing the latest fad doesn't prove anything. Also they don't care if chinese hackers are vacuuming data cause ballstreet doesn't care about sustainability. But I grant you that secure and trust are just words that don't mean anything anymore anyhow.

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#44
post #34

> The most important AI applications being deployed in enterprise today—agents, code generation, and complex reasoning—are bottlenecked by inference latency Is this really true today? I don't work in enterprise, so don't know how things look like, but I'm sure lots of people here do, and it feels unlikely that inference latency is the top bottleneck, even above humans or waiting for human input? Maybe I'm just using…

True, the biggest bottleneck is formulating the right task list and ensuring the LLM is directed to find the relevant context it needs. I feel LLMs in their instruction following are often to eager to output rather than using tools (read files) in their reasoning step.

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#45
post #43

Earlier quoted context omitted.

I have worked for a dozen companies all earnt more than $20b a year in revenue. That includes two banks and a hedge fund. All use the cloud. You must be living under a rock if you think the cloud isn't secure enough for the enterprise.

Some quant-heads endorsing the latest fad doesn't prove anything. Also they don't care if chinese hackers are vacuuming data cause ballstreet doesn't care about sustainability. But I grant you that secure and trust are just words that don't mean anything anymore anyhow.

LOL, all fintech are using or entering the "cloud" very heavily. Cloud is here for long enough that claiming it's insecure shows only the immense ignorance.

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#46
post #35
post #34

> The most important AI applications being deployed in enterprise today—agents, code generation, and complex reasoning—are bottlenecked by inference latency Is this really true today? I don't work in enterprise, so don't know how things look like, but I'm sure lots of people here do, and it feels unlikely that inference latency is the top bottleneck, even above humans or waiting for human input? Maybe I'm just using…

It is if you want good results. I’ve been giving Gemini pro prompts for 200+ seconds multiple times per day this week and for such tasks I really like to make it double/triple check and sometimes give the results to Claude for review, too (and vice versa). Ideally I can just run the prompt 100x and have it pick the best solution later. That’s prohibitively expensive and a waste of time today.

How do you create a prompt for Gemini to spend 200 seconds and review multiple times.

Is it as simple as stating in the prompt:

  Spend 200+ seconds and review multiple times 

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#48
post #43

Earlier quoted context omitted.

Some quant-heads endorsing the latest fad doesn't prove anything. Also they don't care if chinese hackers are vacuuming data cause ballstreet doesn't care about sustainability. But I grant you that secure and trust are just words that don't mean anything anymore anyhow.

LOL, all fintech are using or entering the "cloud" very heavily. Cloud is here for long enough that claiming it's insecure shows only the immense ignorance.

https://www.bleepingcomputer.com/news/security/oracle-custom...

Just one of the later examples of a very long list of cloud data breaches affecting millions of users. But hey who cares as long as it does not affect your own bottom line.

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#49
post #40

Earlier quoted context omitted.

Only an insignificant minority of companies are running their own AI LLM models. Everyone else is perfectly fine using whatever Azure, GCP etc provide. Enterprise companies don't need to be the fastest or have the best user experience. They need to be secure, trusted and reliable. And you get that by using cloud offerings by default and only going third party when there is a serious need.

If you think that cloud offerings are secure and trustworthy by default you truly must be living under a rock.

I feel a lot companies do it to reduce liability. It may not be more secure, but it is not their problem.

Re: Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)

#50
post #40

Earlier quoted context omitted.

If you think that cloud offerings are secure and trustworthy by default you truly must be living under a rock.

I have worked for a dozen companies all earnt more than $20b a year in revenue. That includes two banks and a hedge fund. All use the cloud. You must be living under a rock if you think the cloud isn't secure enough for the enterprise.

I think the key here is twofold. First “the cloud” as commonly understood isn’t what anyone here is talking about. The subject is commercial inference providers.

The “cloud”, or Commercial offerings in storage, VMs, etc are reasonably “secure” in a very general context these days, that is generally true.

OTOH “cloud” AI (commercial inference) is going to use your data for training, incorporating your business processes and domain specific competencies into its innate capabilities, which could eventually impact your value proposition. Empirically, this will happen, eventually, regardless of the user agreement that you signed.

Leakage of proprietary competencies is what is meant by being insecure, in this context.

Second, “cloud isn't secure enough for the enterprise” should be replaced with “enterprise actually cares about security except as a cost/benefit analysis”.

Sending your data to someone else’s data center is a really good way for your data to potentially end up on someone else’s computer. In fact, it’s pretty much the point. If security was the priority, they wouldn’t do that.

Post reply on HN