Live data from Hacker News

Cerebras Inference now 3x faster: Llama3.1-70B breaks 2,100 tokens/s

cerebras.ai

61–70 of 87 posts

Re: Cerebras Inference now 3x faster: Llama3.1-70B breaks 2,100 tokens/s

#61

Earlier quoted context omitted.

If you're using an LLM as a compressed version of a search index, you'll be constantly fighting hallucinations. Respectfully, you're not thinking big-picture enough. There are LLMs today that are amazing at coding, and when you allow it to iterate (eg. respond to compiler errors), the quality is pretty impressive. If you can run an LLM 3x faster, you can enable a much bigger feedback loop in the same period of time.…

To be honest most LLM's are reasonable at coding, they're not great. Sure they can code small stuff. But the can't refactor large software projects, or upgrade them.

Honestly, most software tasks aren’t refactoring large projects, so it’s probably OK.

As the world gets more internet connected and more online, we’ll have an ever expanding list of “small stuff” - glue code that mixes and ever growing list of data sources/sinks and visualizations together. Many of which are “write once” and leave running.

Big companies (eg google) have built complex build systems (eg bazel ) to isolate small reusable libraries within in a larger repo. Which was a necessity to help unbelievably large development teams to manage a shared repository. An LLM acting in its small corner of the wold seems well suited to this sort of tooling, even if it can’t refactor large projects spanning large changes.

I suspect we’ll develop even more abstractions and layers to isolate LLMs and their knowledge of the wold. We already have containers and orchestration enabling “serverless” applications, and embedded webviews for GUIs.

Think about ChatGPT and their python interpreter or Claude and their web view. They all come with nice harnesses to support a boilerplate-free playground for short bits of code. That may continue to accelerate and grow in power.

Re: Cerebras Inference now 3x faster: Llama3.1-70B breaks 2,100 tokens/s

#62
I wonder at what point does increasing LLM throughput only start to serve negative uses of AI. This is already 2 orders of magnitude faster than humans can read. Are there any significant legitimate uses beyond just spamming AI-generated SEO articles and fake Amazon books more quickly and cheaply?

Re: Cerebras Inference now 3x faster: Llama3.1-70B breaks 2,100 tokens/s

#63
post #62

I wonder at what point does increasing LLM throughput only start to serve negative uses of AI. This is already 2 orders of magnitude faster than humans can read. Are there any significant legitimate uses beyond just spamming AI-generated SEO articles and fake Amazon books more quickly and cheaply?

How about just serving more clients in parallel? I don't see why human reading-speed should pose any kind of upper bound.

And then there are use cases like OpenAI's o1, where most tokens aren't even generated for the benefit of a human, but as input for itself.

Re: Cerebras Inference now 3x faster: Llama3.1-70B breaks 2,100 tokens/s

#64
post #49
post #24

Earlier quoted context omitted.

By comparison with reality. The initial LLMs had "reality" be "a training set of text", when ChatGPT came out everyone rapidly expanded into RLFH (reinforcement learning from human feedback), and now there's vision and text models the training and feedback is grounded on a much broader aspect of reality than just text.

Given that there are more and more AI generated texts and pictures that ground will be pretty unreliable.

Perhaps. But CCTV cameras and smartphones are huge sources of raw content of the real world.

Unless you want to take the argument of Morpheus in The Marix and ask "what is real?"

Re: Cerebras Inference now 3x faster: Llama3.1-70B breaks 2,100 tokens/s

#65
post #50
post #37

Earlier quoted context omitted.

I don't understand your question. This isn't turtles all the way down, it's grounded in real world data, and increasingly large varieties of it.

How does the AI know it’s reality and not a fake image or text fed to the system?

I refer you to Wachowski & Wachowski (1999)*, building on previous work including Descartes and A. J. Ayer.

To whit: humans can't either, so that's an unreasonable question.

More formally, the tripartite definition of knowledge is flawed, and everything you think you know has a Munchausen trilemma.

* Genuinely part of my A-level in philosophy

Re: Cerebras Inference now 3x faster: Llama3.1-70B breaks 2,100 tokens/s

#66
post #65
post #50

Earlier quoted context omitted.

How does the AI know it’s reality and not a fake image or text fed to the system?

I refer you to Wachowski & Wachowski (1999)*, building on previous work including Descartes and A. J. Ayer. To whit: humans can't either, so that's an unreasonable question. More formally, the tripartite definition of knowledge is flawed, and everything you think you know has a Munchausen trilemma. * Genuinely part of my A-level in philosophy

I wasn’t expecting your response to be “the truth is unknowable”, but was hoping for something of more substance to discuss.

Re: Cerebras Inference now 3x faster: Llama3.1-70B breaks 2,100 tokens/s

#67
post #65
post #50

Earlier quoted context omitted.

How does the AI know it’s reality and not a fake image or text fed to the system?

I refer you to Wachowski & Wachowski (1999)*, building on previous work including Descartes and A. J. Ayer. To whit: humans can't either, so that's an unreasonable question. More formally, the tripartite definition of knowledge is flawed, and everything you think you know has a Munchausen trilemma. * Genuinely part of my A-level in philosophy

So we get the same flaws as before with a higher power consumption.

And because it’s fast and easy we now get more fakes, scams and disinformation.

That makes AI a lose-lose not to mention further negative consequences.

Re: Cerebras Inference now 3x faster: Llama3.1-70B breaks 2,100 tokens/s

#68
post #44
post #34

Earlier quoted context omitted.

I filled out a lengthy prompt in the demo. submitted it. an auth window pops up. I don't want to login. I want the demo. such a repulsive approach.

chill with the emotionally charged words. their hardware, their rules. if this upsets you you will not have a good time on the modern internet.

You're not wrong, but how it is currently implemented is pretty deceptive. I would have appreciated knowing the login prompt before interacting with the page. I am curious how many bounces they have because of this one dark pattern.

Re: Cerebras Inference now 3x faster: Llama3.1-70B breaks 2,100 tokens/s

#69
post #64
post #49

Earlier quoted context omitted.

Given that there are more and more AI generated texts and pictures that ground will be pretty unreliable.

Perhaps. But CCTV cameras and smartphones are huge sources of raw content of the real world. Unless you want to take the argument of Morpheus in The Marix and ask "what is real?"

So let’s crank up total surveillance for better auto descriptions of a picture.

We aren’t exchanging freedom for security anymore, what could be reasonable under certain conditions, we just get convenience. Bad deal.

Post reply on HN