Live data from Hacker News

SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

infini-ai-lab.github.io

51–60 of 65 posts

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#51
post #34

Earlier quoted context omitted.

I am curious whether this is true - OAI at least has the reputation in the industry of caring the least about safety of the major labs

If they don’t care about safety (or perceived safety), why do they spend so much time lobotomizing models for safety reasons?

market reach e.g. ability to have chat app on iOS (the API is less limited)

public relations, limit the edge case nonsense 'journalists' hype so corporate execs aren't terrorized into avoiding buying

doesn't have to be as smart as it could be, it just has to be smarter than other models, so might as well file down some sharp edges for sake of above

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#52
post #4

this is quite worrying for OpenAI as the rate token prices have been plummeting thanks to Meta and its going to have to keep cutting its prices while capex remains flat. whatever Sam says in interviews just think the opposite and the whole picture comes together. It's almost a mathematical certainty that people who invested in OpenAI will need to reincarnate in multiple universes to ever see that money again but no b…

I agree with you somewhat. You are correct unless they have a much better GPT model that have not released for whatever reason. They are a year ahead than competitors and GPT4 is pretty old now. I find it hard to believe they don’t have much more capable models now. We Will see though

GPT-4 is not a single model. The GPT-4 that was released initially a year ago is way worse in benchmarks than the newest versions of it and the original version has been beat by quite a lot of other models by this point.

The newest version of GPT-4 is probably still overall the best model currently, but it is only a few months old, and the picture depends a lot on what benchmarks you are looking at.

E.g. for what we are doing at our company (document processing, etc.) Claude-3 Opus and Gemini-1.5 Pro are currently the better models. The newest GPT-4 even performed worse than a previous version.

So to me it def. seems like the gap is getting smaller. Of course, OpenAI could be coming out with GPT-5 next week and it could be vastly better than all other current models.

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#53
post #36

I don't need exact results. FP8 quantization is almost lossless and even 6-bit quantization is usually acceptable. Can this be combined with quantization?

> Can this be combined with quantization? It is in their TODO part in https://github.com/Infini-AI-Lab/Sequoia/tree/main

INT8, not FP8

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#54

Earlier quoted context omitted.

Even more important, you know that GPT-4 will probably also not crack it. Which is why the SOTA is not terribly interesting. The delta between GPT-4 and the competition has been closing but why anyone would assume that this is a trend and that it would continue with GPT-4.5 to competition, or GPT-5 to competition instead of the other way around is a mystery to me. I am not saying it could not be true. But extrapolati…

There’s a scatterplot that’s been circulating on Twitter. The trend lines show that since the time of GPT-2, open weights models have improved at a steeper rate than proprietary models, with the two on a path to intersect.

I would argue that's to be expected after the first generally accepted POC (GPT-3.5) was released, with it an entire industry created, and other companies actually started copying/competing in a big way.

It seems a stretch to read this as a continuing trend, when (from what I gather everyone agrees on) the way to better models seems to be ever more efficient handling of ever larger amounts of money, compute and data, with no reasonable limits in sight on any of the three.

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#55

Earlier quoted context omitted.

I disagree. a) A year after GPT-4 set the bar, it's still the best model, despite everyone else not having to do it first. Just copy, and just software. And that's not for lack of trying by every other viable prime player on the planet with unprecedented acceleration. Imagine any other piece of software, where the incumbent has a mere 2-3 year head start, in which they had to work out the entire product that everyone…

> A year after GPT-4 set the bar, it's still the best model Debatable. Many people find Claude Opus superior, and I know I've found it consistently better for challenging coding questions. More importantly, the delta between GPT-4 and everything else is getting smaller and smaller. Llama 3 is basically interchangeable with GPT-4 for a huge number of tasks, despite its smaller size.

> Many people find Claude Opus superior

Many more do not, according to the LMSYS leaderboard.

> Llama 3 is basically interchangeable with GPT-4 for a huge number of tasks

Sure. I am sure the number approaches infinity, if you are willing to let the model inform the task. That's usually not what most people are looking for in a tool.

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#56

Earlier quoted context omitted.

Disclaimer: not a fan of "Open"AI Everyone could say anything about open source models, but they're comparing themselves to what OpenAI released a year ago. They haven't shown all of their cards yet and they have a decent moat already in place; some say they have no moat, I disagree, they have one of the best moats possible which is brand awareness. Sora on its own could bring in billions in revenue; an open-source S…

> they have one of the best moats possible which is brand awareness. Do they though? If you talk to "regular people", everybody knows ChatGPT, but nobody knows or cares about OpenAI. And most of them don‘t even really know that name. They call it ChatUuuuhm, ChatThingy, Chad Gippity, or similar. I think they will just switch, when something better comes along.

I don't really follow your logic ...

ChatGPT is a brand that belongs to OpenAI, that's ... not really hard to understand.

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#57

Earlier quoted context omitted.

There’s a scatterplot that’s been circulating on Twitter. The trend lines show that since the time of GPT-2, open weights models have improved at a steeper rate than proprietary models, with the two on a path to intersect.

I would argue that's to be expected after the first generally accepted POC (GPT-3.5) was released, with it an entire industry created, and other companies actually started copying/competing in a big way. It seems a stretch to read this as a continuing trend, when (from what I gather everyone agrees on) the way to better models seems to be ever more efficient handling of ever larger amounts of money, compute and data,…

Scaling up LLMs is only going to go so far, and it will yield diminishing marginal returns on all of that money, compute, and data. It’s a regime of exponential increases in inputs for linear gains in the outputs - barring some technological breakthroughs which could come from anywhere, not just from OpenAI.

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#58

Earlier quoted context omitted.

Disclaimer: not a fan of "Open"AI Everyone could say anything about open source models, but they're comparing themselves to what OpenAI released a year ago. They haven't shown all of their cards yet and they have a decent moat already in place; some say they have no moat, I disagree, they have one of the best moats possible which is brand awareness. Sora on its own could bring in billions in revenue; an open-source S…

MS had yet to fully stabilize that lead a full decade after they had won the os platform standard for ibm compatible pcs. A platform standard moat goes way way beyond a brand advantage. Azure, while significant, has no similar monopoly to support OpenAI. Do you really see a structural advantage to openAI beyond the Microsoft products integrating it?

Can you clarify what is meant by "structural advantage"?

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#59
post #21

Earlier quoted context omitted.

I agree with you somewhat. You are correct unless they have a much better GPT model that have not released for whatever reason. They are a year ahead than competitors and GPT4 is pretty old now. I find it hard to believe they don’t have much more capable models now. We Will see though

The polish of OpenAI stuff when released has been quite mature since gpt4 or even 3.5. They are no doubt sitting on ultra polished stuff. When you are the tip of the arrow though and the cutting edge itself it might not be as efficient but does it ever show you things you can’t unsee. When OpenAI can launch a video thing a day after because it’s ready to go. I am less and less skeptical e dry time they ship because t…

Also, what will the effect of open models be on the LLM provider industry? What effect will Meta’s scorched earth policy of killing markets by releasing very good open models have?

I use LLMs constantly, but no longer in a commercial environment (I am retired except for writing books, performing personal research projects, and small consulting tasks). I now usually turn first to local models for most things: ellama+Emacs is good enough for me to mostly stop using GPT-4+Emacs or GitHub Copilot, the latest open 7B, 8B, 30B models running on my Mac seem sufficient for most of the NLP and data manipulating things I do.

However, it is also fantastic to have long context Gemini, OpenAI APIs, Claude, etc. available when needed or just to experiment with.

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#60
post #41

Earlier quoted context omitted.

There’s a Pareto frontier where Meta is pushing out the boundaries along the “private” and “cheap” axes. Open AI can release GPT 4.5 or 5 and push out the boundary in the direction of “correctness” and “multimodality”. Either way, we win as customers while the the level of competition remains this hot. I personally want a smart AI much more than a cheap or fast one. Your mileage may vary.

Well, Pareto is about optimisation, not either/or. I want a model that’s smart enough, while also being locally-executable. I don’t know whether/when we’ll get there, and whether it will be improvements in models, or underlying model technology, or GPU/TPUs with larger memory at a consumer price point, or something else, that will deliver it.

That's just the middle of the Pareto frontier. Some people want one corner, others the other corner. It's like compression. You can have light compression and high speed, vice-versa, or a balance. You want balance. I want one of the corners.
Post reply on HN