Live data from Hacker News

SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

infini-ai-lab.github.io

31–40 of 65 posts

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#31

Earlier quoted context omitted.

> A year after GPT-4 set the bar, it's still the best model Debatable. Many people find Claude Opus superior, and I know I've found it consistently better for challenging coding questions. More importantly, the delta between GPT-4 and everything else is getting smaller and smaller. Llama 3 is basically interchangeable with GPT-4 for a huge number of tasks, despite its smaller size.

GPT-4 was released in March 2023. Which means the research that went into it would've been finalised quite some time prior. Meaning that you're getting close to a 2 year head start.

While they still call it GPT-4, the one topping the rankings are newer iterations of it despite still retaining the same name. The latest one is from 2024-04-09. Sure that one probably finished training a few months ago but it is by no means a 2 year head start.

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#33

Earlier quoted context omitted.

There's wide speculation that what will be branded as either GPT-4.5 or GPT-5 has finished pretraining now and is undergoing internal testing for a fairly near-term release.

My speculation is that internally they have much stronger models like Q* but they won’t be able to release them to public even if they want to for lack of compute and safety and other reasons they see probably…

I am curious whether this is true - OAI at least has the reputation in the industry of caring the least about safety of the major labs

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#34

Earlier quoted context omitted.

My speculation is that internally they have much stronger models like Q* but they won’t be able to release them to public even if they want to for lack of compute and safety and other reasons they see probably…

I am curious whether this is true - OAI at least has the reputation in the industry of caring the least about safety of the major labs

If they don’t care about safety (or perceived safety), why do they spend so much time lobotomizing models for safety reasons?

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#35
post #4

this is quite worrying for OpenAI as the rate token prices have been plummeting thanks to Meta and its going to have to keep cutting its prices while capex remains flat. whatever Sam says in interviews just think the opposite and the whole picture comes together. It's almost a mathematical certainty that people who invested in OpenAI will need to reincarnate in multiple universes to ever see that money again but no b…

I agree with you somewhat. You are correct unless they have a much better GPT model that have not released for whatever reason. They are a year ahead than competitors and GPT4 is pretty old now. I find it hard to believe they don’t have much more capable models now. We Will see though

I'm not saying Claude 3 and Gemini are better than GPT4 in every aspect, but those two models can at least perform addition on arbitrarily long numbers, meanwhile GPT4 struggles.

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#36

I don't need exact results. FP8 quantization is almost lossless and even 6-bit quantization is usually acceptable. Can this be combined with quantization?

> Can this be combined with quantization?

It is in their TODO part in https://github.com/Infini-AI-Lab/Sequoia/tree/main

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#38

Other than portability and privacy, are there any benefits to running a local model with a 4090, versus running the same model on-demand on a cloud service with the same or more powerful card?

If you have a robot or self driving car, you're going to want on device inference for your vision language models.

For video games, being locked to a cloud service means the feature will disappear when the servers are shut down.

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#39
post #4

this is quite worrying for OpenAI as the rate token prices have been plummeting thanks to Meta and its going to have to keep cutting its prices while capex remains flat. whatever Sam says in interviews just think the opposite and the whole picture comes together. It's almost a mathematical certainty that people who invested in OpenAI will need to reincarnate in multiple universes to ever see that money again but no b…

Disclaimer: not a fan of "Open"AI Everyone could say anything about open source models, but they're comparing themselves to what OpenAI released a year ago. They haven't shown all of their cards yet and they have a decent moat already in place; some say they have no moat, I disagree, they have one of the best moats possible which is brand awareness. Sora on its own could bring in billions in revenue; an open-source S…

> they have one of the best moats possible which is brand awareness.

Do they though?

If you talk to "regular people", everybody knows ChatGPT, but nobody knows or cares about OpenAI. And most of them don‘t even really know that name. They call it ChatUuuuhm, ChatThingy, Chad Gippity, or similar.

I think they will just switch, when something better comes along.

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#40

Earlier quoted context omitted.

There's wide speculation that what will be branded as either GPT-4.5 or GPT-5 has finished pretraining now and is undergoing internal testing for a fairly near-term release.

My speculation is that internally they have much stronger models like Q* but they won’t be able to release them to public even if they want to for lack of compute and safety and other reasons they see probably…

Honestly I'm pretty puzzled by this mystical fog that hangs over OpenAIs skunkworks projects - don't people leave for other jobs/go to conferences etc.?

I'm surprised that nobody call tell what they infact do or do not have.

Post reply on HN