Live data from Hacker News

SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

infini-ai-lab.github.io

61–65 of 65 posts

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#61
post #34

Earlier quoted context omitted.

I am curious whether this is true - OAI at least has the reputation in the industry of caring the least about safety of the major labs

If they don’t care about safety (or perceived safety), why do they spend so much time lobotomizing models for safety reasons?

Because of PR reasons. They want to avoid government legislations and pretending that they care helps

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#62
post #34

Earlier quoted context omitted.

I am curious whether this is true - OAI at least has the reputation in the industry of caring the least about safety of the major labs

If they don’t care about safety (or perceived safety), why do they spend so much time lobotomizing models for safety reasons?

I didn’t say they don’t care about safety, merely that of the big labs they care the least or close to the least

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#63

Earlier quoted context omitted.

There's wide speculation that what will be branded as either GPT-4.5 or GPT-5 has finished pretraining now and is undergoing internal testing for a fairly near-term release.

My speculation is that internally they have much stronger models like Q* but they won’t be able to release them to public even if they want to for lack of compute and safety and other reasons they see probably…

> My speculation is that internally they have much stronger models like Q*

People used to speculate the same about Google. Everyone hypes up their “secret, too powerful to release” models. Remember the dude who was convinced that there was a sentient AI in the machine? The light of actual public release tends to expose a lot of the hype.

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#64

Earlier quoted context omitted.

My speculation is that internally they have much stronger models like Q* but they won’t be able to release them to public even if they want to for lack of compute and safety and other reasons they see probably…

> My speculation is that internally they have much stronger models like Q* People used to speculate the same about Google. Everyone hypes up their “secret, too powerful to release” models. Remember the dude who was convinced that there was a sentient AI in the machine? The light of actual public release tends to expose a lot of the hype.

That would be a reasonable assumption if OpenAI did not already have an established track record of repeatedly re-defining our fundamental expectations of what technology can do.

GPT-4 was already completed and secretly being tested on Bing users in India in mid-2022 (there were even Microsoft forum posts asking about the funny chatbot). Even after heavy quantization and the alignment tax GPT-4 is still the bar to beat. It's been two years and their funding has increased over 10x since then.

Short of a fundamental Hard Problem that they cannot overcome, their internal bleeding edge models can reasonably be assumed to possess significantly greater capabilities.

Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency

#65
post #23

Earlier quoted context omitted.

There are always going to be pros and cons. That's why solutions like managed databases are reality. From an expert perspective it seems like there's more to lose but from the perspective of a company with employee turn over, possible data loss, security etc. the benefits start to far outweigh the costs. This reasoning can mostly be applied here. If you want to learn about and pull the LLM apart. Perhaps fine-tune an…

What Meta is doing is very nice and differentiates them. I also hope that it ought not change if it became more palatable to not be open.

Meta seem to be thinking 10 years ahead where anyone can run these models at the edge.

Perhaps it's not about where the model is hosted but what can be built on it.

Meta have added Llama3 across the board on all their apps.

Chat is fun but in the wrong context it's useless. However the training data return on millions of users is something interesting to pay attention to.... Llama4 might be a significant jump!

Increased model intelligence and innovative applications of language technology will be where the real value appears. Open-sourcing and allowing public amplification of abilities and enhancements is a very smart move.

The marketing department is also commendable. What happened to Grok? LLMs are everywhere - we're running them on home computers, that's where we should be pondering the next moves.

Post reply on HN