Live data from Hacker News

Fork of Facebook’s LLaMa model to run on CPU

github.com

151–160 of 178 posts

Re: Fork of Facebook’s LLaMa model to run on CPU

#151

Earlier quoted context omitted.

One is a contest to waste the most resources, one has potential to actually have useful results

Why do you think Bitcoin does not have useful results?

Bitcoin has a use, but there are other options for consensus algorithms that don't waste as much energy as the citizens of a medium sized country and fill the same user case (and other expanded use cases). Why not just do that?

Re: Fork of Facebook’s LLaMa model to run on CPU

#152

Earlier quoted context omitted.

I don't work in the field but just to kind of put it into perspective, a 12v 100A LiFePO4 battery has 1200 Watts capacity and weighs 30 pounds. A typical gaming PC (which to be fair, is more willing to trade power for performance) consumes about 600 Watts per hour. Problem for a Tesla? Not so much. Problem for a lightweight drone? Definitely.

ahhh the units in this post are making my eye twitch.

I would be mad too, if my gaming PC was demanding 300 Watts of power but it took half an hour to ramp up ;)

Re: Fork of Facebook’s LLaMa model to run on CPU

#153
post #96

Earlier quoted context omitted.

That's genuinely surprising. What sort of on-board compute do you typically have today?

a common example from my robotics experience (mainly mobile robots) has been getting something powerful enough to run our image recognition/interpreting sensor data. We often have something like several microprocessors (think:arduino equivalent running c++ or c) which run all the motor control etc and a high level system (used to often be raspberry pi, now more often nvidia jetson nano) listening to all of those and…

Limiting ourselves to onboard compute available on mobile robots is one thing, but even for fixed installation robots, aka an arm in a factory where space and power aren't limited, we're very much still limited by compute capacity. Trying to use robots to do something as simple as folding clothes still cannot be done at a reasonable speed. Yeah, on a personal level, just buck up and spend the 20 minutes folding your clothes, or hire a maid to do it for you, but the complexity of automating the task of folding clothes by a robot is a stand in for other tasks in industry that we still can't automate because the complexity is still too high for our current computing power, and have to hire a human for.

Researchers at US Berkeley came out with the algorithm they named SpeedFolding in October of last year. Watch https://youtu.be/UTMT2WAUlRw?t=511 and then realize that linked excerpt is sped up 9x.

If we had 9x faster compute we could have laundry folding robots which is one thing, but that amount of compute would enable robots to do tons more tasks in industry.

Re: Fork of Facebook’s LLaMa model to run on CPU

#154

Earlier quoted context omitted.

> running SHA-256 18 quintillion times or games. People could have been studying or doing something more important than wasting time and energy. I get that it is entertainment, but so are board games and that don't require mining rare earth minerals or putting pressure on the grid as you can always play board games with candles on.

Monopoly by candlelight, just the future I had always envisioned.

Why are you wasting candlestick on playing games? It's such a waste! Don't you know bees died to make that candle?

The most ecologically friendly thing you can do is go to sleep. If you want to play games, do it while the sun is out!

(/s, just in case)

Re: Fork of Facebook’s LLaMa model to run on CPU

#155

Earlier quoted context omitted.

John Hopkins are working on organoids that will replace silicon GPUs for AI.

If they can get pass these new ethical committees...

I'm sure there are rat and mouse brain cells free for the taking from almost any pharmaceutical testing lab.

If organics is the only factor, I don't know why those wouldn't perform as well as human or ape brain cells.

Re: Fork of Facebook’s LLaMa model to run on CPU

#156
post #120

Would it be possible to run the 65B one like this as well? Is the bottleneck just the RAM, or would I need an absurd number of CPUs as well? It's not that hard to create a consumer-grade desktop with 256GB in 2023.

You don't need 256 GB. A pair of the new 48GB DDR5 will work along with a pair of 32GB sticks should work in a consumer DDR5 MB to fit the weights. It does burst when initially loading. So, a fast disk with about the same swap size as RAM seems necessary. It took about 25 mins to generate a single 500 character response using a 5800X & 32 GB DDR4, but I was not able to get to it to run on more than 1 thread with the…

Follow up: https://github.com/facebookresearch/llama/issues/79#issuecom... claims 65B was able to fit in 128 GB by unsharding & merging weights into a single file instead of the multiple pth with 172Gb max swap file usage & appears to stream to GPU.

Re: Fork of Facebook’s LLaMa model to run on CPU

#157
post #153

Earlier quoted context omitted.

a common example from my robotics experience (mainly mobile robots) has been getting something powerful enough to run our image recognition/interpreting sensor data. We often have something like several microprocessors (think:arduino equivalent running c++ or c) which run all the motor control etc and a high level system (used to often be raspberry pi, now more often nvidia jetson nano) listening to all of those and…

Limiting ourselves to onboard compute available on mobile robots is one thing, but even for fixed installation robots, aka an arm in a factory where space and power aren't limited, we're very much still limited by compute capacity. Trying to use robots to do something as simple as folding clothes still cannot be done at a reasonable speed. Yeah, on a personal level, just buck up and spend the 20 minutes folding your…

Robotics is a double whammy, you have compute problems but you also have actuation.

Getting robots to move quickly is easy; getting them to move quickly to exactly where you want them, or with exactly as much force... that is much, much more difficult. Double for mobile robots where you don't have a good energy source. If cost is an issue that is another dimension -- powerful and accurate actuators are extremely expensive.

Re: Fork of Facebook’s LLaMa model to run on CPU

#158

Earlier quoted context omitted.

Oh yes “”” Hackernews senator: “”Someone on the internet said meta aka Facebook is not considered a real data native, clean coder and high IQ company unless your new language model exceeds the elegance and slipperiness of mark Zuckerbergs (you) language output in senate hearings. he is smoother than a lake in the metaverse.“” Mark LLM: “ Yes, unfortunately, the media and our competitors are all over the idea that Met…

I have to say "he is smoother than a lake in the metaverse" is presumably accidental, based on the quality of the rest of that text, but it has to be one of the wittiest phrases ive seen LLMs come out with to date

That was my prompt, I am hackernews senator. People do sometimes ask how many A100s it takes to run me.

Re: Fork of Facebook’s LLaMa model to run on CPU

#159

Earlier quoted context omitted.

Why do you think Bitcoin does not have useful results?

Can you cite an useful result? I can't but I don't think that some people getting richer is useful.

Sure: I can buy servers and domains anonymously and buy drugs online, which are illegal in my country.

Re: Fork of Facebook’s LLaMa model to run on CPU

#160

Earlier quoted context omitted.

I don't work in the field but just to kind of put it into perspective, a 12v 100A LiFePO4 battery has 1200 Watts capacity and weighs 30 pounds. A typical gaming PC (which to be fair, is more willing to trade power for performance) consumes about 600 Watts per hour. Problem for a Tesla? Not so much. Problem for a lightweight drone? Definitely.

ahhh the units in this post are making my eye twitch.

I know Watts per hour is not the right way to phrase it but I feel it helps for those that don't know. Also, I just don't like saying Amp. Hours :)
Post reply on HN