Live data from Hacker News

YaLM-100B: Pretrained language model with 100B parameters

github.com

351–360 of 666 posts

Re: YaLM-100B: Pretrained language model with 100B parameters

#351
post #146

Earlier quoted context omitted.

Well you can have it in the West. We'd prefer something separate from Kremlin.

We can have what, a great search engine? Maybe if you have a time machine to 2003

2003 certainly didn't have a better search engine. It only had a much smaller, open and un-SEO-biased Internet, making the indexing job correspondingly easier.

Re: YaLM-100B: Pretrained language model with 100B parameters

#352
post #83

Earlier quoted context omitted.

Plus, that's the energy costs involved when running a computer now worth 60k, I'm pretty sure that in the current socio-economic climate those power costs will surpass the initial acquisition cost (those 60k, that is) pretty easily.

An 80GB nvidia A100 goes for $20k and uses 300 watts, the energy costs of using one (or three) isn’t going to surpass the hardware costs for… a while.

I wanted to add that I was writing it metaphorically in a way, as in, seeing as how high those energy bills will be they might as well all add up to 60k.

Not sure about most of the people in here, but I would get really nervous at the thought of running something that eats up 3x300 watts per hour, for 24/7, just as part of a personal/hobby project. The incoming power bills would be too high, you have to be in the wage-percentile for which dropping 60k on a machine just to carry out some hobby project is ok, i.e. you’d have to be “high-ish” middle-class at least.

The recent increases in consumer power prices are a heavy blow for most of the middle-class around Europe (not sure about how things are in the States), so a project like this one is just a no-go for most of middle-class European programmers/computer people.

Re: YaLM-100B: Pretrained language model with 100B parameters

#353

Earlier quoted context omitted.

You brought two links on: - phone calls surveillance in Venezuella: no Google no FB mentioned - plain words of some reporter without any evidence provided, no Google no FB mentioned

Weird how there is limited hard evidence of a secret, illegal government program... It's a lot more than I've seen than evidence for the claims of Yandex proactively sharing data with the Russian government. Also where do you see Venezuela?

So, no proof, no evidence. Ok.

> It's a lot more than I've seen than evidence for the claims of Yandex proactively sharing data with the Russian government.

The difference is that checks and balances are much stronger in US, and such activities can be successfully investigated and government sued.

As an example, your verizon case was successfully challenged: https://en.wikipedia.org/wiki/Klayman_v._Obama

In Russia, court system works in manual mode from Kremlin.

> Also where do you see Venezuela?

I misread, you are right.

Re: YaLM-100B: Pretrained language model with 100B parameters

#354

What are some use cases for something like this? I understand it says "generating and processing text", but is it a replacement for OCR? Or something else?

No, it is more like generating a conversations, translating text, summarization texts, writing code, etc.

Re: YaLM-100B: Pretrained language model with 100B parameters

#356

Side note: Yandex search is awesome, and I really hope they stay alive forever. It's the only functional image search nowadays, after our Google overlords neutered their own product out of fear over lawyers/regulation and a disdain for power users. You can't even search for images "before:date" in Google anymore.

Yandex Image Search is today is what Google Image Search should have been. End of the day I’ll use what actually gets the job done. Same goes for OpenAI and Google AI. If you don’t actually ever release and let others use your stuff and end paralyzed in fear at what your models may do then someone else is gonna release the same tech, and at this rate it seems like that’ll be Chinese or Russian companies who don’t sha…

The “ethical concerns” thing is just a progressive-sounding excuse for why they’re not going to give their models away for free. I guarantee you those models are going to be integrated into various Google products in some form or another.

Re: YaLM-100B: Pretrained language model with 100B parameters

#357
post #249

Earlier quoted context omitted.

What does it have to do with OpenAI branding? Their "moral" reasoning behind not publishing models is simply laughtable because they do sell API access to them to anyone who can pay. And "bad guys" generally have money.

They can (and do) revoke API access from bad guys. They can't do that to downloaded models. Look, I don't like what OpenAI does, but "API access, but no model download" makes sense if you are worried about misuses.

>if you are worried about misuses

why is morality into this? is this the same discussion of car manufacturers not selling cars to certain people because they are worried about misuse?

Re: YaLM-100B: Pretrained language model with 100B parameters

#358

Earlier quoted context omitted.

Weird how there is limited hard evidence of a secret, illegal government program... It's a lot more than I've seen than evidence for the claims of Yandex proactively sharing data with the Russian government. Also where do you see Venezuela?

So, no proof, no evidence. Ok. > It's a lot more than I've seen than evidence for the claims of Yandex proactively sharing data with the Russian government. The difference is that checks and balances are much stronger in US, and such activities can be successfully investigated and government sued. As an example, your verizon case was successfully challenged: https://en.wikipedia.org/wiki/Klayman_v._Obama In Russia, c…

> The difference is that checks and balances are much stronger in US,

You say that after we were talking about the NSA literally spying on US citizens, and without any proof? C'mon, are you really going to badger me about not having having the exact "hard evidence", and not even read my sources or provide ANY evidence yourself.

edit: Yes, it got challenged AFTER needing to be leaked by a whistleblower that still can't return to his home.

Re: YaLM-100B: Pretrained language model with 100B parameters

#359
post #340

Earlier quoted context omitted.

You misunderstand. The NSA went out of their way to tap Google's lines outside of the US, which made the leadership at Google furious . It accelerated the work to encrypt international fiber (I think many people were really bothered by the tcpdump of a bigtable RPC containing a user ID). I was at a conference shortly after an saw a SVP rip an NSA rep to pieces. If Google is doing anything that is required of them leg…

That's what Google claims, however the leaked slides claimed "direct access". edit: Does it really matter if they setup an FTP server instead of direct access, when we know a request can literally ask for "all" data (see Verizon). > When required to comply with these requests, we deliver that information to the US government — generally through secure FTP transfers and in person," Google spokesman Chris Gaither told…

right, you're discussing the mechanism by which Google shares information with the US government- when required by law.

These systems don't give access to "all" data. Telephone companies are different- AT&T had a long standing, off the books agreement with US intelligence agencies (see Idea Factory for a fact-based discussion of what AT&T did) to share large amounts of information illegally.

Re: YaLM-100B: Pretrained language model with 100B parameters

#360

It's just crazy how much it costs to train such models. As I undestand 800 A100 cards would cost about 25.000.000 without considering the energy costs for 61 days of training.

Lambda labs will rent you an 8xA100 instance for 3 months for $21,900. That would put it at around $2m

Still a bit to expensive for my sideproject ; ) To be honest it seems only big corp can do that kind of stuff. By the way if try to do hyper parameter tuning or some exploration in the architecture it becomes guess 10x or 100x more expensive.
Post reply on HN