Live data from Hacker News

PaLM 2 Technical Report [pdf]

ai.google

161–170 of 297 posts

Re: PaLM 2 Technical Report [pdf]

#161
post #146

Earlier quoted context omitted.

It makes up those numbers, I asked about the difference between the small and large PaLM 2 data set size, and it asserted the small model was trained on 540 billion and the large model was trained on 540 trillion. A different draft instead specified 1.4 trillion for the large.

It even gave me a table with a whole bunch of differences. All thats made up? Here is a table that summarizes the key differences between the two language models: Feature Palm Bard Number of parameters 400 billion 540 billion Vocabulary size 137 billion words 1.5 trillion words

I didn't mean to argue that everything generated is incorrect. But in my experience, the numbers it generates seem closer to random guesses. If you ask it enough times, it sometimes converges on a number, but I don't think that means it's an accurate value. I was able to make it generate a similar table for the different PaLM 2 sizes, and laMDA, and it listed, PaLM 2 Gecko 137 billion, PaLM 2 Otter 540 billion, PaLM 2 Bison 1.8 trillion, PaLM 2 Unicorn 5.4 trillion, LaMDA 137 billion. For Unicorn, it also lists "Still under development."

Edit: Playing around with it more and it listed WuDao 2.0 1.75 Trillion, Chinchilla 175B, Codex 175B, Dalle2 1.3B, GPT4 1.75T, GPT3.5 540B, GPT3 175B, GPT2 1.37B, GPT 1.3B.

But in the previous question it listed GPT4 540 billion and Codex 5.4 trillion among other contradictions.

Re: PaLM 2 Technical Report [pdf]

#162
post #140

I don't understand how this can be considered a technical report. No information on model architecture, distributed training methodology, or optimizations. The "Training dataset" section is a pathetic 0.5 pages long. Come on, Google.

In that sense, it's very similar to the GPT-4 Technical Report. The era of being "open" about LLMs or other "secret sauce" models in published papers may be over, since these things have become existential threats to companies.

Btw I've arrived at a different interpretation of the "Open" in OpenAI. It's open in the sense that the generic LLM is exposed via an API, allowing companies to build anything they want on top.

Companies like Google have been working on language models (and AI more broadly) for years but have hid the generic intelligence of their models, exposing it only via improvements to their products. OpenAI bucked this trend and exposed an API to generic LLMs.

Re: PaLM 2 Technical Report [pdf]

#163
post #84
post #62

personal experience - I'm using GPT4 for writing code especially in python. After using bard today, I feel bard is doing quite well considering its free. I will keep using it and if its keep doing well, I will cancel GPT4 $20/month subscription.

why don't you just use chatGPT? from what i know it's running GPT3.5 and it's not that different (at least in terms of code quality)

When my 25 queries per 3 hours runs out I don't use openai at all. That's how bad chat gpt is in comparison to gpt4 in my use cases.

Re: PaLM 2 Technical Report [pdf]

#164
post #5

> "We then train several models from 400M to 15B on the same pre-training mixture for up to 1 × 1022 FLOPs." Seems that for the last year or so these models are getting smaller. I would be surprised if GPT-4 had > the number of parameters as GPT-3 (i.e. 175B). Edit: Seems those numbers are just for their scaling laws study. They don't explicitly say the size of PaLM 2-L, but they do say "The largest model in the PaLM…

GPT-4 is way slower than GPT-3. Unless they are artificially spiking the latency to hide parameter count, it’s likely around 1trn params

Someone on HN has educated me that gpt4 and 3 should be on a similar param count. This is based on inference times of gpt4 vs gpt3.5 pre-speedup (where distilled version was used only post-speedup in the turbo version).

Re: PaLM 2 Technical Report [pdf]

#165

Earlier quoted context omitted.

In that sense, it's very similar to the GPT-4 Technical Report. The era of being "open" about LLMs or other "secret sauce" models in published papers may be over, since these things have become existential threats to companies.

Btw I've arrived at a different interpretation of the "Open" in OpenAI. It's open in the sense that the generic LLM is exposed via an API, allowing companies to build anything they want on top. Companies like Google have been working on language models (and AI more broadly) for years but have hid the generic intelligence of their models, exposing it only via improvements to their products. OpenAI bucked this trend an…

Sounds like the famous Facebook hoodie. "Open and connected" was one slogan on it. The API can be shut down at any time.

https://venturebeat.com/social/facebook-insignia-hoodie/

In the end they just shat all over RSS etc.

Re: PaLM 2 Technical Report [pdf]

#166

Earlier quoted context omitted.

I think the secret sauce is just bucket loads of cash to spend on compute. And because of this I don’t buy that AI is an existential threat to Google at this point. If they were really worried they could spend a tiny portion of their ~280 billion dollars in revenue to train a bigger model.

I assume this is just a PR/IR-driven project to stay the "Google is Dead" headlines hence the budget, especially considering an oversized chunk was spent on the scaling law, doesn't seem they were serious about building a GPT4-killer. I wasn't aware autoregressive LLMs were still considered an existential threat to Google. What's the threat supposed to be, ChatGPT is just going to keep eating Google search market sha…

I agree that if training data is what matters, it is likely that no one can compete with Google with Google Books, which scanned 25 million volumes (source: http://www.nytimes.com/2015/10/29/arts/international/google-...), which is approximately all the books.

DeepMind's RETRO paper https://arxiv.org/abs/2112.04426 mentions a dataset called MassiveText, which includes 20 million books of 3T tokens. So we know Google is using Google Books, since there is simply no other source of 20 million books. Also as far as I know 3T tokens is more than publicly known to be used by anyone so far: Google could train on more data than anyone else, solely from Google Books, even without using its web crawl.

Edit: it was 2005(!), so it is possible that many of you haven't heard of this. George Dyson, in Turing's Cathedral written in 2005 says:

> My visit to Google? Despite the whimsical furniture and other toys, I felt I was entering a 14th-century cathedral: not in the 14th century but in the 12th century, while it was being built. Everyone was busy carving one stone here and another stone there, with some invisible architect getting everything to fit. The mood was playful, yet there was a palpable reverence in the air. "We are not scanning all those books to be read by people," explained one of my hosts after my talk. "We are scanning them to be read by an AI."

https://www.edge.org/conversation/george_dyson-turings-cathe...

Read the whole thing. It is not an accident Google got Google Books to train AI. That was the plan from the start.

Re: PaLM 2 Technical Report [pdf]

#167
post #95

Earlier quoted context omitted.

it claims to run on LaMDA at the moment

If you mean asking it what it's running on, it just hallucinates. As others have noted in the comments here, you can get it to say that it runs on PaLM 3 quite easily.

In chat history you can see which model generated each request - for me it’s always LaMDA

Re: PaLM 2 Technical Report [pdf]

#168

Earlier quoted context omitted.

In that sense, it's very similar to the GPT-4 Technical Report. The era of being "open" about LLMs or other "secret sauce" models in published papers may be over, since these things have become existential threats to companies.

Btw I've arrived at a different interpretation of the "Open" in OpenAI. It's open in the sense that the generic LLM is exposed via an API, allowing companies to build anything they want on top. Companies like Google have been working on language models (and AI more broadly) for years but have hid the generic intelligence of their models, exposing it only via improvements to their products. OpenAI bucked this trend an…

Guess they should rebrand as AvailableAI then...

Re: PaLM 2 Technical Report [pdf]

#169

Earlier quoted context omitted.

In that sense, it's very similar to the GPT-4 Technical Report. The era of being "open" about LLMs or other "secret sauce" models in published papers may be over, since these things have become existential threats to companies.

Btw I've arrived at a different interpretation of the "Open" in OpenAI. It's open in the sense that the generic LLM is exposed via an API, allowing companies to build anything they want on top. Companies like Google have been working on language models (and AI more broadly) for years but have hid the generic intelligence of their models, exposing it only via improvements to their products. OpenAI bucked this trend an…

> Btw I've arrived at a different interpretation of the "Open" in OpenAI.

I don't understand why people have to keep trying to wrap their head around the word 'Open' in OpenAI. If you ever saw a commercial like a product has a 'great new taste' but then you tried it and it tasted bad, would you twist yourself into knots trying to understand how you went wrong in your interpretation of 'great'? No that's ridiculous. Same with 'Open' in 'OpenAI'. It's just some letters that form part of the name that they chose for themselves when they filled the form to incorporate their company.

Re: PaLM 2 Technical Report [pdf]

#170
post #169

Earlier quoted context omitted.

Btw I've arrived at a different interpretation of the "Open" in OpenAI. It's open in the sense that the generic LLM is exposed via an API, allowing companies to build anything they want on top. Companies like Google have been working on language models (and AI more broadly) for years but have hid the generic intelligence of their models, exposing it only via improvements to their products. OpenAI bucked this trend an…

> Btw I've arrived at a different interpretation of the "Open" in OpenAI. I don't understand why people have to keep trying to wrap their head around the word 'Open' in OpenAI. If you ever saw a commercial like a product has a 'great new taste' but then you tried it and it tasted bad, would you twist yourself into knots trying to understand how you went wrong in your interpretation of 'great'? No that's ridiculous. S…

You mean when they filled out a form to incorporate their non-profit. Which they later turned into a for-profit company after reaping all the goodwill. The “Open” used to mean something.
Post reply on HN