Live data from Hacker News

What we still don’t know about how A.I. is trained

newyorker.com

51–60 of 211 posts

Re: What we still don’t know about how A.I. is trained

#51
post #10

Earlier quoted context omitted.

It’s amazing how one can found a nonprofit with a goal of conducting “open” research and then end up publishing something like this a couple of years later. Greed is good I guess.

We're not publishing any details on the models for safety reasons! Also would be great if the government cracked down on our competitors because they don't care about safety like we do.

"Safety". As if they cared enough about it with ChatGPT. Its purely for competitive reasons because of they hype that was generated.

Re: What we still don’t know about how A.I. is trained

#53
post #4

Say I'm a conspiracy theorist but I'm calling it, that the pentagon isn't letting OpenAI tell the details of GPT-4 even if they wanted to (which they don't except for some of the researchers probably). National security, export restrictions, munitions classification, new Manhattan project etc. EDIT: I know that most people think it's unlikely and I can't give any direct evidence for it. Does that mean it's not welcom…

Just say what's on your mind and don't mind the votes. One thing you'll discover is that you're not alone in your views, whatever they are. Few days ago I came across this bone chilling AI generated Metal Gear Solid 2 meme with Hideo Kojima characters talking about how the purpose of this technology is to make it impossible to tell what's real or fake, leading directly to regulation of information networks with ident…

> Just say what's on your mind and don't mind the votes. One thing you'll discover is that you're not alone in your views, whatever they are.

Preach it!

Re: What we still don’t know about how A.I. is trained

#54
post #47
post #31

Earlier quoted context omitted.

> Imagine the architectural advantage of a big switch statement over models trained in different domains or initial vectors. Given that the emergent abilities come from the large parameter count and massive amount of training data, using smaller models seems like a distinct disadvantage .

> Given that the emergent abilities come from the large parameter count Where can I find evidence of this?

https://openreview.net/pdf?id=yzkSU5zdwD

https://arxiv.org/pdf/2203.15556.pdf

There were also some informal comparisons of GPT models with various parameter counts.

Re: What we still don’t know about how A.I. is trained

#55

Earlier quoted context omitted.

Right, safety is a moat. If you don't meet the standards of the closed model you won't be allowed to exist.

What is OpenAI is right and the risks are real? They likely already have some glimpses of GPT-5 internally. And GPT-4 is closely resembles AGI already.

the risks that a text simulator without gaurdrails will be able to generate text we don't like?

Or that someone will automate cyberattacks, as if the government isn't already doing it?

my greatest fear is that there is only one superintelligence, with access controlled by a monopoly of a few san franciscans deciding what moral boundaries AI will have. I couldn't even get Claude+ to talk to me about houses of zodiac because it insisted it's not an astrologer, it's an AI assistant designed to be helpful blah blah blah, tell me what use is this kind of "safety"?

Re: What we still don’t know about how A.I. is trained

#56
post #4

Say I'm a conspiracy theorist but I'm calling it, that the pentagon isn't letting OpenAI tell the details of GPT-4 even if they wanted to (which they don't except for some of the researchers probably). National security, export restrictions, munitions classification, new Manhattan project etc. EDIT: I know that most people think it's unlikely and I can't give any direct evidence for it. Does that mean it's not welcom…

There has long been a relationship between the tech industry (Silicon Valley definitely included) and the military.

While I can't know what you say is true for sure, given the military's history with things like the internet, GPS, and encryption, I would not be surprised

Re: What we still don’t know about how A.I. is trained

#57
post #10
post #2

The author is right we know almost nothing about the design and training of GPT-4. From the technical report https://cdn.openai.com/papers/gpt-4.pdf : "Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar."

It’s amazing how one can found a nonprofit with a goal of conducting “open” research and then end up publishing something like this a couple of years later. Greed is good I guess.

"Open"AI is how people write it

Re: What we still don’t know about how A.I. is trained

#58
post #47
post #31

Earlier quoted context omitted.

> Imagine the architectural advantage of a big switch statement over models trained in different domains or initial vectors. Given that the emergent abilities come from the large parameter count and massive amount of training data, using smaller models seems like a distinct disadvantage .

> Given that the emergent abilities come from the large parameter count Where can I find evidence of this?

https://ai.googleblog.com/2022/11/characterizing-emergent-ph...

Re: What we still don’t know about how A.I. is trained

#59
post #54
post #47

Earlier quoted context omitted.

> Given that the emergent abilities come from the large parameter count Where can I find evidence of this?

https://openreview.net/pdf?id=yzkSU5zdwD https://arxiv.org/pdf/2203.15556.pdf There were also some informal comparisons of GPT models with various parameter counts.

Excellent info - I did find a bit in the conclusion from the arXiv article:

> While the desire to train these mega-models has led to substantial engineering innovation, we hypothesize that the race to train larger and larger models is resulting in models that are substantially underperforming compared to what could be achieved with the same compute budget.

This mirrors some of my experience. Training/tuning a 7B parameter model feels like goldilocks right now. We are thinking more about 1 specific domain with 3-4 highly-targeted tasks. Do we need 175B+ parameters for that? I can't imagine it would make our lives easier at the moment. Iteration times & cost are a really big factor right now. Being able to go 10x faster/cheaper makes it worth trying to encourage the smaller model(s) to fit the use case.

Re: What we still don’t know about how A.I. is trained

#60
post #41
post #40

Earlier quoted context omitted.

Going to North Korea and assisting them to launder money and bypass sanctions is illegal (aside from being utterly stupid and immoral), which is why he plead guilty and is now in prison.

It is hardly different from saying "you could put cash in duffel bags and the transaction would be hard to trace" Is that assistance? It is just a basic statement of fact. Is wikipedia guilty of providing assistance to NK? They provide far more in depth "assistance" to anyone wanting to perform a Bitcoin transaction. Bringing this back to my original comment, you can see why the federal government would restrict the…

So wait, if the North Koreans can just read all about it on Wikipedia, why did they invite him to the conference?

Also North Korea is a strange hill to die on. It's a brutal dictatorship which represses their own people and threatens to reign nuclear hell on their neighbours and the US. There's a very clear moral line that it's wrong to help them to launder money and evade sanctions, even if it weren't illegal.

Post reply on HN