Live data from Hacker News

What we still don’t know about how A.I. is trained

newyorker.com

71–80 of 211 posts

Re: What we still don’t know about how A.I. is trained

#71

Earlier quoted context omitted.

God bless these saviors for prioritizing our "safety!"

Is this the new "think of the children"?

> Is this the new "think of the children"?

As in it’s not about the children, it’s about control? Yes.

Re: What we still don’t know about how A.I. is trained

#72
post #66

GPT Is Not A.I. We tech people should actively go on the offence and educate whomever we can that text inference is not intelligence.

You say that, but if I'm confused about something and think hard about it, I think in language. If you blinded me, paralyzed me, deafened me and desensitized my olfactions, I could still think, but what I would be doing is feeding one language thought into another. It's not so much different from "text" imho.

Re: What we still don’t know about how A.I. is trained

#73
post #61

Earlier quoted context omitted.

Just for some more perspective, a 747 outputs 12,500 tons of CO2 per year. So training GPT4 is basically of no major CO2 concern, especially when you consider how much CO2 it saves. Saves in the sense that humans no longer have to do the work, GPT-4 can just spit it out in seconds so no need for lights, computers to run, food to be produced for the human to eat, etc.

Except of course, now everyone is going to be doing the same thing as OpenAI, pretty much every day, until forever? We'll want to keep throwing hardware at the problem until who knows when and what happens.

If we agree GPT4 is a net negative compared to the work it's replacing, then the more hardware you throw at it, the less C02 would result. Scale in this case is a Negative

Re: What we still don’t know about how A.I. is trained

#74
post #2

The author is right we know almost nothing about the design and training of GPT-4. From the technical report https://cdn.openai.com/papers/gpt-4.pdf : "Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar."

GPT-4 is an amazing achievement, however, it is just a language model. LLM (large language models) are well documented in literature and GPT-4 is just a much larger version (more parameters) of these LLM models. Training of LLM models is also well documented. GPT-4 just has been trained on a very large subset of the Internet.

Of course there are proprietary models, that will be improved versions of the academic LLM models, however, there are no big secrets or mysteries.

Re: What we still don’t know about how A.I. is trained

#75

Earlier quoted context omitted.

Right, safety is a moat. If you don't meet the standards of the closed model you won't be allowed to exist.

What is OpenAI is right and the risks are real? They likely already have some glimpses of GPT-5 internally. And GPT-4 is closely resembles AGI already.

> GPT-4 is closely resembles AGI already

Thats a very bold statement and goes against everything I've read on it so far, care to backup such a claim with some facts? Of course each of us has their own bar for such things, but for most its pretty darn high

Re: What we still don’t know about how A.I. is trained

#76

Earlier quoted context omitted.

Right, safety is a moat. If you don't meet the standards of the closed model you won't be allowed to exist.

What is OpenAI is right and the risks are real? They likely already have some glimpses of GPT-5 internally. And GPT-4 is closely resembles AGI already.

If they're right that GPT-4 is extremely dangerous, then it's an extraordinarily irresponsible of them to release working implementations as a consumer chat app and integrate it into a search engine.

If they're right that LLMs on that scale are generally dangerous but theirs is the exception as they've very carefully calibrated it to be safe, it's extraordinarily irresponsible of them to withhold all details of steps that might make it safe...

Re: What we still don’t know about how A.I. is trained

#77
post #2

The author is right we know almost nothing about the design and training of GPT-4. From the technical report https://cdn.openai.com/papers/gpt-4.pdf : "Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar."

GPT-4 is an amazing achievement, however, it is just a language model. LLM (large language models) are well documented in literature and GPT-4 is just a much larger version (more parameters) of these LLM models. Training of LLM models is also well documented. GPT-4 just has been trained on a very large subset of the Internet. Of course there are proprietary models, that will be improved versions of the academic LLM m…

I'm not so sure about this. There is speculation that GPT4 may utilize additional specialized models underneath it for specific tasks.

Re: What we still don’t know about how A.I. is trained

#78
"To avoid this problem, according to Time, OpenAI engaged an outsourcing company that hired contractors in Kenya to label vile, offensive, and potentially illegal material that would then be included in the training data so that the company could create a tool to detect toxic information before it could reach the user. Time reported that some of the material "described situations in graphic detail like child sexual abuse, bestiality, murder, suicide, torture, self-harm, and incest." The contractors said that they were supposed to read and label between a hundred and fifty and two hundred and fifty passages of text in a nine-hour shift. They were paid no more than two dollars an hour and were offered group therapy to help them deal with the psychological harm that the job was inflicting. The outsourcing company disputed those numbers, but the work was so disturbing that it terminated its contract eight months early. In a statement to Time, a spokesperson for OpenAI said that it "did not issue any productivity targets," and that the outsourcing company "was responsible for managing the payment and mental health provisions for employees," adding that "we take the mental health of our employees and those of our contractors very seriously.""

Re: What we still don’t know about how A.I. is trained

#79

Earlier quoted context omitted.

Just say what's on your mind and don't mind the votes. One thing you'll discover is that you're not alone in your views, whatever they are. Few days ago I came across this bone chilling AI generated Metal Gear Solid 2 meme with Hideo Kojima characters talking about how the purpose of this technology is to make it impossible to tell what's real or fake, leading directly to regulation of information networks with ident…

MGS2 and MGS4 explore ideas about AI, misinformation, the media, and society that are only now being discussed in the mainstream. The concept of an autonomous AGI that generates and filters news stories to provoke humanity into a state of constant division and war is fascinating and worth exploring IMO. Death Stranding also explores ideas of what it means to find connection in a disconnected world that I think are re…

> The concept of an autonomous AGI that generates and filters news stories to provoke humanity into a state of constant division and war is fascinating and worth exploring IMO.

You don't have to speculate much. Facebook and Twitter's recommendation algorithms are doing that fairly well already.

Re: What we still don’t know about how A.I. is trained

#80
post #66

GPT Is Not A.I. We tech people should actively go on the offence and educate whomever we can that text inference is not intelligence.

You don't own the definition of AI. Whether LLMs are intelligent or just pretending to be doesn't matter to many people, and it's not your place to tell them their opinion is wrong.
Post reply on HN