Live data from Hacker News

GPT-4 details leaked?

threadreaderapp.com

191–200 of 648 posts

Re: GPT-4 details leaked?

#192

Earlier quoted context omitted.

The application in games I'm most excited about is commenters in FIFA career mode that don't have a limited set of prerecorded voice lines, and take your recent games, formation changes etc into account too, like real commentators would. The recent installments already do that to a small degree. Of course this would also easily open the doors to having multiple commentators/analysts to choose from, each with their in…

There is a mod for Skyrim where someone piped together multiple AI models. It goes like this: You speak into your microphone and ask a NPC something. This gets transcribed (voice to text) by Whisper AI. This transcript gets send to eg. GPT-4 with a pre-prompt engineered to give background, current information and the "personality" for the NPC you are talking to. The output of this gets piped back to a Text-to-Speech…

I've seen an example and the weakest link seemed to be the TTS, which sounded several generations behind.

Re: GPT-4 details leaked?

#193

Earlier quoted context omitted.

> But what I find most interesting is that there is absolutely no taking of responsibility of any technological creations. I appreciate your willingness to talk about it, but to be honest it doesn't seem like it matters much what you, or any of us (not singling you out in particular), thinks about it, does it? It probably doesn't even matter who these people are who should take responsibility. This is one genie, like…

It actually does make a difference. The genie is out of the bottle partially but it depends a lot on what we allow it it be used on. If we sit idly and allow for ingesting all what’s written for instance, including whats currently written and let bros make derivative works for a quick buck then we mostly killed the writer’s incentive to write or publish. If we slow down and not allow ripping one another off it could…

You can slow things down, but not by more than a few years, because of the gradual democratization of training foundation models. Right now training a model competitive with chatgpt can be done for $150K (microsoft orca 13b). In a few years the cost will be low enough that individuals can train models. At that point regulating it will require draconian dictatorships.

I’m also very wary of the copyright angle on this, because just like we don’t prohibit people from learning copyrighted materials in their brains, it feels very wrong to regulate how we train digital brains. I’m ok with forbidding the output of copies of individual existing copyrighted works, but we already have laws on the books for that. I find it downright immoral to prohibit the generation of works “in the style of”. That again reeks like the kind of draconian society I don’t want to live in.

People will always be willing to pay for human-made art, just like we pay more for handmade pots, even though machines can make them better, so I think the doomsayers who predict the end of art are flat out wrong. Easy access to mass-generated AI content could be the best thing that happened to true artists, just like chess AI that can beat every human player was the best thing that happened to the chess world. We need labeling laws that show the origin of works so people can choose whether they want artificial or human-made, but please not another extension of the copyright regime to be even more suffocating and hostile of cultural flourishing.

Re: GPT-4 details leaked?

#194
post #62

Earlier quoted context omitted.

If it involves Elon being a bad guy, it is certain to have HNers salivating at the thought.

At this point, after attempting to start a literal fight with Zuckerberg, Musk is proposing a penis-measuring contest. He appears to be decompensating in real time. If that's salivation fodder, so be it, but it just makes me sad. You hate to see it happen... or at least I do.

Maybe Russia has been injecting lead into his water supply lol.

I’m also saddened.

Musk is not really a hero or a villain, but his manic stages have given us our first realistic shot at becoming a spacefaring civilization, and moved the needle big time on the lock that the perro cartels had on the automotive industry vis-a-vis electric cars.

I hope elon gets better. Losing a billionaire tech maximalists manic episodes is going to set us back decades as more reasonable people chase profit instead of dreams.

Re: GPT-4 details leaked?

#195

Earlier quoted context omitted.

Every company who promotes and develops AI is morally responsible for the coming disaster that it will bring on us. If I could have one wish it would be that every trace of AI research is destroyed.

I recommend reading comments like this and substituting "a baby" for AI. A baby also can't be aligned and is capable of deciding to destroy the world. It's not gonna do it though.

A baby doesn't output propaganda at 6GB/s in computer readable text.

Re: GPT-4 details leaked?

#196

Earlier quoted context omitted.

The real moat is an abundance of high quality data.

Well open AI raised eye brows by crawling the internet and using everyone's data to make a commercial product One day some new startup will train on all of libgen and torrent networks, but it will be very hard to prove. You'll keep getting these gaps up in questionable morality and legality, and even openai will complain about playing fair

Many people train on libgen/torrent in the form of books3 (e.g. LLaMa does this).

Re: GPT-4 details leaked?

#197

Earlier quoted context omitted.

>There are no license issues like with the facebook llama. OpenLLaMa uses a dataset which does not seem to have gotten propper commercial licensing for the training data. There is potential licensing issues because the copyright situation is not well defended.

So far as anyone knows, this is not a derivative work, its transformative, and therefore not subject to any licensing requirement. You're right though, that's arguably still up for debate, but I think the precedent of transformative work is pretty well attested.

is it?

Because condensation is literally part of the definition of derivative, and basically the weights are a condensed form of the input data. It's some sort of lossy compression, when looking at it from the right point of view.

Summarization and translation are also clearly derivative.

The definition of transformative I found:

  - add something new (context of other books I guess, this one might pass)
  - with a further purpose or different character (further purpose clearly yes)
  - do not substitute for the original use of the work (this one I find difficult. In the case of books, probably. In the case of github, it aims to replace quite some aspects of it)

Re: GPT-4 details leaked?

#199
post #82

The tweet is gone. What was in it? Also, I'm dubious about this unsubstantiated claim. The biggest past innovation (training with human feedback) actually shrunk the size of a model. Compare Bloom-366B with falcon-40B (much better). I would be mildly surprised if it turned out Gpt4 has 1.8T parameters. (even if it's a composite model as they say) The article says they use 16 experts 111B each. So the best thing to as…

As a note the 366B in Bloom-366B refers to the number of tokens, not the number of parameters. Bloom had 176B parameters (still many more than Falcon)

Re: GPT-4 details leaked?

#200
post #108

I wonder what the legal implications of them using SciHub and Libgen would be if that's true. I'd imagine OpenAI is big enough to make deals with publishers.

Libgen / Scihub or not, if the model can provide details about the book other than just high level info like the summary and no explicit deal with the publisher has been made, you can make a strong argument that it is plagiarism. Even if bits and pieces of the book text are distributed across the internet and you end up picking up portions of the book, you still read the book. It is extremely sad but ChatGPT will be…

If I read a book and then write a summary, is that plagiarism? What's the difference? I am legitimately not familiar with copyright law, but real lawyers seem to think it is unclear whether training on copyrighted data is illegal (in Japan it's definitely not).
Post reply on HN