Live data from Hacker News

GPT-4 details leaked?

threadreaderapp.com

571–580 of 648 posts

Re: GPT-4 details leaked?

#571
post #566

Earlier quoted context omitted.

[flagged]

Can we please stop trying to diagnose mental health issues when we have no background to do so and don’t actually know the patient

My apologies, however Elon has identified as bipolar publicly.

https://www.dailymail.co.uk/health/article-4746914/Elon-Musk...

As someone who has bipolar as well, I think it’s an important thing to talk about and glad he has. Now that he has, I hope we can too without being shut down. It’s not uncommon in our field but it has a huge stigma attached to it that’s unhelpful to the people with it, or to the people are are affected by a friends, coworkers, or loved ones mania or depression. When I’m under a lot of stress I tend towards mania, especially when it’s “good stress” like an achievement or great new job or something really exciting to work on. Inexorably I get drawn into a pit of despair, especially as I start to realize the impact my mania has had on my relationships and reputation. I have a good network and good self awareness built over 30 years of meditation and Buddhist study, so the impacts are mitigated.

I suspect if people understood bipolar and were willing to discuss and learn, we might understand better what Elon does and why. He’s not bad. He’s just different. Bipolar is considered an dimension of neurodiversity, and like other aspects like autism, is nothing to be ashamed of.

Re: GPT-4 details leaked?

#572

Earlier quoted context omitted.

This is common behavior for inference based learners who don’t hail from strong academic backgrounds. Many developers who are self taught utilize a similar method of learning, essentially using pattern recognition to make “educated guesses” that are then internalized as potential facts and tested at the earliest opportunity. In this instance the test was to project the incorrect information out onto a public forum co…

> Yes, this is done in lieu of actually looking up extended details on what something means. If they were capable of understanding the extended details, they would already have an academic background in the subject. Laymen aren't going to have a clue what MoE means even if they went to the trouble of digging up the paper. > Many developers who are self taught utilize a similar method of learning, essentially using pa…

> If they were capable of understanding the extended details, they would already have an academic background in the subject. Laymen aren't going to have a clue what MoE means even if they went to the trouble of digging up the paper.

This is, hopefully, an accidental thought experiment gone awry. "IF THEY WERE CAPABLE of understanding the extended details, they would already have an academic background in the subject" can and should == "I spent a ton of time in the library", and an follow-up apology for putting "capable" and "academic background" in the same sentence.

The whole friggin' point of this glorified LAN is that we can break down those dumb walled gardens and let kids learn from random BBS textfiles, MIT YT videos and the gathered wisdom of HN.

If you are going to just dismiss auto-didacts, you're going to have to re-write the complete higher education History in Western society. I won't even begin to try and validate how wrong this is for Eastern History as well.

Re: GPT-4 details leaked?

#573

> The conspiracy theory that the new GPT-4 quality had been deteriorated might be simply because they are letting the oracle model accept lower probability sequences from the speculative decoding model. In other words: the speculation was likely right, I'll propose a specific mechanism explaining it, but then still insult the people bringing it up and keep gaslighting them.

Calling something a conspiracy theory is not an insult against anybody. It's a theory because it's unproven and it's a conspiracy because people think OpenAI purposely degraded their own service, hence conspiracy theory.

Re: GPT-4 details leaked?

#574

Earlier quoted context omitted.

The advantages in a social setting lie in the introduction of entropy, that is _creativity_, to a community. In a rigorous academic setting and with proper training these individuals are more likely identify links between ideas or information that may not seem obvious at first, and tend to be your more 'eccentric' academics. For the interests of the wider group, the best outcome is to help these individuals refine th…

Given that this is an online forum, another advantage is that a conversational trail is left for others to discover. The inferences these types of individuals make are often based on a structure of knowledge and reality that others share, so the most common preconceived and incorrect notions tend to have the most documentation on how to ameliorate the incorrectness (given that these individuals are allowed to state t…

This had got to be the best thread I’ve ever (inadvertently) started.

Re: GPT-4 details leaked?

#575

Google has been doing research into mixture of experts for scaling LLMs. Their GLaM model published in 2022 has 1.7 trillion parameters and 64 experts. https://icml.cc/media/icml-2022/Slides/17378.pdf

Google is jokingly behind in terms of LLMs. They've done a pretty good job at incorporating vision and audio ML models into their ecosystem, but they underestimated language.

Maybe, but they're relatively open with their research, which is great. They also made BERT and released it for free.

Re: GPT-4 details leaked?

#576

Earlier quoted context omitted.

Not sure where you picked nihilism in my reply, referencing the "we live in a society" meme I actually smiled. Nevertheless, stopping the pretence today, painful as it is, realizing the state of the world is far from perfect, is the first step for a better tomorrow.

There is no political debate in which increasing the opposition’s nihilism is not advantageous. Nhilihism keeps people from thinking. It keeps people at home, away from voting booths. There is daylight between Panglossian utopia and dystopia. If you’re cynical, one of the most productive uses of your time might be engaging the other side and convincing them civic engagement is worthless.

Again, haven't brought up nihilism or cynicism.

Incidentally, just to give you some more to rummage about, staying away from the voting booth seems like the best thing to do and civic engagement is indeed worthless. Unless you are able to say and vote "No", effectively banning all the "sides" from political action, forcing new sides to emerge. No people will ever be free, representational democracy or not, if they can't even say "No". Not being able to say "No" is also the root of the cynicism.

Re: GPT-4 details leaked?

#577

Earlier quoted context omitted.

This is common behavior for inference based learners who don’t hail from strong academic backgrounds. Many developers who are self taught utilize a similar method of learning, essentially using pattern recognition to make “educated guesses” that are then internalized as potential facts and tested at the earliest opportunity. In this instance the test was to project the incorrect information out onto a public forum co…

> Yes, this is done in lieu of actually looking up extended details on what something means. If they were capable of understanding the extended details, they would already have an academic background in the subject. Laymen aren't going to have a clue what MoE means even if they went to the trouble of digging up the paper. > Many developers who are self taught utilize a similar method of learning, essentially using pa…

> At least when humans do it, we bother with the verification step instead of just acting like we know what we're talking about

We do??

Re: GPT-4 details leaked?

#578

Earlier quoted context omitted.

I think we're far from that point though. For the vast majority of use cases, I always wish that the answers could be more accurate. Sure - they might be 'good enough' to build a business on. But if a competitor builds their business on top of a more accurate model, their product will work better, and they will win the market.

Yea but the bench being discussed here is FOSS. Which for me, and many, translates to can i run something useful in my closet or on my phone. I've found LLaMA neat and yea, some FOSS models are getting decent - but they're a far cry from GPT4. I pay for GPT4, use it almost daily and that's my bench. Yes, when i can run GPT4 in my closet, OpenAI will have GPT7 or w/e - but it doesn't change the fact that i have someth…

My guess is you'll be running GPT4 equivalent in your closet, but with a 4K context window.

Where the big guys will have GPT-who-cares-what-version with a 100K context window.

Context size is as much of a big deal as newer generations of models imo.

Re: GPT-4 details leaked?

#579
post #260

Earlier quoted context omitted.

>> The training cost of GPT-4 is now only 1/3 of what it was about a year ago. It is absolutely staggering how quickly the price of training an LLM is dropping, which is great news for open source. The google memo was right about the lack of a moat. That really doesn't change anything at all. The more training large models gets cheaper, the more large corporations are able to train larger models than everyone else. S…

Surely there are diminishing returns for the AI computing though? I mean, is a model with 10x the parameter count 10x better? I think it is still possible that the training costs will be irrelevant for all players at some point with this non-linear scale. Access to data is another story

It's still SO early. We are in the "640K [of memory] ought to be enough for anybody" phase of LLMs. So much more to go.

Re: GPT-4 details leaked?

#580
post #577

Earlier quoted context omitted.

> Yes, this is done in lieu of actually looking up extended details on what something means. If they were capable of understanding the extended details, they would already have an academic background in the subject. Laymen aren't going to have a clue what MoE means even if they went to the trouble of digging up the paper. > Many developers who are self taught utilize a similar method of learning, essentially using pa…

> At least when humans do it, we bother with the verification step instead of just acting like we know what we're talking about We do??

We ask Google (and now ChatGPT) if it's true. It goes round.
Post reply on HN