Live data from Hacker News

GPT-4 details leaked?

threadreaderapp.com

541–550 of 648 posts

Re: GPT-4 details leaked?

#541
post #193

Earlier quoted context omitted.

You can slow things down, but not by more than a few years, because of the gradual democratization of training foundation models. Right now training a model competitive with chatgpt can be done for $150K (microsoft orca 13b). In a few years the cost will be low enough that individuals can train models. At that point regulating it will require draconian dictatorships. I’m also very wary of the copyright angle on this,…

>> You can slow things down, but not by more than a few years, because of the gradual democratization of training foundation models. Just to be clear, what's being "democratised" is the fine-tuning of second-tier, inferior-performance models; or pre-training of third-tier ones. In the game of training large neural nets, the players that can afford to train the largest models with the most amount of data and compute a…

Capabilities of the open-source models are only increasing over time by objective measurement. Yes, every one of them is demonstrably inferior to GPT-4, but we have historical precedent that the cost of compute only ever goes down.

Additionally, assuming the leaked details given here are accurate, there might not be a GPT-6. This entire approach of AI via language models very well could be approaching a local maximum and/or have already reached the point of diminishing returns.

If that is the case, OpenAI's moat is guaranteed to run dry. It should be telling that very few of the improvements over the past few months involve the base model, rather they are value-adds like plug-ins and hooking it up to a VM, things that are not protected by training difficulty.

Re: GPT-4 details leaked?

#542

Earlier quoted context omitted.

Not sure if you're being sarcastic but I agree. Humans should not have AI research or advanced AI at all. It (a) removes purpose from people, (b) presents a situation that is too alien for human minds to handle, (c) increases the addictiveness of technology and thereby pushes us further into growing the technological system, (d) crosses the "adaptability threshhold", i.e. the point at which the PACE of technological…

It’s ironic you say: “we are playing with fire.” Playing with fire is, in large part, literally how humans have come to dominate this planet. Why stop now?

Are you serious? Us dominating the planet is NOT a good thing.

Re: GPT-4 details leaked?

#543

Earlier quoted context omitted.

I would love to see your rebuttals, especially since I have never seen any strong arguments in favour of AI being a net benefit to society, and I have thought and read about this at lenght. Of course, I always expect downvotes on my posts here since there is a strong tendency towards loving technology here. But what I find most interesting is that there is absolutely no taking of responsibility of any technological c…

If we were a rational civilization we'd stop all scientific research immediately. First there's a good chance the great filter is ahead of us and will be triggered by a technology break-trough. Second with nuclear weapons we got lucky in that it's extremely hard to separate fission capable isotope of uranium from mineral ores; if in the future we invent a powerful weapon that's easy to produce organizations like al-Q…

I agree with you that we'd stop scientific research immediately, or at least most of it.

Re: GPT-4 details leaked?

#544

Earlier quoted context omitted.

Corporate wants you to find the difference... That is to say, the two things you mention are the same process. "Identifying the best ways of smelting ore to obtain the metal through trial and error" is the easy part, when you get to pick low-hanging fruits in a field. But as the easy options get cleared out, continuing improvements requires increasingly complex, sophisticated methods - that's where the process transi…

The difference, while moving in "interdependent" directions, is in the purpose: obtaining some sufficient information on how things work versus an actual consideration of the nature of things. It is not really (fully the same process), because you could (in theory) "early stop" when you have achieved technically sufficient competence - the description of the optimal process -, before the jump to the understanding. So…

I wanted to address that in the second paragraph, which I ultimately deleted before submitting, because I couldn't phrase it right. But since you brought it up: I'd consider this an effect of specialization.

In the process of improving your object-level "how things work" goals, you end up generalizing and stacking increasingly complex theoretical models. Soon enough, you end up with people working high up the stack - not knowing or caring about the initial goals. Those people end up growing the "mound of knowledge" both upwards and sideways. The work is sort of self-justifying, but really, it's also self-similar. Where early materials science may have been driven by, say, desire for better/cheaper weapons, soon enough, you have people doing materials science because they desire to solve a puzzle. Whether the practitioners are smelting different combinations of ores to find one that will win them the war, or they're mixing up different kinds of equations to figure out a clean solution to a theoretical conundrum - it's the same process, same motivation. And it always involves play.

The kind of methodical, boring approach, with hypotheses and control groups and peer-reviewed papers? That's the boring part you have to do after play.

See also (with no implied judgement in this context): software developers that lose sight of (or care little about) business goals, and instead aim for theoretical markers of what "good code" is, and/or solve abstract puzzles of algorithms and architecture. Or the MBAs that view companies as abstract money-printing processes, running them by the rulebook that's entirely independent of whatever it is the business is actually doing or selling. Both are cases of growing complexity creating a new field of work that's independent of what brought it into existence.

Re: GPT-4 details leaked?

#546

Earlier quoted context omitted.

> It has its advantages though. Seems the advantage is somewhat localized to the individual inference-based learner; it doesn't seem like a pro-social strategy which would optimize benefit to the group. Overall this seems like it would generalize to widespread misinformation if the majority of uses adopted this behavior. I'm guessing it's in the best interests of the wider group to try to minimize the occurrence of t…

The advantages in a social setting lie in the introduction of entropy, that is _creativity_, to a community. In a rigorous academic setting and with proper training these individuals are more likely identify links between ideas or information that may not seem obvious at first, and tend to be your more 'eccentric' academics. For the interests of the wider group, the best outcome is to help these individuals refine th…

Given that this is an online forum, another advantage is that a conversational trail is left for others to discover. The inferences these types of individuals make are often based on a structure of knowledge and reality that others share, so the most common preconceived and incorrect notions tend to have the most documentation on how to ameliorate the incorrectness (given that these individuals are allowed to state their inferences out loud).

Re: GPT-4 details leaked?

#547

Previously posted about here: https://news.ycombinator.com/item?id=36671588 and here: https://news.ycombinator.com/item?id=36674905 With the original source being: https://www.semianalysis.com/p/gpt-4-architecture-infrastruc... The twitter guy seems to just be paraphrasing the actual blog post? That's presumably why the tweets are now deleted. --- The fact that they're using MoE was news to me and very interesting. I…

Yeah the Mixture of Experts might have not been called out by name, but it was pretty obvious you were getting different models depending on the question. It goes to show how LLMs are nothing like AGI. I think combining it with a calculator is just a bandaid. A useful bandaid, but its not going to be able to do science ever.

I dont see how LLMs using many experts means it's very different from AGI. Why would anyone assume that human AGI isn't based on multiple models running in a similar architecture? At minimum humans are operating with a left and right brain, which process data very differently.

Re: GPT-4 details leaked?

#548

Earlier quoted context omitted.

Sorry, I do not understand what you mean...

I think they read the first line of your reply and took it as you arguing that it's nearly impossible for an LLM to give output that would get someone killed.

I was agreeing with them. Why assume wilful ignorance? Have we become the new reddit? People just yelling in disagreement? Maybe it's time to log out and delete this password...

Re: GPT-4 details leaked?

#549
post #276

Earlier quoted context omitted.

> And safely here means non disruptive to established businesses Why would OpenAI care about that? No, it means safety, as in not giving out dangerous answers that get people killed.

Because they belong to Microsoft. Because they are funded by rich people who own established businesses. Because they are the privileged .01% who are invested in the status quo. Because their whole networks consist of people with entrenched business interests. Pick one or more. Why would you think rich American engineers are in general any more worried about the overall safety of the world then their own self interes…

Eh? People as a rule care about others and this tendency actually INCREASES as they get wealthier and safer as are less worried about themselves.

Your average OpenAI machine learning expert cares plenty about not killing people, just like you do.

Re: GPT-4 details leaked?

#550
post #486

Earlier quoted context omitted.

Just theorizing from the top-level post here. No clue if it's legitimate.

To be clear, you just made up MoE details while MoE is actually well established and hails from decades old research?

Can you detail what mistakes were made? I’m brushing up on it currently but having trouble grokking it.
Post reply on HN