Live data from Hacker News

GPT-4 details leaked?

threadreaderapp.com

621–630 of 648 posts

Re: GPT-4 details leaked?

#621
post #486

Earlier quoted context omitted.

To be clear, you just made up MoE details while MoE is actually well established and hails from decades old research?

This is common behavior for inference based learners who don’t hail from strong academic backgrounds. Many developers who are self taught utilize a similar method of learning, essentially using pattern recognition to make “educated guesses” that are then internalized as potential facts and tested at the earliest opportunity. In this instance the test was to project the incorrect information out onto a public forum co…

I don’t appreciate the many bad-faith assumptions made. in particular the assertion that I’m not from a “strong academic background”. For what it’s worth, I made a mistake in a public forum, admitting to this twice. I’ve sought accurate versions of my response which no one provided. Nevertheless I have continued to educate myself on the subject and only feel more confident that I wasn’t misinformed to the degree you all indicated.

You may not realize it but not everyone on the internet is nefarious and if you were to speak in this analytical way about say a classmate while they were in ear shot - that person would likely be quite upset.

Re: GPT-4 details leaked?

#622

Earlier quoted context omitted.

>> RE Chomsky: You can see it's like epicycles: with enough parameters, an LLM is like a numerical method for curve fitting, that doesn't explain the data (any more than a fourier transform does). Curiously, they do seem to predict very accurately... yet also generalize strangely ("hallucinate"). What to think? Well, that's the fundamental problem of modelling: that for any set of observations there's an arbitrary nu…

> Chomsky used it to support his argument about the poverty of the stimulous but linguistics Do you know where Chomsky refers (directly or indirectly) to Gold? I've been searching for a reference for some time.

No, I'm sorry. I'm not a linguist so I only know the relation between Gold's result and linguistics second-hand. I'm more interested in it from the point of view of inductive generalisation in machine learning; that's my schtick.

Just to make sure I didn't hallucinate all that, I had an admittedly perfunctory search online and I could find this paper:

https://proceedings.neurips.cc/paper/2002/file/04ad5632029cb...

Whose introduction describes how Gold's result is considered to support the arguments for linguistic nativism from the poverty of the stimulus. Then again, the author doesn't seem to be a linguist himself and he doesn't give any more specific references, so I'm now a little worried; and your question remains un-answered.

Have you tried wading through Chomsky's early work on linguistics? I don't have the courage to. The closest I've got to is I have a friend who has read a couple of Chomsky's linguistics books. My friend is making a living as an astrologist now so maybe that's a bit of a warning there :P

Re: GPT-4 details leaked?

#623

Earlier quoted context omitted.

>> RE Chomsky: You can see it's like epicycles: with enough parameters, an LLM is like a numerical method for curve fitting, that doesn't explain the data (any more than a fourier transform does). Curiously, they do seem to predict very accurately... yet also generalize strangely ("hallucinate"). What to think? Well, that's the fundamental problem of modelling: that for any set of observations there's an arbitrary nu…

The gravity model is similar though: we posit a force that pulls things together, but we don't know /why/ that force seems to exist, no more than the ancients knew /why/ the planets seemed to move in smaller circles along their circular paths. We're really not /that/ enlightened, after all.

I think that's right, but ultimately all explanations we have are based on prior knowledge that is itself not necessarily complete. It's explanations all the way down, until we hit some primary observations or axiomatic assumptions that are the hardest to get rid of.

"Enlightened" was my bad choice of a word. I get overexcited when I think of how much we have learned in the past couple thousand years and I forget that we mainly learned how little we know. Or can explain!

Re: GPT-4 details leaked?

#624
post #41

If it was trained on CS textbooks, they weren't very good ones. I asked it (GPT4) to write a quantum computer algorithm to square a number. It very confidently told me that to simplify the problem it would use two bits. Okay, fine. But then the algorithm it (again confidently) implemented did a left shift (which it reminded me was multiplying by 2, so it definitely intended this!) and then add the number to itself. I…

It's arguably the first useful general purpose AI. Claiming it is not worth anything at all because it can't solve a problem that 99.999% of humans would not be able to solve is a pretty ridiculous definition of 'worth'.

You (and everyone else) is missing the point of my post. (I admit to having thrown it poorly.) Forget the QC part; It confidently described an algorithm to square a number, which literally any beginning CS student could do, that didn’t even come close.

I will admit to using it all the time for simple programming tasks, and it occasionally does them correctly. Often it comes close enough that I can fix them. (Interestingly in most of these cases I can’t talk it into fixing itself. It kinda gets into wrong-approach ruts), and sometimes it’s horribly wrong (like here).

I find the horribly wrong cases funny.

Re: GPT-4 details leaked?

#625

Earlier quoted context omitted.

If a child watching an adult is "science", then when I brush my teeth I'm a dentist. Words have meaning and when you dilute them down this far the only result is the destruction of communication.

Words do have meaning. den· tist ˈden-təst : one who is skilled in and licensed to practice the prevention, diagnosis, and treatment of diseases, injuries, and malformations of the teeth, jaws, and mouth and who makes and inserts false teeth Brushing your teeth does not imply you are a dentist. sci·ence noun 1. the systematic study of the structure and behavior of the physical and natural world through observation, e…

[deleted]

Re: GPT-4 details leaked?

#626

Earlier quoted context omitted.

This is common behavior for inference based learners who don’t hail from strong academic backgrounds. Many developers who are self taught utilize a similar method of learning, essentially using pattern recognition to make “educated guesses” that are then internalized as potential facts and tested at the earliest opportunity. In this instance the test was to project the incorrect information out onto a public forum co…

I don’t appreciate the many bad-faith assumptions made. in particular the assertion that I’m not from a “strong academic background”. For what it’s worth, I made a mistake in a public forum, admitting to this twice. I’ve sought accurate versions of my response which no one provided. Nevertheless I have continued to educate myself on the subject and only feel more confident that I wasn’t misinformed to the degree you…

I take it all back. The comment I responded to appears to be both correct in assertion and implication.

Re: GPT-4 details leaked?

#627

Earlier quoted context omitted.

Given that this is an online forum, another advantage is that a conversational trail is left for others to discover. The inferences these types of individuals make are often based on a structure of knowledge and reality that others share, so the most common preconceived and incorrect notions tend to have the most documentation on how to ameliorate the incorrectness (given that these individuals are allowed to state t…

This had got to be the best thread I’ve ever (inadvertently) started.

I have quite deeply enjoyed this thread myself. Thanks :)

Re: GPT-4 details leaked?

#628

Earlier quoted context omitted.

I'm not a lawyer and obviously we won't get any definite answer unless it actually goes to court, all of this is just hand waving and guessing. But I think that unless GPT starts reciting large parts outside of the context of learning/education/research, reciting smaller snippets would fall into "fair use" and not be illegal.

For it to be fair use, they still have to have legally owned the book (as far as I understand). You can't steal a book, photocopy some pages, then claim the photocopied pages are fair use.

I think you can. It is a separate "crime". You would get 2 cases one for fair use (which if you are quoting, commenting, reviewing, generally repurposing content and it is in fact fair) and second case for license/terms breach and/or illegally obtaining this piece of work(for example if you stolen it from bookstore).

Re: GPT-4 details leaked?

#629
post #591

Earlier quoted context omitted.

The advantages in a social setting lie in the introduction of entropy, that is _creativity_, to a community. In a rigorous academic setting and with proper training these individuals are more likely identify links between ideas or information that may not seem obvious at first, and tend to be your more 'eccentric' academics. For the interests of the wider group, the best outcome is to help these individuals refine th…

You have clarified something I have always thought about very intensely and deeply but haven’t really ever read anyone else who understands that so well or rather put it into words so clearly. I’m an inferenced based learner to an extreme and it definitely has many upsides and also downsides. The upsides are being able to learn extremely rapidly by making connections between pieces of information where there’s gaps a…

I really appreciate your response, thank you for sharing your perspective and experiences.

When I read message I can't help but picture you as the storybook 'inventor' who is locked away inside his house with strange colored smoke coming out the chimney & weird noises heard from the street, yet when the doors open the whole town would gather to see what you made.

Re: GPT-4 details leaked?

#630
post #591

Earlier quoted context omitted.

The advantages in a social setting lie in the introduction of entropy, that is _creativity_, to a community. In a rigorous academic setting and with proper training these individuals are more likely identify links between ideas or information that may not seem obvious at first, and tend to be your more 'eccentric' academics. For the interests of the wider group, the best outcome is to help these individuals refine th…

You have clarified something I have always thought about very intensely and deeply but haven’t really ever read anyone else who understands that so well or rather put it into words so clearly. I’m an inferenced based learner to an extreme and it definitely has many upsides and also downsides. The upsides are being able to learn extremely rapidly by making connections between pieces of information where there’s gaps a…

This was very well said and I don't know if you could have said it any better. FWIW, I'm a person with multiple degrees in CS, but the best programmers I've worked with and who get stuff done have zero degrees. I have eight years of hardcore programming experience to include professional and side project stuff - I've learned more actually doing than in any classroom. Yeah it's cool to know what a bubble sort is and how it compares to a merge sort, but knowing all the fine details isn't really needed for actually building things, especially now that we're at the point where an AI can give you the code along with complete instruction.

It sounds like you've done completely fine for yourself and built things that people want, so I would try not to be too hard on yourself.

Post reply on HN