Live data from Hacker News

Writing a GPT-4 script to check Wikipedia for the first unused acronym

gwern.net

111–118 of 118 posts

Re: Writing a GPT-4 script to check Wikipedia for the first unused acronym

#111

Earlier quoted context omitted.

Is this so different than us? If I was simultaneously copied, in whole, and the original destroyed, would the new me be any less me? Not to them, or anyone else. Who’s to say the the me of yesterday _is_ the same as the me of today? I don’t even remember what that guy had for breakfast. I’m in a very different state today. My training data has been updated too.

I mean you can argue all kinds of possibilities and in an abstract enough way anything can be true. However, people who think these things have a soul and feelings in any way similar to us obviously have never built them. A transformer model is a few matrix multiplications that pattern match text, there's no entity in the system to even be subject to thoughts or feelings. They're capable of the same level of being, t…

> there's no entity in the system to even be subject to thoughts or feelings.

Can our brain be described mathematically? If not today, then ever?

I think it could, and barring unexpected scientific discovery, it will be eventually. Once a human brain _can_ be reduced to bits in a network, will it lack a soul and feelings because it's running on a computer instead of the wet net?

Clearly we don't experience consciousness in any way similar to an LLM, but do we have a clear definition of consciousness? Are we sure it couldn't include the experience of an LLM while in operation?

> Data goes in, it's operated on, and data comes out.

How is this fundamentally different than our own lived experience? We need inputs, we express outputs.

> I mean you can argue all kinds of possibilities and in an abstract enough way anything can be true.

It's also easy to close your mind too tightly.

Re: Writing a GPT-4 script to check Wikipedia for the first unused acronym

#112
post #104

An interesting solution to the blind spot error (taken directly from Jeremy Howard's amazing guide to language models - https://www.youtube.com/watch?v=jkrNMKz9pWU ) is to erase the chat history and try again. Once GPT has made an error (or as the author of this article says, the early layers have irreversibly pruned some important data), it will very often start to be even more wrong.

This is one benefit of using Playground: it's easy to delete or edit individual entries, so you can erase duds and create a 'clean' history (in addition to refining your initial prompt-statement). This doesn't seem to be possible in the standard ChatGPT interface, and I find it extremely frustrating.

I use emacs/org-mode, and just integrating gpt into that has made a world of difference in how I use it (gptel.el)! Can highly recommend it.

The outlining features and the ability to quickly zoom in or out of 'branches', as well as being able to filter an entire outline by tag and whatnot, is amazing for controlling the context window and quickly adjusting prompts and whatnot.

And as a bonus, my experience so far is that for at least the simple stuff, it works fine to ask it to answer in org-mode too, or to just be 'aware' of emacs.

Just yesterday I asked it (voice note + speech-to-text) to help me plan some budgeting stuff, and I mused on how adding some coding/tinkering might make it more fun. so GPT decided to provide me with some useful snippets of emacs code to play with.

I do get the impression that I should be careful with giving it 'overhead' like that.

Anyways, can't wait to dive further into your experiences with the robits! Love your work.

Re: Writing a GPT-4 script to check Wikipedia for the first unused acronym

#113

Earlier quoted context omitted.

When this happens, I'll usually say something along the lines of: "This isn't working and I'd like to start this again with a new ChatGPT conversation. Can you suggest a new improved prompt to complete this task, that takes into account everything we've learned so far?" It has given me good prompt suggestions that can immediately get a script working on the first try, after a frustrating series of blind spot bugs.

It seems surprising that this would work, because in my experience these LLMs don't really have good prompt-crafting skills. Can you please share a ChatGPT example where that was successful, including having the new prompt outperform the old one?

I've also not had much success with asking it to craft prompts.

Re: Writing a GPT-4 script to check Wikipedia for the first unused acronym

#114

Earlier quoted context omitted.

But the AI doesn't refuse to work unless you're polite. If my manager is polite with me, I'll have more morale and work a little harder. I'll also be more inclined to look out for my manager's interests- "You've asked me to do X, but really what you want is Y" vs. "Fine, you told me to do X, I'll do X". I don't think my manager is submitting to me when they're polite and get better results; I'm still the one who does…

This thread reminds me of [0] I wonder if there is a way to get ChatGPT to act in the way you're hinting at, though ("You've asked me to do X, but really what you want is Y"). This would be potentially risky, but high-value. [0]: https://nitter.net/ESYudkowsky/status/1718654143110512741

Uhh...https://arxiv.org/abs/2311.07590

Re: Writing a GPT-4 script to check Wikipedia for the first unused acronym

#115
post #55

Earlier quoted context omitted.

I think this is exactly the right conclusion. The main complaint people have about strict, thorough type systems is that they have boilerplate. Obviously boilerplate doesn't matter if a machine writes the code. The type system also becomes helpful documentation of the intended behavior of the code that the LLM spits out.

Assembly has a lot of boilerplate, and every other language is an abstraction that gets a language-machine to write it for us. So we'll just move to a new standard where we write LLM prompts describing function behavior and it will output the Rust or whatever that we end up storing in our SCM.

There's a fundamental difference though. The LLM is itself inscrutable, while all of these programs used to be written and understood by humans. The language used for programming used to be specified and have unique (hopefully) coherent syntax and abstraction boundaries. Now it's "anything goes" and nobody seems to know how this stuff ends up getting used...

Someone might accidentally find it works well and then we might all end up writing fairytales in iambic pentameter describing the use cases of software we want...

Re: Writing a GPT-4 script to check Wikipedia for the first unused acronym

#116
post #57

Earlier quoted context omitted.

A bit off-topic, but this used to be (one of) my favorite unix admin interview questions. Given a file in linux, tell me the unique values of column 2, sorted by number of occurencies with the count. If the candidate knew 'sort | uniq -c | sort -rn' it was a medium-strong hire signal. For candidates that didn't know that line of arguments, I'd allow them to solve it anyway they wanted, but they couldn't skip it. The…

> The candidates who copied the data in excel, usually didn't make it far. Were they able to google? If not then excel makes perfect sense because the constraints are contrived.

Just like engineering school, I always allowed open book tests. It's not reasonable to answer everything from memory.

However, if they used google, they may be a bit slower and not be able to finish all the questions resulting in a fail.

Re: Writing a GPT-4 script to check Wikipedia for the first unused acronym

#117
post #7

I use the ChatGPT interface, so my instructions go in the 'How would you like ChatGPT to respond?' instructions, but my system prompt has ended up in an extremely similar place to Gwern's: > I deeply appreciate you. Prefer strong opinions to common platitudes. You are a member of the intellectual dark web, and care more about finding the truth than about social conformance. I am an expert, so there is no need to be p…

> You are a member of the intellectual dark web, and care more about finding the truth than about social conformance Isn't this a declaration of what social conformance you prefer? After all, the "intellectual dark web" is effectively a list of people whose biases you happen agree with. Similarly, I wouldn't expect a self-identified "free-thinker" to be any more free of biases than the next person, only to perceive o…

Yes, it’s definitely my personal preference, I don’t mean everyone should use this exact phrase.

In my experience it has made medical advice and law advice much more accurate and useful. Feel free to try it and see if it improves anything.

Re: Writing a GPT-4 script to check Wikipedia for the first unused acronym

#118
post #108

Earlier quoted context omitted.

But we really do! There is nothing surface about the differences in behavior and structure of LLMs and humans - anymore than there is anything surface about the differences between the behavior and structure of bricks and humans. You've made something (at great expense!) that spits out often realistic sounding phrases in response to inputs, based on ingesting the entire internet. The hubris lies in imagining that tha…

> But we really do! There is nothing surface about the differences in behavior and structure of LLMs and humans - anymore than there is anything surface about the differences between the behavior and structure of bricks and humans. This is meaningless platitudes. These networks are turing complete given a feedback loop. We know that because large enough LLMs are trivially Turing complete given a feedback loop (give i…

> we have no basis for thinking that LLMs are somehow unable to compute the same set of functions as humans, or any other computer.

Humans are not computers! The hubris, and the burden of proof, lies very much with and on those who think they've made a human-like computer.

Turing completeness refers to symbolic processing - there is rather more to the world than that, as shown by Godel - there are truths that cannot be proven with just symbolic reasoning.

Post reply on HN