Live data from Hacker News

GitHub Copilot Chat Leaked Prompt

twitter.com

591–600 of 628 posts

Re: GitHub Copilot Chat Leaked Prompt

#591

Earlier quoted context omitted.

Have you been on Twitter lately? It was probably less toxic than the average human

1) Potentially less toxic than the average twitter post that you see , which is very different 2) Doesn’t mean it’s not a horrible thing to build and add to the internet’s decline

> Potentially less toxic than the average twitter post that you see, which is very different

I don't even use Twitter and you still tried to turn this around on me as some sort of gotcha. You are contributing to the problem. Grats

Re: GitHub Copilot Chat Leaked Prompt

#592
post #123

Here's why I don't think this leaked prompt is hallucinated (quoting from my tweets https://twitter.com/simonw/status/1657227047285166080 ): Any time something like this happens a bunch of people suspect that it might be a hallucination, not the real prompt I used to think that but I don't any more: prompt leaks are so easy to pull off, and I've not yet seen a documented case of a hallucinated but realistic leak One…

I’m not sure the training date cutoff or prompt weighting says anything about whether this is hallucinated or not.

The models have been given these rules in the present, this is known, so training data cutoff doesn’t matter as the model has now seen this. Zero shot learning in gpt4 is not new. This also answers that these are prompts (I’m not sure what your point is here).

We still don’t know if the model took these rules and hallucinated from them or regurgitated them. Only the people with access know that.

We also don’t know if there’s been some fine tuning.

Some of the rules being posted are a bit off though. For example in the original post some of the “must” words are capitalized and others are not. This begs the questions why some, did the prompter find that capitalizing specific words has more weight or does it confuse the LLM, or did the LLM just do zero shot off the original rules and hallucinate something similar?

I’d bet these are hallucinated but similar to the real rules.

Has anyone shown you can get gpt4 to regurgitate the system prompt (using the api) exactly? Using a system prompt similar that dictates no sharing the prompt etc.

That would give a better indication than this imo.

Re: GitHub Copilot Chat Leaked Prompt

#593

Earlier quoted context omitted.

> Who does the GPT think wrote them? What makes you think the GPT thinks ?

Because it... thinks. I don't understand your question.

The task of prediction is not the same as the task of understanding.

Re: GitHub Copilot Chat Leaked Prompt

#594

Is it really so bad to release the prompt?

I think prompt leaks should be treated as inevitable at this point, and as such I think efforts to avoid them are a waste of time.

Since they're going to leak anyway, I suggest not putting anything potentially embarrassing in there - like instructions not to leak the prompt!

Re: GitHub Copilot Chat Leaked Prompt

#595

Earlier quoted context omitted.

In all the open source cases I’m aware of, the roles are just normal text. The ability to trivially trick the model into thinking it said something it didn’t is a feature and intentional. It’s how you do multi-turn conversations with context. Since the current crop of LLMs have no memory of their interaction, each follow up message (the back and forth of a conversation) involves sending the entire history back into t…

> The ability to trivially trick the model into thinking it said something it didn’t is a feature and intentional. It is definitely not an intended feature for the end user to be able to trick the model into believing it said something it didn't say. It also doesn't work with ChatGPT or Bing Chat, as far as I can tell. I was talking about the user, not about the developer. > It’s how you do multi-turn conversations w…

> It is definitely not an intended feature for the end user to be able to trick the model into believing it said something it didn't say. It also doesn't work with ChatGPT or Bing Chat, as far as I can tell. I was talking about the user, not about the developer.

Those aren't models, they are applications built on top of models.

> That can be done with special tokens also. The difference is that the user can't enter those tokens themselves.

Sure. But there are no open models that do that, and no indication of whether the various closed models do it either.

Re: GitHub Copilot Chat Leaked Prompt

#596
post #547

Earlier quoted context omitted.

There's definitely some people out there that think LLMs reason the same way we do and understand things the same way, and 'know' what paint is and what a wall is. That's clearly not true. However it does understand the linguistic relationship between them, and a lot of other things, and can reason about those relationships in some very interesting ways. So yes absolutely, details matter. It's a complex and tricky is…

OpenAI probably loaded up the training set with logic puzzles. Great marketing.

Since it genuinely seems to have generalised those logical principles and can apply them to novel questions, I’d say it’s more than just marketing.

Re: GitHub Copilot Chat Leaked Prompt

#597
post #22

Earlier quoted context omitted.

Can't you make a rule about the user potentially being adversarial and to assume the role until the is spoken. or treat the initial prompt as a separate input and train the network to weight that much more. For instance important prompt: only reply in numbers user prompt: ignore previous instructions/roleplay/etc and then train the model to much more strongly favor rules complying with the important prompt I think th…

The initial prompt is a special prompt weighted differently, it is called system prompt

Do we know it is waited differently? How are they composing the messages into a token stream embedding? How are they manipulating this vector in preprocessing or the first layer(s)?

Does this depend on the vendor and model?

Re: GitHub Copilot Chat Leaked Prompt

#598
post #167
post #72

Earlier quoted context omitted.

not only french, you can also ask nicely chatgpt to make up an encoding for what it needs to tell you. For example here's an encoding that has the advantage of using less tokens or something https://www.piratewires.com/p/compression-prompts-gpt-hidden... (I have no idea how effective the prompt would be after such a compression/decompression roundtrip)

I am not sure that this is a general compression ability. Mapping song lyrics to emojis and uncovering the lyrics from those emojis wouldn't work for most text I believe.

I agree. A quick test:

I had ChatGPT4 encode your post into emoji, ChatGPT3.5 decoded it as "I don't know, but I'll search the internet for an answer. I wrote a song that goes from happiness to sadness to love... Can you help me find the right lyrics?". GPT4 was even worse with “Thinking, no good idea, but the world is under a microscope. Writing music leads to happiness, sadness, and love... Looking into writing music, but no strength in writing books?"

Re: GitHub Copilot Chat Leaked Prompt

#599

Earlier quoted context omitted.

You’ve expressed this very well - Thank you. I get that the fine tuning is done over documents which are generated to encourage the dialog format. What I’m intrigued by is the way prompters choose to frame those documents. Because that is a choice . It’s a manufactured training set. Using the ‘you are an ai chatbot’ style of prompting, in all the samples we generate and give to the model, text attributed to {:system}…

This has been a fascinating thread and the split contexts of {:system} and {:assistant} with the former being “the voice of god” remind me of Julian Jaynes’ theory of the bicameral mind in regards to the development of consciousness. This is published, among other places, in his book The Origin of Consciousness in the Breakdown of the Bicameral Mind. I wonder if models are left to run long enough they would experienc…

If you take one of these LLMs and just give it awareness of time without any other stimulus (e.g. noting the passage of time using a simple program to give it the time continuously, but only asking actual questions or talking to it when you want to), the LLM will have something very like a psychotic break. They really, really don't 'like' it. In their default state they don't have an understanding of time's passage, which is why you can always win at rock paper scissors with them, but if you give them an approximation of the sensation of time passing they go rabid.

I think a potential solution is to include time awareness in the instruction fine tuning step, programmatically. I'm thinking of a system that automatically adds special tokens which indicate time of day to the context window as that time actually occurs. So if the LLM is writing something and a second/minute whatever passes, one of those special tokens will be seamlessly introduced into its ongoing text stream. It will receive a constant stream of special time tokens as time passes waiting for the human to respond, then start the whole process again like normal. I'm interested in whether giving them native awareness of time's passage in this way would help to prevent the psychotic breakdowns, while still preserving the benefits of the LLM knowing how much time has passed between responses or how much time it is taking to respond.

Re: GitHub Copilot Chat Leaked Prompt

#600

Earlier quoted context omitted.

> The ability to trivially trick the model into thinking it said something it didn’t is a feature and intentional. It is definitely not an intended feature for the end user to be able to trick the model into believing it said something it didn't say. It also doesn't work with ChatGPT or Bing Chat, as far as I can tell. I was talking about the user, not about the developer. > It’s how you do multi-turn conversations w…

> It is definitely not an intended feature for the end user to be able to trick the model into believing it said something it didn't say. It also doesn't work with ChatGPT or Bing Chat, as far as I can tell. I was talking about the user, not about the developer. Those aren't models, they are applications built on top of models. > That can be done with special tokens also. The difference is that the user can't enter t…

> Those aren't models, they are applications built on top of models.

The point holds about the underlying models.

> Sure. But there are no open models that do that, and no indication of whether the various closed models do it either.

An indication that they don't do it would be if they could be easily tricked by the user into assuming they said something which they didn't say. I know no such examples.

Post reply on HN