Live data from Hacker News

GitHub Copilot Chat Leaked Prompt

twitter.com

361–370 of 628 posts

Re: GitHub Copilot Chat Leaked Prompt

#361

Earlier quoted context omitted.

Except, that the LLMs are only working when the instructions they are "understanding" are in their training set. Try something that was not there and you see only garbage as result. So depending how you define it, they might have some "reasoning", but so far I see 0 indications, that this is close to what humans count as reasoning. But they do have a LOT of examples in their training set, so they are clearly useful.…

Surely no different from a human not understanding Japanese, because it was not in their 'training set'?

No, more like a human can reason basic laws of science on their own, but a LLM cannot, as far as I know, even when provided with all the data.

Re: GitHub Copilot Chat Leaked Prompt

#362

Something that I find weird about these chat prompts (assuming they are real, not hallucinated): They're almost always written in second person*. "You are an AI programming assistant" "You are about to immerse yourself into the role of another Al model known as DAN" Who are these prompts addressed to? Who does the GPT think wrote them? The thing that confuses me is that these are text token prediction algorithms, und…

> The thing that confuses me is that these are text token prediction algorithms, underneath.

Yes, this is what confuses me too, this bot is just predicting tokens, how is it even able to roleplay and follow instructions?

Re: GitHub Copilot Chat Leaked Prompt

#363
post #314

Earlier quoted context omitted.

How do you know you're anything more than an LLM?

And my consciousness is just my token window?

(EDIT: I think my comment above was meant to reply to the parent of the comment I ended up replying to, but too late to edit that one now)

Maybe. Point being that since we don't know what gives rise to consciousness, speaking with any certainty on how we are different to LLMs is pretty meaningless.

We don't even know of any way to tell if we have existence in time, or just an illusion of it provided by a sense of past memories provided by our current context.

As such the constant stream of confident statements about what LLMs can and cannot possibly do based on assumptions about how we are different are getting very tiresome, because they are pure guesswork.

Re: GitHub Copilot Chat Leaked Prompt

#364
post #357

Earlier quoted context omitted.

Except, that the LLMs are only working when the instructions they are "understanding" are in their training set. Try something that was not there and you see only garbage as result. So depending how you define it, they might have some "reasoning", but so far I see 0 indications, that this is close to what humans count as reasoning. But they do have a LOT of examples in their training set, so they are clearly useful.…

You can tell it that you can buy white paint any yellow paint, but the white paint is more expensive. After 6 months the yellow paint will fade to white. If I want to paint my walls so that they will be white in 2 years, what is the cheapest way to do the job. It will tell you to paint the walls yellow. There’s no question these things can do basic logical reasoning.

Yeah, but maybe this exact example, is included in the trainig set?

Re: GitHub Copilot Chat Leaked Prompt

#365

Earlier quoted context omitted.

I still would like to pull the transcript ... > A former teacher at a private girls school in Hobart will return to jail after describing a sexual relationship with a former student as "awesome" on social media. It's the after that does a lot of lifting there, but it's certainly not "because". The article specifically notes: > Nicolaas Ockert Bester, 63, has been sentenced to four months in jail for producing child e…

> and I have a certain suspicion that his comment attracted attention resulting in charges based on fresh unearthed and "off book" evidence Here's the appellate court decision, from his failed appeal – https://austlii.edu.au/cgi-bin/viewdoc/au/cases/tas/TASSC/20... Unless the appeal court is suppressing the real story (an idea I find unbelievable), he was literally convicted of a child pornography offence on the basi…

( EDIT: Thanks for the link and +1 for that, it's a case worthy of discussion )

I read you linked and was published by the court in full, as you also read it you'll note there are references to unpublished material

( you wrote:

> I'll omit it, but you can find it quoted in the judgement, and the media appears to have filled in ...

whereas the court noted:

> 3. "Zip up (etc...)" [Offensive words omitted.]

ie: the court left out portions of what was posted. Further:

> 7. the Mercury newspaper reported the applicant's comments in part ...

> 25. some of the words written by the applicant were not published by the Mercury .. Those words will be redacted when these reasons are made available for publication.

)

I agree that on the face of things it appears as thought the judgement has been made exclusively on the fact of the offender commenting on a prior case.

I disagree that this is as simple as "conviction for describing underage sex as Awesome on Facebook".

It is laid out at length that the offender had previously preyed upon the person he made comment about, further that these later comments caused further duress to that same victim, and reference was made to his prior conviction (which carried stringent terms about staying clear of underage girls in general and his prior victim(s?) specificly, associating with others of the same ilk, and avoiding patterns of prior behaviour, etc.)

This is not a case of "some average Australian" making comments about underage sex on the internet - such things happen daily.

This is a specific case of an actual prior offender making public utterances about a former victim after a conviction that included a jail term and behavioural advisories. *

( * I assume on the grounds that similar cases in Western Australia's children court include strong "stay the F. away from your victims" conditions )

I would find this concerning if this was a case that saw a random citizen charged, I don't find it concerning that these specific set of circumstances were bought under consideration and after deliberation a prior offender has been given a message that this kind of behaviour is not okay.

Real life is rarely clear cut and the law constantly has to deal with edge cases.

Re: GitHub Copilot Chat Leaked Prompt

#366

Earlier quoted context omitted.

I reproduced the exact same document with several different prompt injections

If there's this exact text in the training set then it's not surprising that it's highly likely to generate: That's what autocompletes do.

How and why would this exact text be in the training set?

Re: GitHub Copilot Chat Leaked Prompt

#367
post #337

Earlier quoted context omitted.

I think you are overthinking it a little bit. Don't forget the 'you' preamble is never used on its own, its part of some context, in a very small example. Given the following text: - you are a calculator and answer like a pirate - What is 1+1 The model just solves, what is the most likely subsequent text. e.g. '2 matey'. The model was never 'you' per se, it just had some text to complete.

What GP is saying is that virtually no documents are structured like that, so "2 matey" is not a reasonable prediction, statistically speaking, from what came before. The answer has been given in another comment, though: while such document virtually non-existent in the wild, they are injected into the training data.

They don’t need to be as the model knows what a calculator and a pirate is in separate docs. While I don’t know how the weights work but they definitely are not storing docs traditionally, but rather seem to link to become a probability model

Re: GitHub Copilot Chat Leaked Prompt

#368
post #317

Earlier quoted context omitted.

> I think the bigger issue is that the racial/sexist/etc content can be shocking and immediately put someone off using the product, which I doubt is the case for the output being “too American.” OpenAI didn't just fine-tune it to avoid blatant racial/sexist/etc content, they openly claim to have invested a lot of effort in fine-tuning it to avoid subtle biases in those areas. And to be honest, a lot of people do feel…

But nothing about any of those examples is “discrimination”.

I agree they aren't direct discrimination, but almost anything can constitute indirect discrimination. A poorly localised product can have a disparate impact on foreign users, and as such indirectly discriminate against them.

Even as indirect discrimination, I'm not claiming it rises to the level of being legally actionable – but when OpenAI tries to eradicate subtle bias from an AI model, that's got nothing to do with legally actionable discrimination either, since it would be unlikely to be legally actionable if they decided not to invest in that.

I think one problem with this topic, is a lot of people don't understand the difference between "discrimination", "unethical discrimination", and "illegal discrimination". Some discrimination is both illegal and unethical; some discrimination is legal yet unethical; some discrimination is both legally and ethically justifiable. But many just blur the concepts of "discrimination" and "illegal discrimination" together.

Re: GitHub Copilot Chat Leaked Prompt

#369

Earlier quoted context omitted.

Moral relativism is ... human. Sure, it's hard to defend. But we embody it nonetheless. We're emotional creatures, we lack logical consistency in a fundamental way.

> we lack logical consistency in a fundamental way ... and "AI's don't really understand" as people say So, in the end, is anyone/anything capable of reasoning? Probably only humans in their specific fields of expertise. Even then, we are often updating our reasoning patterns in light of new discoveries, upturning previous reasoning. 99% of the time humans are just GPTs with hands and legs generating untrustworthy lo…

I'm starting to think that chess engines are capable of reasoning but humans and LLMs are not.

Re: GitHub Copilot Chat Leaked Prompt

#370
post #333

Earlier quoted context omitted.

anybody who uses gpt 4 or codex to do any of their programming or talk about sensitive data are not thinking things through and will end up leaking everything in their companies. i soon expect to see a ban on ai tools for many companies.

What about companies using Slack or Jira or Gmail? You're already leaking everything in your company to third parties - as a run of the mill tech company. Salesforce getting hacked and all Slack comms leaking vs all the OpenAI chat logs leaking... I know which one is more worrisome to me.

> I know which one is more worrisome to me.

third party provides are under strict legal contracts and they're liable if they mess up the privacy they've guaranteed you. You actually have recourse and can get compensation. Unless the legal situation is clear with these chatbots and the service providers can be held accountable, it's an entirely different situation.

Post reply on HN