Live data from Hacker News

Google warns its own employees: Do not use code generated by Bard

theregister.com

151–160 of 166 posts

Re: Google warns its own employees: Do not use code generated by Bard

#151

Earlier quoted context omitted.

Everything you do or know is a derivative of your own training set. True original thoughts without context don’t exist, well at least in any way we can find out as every human is the product of their own training set. Just out curiosity, what makes the code you write, more original than an LLMs who’s training set is way bigger than yours and likely will more variance in how to achieve same goal. I’m not trying to be…

This is an interesting thought. How much do we really owe to our teachers? If it wasn't for the school that taught me to read all those years ago I wouldn't have a job!

Yes deduced to it’s lowest form but there’s billions of inputs that go into your own training set to produce the outcome you’re at.

Along the way, you’ve experienced trauma which nudged you a direction, you’ve been inspired by other people which left an impression in your mind, you’ve had people teach you 1 skill and another person teach you a different skill, and you combined them, life experiences that changed your outlook etc.

It’s less about the specific teachers and more about every single person and event that has happened to you is your own ‘training set’. But sometimes we can definitely point to a specific event or person who influenced us to where we are at today and I like to think of that just like model weights in current LLMs.

Re: Google warns its own employees: Do not use code generated by Bard

#152

Earlier quoted context omitted.

On the other hand, somewhat in the style of NP-Completeness, it is often easier to verify that something is correct than it is to generate it from scratch. Even if I can't cook, I can say it's a tasty meal (:

It's usually somewhat easy to verify a piece of code is not obviously wrong, what's much harder is proving that a piece of code is not subtlety wrong. When given a complete piece of code that appears to work, it can be very easy to convince yourself that you understand it well enough to know that it is correct, even when it's not. This problem isn't unique to LLMs, refer to the case of programmers copying binary sear…

I'm reminded of a friend who worked in radio hardware design. They'd use simulation and fuzzy/genetic algorithms to create a circuit, and then verify its performance with experiments. But they couldn't always say exactly why the circuit worked, just that it met the performance criteria.

It's an interesting divergence in software, between those who manage complexity by adding more human-understandable abstraction, and those who manage it by just verifying the results, letting the complexity fly free. All the ML stuff is definitely taking big steps down the latter path.

Re: Google warns its own employees: Do not use code generated by Bard

#153

Earlier quoted context omitted.

I don't think the comparison with "employees must wash hands" signs is apt. Should we see legal action against large language model of questionable heritage produced code, the consequences will be dire for companies . Being able to say to their (probably then ex-) employees "I told you so" will gain them nothing. There will not be nearly enough to get from the breaching ex-employee to compensate for any damages.

Let's say I go to a restaurant, get food poisoning, and sue the restaurant. Turns out the employee didn't wash their hands (or at least, that's the legal theory). If the restaurant has signs like that up, and consistent messaging to their employees, and new employee training, and other "best practices", then I may still be able to sue them and win. But if they don't have all that stuff, I may be able to sue them for…

I agree with that. I think the fundamental difference in Google's case is that the suing party is different from the addressee of the warning. Moreover the suing party is powerful and the addressee is completely negligible. Hopefully courts sees that, but who knows?

Re: Google warns its own employees: Do not use code generated by Bard

#154

Earlier quoted context omitted.

Can you quote where he said he fed it internal code?

“ as a google swe I find bard almost useless. it doesnt understand any of our internal tooling or context.” How would he know it doesn’t understand if he hasn’t used it?

He likely asked it a question.

Re: Google warns its own employees: Do not use code generated by Bard

#155
post #21

This is Google, who have a monorepo, who are afraid of the legal risks to the monorepo. The copyright of the output of these tools is not yet determined, it's a risk. The advice not to use the tools in a risky way not only directly mitigates the risk but like an "employees must wash hands" sign in a restaurant restroom, it can be seen as a reasonable step that transfers some of the liability to the staff member. It's…

I'm waiting for the patent troll equivalent of training data

Sounds like Reddit CEO's next pivot...

Re: Google warns its own employees: Do not use code generated by Bard

#156

Earlier quoted context omitted.

Can you quote where he said he fed it internal code?

“ as a google swe I find bard almost useless. it doesnt understand any of our internal tooling or context.” How would he know it doesn’t understand if he hasn’t used it?

It's really easy to use it without violating policies.

In fact, we were literally asked to.

Re: Google warns its own employees: Do not use code generated by Bard

#157

Earlier quoted context omitted.

Everything you do or know is a derivative of your own training set. True original thoughts without context don’t exist, well at least in any way we can find out as every human is the product of their own training set. Just out curiosity, what makes the code you write, more original than an LLMs who’s training set is way bigger than yours and likely will more variance in how to achieve same goal. I’m not trying to be…

This seems absurdly reductionist to me. Wouldn't all work simply be attributable to the first proto-human who grabbed a stick or rock and used it as a tool? Surely even proto-humans communicated knowledge by demonstrations amongst themselves, even if they lacked complicated language.

Humans are just a collection of experiences that are using past data to make future assumptions. So yes, thank you Mr. Caveman because of your ingenuity of using a rock and a stick to create a hammer, we now have a full modern society built off the back of tools. But that seems ridiculous doesn’t it?

So at what point do we stop attributing the previous innovations that led to our current innovations? And why do we stop there? Would Mr. Caveman be the only special human to ever figure that out or could the argument be made that eventually someone else would have figured out how to make tools and therefore attribution is just pointless?

What I am trying to get at is, everything you do or create, is because of work the all of humanity has done. So pertaining to copyright, why should 1 single person claim an idea as solely theirs when that idea was not created in a vacuum.

I should also say, I am not discounting anyone’s work, but rather if the monetary reasons for creating became secondary, would the need for copyright even exist?

Re: Google warns its own employees: Do not use code generated by Bard

#158

Earlier quoted context omitted.

Everything you do or know is a derivative of your own training set. True original thoughts without context don’t exist, well at least in any way we can find out as every human is the product of their own training set. Just out curiosity, what makes the code you write, more original than an LLMs who’s training set is way bigger than yours and likely will more variance in how to achieve same goal. I’m not trying to be…

> Just out curiosity, what makes the code you write, more original than an LLMs Indeed in many ways LLM code is more original than mine, as I limit myself to calling only functions that exist and producing code that will compile, unlike LLMs which have no such limitations. /s

Maybe, but it sounds like you’re operating within pretty restricted walls. Sometimes there are better ways to do things that you haven’t thought of but it requires experimentation.

That would be like assuming every piece of code you’ve ever written compiled perfectly on every run. Do you hold yourself to the same standard? Or do you give yourself some leeway because you walk through it, talk through, think through it and then come to a solution?

Re: Google warns its own employees: Do not use code generated by Bard

#159

Earlier quoted context omitted.

I'm waiting for the patent troll equivalent of training data

Sounds like Reddit CEO's next pivot...

next pivot? they ain't never made money off of ads son. datamining and consensus building were always an aim.

Re: Google warns its own employees: Do not use code generated by Bard

#160

Earlier quoted context omitted.

I was pointing out that the case where the network produces exact copies of copyrighted code (or very very close to exact copies, such as changing a variable name or removing a comment) is clear, and requires no new litigation. While I also think there is a good chance that what you are claiming (that any output is a derived work of the training set), this is definitely not settled, and will require someone going to…

Worse: The US Supreme Court will reach a verdict. The European equivalent will also reach a verdict, which may be different. And the Australians. And the British. And the Chinese. And... It won't be over when the first top-level court decides.

Hopefully it'll over when the first top-level court of any country too big to ignore says "no", at which point anyone trying to be safe will stop allowing LLM-generated code in their codebase.
Post reply on HN