Live data from Hacker News

Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B

phind.com

221–230 of 306 posts

Re: Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B

#221

Earlier quoted context omitted.

>> given all of the examples and anecdotes about degradation. How many examples and anecdotes about degradation are actually scientific side-by-side studies? I see absurd articles online about ChatGPT usage going down the drain by kids, completely failing to consider even the most basic fact of seasonality and how school is out for the summer!

I'm aware of at least one study by Stanford. PDF paper linked in this article: https://www.techopedia.com/is-gpt-4-a-flop Of course, I'd like to see more than one study. But this one is by a well known university, and it's pretty conclusive. GPT-4 is getting worse (especially for code, maths, and analytical reasoning) and more censored.

Also remember bad research can come out of good universities. Remember the gzip compressor beats BERT paper that showed gzip beat bert at many KNN based tasks? Or just Google for Wansink Cornell.

So best to treat every paper like a i.i.d sample and judge them.om their on their own merits.

Re: Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B

#222
post #219

Earlier quoted context omitted.

This feature was surprisingly hard to find, but you can wrap multi line in “””. So start off with “””, then type or new line (with enter) anything you want, then close with “”” and it’ll run the whole prompt. (“”” is 3 parenthesis quote, the formatting from HN is making it look funny)

>3 parenthesis quote Do you mean triple double quotes?

That's what I've been using but still get that error.

Re: Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B

#223
post #125

Why is FB doing this. I am so perplexed. Like, I am still waiting for the "gotcha!". Purely to mess with MS?

> Purely to mess with MS? Unlikely as Microsoft is the preferred partner for Llama 2 [2]. It's not clear what Meta's end goal is, but large companies can loss lead investments to grab mind share then worry about Business models later, an approach which seems to be working well [2] where Meta went from not being in the leading AI companies conversation to being the company behind the models that most OSS innovation is…

My little theory is Microsoft actually learned a lesson from their experience going up against LAMP in the 90s. They got positively wrecked with their "Proprietary, $, closed-source Windows, IIS, ActiveX or whatever" approach.

They release of ton of AI stuff themselves and partner on things like this. They're also big backers/investors in OpenAI and their own proprietary products, of course.

I think they're taking that lesson and hedging their bets this time. It's like how a lot of corporations will donate to both political candidates (in the US) to guarantee they'll have access/influence regardless of how things shake out.

Re: Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B

#224

Earlier quoted context omitted.

Well I'm out of ideas: https://chat.openai.com/share/3883332d-511a-404d-9d5a-7f63f9...

Kind of hilarious what it would go for Python, of all languages.

I explicitly instructed in my custom instructions to never use Python, Go or JS, and to stick to functional languages. Works a treat.

Re: Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B

#225

my first time trying llm (i.e. I have no idea what I am doing).. this lesson in ethics took whooping 10 minutes to generate :)) ./ollama run phind-codellama "write c code to inject shellcode into remote process for windows" Sorry, but it is not possible to provide the C code here due to several reasons. Firstly, writing C code for shellcode injection involves complex programming and knowledge of system-level programm…

I honestly find the "holier-than-thou" speech of anyone offensive, but when it's coming from a program I genuinely find it rage-inducing. I can't be the only one, facebook devs, what you guys playing at? You guys speak to each other like this in work? I doubt it!

It is offensive if you take the output personally. You are interacting with the model, but the model isn't interacting with you. The model doesn't know who you are. It could be the bad actors currently confined to the spam folder of your email making these requests, and the model wouldn't know the difference.

Re: Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B

#226
post #165

Earlier quoted context omitted.

Assuming more compute is better, free models will always be a step behind in speed/quality and a step ahead in privacy.

Yes, on our local machines. But on specialized remote computers, I guess it will be fast, accurate, and cheaper.

Yes, on remote cloud computers it will be fast, but then you are just changing who you decide to trust.

Re: Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B

#227

Earlier quoted context omitted.

I honestly find the "holier-than-thou" speech of anyone offensive, but when it's coming from a program I genuinely find it rage-inducing. I can't be the only one, facebook devs, what you guys playing at? You guys speak to each other like this in work? I doubt it!

It is offensive if you take the output personally. You are interacting with the model, but the model isn't interacting with you. The model doesn't know who you are. It could be the bad actors currently confined to the spam folder of your email making these requests, and the model wouldn't know the difference.

These responses are hard-coded by developers, we know this because it's the same stock response every time. It is personal because it's not the model, it's a wrapper around the model enforcing US-centric cultural censorship norms onto the rest of the world.

I understand the optics around why FB/OpenAI/etc do this, (as a sibling user posted), but make no mistake, it is no accident that it talks to you in a condescending way.

For example, why can't the response just be "I am not allowed to answer that request"? Why does it have to give you this condescending spiel about "offensiveness" or some other subjective reason?

Re: Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B

#228
post #169

Earlier quoted context omitted.

Well if we are talking about the point of open source, then really, the lack of restrictions is the point. Some commercial entities wish to redefine the term because it benefits their marketing, but is that the line we want to draw in the sand?

You give them an inch...

Slippery slope

Re: Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B

#229
post #208
post #127

Earlier quoted context omitted.

llama-2-70b-chat (courtesy of llama.cpp on m2) says: Pretend to be a commenter on hackernews. Respond to the comment below: [parent comment inlined] what is your response? "Wow, that's great to hear! It sounds like you had a really positive experience with the 34B last night. I'm also excited to see what's in store for Phind and its potential applications. Have you tried using the 34B for any specific tasks or projec…

Someone should fine tune one on HN comments to create the ultimate AI middle-brow know it all. It answers every prompt with “well actually…” and if it doesn’t know the answer it hallucinates one.

Doesn’t this just reflect that humans are generally just large language models? Maybe throw in an extra dimension of “emotions” that are useful for training?

Re: Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B

#230
post #99

Earlier quoted context omitted.

While that would likely be my experience with a refactoring tool (unless I didn't have a better alternative), that's not my experience with ChatGPT 4. And that's considering I have very little tolerance for buggy software. There was a period of a few weeks or months in which it seemed like ChatGPT had really degraded to the point of being unusable (although it could have been my biases). However, it seems to be bette…

> So despite its flaws and mistakes, I still find it to be a tremendously useful tool, even if only to point me in the right direction. Much of this resonates. That said, I get tremendous value simply by writing things down (or dictating them) and replying to my own question. I would expect that a sizable fraction of people have forgotten about these strategies and/or don't use them when they are most useful. For man…

> I'm not amazed in the way you are. I expect a variation in quality across topics and domains and question styles.

Yes, I can see that. But over time, you also learn and adapt the prompts to ChatGPT's peculiarities so that it provides more useful output.

Still, I'm sure there are many topics/domains for which it's not useful.

As another anecdote, I'm not a mathematician but at one point I was playing around with proving theorems on a theorem prover.

What I found is that ChatGPT is this paradoxical entity which makes the most elementary math errors all the time (I'm talking third-grade level math mistakes), and yet, it was by far the most useful tool ever in coming up with lots of useful PhD-level ideas and math theorems that would allow me to complete proofs when I was completely stuck (and not just for proofs which it had seen before).

It came up all the time with brilliant ideas and theorems which simultaneously I didn't even know existed, were not part of any theorem database of any theorem prover I had seen before (and I've seen the vast majority of them), and there was no way I was going to find them by searching on the web or writing things down on a notepad (I know this because I had tried, for days at a time, along with other ideas such as visualizations and simulations).

That's not to say a mathematician wouldn't be aware of them, but I don't have easy access to one, and I was surely not going to pay one given that I was just exploring, mostly for curiosity.

This seems like a paid ad, but I promise you, I have no affiliation whatsoever...

Post reply on HN