Live data from Hacker News

Phind Model beats GPT-4 at coding, with GPT-3.5 speed and 16k context

phind.com

341–350 of 358 posts

Re: Phind Model beats GPT-4 at coding, with GPT-3.5 speed and 16k context

#341

I tried my standard "trick" question I use for LLMs: "Give me five papers with code demonstrating the state of the art of machine learning which uses geospatial data (e.g. GeoJSON) as both input and output." There is no such state of the art. My hand-wavey understanding is that GIS data is non-continuous, which makes it useless for transformers, and also contextual, which makes it useless for anything else. Will defe…

I don't see how this would be relevant for a code model? The code model isn't trained to retrieve papers/articles, it's meant to complete code. Whether or not you find hallucination in a unrelated task isn't particularly interesting.

[deleted]

Re: Phind Model beats GPT-4 at coding, with GPT-3.5 speed and 16k context

#342

I tried my standard "trick" question I use for LLMs: "Give me five papers with code demonstrating the state of the art of machine learning which uses geospatial data (e.g. GeoJSON) as both input and output." There is no such state of the art. My hand-wavey understanding is that GIS data is non-continuous, which makes it useless for transformers, and also contextual, which makes it useless for anything else. Will defe…

> the state of the art of machine learning which uses geospatial data (e.g. GeoJSON) as both input and output > There is no such state of the art Some GIS work uses vector data: points/lines/polygons representing features (e.g., the location of roads or the outlines of buildings), which can be stored in formats like GeoJSON or WKT. But other work uses remote sensing data/satellite imagery that can be stored in raster…

[deleted]

Re: Phind Model beats GPT-4 at coding, with GPT-3.5 speed and 16k context

#343

I tried my standard "trick" question I use for LLMs: "Give me five papers with code demonstrating the state of the art of machine learning which uses geospatial data (e.g. GeoJSON) as both input and output." There is no such state of the art. My hand-wavey understanding is that GIS data is non-continuous, which makes it useless for transformers, and also contextual, which makes it useless for anything else. Will defe…

I don't see how this would be relevant for a code model? The code model isn't trained to retrieve papers/articles, it's meant to complete code. Whether or not you find hallucination in a unrelated task isn't particularly interesting.

Damn, this is how I learn that HN doesn't have a block function. What a shame.

My friend, can you do me a favour and actually click the link and have a play with the app? If you do, you will discover that what you're dealing with there is an LLM. That's literally why it's being compared to other LLMs.

No idea what you were trying to achieve with this comment. "The code model isn't trained to retrieve articles." a) neither is any other LLM, what's your point? and b) the app on the other end of that URL retrieves articles - it's not even tangential to the app, it's key functionality.

Re: Phind Model beats GPT-4 at coding, with GPT-3.5 speed and 16k context

#344

Earlier quoted context omitted.

Well no we have the anecdotes of all the HN folks which I trust many, many times more than a benchmark.

lol, you can continue trusting anecdotes from internet. Industry prefers more scientific methods.

So Paul Graham posted that Phind is better and got absolutely destroyed in the comments

https://twitter.com/paulg/status/1719657855240815026

No, I do not take these benchmarks seriously and for good reason. They're benchmarks. The only thing that matters is the user's direct experience of the product. And Phind isn't there.

Re: Phind Model beats GPT-4 at coding, with GPT-3.5 speed and 16k context

#345

Earlier quoted context omitted.

lol, you can continue trusting anecdotes from internet. Industry prefers more scientific methods.

So Paul Graham posted that Phind is better and got absolutely destroyed in the comments https://twitter.com/paulg/status/1719657855240815026 No, I do not take these benchmarks seriously and for good reason. They're benchmarks. The only thing that matters is the user's direct experience of the product. And Phind isn't there.

> got absolutely destroyed in the comments

by tweeter trolls?..

Re: Phind Model beats GPT-4 at coding, with GPT-3.5 speed and 16k context

#347
post #138

It failed for me at a much more basic level. I asked 5 different, and increasing explicit, variations of the following question: "Can you generate HTML and CSS for a JPG mockup I'm going to give you?" Each time it answered along the following lines: "Sure, here is how you can create HTML and CSS from a JPG mockup. Follow this process..." In my experience this never happens with GPT-4.

I've not seen that anywhere, ChatGPT does image input now? Do you have examples of the output from feeding it a JPEG?

Yep, ChatGPT did a pretty impressive job when I tested a few days ago. Just grab a mockup from a Google search and prompt ChatGPT 4 to generate a web page. I'm sure your milage may vary.

However, my point was that Phind's answer was worse than a No, or a hallucinated attempt would've been. By saying "Yes, here is how YOU can do it...", it left the impression that it didn't even understand the question.

Re: Phind Model beats GPT-4 at coding, with GPT-3.5 speed and 16k context

#348
post #347

Earlier quoted context omitted.

I've not seen that anywhere, ChatGPT does image input now? Do you have examples of the output from feeding it a JPEG?

Yep, ChatGPT did a pretty impressive job when I tested a few days ago. Just grab a mockup from a Google search and prompt ChatGPT 4 to generate a web page. I'm sure your milage may vary. However, my point was that Phind's answer was worse than a No, or a hallucinated attempt would've been. By saying "Yes, here is how YOU can do it...", it left the impression that it didn't even understand the question.

Thanks for expanding on your answer.

Re: Phind Model beats GPT-4 at coding, with GPT-3.5 speed and 16k context

#349

Earlier quoted context omitted.

They are. Moreover, the idea that AI companies are missing and/or not implementing this “obvious” tactic is hilarious. Folks, these approaches have profound consequences for training and inference performance. Y’all aren’t pointing out some low hanging fruit here, lol

Actually, yes I am pointing out low hanging fruit here. These approaches do not have "profound consequences" for inference or training performance. In fact, sentence transformer models run orders of magnitude more quickly. Performance penalties will be small. Also, I actually have several top NLP conference publications, so I'm not some charlatan when I say these things. I've actually physically used and seen these t…

> In fact, sentence transformer models run orders of magnitude more quickly. Performance penalties will be small.

They do not. Sentence transformers aren't new, and have well-known trade offs. What source or line of reasoning misled you to believe otherwise?

> Here's more examples of low hanging fruit. The proof in that they work is in the implementations which I provide. You can run them, they work!: https://gist.github.com/Hellisotherpeople/45c619ee22aac6865c...

This...is your blog about prompt engineering. What do you believe this "proves"? How have you blown away current production encoding or attention mechanisms?

Re: Phind Model beats GPT-4 at coding, with GPT-3.5 speed and 16k context

#350

> You can now get high quality answers for technical questions in 10 seconds instead of 50. ChatGPT 4 does not take 50 seconds to answer, so I don't understand this comparison.

Recently I've used gpt 4 and yes it does take up to a minute even for easy questions. I've asked it how to scp a file on Windows 11 and it'll take a minute to tell me all the options possible. If this takes 1/5th the time for equivalent questions, I'd consider switching

Take a look at the AutoExpert custom instructions: https://github.com/spdustin/ChatGPT-AutoExpert

It lets you specify verbosity from 1 to 5 (e.g. "V=1" in the prompt). Sometimes the model will just ignore that, but it actually does work most of the time. I use a verbosity of 1 or 2 when I just want a quick answer.

Post reply on HN