Live data from Hacker News

OpenAI GPT-4 vs. Groq Mistral-8x7B

serpapi.com

41–50 of 139 posts

Re: OpenAI GPT-4 vs. Groq Mistral-8x7B

#41
post #34
post #24

Earlier quoted context omitted.

They already have superhuman image classification performance.

I remember talking to a radiologist who said he was sure something like this was coming like ten years ago where instead of a radiologist looking at scans manually, a machine would go through a lot of images and flag some for manual review. We haven't even gotten there yet, have we?

Yes, we absolutely are there: https://youtu.be/D3oRN5JNMWs?feature=shared

My professor (Sir Michael Brady) at university 14 years ago set up a company to do this very thing, and he already had reliable models back before 2010. I believe their company was called Oxford Imaging or something similar.

Re: OpenAI GPT-4 vs. Groq Mistral-8x7B

#42
post #15

Earlier quoted context omitted.

I hear it should be dropped this summer

According to Sam Altman in a podcast with Lex Fridman this week, there is no real indication that it will be dropped this year. They will release a new model, but it might not be GPT-5

Fair enough, I got the info from this article

https://web.archive.org/web/20240319224624/https://www.busin...

Re: OpenAI GPT-4 vs. Groq Mistral-8x7B

#43
post #22
post #13

Earlier quoted context omitted.

If you have speed you can generate multiple answers and have another model pick the best one.

If I ask an LLM a very complex and specific question 500 times, if it just doesn't know the facts you'll still get the wrong answer 500 times. That's understandable. The real problem is when the AI lies/hallucinates another answer with confidence instead of saying "I don't know".

The problem is asking for facts, LLM are not a database so they know stuff but it is compressed so expect wrong facts, wrong names, dates, wrong anything.

We will need an LLM as a front end then it will generate a query to fetch the facts from the internet or a database , then maybe format the facts for your consumption.

Re: OpenAI GPT-4 vs. Groq Mistral-8x7B

#45
post #15

A bit off-topic but maybe not? Any words on GPT-5? Is that coming? Or is OpenAI just focusing on the Sora model?

I hear it should be dropped this summer

My understanding from the lex podcast: they will release a lot of new models this year, but they will release intermediate models first before gpt-5

Re: OpenAI GPT-4 vs. Groq Mistral-8x7B

#46

A bit off-topic but maybe not? Any words on GPT-5? Is that coming? Or is OpenAI just focusing on the Sora model?

There's no reason for OpenAI to release the model. They have close to 100% market anyways and releasing GPT-5 likely won't increase the total market as it is a incremental leap. And it's a open secret that most other models used GPT-4 synthetic data for training to come close to it. They would likely wait till any model performs better than GPT 4 for the same price

There is reason to release new models if said models would be capable of grabbing a significant portion of job market currently occupied by humans.

Re: OpenAI GPT-4 vs. Groq Mistral-8x7B

#47

Earlier quoted context omitted.

If you need a perfect score, don't use LLMs. This seems obvious to me, even given the state of the art LLMs. I am a heavy user of GPT4 and I wouldn't bet $1000 bucks on it being 100% reliable for any non-trivial task.

They'll get better. Humans are far from perfect, and I have no doubt that LLMs will eventually outperform them for non-trivial tasks consistently.

Or they might not get better. It could be that we are at a local optimum for that sort of thing, and major improvements will have to wait (perhaps for a very long time) for radical new technologies.

Re: OpenAI GPT-4 vs. Groq Mistral-8x7B

#48

Earlier quoted context omitted.

If you need a perfect score, don't use LLMs. This seems obvious to me, even given the state of the art LLMs. I am a heavy user of GPT4 and I wouldn't bet $1000 bucks on it being 100% reliable for any non-trivial task.

They'll get better. Humans are far from perfect, and I have no doubt that LLMs will eventually outperform them for non-trivial tasks consistently.

Machine learning models will get better for sure. We don't know if LLM are the end game though and it's not sure if this particular technique is what we'll need to reach the next level.

Re: OpenAI GPT-4 vs. Groq Mistral-8x7B

#49
post #22
post #13

Earlier quoted context omitted.

If you have speed you can generate multiple answers and have another model pick the best one.

If I ask an LLM a very complex and specific question 500 times, if it just doesn't know the facts you'll still get the wrong answer 500 times. That's understandable. The real problem is when the AI lies/hallucinates another answer with confidence instead of saying "I don't know".

The weird problem is with LLM hallucinations is that it usually will acknowledge its mistake and correct itself if you call it out. My question is why can't LLMs included a sub-routine to check itself before answering. Simply asking itself something like "this answer may not be correct, are you sure you're right?"

Re: OpenAI GPT-4 vs. Groq Mistral-8x7B

#50
post #41
post #34

Earlier quoted context omitted.

I remember talking to a radiologist who said he was sure something like this was coming like ten years ago where instead of a radiologist looking at scans manually, a machine would go through a lot of images and flag some for manual review. We haven't even gotten there yet, have we?

Yes, we absolutely are there: https://youtu.be/D3oRN5JNMWs?feature=shared My professor (Sir Michael Brady) at university 14 years ago set up a company to do this very thing, and he already had reliable models back before 2010. I believe their company was called Oxford Imaging or something similar.

Yep, everyone seems to forget that ML was available before 2021. Had a conversation recently with my former colleague who learned about some plastic packaging company which used "AI" to predict client orders and inform them about scheduling implications. When I told him that you don't need Transformers and 30GB models for that, he was quasi-confused, cause he kinda knew it but the hype just overtook his knowledge.
Post reply on HN