Earlier quoted context omitted.
We do the same (all requests go to o1, sonnet and gemini and we store the results for later to compare) automatically for our research: Claude always wins. Even with specific prompting on both platforms. Especially frontend it seems o1 really is terrible.
Claude is trained on principles. GPT is trained on billions of edge cases. Which student do you prefer?
GPT-5 is behind schedule
641–650 of 1001 posts
Re: GPT-5 is behind schedule
#642Earlier quoted context omitted.
It would seem you don't care too much about verifying its output or about its correctness. If you did, it wouldn't take you just an hour. I guess you'll let correctness be someone else's problem.
Your wild assumptions and snarky accusations are unnecessary. The library is for me to use; there isn't a "someone else" for me to pass problems onto. I then did what I usually do — start writing real code with it ASAP, because real code is how you find real problems. I developed the library interactively, one API call at a time, in a manner akin to pair programming. Code quality was significantly better than I'd exp…
Re: GPT-5 is behind schedule
#643Earlier quoted context omitted.
That's the point being made. It's transformed robotics research, yes, but it both remains to see whether it will have a truly transformative effect on the field as experienced by people outside academia (I think this is quite probable) and more pointedly when .
I think it's impossible to spend a lot of time with these models without believing robotics is fundamentally about to transform. Even the most sophisticated versions of robotic logic pre-LLM/VLM feel utterly trivial compared to what even rudimentary applications of these large models can accomplish.
When questioned:
> believing robotics is fundamentally about to transform
These are not even remotely the same thing. Something that has happened already and is verifiable fact is not the same thing as your opinion, even if your opinion is based on a lot of sound arguments and reasoning.
Very tiresome to read so many claims of fact based on opinion of what will happen in the future.
Re: GPT-5 is behind schedule
#644"OpenAI’s is called GPT-4, the fourth LLM the company has developed since its 2015 founding." - that sentence doesn't fill me with confidence in the quality of the rest of the article, sadly.
When I read this I was honestly confused. I had never heard of NotebookLM before.
Re: GPT-5 is behind schedule
#645I'm sure the debate over the definition of AGI is important and will continue for a while, but... I can't care about it anymore. Between Perplexity searching and summarizing, Claude explaining, and qwen (and other tools) coding, I'm already as happy as can be with whatever you want to call this level of intelligence. Just today I used a completely local AI research tool, based on Ollama. It worked great. Maybe it won…
Same here. The ability to “talk to an expert” about any topic I’m curious about and ask very specific questions has been invaluable to me. It reminds me of being a kid and asking my grandpa a million questions, like how light bulbs worked, or what was inside his radio, or how do we have day and night. And before anyone talks about accuracy or hallucinations, these conversations usually are treated as starting off poi…
Re: GPT-5 is behind schedule
#646Earlier quoted context omitted.
Same here. The ability to “talk to an expert” about any topic I’m curious about and ask very specific questions has been invaluable to me. It reminds me of being a kid and asking my grandpa a million questions, like how light bulbs worked, or what was inside his radio, or how do we have day and night. And before anyone talks about accuracy or hallucinations, these conversations usually are treated as starting off poi…
(throwaway account because of what I'm about to say, but it needs to be said) While my main use case for LLMs is coding just like most people here, there are lots of areas that are being ignored. Did you know llama 3.X models have been trained as psychotherapists? It's been invaluable to dump and discuss feelings with it in ways I wouldn't trust any regular person. When real therapists also cost more than what people…
It's in our nature to crave outside acceptance of who we are. But may be taken to extreme, when we stop being wanting to be challenged at all we could lose touch with reality, society...
Re: GPT-5 is behind schedule
#647Earlier quoted context omitted.
It has replaced ~50% of my Google searches.
Yes but it also hasn't been attacked by ads yet. Google doesn't suck for lack of search results, it sucks because of ads. Imagine asking chatgpt to tell you about slopes in Colorado, and the first five answers are about how awesome North Face is and how you can order from them. You probably wouldn't use it as much.
Re: GPT-5 is behind schedule
#648So the team I lead does a lot of research around all the “plumbing” around LLMs. Both technical and from a product-market perspectives. What I’ve learned is that for the most part that AI revolution is not going to be because of PHD-level LLMs. It will be because people are better equipped to use the high-schooler level LLMs to do their work more efficiently. We have some knowledge graph experiments where LLMs contin…
Re: GPT-5 is behind schedule
#649Earlier quoted context omitted.
When you think about it it's astounding how much energy this technology consumes versus a human brain which runs at ~20W [1]. [1] https://hypertextbook.com/facts/2001/JacquelineLing.shtml
20w for 20 years to answer questions slowly and error-prone at the level of a 30B model. An additional 10 years with highly trained supervision and the brain might start contributing original work.
Re: GPT-5 is behind schedule
#650Earlier quoted context omitted.
Same here. The ability to “talk to an expert” about any topic I’m curious about and ask very specific questions has been invaluable to me. It reminds me of being a kid and asking my grandpa a million questions, like how light bulbs worked, or what was inside his radio, or how do we have day and night. And before anyone talks about accuracy or hallucinations, these conversations usually are treated as starting off poi…
> The ability to “talk to an expert” about any topic I’m curious about and ask very specific questions has been invaluable to me. It is dangerous to assume that LLMs are experts on any topic. With or without quotes. You are getting a super fast journalist intern with a huge memory but inability to reason critically, lacking understanding about anything and huge unreliability when it comes to answering questions (you…
And all in the hand of a few big tech corporations...