Live data from Hacker News

GPT-5.1: A smarter, more conversational ChatGPT

openai.com

401–410 of 766 posts

Re: GPT-5.1: A smarter, more conversational ChatGPT

#401

Having gone through the explainations of the Transformer Explainer [1], I now have a good intuition for GPT-2. Is there a resource that gives intuition on what changes since then improve things like more conceptually approaching a problem, being better at coding, suggesting next steps if wanted etc? I have a feeling this is a result of more than just increasing transformer blocks, heads, and embedding dimension. [1]…

Most improvements like this don't come from the architecture itself, scale aside. It comes down to training, which is a hair away from being black magic.

The exceptions are improvements in context length and inference efficiency, as well as modality support. Those are architectural. But behavioral changes are almost always down to: scale, pretraining data, SFT, RLHF, RLVR.

Re: GPT-5.1: A smarter, more conversational ChatGPT

#403

Seems like people here are pretty negative towards a "conversational" AI chatbot. Chatgpt has a lot of frustrations and ethical concerns, and I hate the sycophancy as much as everyone else, but I don't consider being conversational to be a bad thing. It's just preference I guess. I understand how someone who mostly uses it as a google replacement or programming tool would prefer something terse and efficient. I fall…

Ideally, a chatbot would be able to pick up on that. It would, based on what it knows about general human behavior and what it knows about a given user, make a very good guess as to whether the user wants concise technical know-how, a brainstorming session, or an emotional support conversation.

Unfortunately, advanced features like this are hard to train for, and work best on GPT-4.5 scale models.

Re: GPT-5.1: A smarter, more conversational ChatGPT

#404
it feels incredibly dumb now, getting some really basic questions wrong and just throwing nuance to the wind. for claiming to be more human, it understands far less. for example: if I start at a negative net worth how long until I am a millionaire if I consistently grow 2.5% each month? Anyone here would have a basic understand the premise and be able to start answering, 5.1 says it's impossible, with hand holding it will insist you can only reach 0 but that growth isn't the same as a source of income. further hand holding gets it to the point of insisting it cannot continue without making assumptions, goading it will have it arrive at the incorrect value of 72 months, further goading will get 240 months, it took the lazy way out and assumed a static inflation from 2024, then a static income.

o3 is getting it no problem, first try, a simple and reasonable answer, 101 months. claude (opus 4.1) does as well, 88-92 months, though it uses target inflation numbers instead of something more realistic.

Re: GPT-5.1: A smarter, more conversational ChatGPT

#405
post #345

Earlier quoted context omitted.

> All the examples of "warmer" generations show that OpenAI's definition of warmer is synonymous with sycophantic, which is a surprise given all the criticism against that particular aspect of ChatGPT. Have you considered that “all that criticism” may come from a relatively homogenous, narrow slice of the market that is not representative of the overall market preference? I suspect a lot of people who are from a very…

I'll be honest, I like the way Claude defaults to relentless positivity and affirmation. It is pleasant to talk to. That said I also don't think the sycophancy in LLM's is a positive trend. I don't push back against it because it's not pleasant, I push back against it because I think the 24/7 "You're absolutely right!" machine is deeply unhealthy. Some people are especially susceptible and get one shot by it, some pe…

I hate NOTHING quite the way how Claude jovially and endlessly raves about the 9/10 tasks it "succeeded" at after making them up, while conveniently forgetting to mention it completely and utterly failed at the main task I asked it to do.

Re: GPT-5.1: A smarter, more conversational ChatGPT

#406
post #373
post #359

Earlier quoted context omitted.

ChatGPT nowdays gives the option of choosing your preferred style. I have choosen "robotic" and all the ass kissing instantly stopped. Before that, I always inserted a "be conciseand direct" into the prompt.

i found robotic consistenly underperformed in tasks and it also drastically reduced the temperature, so connecting suggestions and ideas basically disappeared. I just wanted it to not kiss my ass the whole time

Did you made a comparison?

I got did not and also had the impression it performed lower, but it still solved the things I told it to do and I just switched very recently.

Re: GPT-5.1: A smarter, more conversational ChatGPT

#407
post #345

Earlier quoted context omitted.

I'll be honest, I like the way Claude defaults to relentless positivity and affirmation. It is pleasant to talk to. That said I also don't think the sycophancy in LLM's is a positive trend. I don't push back against it because it's not pleasant, I push back against it because I think the 24/7 "You're absolutely right!" machine is deeply unhealthy. Some people are especially susceptible and get one shot by it, some pe…

I hate NOTHING quite the way how Claude jovially and endlessly raves about the 9/10 tasks it "succeeded" at after making them up, while conveniently forgetting to mention it completely and utterly failed at the main task I asked it to do.

An old adage comes to my mind: If you want something to be done the way you liked, do it yourself.

Re: GPT-5.1: A smarter, more conversational ChatGPT

#408
post #191

I've switched over to https://thaura.ai , which is working on being a more ethical AI. A side effect I hadn't realized is missing the drama over the latest OpenAI changes.

What a bizarre product. Weirdly political message and ethnic branding. I suppose "ethical AI" means models tuned to their biases instead of "Big Tech AI" biases. Or probably just a proxy to an existing API with a custom system prompt. The least they could've done is check their generated slop images for typos ("STOP GENCCIDE" on the Plans page). The whole thing reeks of the usual "AI" scam site. At best, it's profiti…

I assure you it's not a scam. We work with them heavily at Tech for Palestine. Will send over your feedback, thanks!

What would be helpful to assuage your fears? Would you like more technical info, or perhaps a description of the "biases" used?

Re: GPT-5.1: A smarter, more conversational ChatGPT

#409

Seems like people here are pretty negative towards a "conversational" AI chatbot. Chatgpt has a lot of frustrations and ethical concerns, and I hate the sycophancy as much as everyone else, but I don't consider being conversational to be a bad thing. It's just preference I guess. I understand how someone who mostly uses it as a google replacement or programming tool would prefer something terse and efficient. I fall…

A chatbot that imitates a friendly and conversational human is awesome and extremely impressive tech, and also horrifyingly dystopian and anti-human. Those two points are not in contradiction.

Re: GPT-5.1: A smarter, more conversational ChatGPT

#410

All the examples of "warmer" generations show that OpenAI's definition of warmer is synonymous with sycophantic , which is a surprise given all the criticism against that particular aspect of ChatGPT. I suspect this approach is a direct response to the backlash against removing 4o.

I know it is a matter of preference, but I loved the most GPT-4.5. And before that, I was blow away by one of the Opus models (I think it was 3).

Models that actually require details in prompts, and provide details in return.

"Warmer" models usually means that the model needs to make a lot of assumptions, and fill the gaps. It might work better for typical tasks that needs correction (e.g. the under makes a typo and it the model assumes it is a typo, and follows). Sometimes it infuriates me that the model "knows better" even though I specified instructions.

Here on the Hacker News we might be biased against shallow-yet-nice. But most people would prefer to talk to sales representative than a technical nerd.

Post reply on HN