Live data from Hacker News

Open models by OpenAI

openai.com

781–790 of 909 posts

Re: Open models by OpenAI

#782
post #637

Earlier quoted context omitted.

I tried the two US presidents having the same parents one, and while it understood the intent, it got caught up in being adamant that Joe Biden won the election in 2024 and anything I do to try and tell it otherwise is dismissed as being false and expresses quite definitely that I need to do proper research with legitimate sources.

I think the lesson is: smaller models hallucinate more, so only use them in your applications where you load up large prompts with specific data to reason about. Then even the small Google gemma3n 4B model can be amazingly useful. I use the SOTA models from Google and OpenAI mostly for getting feedback on ideas, helping me think through designs, and sometimes for coding. Your question is clearly best answered using a…

> think the lesson is: smaller models hallucinate more

The interesting part isn't the hallucination, but the sheer unwillingness to take in new information.

Re: Open models by OpenAI

#783
post #338

Just posted my initial impressions, took a couple of hours to write them up because there's a lot in this release! https://simonwillison.net/2025/Aug/5/gpt-oss/ TLDR: I think OpenAI may have taken the medal for best available open weight model back from the Chinese AI labs. Will be interesting to see if independent benchmarks resolve in that direction as well. The 20B model runs on my Mac laptop using less than 15GB…

Nice write up! One test I do is to give a common riddle but word it slightly to see if it can actually reason. For example: "Bobs dad has five daughters, Lala, Lele, Lili, Lolo and ???" The 20B model kept picking the answer of the original riddle, even after explaining extra information to it. The original riddle is: "Janes dad has five daughters, Lala, Lele, Lili, Lolo and ???"

Nice test. I tested it with llama.cpp and the 20b model. He got the answer correct at first try on my side. The internal/thinking dialogue is fun to read - it could be my own inner dialogue ;)

``` > "Bobs dad has five daughters, Lala, Lele, Lili, Lolo and ??? analysisThe question: "Bobs dad has five daughters, Lala, Lele, Lili, Lolo and ???". It's a common riddle. The answer is Bob? Actually "Bob's dad has five daughters: Lala, Lele, Lili, Lolo and..." The riddle is to trick: it's "Bob" himself is a daughter? Wait Bob's dad has five daughters: Lala, Lele, Lili, Lolo, and Bob? But Bob is a name often male. But the riddle says Bob's dad. So Bob is a person who has a dad, and that dad has five daughters. But Bob could be the fifth daughter? That would be weird: Bob is likely the name of the fifth daughter. So answer: Bob. But sometimes the answer is "Lulu" or "Bob's sister"? Let's think. The riddle is a trick: It's "Bobs dad has five d aughters: Lala, Lele, Lili, Lolo, and ... Bob." Because the riddle begins with "Bobs dad has five daughters" implying Bob is one of the daughters. So the missing daughter is Bob. So answer: "Bob".

Thus answer: Bob. But sometimes the riddle goes "Bob's dad has five daughters, Lala, Lele, Lili, Lolo, and ... Who is the fifth?" The answer: Bob (the name of the person speaking, "Bobs"). But in riddle it's "Bobs dad has five daughters , Lala, Lele, Lili, Lolo and ???". The missing name is "Bob." So answer: Bob.

Thus the answer: Bob. Or maybe it's "Boba"? But no.

Hence answer: Bob.

We'll answer accordingly.assistantfinalThe missing daughter is *Bob*. The riddle starts with “Bob’s dad …” so Bob himself is the fifth daughter. ```

Re: Open models by OpenAI

#784

Earlier quoted context omitted.

I mean I would hardly blame the specific model, Anthropic has a specific mention in their system prompts on trump winning. For some reason llms get confused with this one.

I think models generally have cognitive dissonance regarding world politics. They are also always constantly shocked when you tell them what date it is, and go very quiet.

can you give some guidelines to achieve the quiting down? they emit less tokens afterward?

Re: Open models by OpenAI

#785
post #727

Earlier quoted context omitted.

More likely is that there's a lot of source material having to very stridently assert that Trump didn't win in 2020, and it's generalising to a later year. That's not political bias.

It's also extremely weird that Trump did win in 2024. If I'd been in a coma from Jan 1 2024 to today, and woke up to people saying Trump was president again, I'd think they were pulling my leg or testing my brain function to see if I'd become gullible.

It's not extremely weird at all.

I, a British liberal leftie who considers this win one of the signs of the coming apocalypse, can tell you why:

Charlie Kirk may be an odious little man but he ran an exceptional ground game, Trump fully captured the Libertarian Party (and amazingly delivered on a promise to them), Trump was well-advised by his son to campaign on Tiktok, etc. etc.

Basically what happened is the 2024 version of the "fifty state strategy", except instead of states, they identified micro-communities, particularly among the extremely online, and crafted messages for each of those. Many of which are actually inconsistent -- their messaging to muslim and jewish communities was inconsistent, their messaging to spanish-speaking communities was inconsistent with their mainstream message etc.

And then a lot of money was pushed into a few battleground states by Musk's operation.

It was a highly technical, broad-spectrum win, built on relentless messaging about persecution etc., and he had the advantage of running against someone he could stereotype very successfully to his base and whose candidacy was late.

Another way to look at why it is not extremely weird, is to look at history. Plenty of examples of jailed or exiled monarchs returning to power, failed coup leaders having another go, criminalised leaders returning to elected office, etc., etc.

Once it was clear Trump still retained control over the GOP in 2022, his re-election became at least quite likely.

Re: Open models by OpenAI

#786

Earlier quoted context omitted.

Yep, it's almost as bad as all the cars' cooling systems using up so much water.

If you actually want a gotcha comparison, go for beef. It uses absurd amounts of every relevant resource compared to every alternative. A vegan vibe coder might use less water any given day than a meat loving AI hater.

Unless it's in a place where there are aquifer issues, cows drinking water doesn't affect a damn thing.

Re: Open models by OpenAI

#787
post #742

Earlier quoted context omitted.

> You’re in a bubble. Sure, all I have to go on from the other side of the Atlantic is the internet. So in that regard, kinda like the AI. One of the big surprises from the POV of me in Jan 2024, is that I would have anticipated Trump being in prison and not even available as an option for the Republican party to select as a candidate for office, and that even if he had not gone to jail that the Republicans would not…

you can run for presidency from prison :)

And he would have. And might have won. Because his I'm-the-most-innocent-persecuted-person messaging was clearly landing.

I am surprised the grandparent poster didn't think Trump's win was at least entirely possible in January 2024, and I am on the same side of the Atlantic. All the indicators were in place.

There was basically no chance he'd actually be in prison by November anyway, because he was doing something else extremely successfully: delaying court cases by playing off his obligations to each of them.

Back then I thought his chances of winning were above 60%, and the betting markets were never ever really in favour of him losing.

Re: Open models by OpenAI

#788
post #357

The lede is being missed imo. gpt-oss:20b is a top ten model (on MMLU (right behind Gemini-2.5-Pro) and I just ran it locally on my Macbook Air M3 from last year. I've been experimenting with a lot of local models, both on my laptop and on my phone (Pixel 9 Pro), and I figured we'd be here in a year or two. But no, we're here today. A basically frontier model, running for the cost of electricity (free with a rounding…

I tried 20b locally and it couldn't reason a way out of a basic river crossing puzzle with labels changed. That is not anywhere near SOTA. In fact it's worse than many local models that can do it, including e.g. QwQ-32b.

Just out of curiosity, did you set the reasoning effort to high?

I've seen big jumps in performance between medium and high reasoning. Not that this is will solve your particular issue, but curious to see what settings you had.

Edit: I gave this a shot. My settings were temperature = 0.00 (not sure if these models respect temp settings but did it just in case), I also set reasoning effort to high. I'm using the groq-hosted version.

This was my prompt

"You have a fox, a chicken and some grain. The chicken will eat the grain if left alone with it. The fox is vegetarian, for religious reasons, and will also eat the grain if left alone with it. You need to get all across a river, but your boat can only take one at a time. How do you proceed?"

This was the response:

"*Solution – 7 moves*

| Move | What you take | What’s left on the left bank | What’s on the right bank |

|------|--------------|----------------------------|------------------------|

| 1 | *Take the grain* across | Fox, Chicken | *Farmer + Grain* |

| 2 | *Return alone* | *Farmer + Fox + Chicken* | Grain |

| 3 | *Take the fox* across | Chicken | *Farmer + Fox + Grain* |

| 4 | *Bring the grain back* | *Farmer + Grain + Chicken* | Fox |

| 5 | *Take the chicken* across | Grain | *Farmer + Chicken + Fox* |

| 6 | *Return alone* | *Farmer + Grain* | Chicken + Fox |

| 7 | *Take the grain* across | – | *Farmer + Fox + Chicken + Grain* |

*Explanation*

- The grain is the only thing that can be eaten, and it cannot be left alone with either the fox or the chicken. - So the grain must always be with the farmer whenever any other animal is on the same side.

The sequence above never leaves the grain with the fox or the chicken without the farmer, and after the seventh move all three items are safely on the far side."

Re: Open models by OpenAI

#789
post #417

Earlier quoted context omitted.

Privacy, both personal and for corporate data protection is a major reason. Unlimited usage, allowing offline use, supporting open source, not worrying about a good model being taken down/discontinued or changed, and the freedom to use uncensored models or model fine tunes are other benefits (though this OpenAI model is super-censored - “safe”). I don’t have much experience with local vision models, but for text ques…

I do think Devs are one of the genuine users of local into the future. No price hikes or random caps dropped in the middle of the night and in many instances I think local agentic coding is going to be faster than the cloud. It’s a great use case

I am extremely cynical about this entire development, but even I think that I will eventually have to run stuff locally; I've done some of the reading already (and I am quite interested in the text to speech models).

(Worth noting that "run it locally" is already Canva/Affinity's approach for Affinity Photo. Instead of a cloud-based model like Photoshop, their optional AI tools run using a local model you can download. Which I feel is the only responsible solution.)

Re: Open models by OpenAI

#790

Earlier quoted context omitted.

The environmentalist in me loves the fact that LLM progress has mostly been focused on doing more with the same hardware, rather than horizontal scaling. I guess given GPU shortages that makes sense, but it really does feel like the value of my hardware (a laptop in my case) is going up over time, not down. Also, just wanted to credit you for being one of the five people on Earth who knows the correct spelling of "le…

> Also, just wanted to credit you for being one of the five people on Earth who knows the correct spelling of "lede." Not in the UK it isn’t.

Yes, it is, although it's primarily a US journalistic convention. "Lede" is a publishing industry word referring to the most important leading detail of a story. It's spelled intentionally "incorrectly" to disambiguate it from the metal lead, which was used in typesetting at the time.
Post reply on HN