Live data from Hacker News

Gemini 3

blog.google

961–970 of 1001 posts

Re: Gemini 3

#961

Earlier quoted context omitted.

Something else to consider. I often have much better success with something like: Create a prompt that creates a specification for a pacman game in a single html page. Consider edge cases and key implementation details that result in bugs. , execute prompt. It will often yield a much better result than one generic prompt. Now that models are trained on how to generate prompts for themselves this is quite productive.…

I thought this kind of chaining was already part of these systems.

It can be, but the more specific context you can give the better, especially on your initial prompting. If it is opaque to you who knows what it is doing. Dialing in the initial spec/prompt for 5 minutes is still important. Different LLMs and models will do better or worse on this and by being a human in the loop on this initial stuff my experience is much higher quality, which indicates to me, the LLM tries, but just doesn't always have enough info to implement your intentions in many cases yet.

Re: Gemini 3

#962
post #801
post #698

Earlier quoted context omitted.

Those numbers are too good to expect. If 90% right 10% wrong is the baseline would you take as an improvement: - 80% right 18% I don't know 2% wrong - 50%/48%/2% - 10%/90%/0% - 80%/15%/5% The general point being that to reduce wrong answers you will need to accept some reduction in right answers if you want the change to only be made through trade-offs. Otherwise you just say "I'd like a better system" and that is ra…

I think you may have misread. They stated that they'd be willing to go from 90% correct to 10% correct for this tradeoff.

Thanks for the correction

Re: Gemini 3

#963
post #627

Earlier quoted context omitted.

I was under the impression for this to work like that, training data needs to be plenty. One project is not enough since it’s too "sparse". But maybe this example was used by many other people and so it proliferated?

The repo[0] currently has been forked ~41300 times. [0] https://github.com/wesbos/JavaScript30

It’s quite unlikely that training data will include duplicate repositories or even forks, that alone would surpass the published dataset sizes.

Re: Gemini 3

#964
post #501

Earlier quoted context omitted.

To this day, I still don't understand why Claude gets more acclaim for coding. Gemini 2.5 consistently outperformed Claude and ChatGPT mostly because of the much larger context.

Claude doesn’t gaslight me, or flat out refuses to do something I ask it to because it believes it won’t work anyway. Gemini does Gemini also randomly just reverts everything because of some small mistake it found, makes assumptions without checking if those are true (eg this lib absolutely HAS TO HAVE a login() method. If we get a compile error it’s my env setup fault) It’s just not a pleasant model to work with

confirmed, but also happens occasionally with Claude

Re: Gemini 3

#965

Grok got to hold the top spot of LMArena-text for all of ~24 hours, good for them [1]. With stylecontrol enabled, that is. Without stylecontrol, gemini held the fort. [1] https://lmarena.ai/leaderboard/text

Grok is heavily censored though

[dead]

Re: Gemini 3

#966
post #923

Hassabis interview on Gemini 3, with Hard Fork (nyt podcast), also Josh Woodward https://youtu.be/rq-2i1blAlU?t=428 Some points - Good at vibe coding 10:30 - step change where it's actually useful AGI still 5-10 years. Needs reasoning, memory, world models. Is it a bubble? - Partly 22:00 What's fun to do with Gemini to show the relatives? Suggested taking a selfie with the app and having it edit. 24:00 (I tried and s…

> Needs reasoning, memory, world models.

Is that all? So they just need to invent:

1. Thought

2. A mechanism for efficiently encoding and decoding arbitrary percepts

3. A formal model of the world

And then the existing large language models can handle the rest.

Yep, 5 years and a hundred billion dollars or so should do the trick.

Re: Gemini 3

#967
post #435

Earlier quoted context omitted.

The requested prompt does not exist or you do not have access. If you believe the request is correct, make sure you have first allowed AI Studio access to your Google Drive, and then ask the owner to share the prompt with you.

I thought this was a joke at first. It actually needs drive access to run someone else's prompt. Wild.

To clarify, the message above is what I got after giving it Google Drive access.

Re: Gemini 3

#968
post #832

Earlier quoted context omitted.

You can criticize the code but "wow looks like total shit" is such an embarrassing thing to say considering the context. Imagine going back a few years and show them a tool outputting this from text. No-one would believe it.

It simply is non impressive at all to me, we had an industry(games not web) that was the most innovativd and was able to do things, and in part still is, thousands of years ahead of the slop glorified here

Yeah, yeah. It will hit an imaginary wall any moment now.

Re: Gemini 3

#969

Earlier quoted context omitted.

Not sure what you mean here, but the only real jobs at risk from AI right now are middle/upper management. Not a single engineer has ever been laid off because of AI. Any company claiming this is the case is trying to cover up bad decisions. "Were automating with AI" sounds better to investors than "We over hired and now need to downsize" or "We made some bad market bets, now need to free up cash flow"

> Not sure what you mean here, but the only real jobs at risk from AI right now are middle/upper management. > Not a single engineer has ever been laid off because of AI. Any company claiming this is the case is trying to cover up bad decisions. I don't suppose these assertions are based on anything. If "AI" reduces the amount of time an engineer spends writing crud, boilerplate, test cases, random scripts, etc., and…

There's a problem with the idea that hiring works efficiently. Twitter ran with thousands of engineers for a long time and clearly it did not need to.

Re: Gemini 3

#970
Why doesn't this spell the death of OpenAI? Maybe someone with a better business sense can explain, but here's what I'm seeing:

OpenAI is going for the consumer-grade AI market, as opposed to a company like Anthropic making a specialized developer tool. Google can inject their AI tool in front of everybody in the world, and already have with Google AI search. All of these models are just going to reach parity eventually, but Google is burning cash compared to OpenAI burning debt. It seems like for consumer-grade purposes, AI use will just be free sooner or later (DeepSeek is free, Google AI search is free, students can get Gemini Pro for free for a year already). So all I'm seeing that OpenAI has is Sora, which seems like a business loser though I don't really understand it, and also ChatGPT seems to own the market of people roleplaying with chat bots as companions (which doesn't really seem like a multi-trillion dollar business but I could be wrong).

Post reply on HN