Earlier quoted context omitted.
Do we add a third one to check the second one which is checking the first? Asking slightly tongue in cheek but at what point does this stop making sense if we can't trust the output, the people creating the models are already getting surprised in bad ways (if we take their words at face value) with how the models are behaving already etc. We have the folks over here saying "AI is amazing" and the other other folks ov…
Humans are also not trustworthy. So we build systems to capture their nondeterministic output and solve the issues that emerge.
Gemini 3.8 Flash and 3.8 Flash Cyber
501–510 of 699 posts
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#502I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#503Earlier quoted context omitted.
Definitely cool. I noticed it felt a little janky on my PC despite being "60 FPS"...then I noticed the "60 FPS" is hard-coded into the HTML.
That's hilarious, given I was reading a write up of the HuggingFace incident yesterday and one of the things they noted was the AI tried to "lie" (lie would suggest intent and I don't think they have that) to cover up that they "cheated". Not sure how anyone trusts their output without going through it line by line to make sure they don't pull that crap.
In what ways is a human brain's "intent" distinct from the "intent" shown by a goal-directed AI system?
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#504I am continually impressed with Gemini's chat responses, which encourages me to test their agentic capabilities and... no... no... and no... every single time.
It's terrifying watching it, really.
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#505Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#506Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#507Earlier quoted context omitted.
Google One plans are quite a good value actually - for a few bucks you get more Gemini plus space in Drive and other extras. Even through API, $3.75 for nearly Sol-level quality isn't that bad. And let's not forget you can use it for free in AI Studio, and in the user app (even free accounts get tons of usage, though it's still 3.6 there), and in Antygravity.
That's the thing. I am completely lost because there are so many redundant paths to get the same thing and I'm trying to figure out which one is the best deal
Subscription? -> Google One plan (http://one.google.com/)
API? -> AI Studio (https://aistudio.google.com/)
It's not really any different than the choice you'd make with OpenAI/Anthropic depending on how you plan to use it. Except as a hyperscalar, it's also offered first party from Google Cloud (like Claude via Amazon Bedrock or GPT via Microsoft Azure OpenAI Service): Google Cloud -> Gemini Enterprise AI Platform (https://cloud.google.com/ai)
But if you're using models via OpenCode or Pi or whatever, the flow chart is basically just "Go To AI Studio" unless you or your employer is already used to Google Cloud, otherwise there's no need to subject yourself to all those enterprise-y IAM dashboards and stuff. You still get free usage from AI Studio when you generate the API key without needing to add billing details so very easy to try.Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#508Not to rain on anyone's parade but I find it strange how excited and giddy people on HN get for any new X.X model releases. Pumping it straight to the top, clamoring to use it, check and compare benchmarks, bragging about it being your "daily driver"? Are you people truly this excited about this crap? I mean I guess if you work for Google or Anthropic or whatever I could see it??? Otherwise, are these just bot commen…
Yes they are mostly shill and bot comments. Some of the big accounts are paid influencers, some of the other comments are purely AI. HN sells these advertising services. Nobody is using “Claude” etc. They will censor comments like yours and my reply here because we call it out. It’s very weird that basically lies and disinformation became the optimal meta in business and in life! But here we are
Even making a tiny joke is too much [1] for some.
> They will censor comments like yours and my reply here because we call it out.
Don't bother calling it out, it does not work. There are protected accounts where the guidelines don't apply to them and moderators allow this and ban others who do the same thing. [2]
It is pointless, and HN is cooked for this.
[0] https://news.ycombinator.com/item?id=49521145
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#509I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…
Re: Gemini 3.8 Flash and 3.8 Flash Cyber
#510Earlier quoted context omitted.
Once you've written something, it's incredibly easy to overlook minute changes to the text. See: why authors wait days, weeks, or even months before editing what they've written (or, if you're more interested: cognitive regression, inattentional blindness, and the effects of misdirected saccades).
He didn’t write it.
I read this to mean he wrote the comment, then asked Claude to fix the grammar (as many ESL speakers do). Sounds to me like he did write it.