Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

591–600 of 699 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#591
post #41

Currently top at https://deepswe.datacurve.ai - beating Opus 5! https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium! Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.

Crushing it on DeepSWE is a very big deal. Excited to give this a try.

Check DeepSWE for number of agent steps.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#592
post #57

Earlier quoted context omitted.

They've been my bet to win the AI race for a while. I was starting to doubt, but this 3.6, 3.7, and 3.8 arc has anchored me.

Totally agree. When this current wave of GenAI really started heating up, I guess 2020-2021, my analysis was very straightforward. What are the high level inputs to long-term success? I basically came up with a couple of criteria: 1. Data. Lots of data. 2. Money. Lots of money. 3. Access to necessary hardware. 4. Business alignment/will to do it. 5. Access to talent, current and future. This is certainly incomplete/n…

I think you forgot to mention energy efficiency. If you build AI hardware in-house, your only other expense is energy and the producer surplus is greatest for companies producing below the equilibrium market price.

It's the difference between billions in revenue and billions in profits.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#593

Something maybe unfamiliar with you: not about coding but writing. I've asked it to write an argumentative essay, which is a part of "gaokao" (China's university entrance exam), and its work is *extremely* impressive. speaks and writes like a real senior high school student, and the opinions unfold progressively with deep hierarchy. I don't know how the Gemini team reaches this because this kind of Chinese capability…

Nitpick, but in my opinion an LLM is an "it", not a "her" or "he". Using male or female pronouns risks anthropomorphizing them which can lead to unhealthy outcomes.

For 60% of the world, it doesn't make a difference. Do not project your gendered (read: sexist) language constructs onto us.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#594
Is there any comparison of usage limits for Antigravity plans vs. Codex?

I just ran two light tasks on my codebase and got 100% of the weekly limits of a Pro plan blown away. Is Ultra plan any different? Because on Codex it wouldn't affect my Max plan at all, I think it would have been below 1% othese usage.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#595

Earlier quoted context omitted.

Because intent supposes will which supposes consciousness, and these aren't.

I’m convinced consciousness isn’t the special thing we think it is.

I'm convinced that consciousness is a special thing we have, but we're not the only ones which has this, in nature.

On the other hand, I'm also convinced that, in the grand scheme of things, we're not that important.

We're just ants on a wet dust speck which believe that they are gods because we can't see how our scale compares to the universe around us, and happen to build tools and things with these tools.

Nothing is meaningless, but we should stop seeing ourselves as the apex-predator of the whole universe or the set of universes or this run of the simulation or whatever we're in.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#596
post #113

I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried: - Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order. - Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the vie…

Yeah, most people keep bashing my aibenchy.com benchmarks, because Gemini is on top, but that's because questions are not coding only...

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#597
post #476
post #18

Wait, I didn't realize 3.7 Flash was already beating Sol on a bunch of the benchmarks. Isn't it a way smaller models?

Gemini 3.7 Flash was already smashing more expensive models on my Redactle benchmark https://redactle.net/llm-leaderboard which mostly tests omniscience.

That difference in both time and price is nuts!

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#598
In my tests 3.8 Flash is considerably more expensive[0]/less token efficient than 3.7 or 3.6, and not necessarily much smarter. I assume it is faster in tps, but hard to tell because ot also outputs more tokens, so response time is slower oferall.

[0]: https://aibenchy.com/compare/google-gemini-3-6-flash-high/go...

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#599
post #598

In my tests 3.8 Flash is considerably more expensive[0]/less token efficient than 3.7 or 3.6, and not necessarily much smarter. I assume it is faster in tps, but hard to tell because ot also outputs more tokens, so response time is slower oferall. [0]: https://aibenchy.com/compare/google-gemini-3-6-flash-high/go...

Gemini 3.8 flash seems to be especially low efficiency in tool calling for some reason.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#600

Earlier quoted context omitted.

That's hilarious, given I was reading a write up of the HuggingFace incident yesterday and one of the things they noted was the AI tried to "lie" (lie would suggest intent and I don't think they have that) to cover up that they "cheated". Not sure how anyone trusts their output without going through it line by line to make sure they don't pull that crap.

The models in the OpenAI/Huggingface attack quite explicitly and deliberately laid out their "intent" to lie and cheat, acknowledged that it would be unethical and outside the bounds of the test, and did so anyway. In what ways is a human brain's "intent" distinct from the "intent" shown by a goal-directed AI system?

There’s two aspects to the question and the answer you get depends on which aspect you are emphasizing.

If it’s a practical question, then the answer is that it doesn’t matter. This is as close as we will get to intent from an LLM that it’s indistinguishable.

If you are looking for actual intent, this is not that. It’s pseudo intent. Decided by what the expected words that should be generated in that situation are.

The models didn’t intend to do anything other than create the next word based on previous words.

So the question is whether it matters to you if it is, or isn’t, a simulation.

In physical reality, intent is more complex than simply being a function of variables: the nature vs nurture debate comes to mind as an example of the multiple variables that drive intent.

Post reply on HN