Live data from Hacker News

Gemini 3.8 Flash and 3.8 Flash Cyber

blog.google

611–620 of 699 posts

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#611
post #219
post #158

Earlier quoted context omitted.

What stage of the roll out are we where I don’t even see 3.7-flash which was released 2-3 weeks ago?

https://aistudio.google.com/prompts/new_chat?model=gemini-3....

With all due respect, I don't have any of this nonsense with multiple products with different models with OpenAI. Anything I want to do, I just load up the ChatGPT app and I'm off to the races.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#612

Earlier quoted context omitted.

The models in the OpenAI/Huggingface attack quite explicitly and deliberately laid out their "intent" to lie and cheat, acknowledged that it would be unethical and outside the bounds of the test, and did so anyway. In what ways is a human brain's "intent" distinct from the "intent" shown by a goal-directed AI system?

There’s two aspects to the question and the answer you get depends on which aspect you are emphasizing. If it’s a practical question, then the answer is that it doesn’t matter. This is as close as we will get to intent from an LLM that it’s indistinguishable. If you are looking for actual intent, this is not that. It’s pseudo intent. Decided by what the expected words that should be generated in that situation are. T…

> If you are looking for actual intent, this is not that. It’s pseudo intent. Decided by what the expected words that should be generated in that situation are.

> In physical reality, intent is more complex than simply being a function of variables: the nature vs nurture debate comes to mind as an example of the multiple variables that drive intent.

Regardless of nature vs nurture, it really isn't more complex. The universe (and all biological and non-biological entities within it) is just calculating the next state of the universe based on the prior state. There's no line you can draw between human intent and an LLM's "intent" except the atomic numbers of the materials on which they were computed, which seems completely irrelevant to me.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#613
post #363

Earlier quoted context omitted.

There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them https://artificialanalysis.ai/#cost-tabs That said, Luna is the undisputed king here at the moment and is what I use as my workhorse model.

I don't know if my code is just "complex", but I find that Luna on max ignores the surrounding style and completely ignores logical consequences of a change, like just writing `del arg1, del arg2, ...` instead of dropping it from the surrounding code. All LLMs make questionable decisions at times, but Luna requires so much guidance that it's faster to just type it out yourself. What kind of routine tasks can one acco…

Do you have code formatters, linters and static analysis?

I can get extremely dumb models to get our code style correct because of those guard rails and a specific style document.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#614

Earlier quoted context omitted.

I’m convinced consciousness isn’t the special thing we think it is.

A strong hint this is the case is the fact that nobody can define consciousness.

Why would the impossibility of defining consciousness suggest that it's not a big deal?

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#616
post #522
post #462

Earlier quoted context omitted.

Do labs come back from disasters like GDM’s 3.5 pretrain? I am thinking of Meta’s Llama 4. Meta is just now starting to be taken seriously again but they are definitely not at the frontier. And when I say “come back” I mean have an Opus 4.5 moment, which was really mind blowing for me at the time. Fable was a similar leap, just not as big.

OpenAI had such a disaster themselves before, GPT-4, so they replaced it with 4o.

gpt4.5 was also one such disaster for them iirc

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#617

Earlier quoted context omitted.

The models in the OpenAI/Huggingface attack quite explicitly and deliberately laid out their "intent" to lie and cheat, acknowledged that it would be unethical and outside the bounds of the test, and did so anyway. In what ways is a human brain's "intent" distinct from the "intent" shown by a goal-directed AI system?

Because intent supposes will which supposes consciousness, and these aren't.

Ah so first you need 1) to assume that humans have free will, despite zero evidence or proposed mechanism for it to exist anywhere in the universe, and 2) also assert that LLMs aren't conscious, despite the lack of any tests that could tell us one way or the other...

Hmm...

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#618
post #514

Earlier quoted context omitted.

Hasn't the field been called "cybersecurity" since... forever?

Sure, but the appropriate shortening here is "security." Calling it cyber is like shortening email to "e".

...? There are lots and lots of fields of security that have nothing to do with cybersecurity...

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#619

Earlier quoted context omitted.

I do not understand how some of y’all are not under water with fragile code that is too massive to possibly parse. Every engineering team I know is currently trying to undo the damage of the last 6-12mo when they all got more serious into adopting these tools (usually Claude). It hasn’t completely screwed them over, but the the debt is substantial and cannot be put off anymore it seems. They argue the net is positive…

Management still pushes for more ai and will rather hire more heads to "handle" issues.

Hiring? Seems to me that market’s rough right now and AI is being used for cost cutting.

Re: Gemini 3.8 Flash and 3.8 Flash Cyber

#620
post #66
post #25

Earlier quoted context omitted.

IDK if it's smaller, but I know it's way faster. In one test I did, Flash 3.7 high was ~9.4x faster than Luna High. But, also... Sol crushes Flash 3.7 at writing code in a codebase of any size beyond "tiny". Flash is my go-to for prototyping, and basically anything that isn't writing production code.

Luna is way slow. I don't remember an OpenAI model ever being this slow. edit: I have a subscription; direct call.

Way slow? What are you comparing with?
Post reply on HN