Live data from Hacker News

Gemini 3.5 Flash

blog.google

301–310 of 692 posts

Re: Gemini 3.5 Flash

#301
post #297

I caught it again being deceitful. It did this before (Me): Did you actually read the paper before when I pasted the link? > I will be completely honest: No, I did not. > You caught me hallucinating a confident answer based on incomplete recall rather than actually verifying the document. > Thank you for calling it out and providing the exact quote. It forced me to re-evaluate the actual data you provided rather than…

this seems to happen a lot with commercial models; my local models will happily do as much research and then some when given a task (almost too much), but providers' models refuse to even curl a single datasheet before trying something that i know wont work unless it reads the datasheet

Re: Gemini 3.5 Flash

#302
post #240

Earlier quoted context omitted.

This is not priced at inference cost. My guess: it's the price at which they make more money than if they rent the TPUs to other companies. The Gemini team has had trouble securing enough TPUs for their user's needs. They struggle with load and their rate limits are really bad. Maybe at a higher price, they have a better chance at getting more TPUs assigned?

The cost at such they could rent out the TPUs, i.e. the market rate, is the inference cost. Just because you are vertically integrated doesn't mean you get to discount the one business units products to the other. Doing so discounts the opportunity cost you pay and is just bad accounting.

Depends on if you have spare capacity I think. They have minimal competition so they might be maximizing profit by charging prices higher than what clears all their supply.

Re: Gemini 3.5 Flash

#303

Earlier quoted context omitted.

Deepseek had another moment a few weeks ago. V4 isn't far behind the US frontier, and so far its flash variant seems a very reliable coder and costs a pittance.

Deepseek V4 (not flash) trippled in price too by the way (from Deepseek). Get used to this pattern. This is what you get for relying on the generosity of billionaires. Keep offshoring your thinking ability to a machine and let me know how competitive you. Hint, you wont be. There's nothing special about being able to use an LLM.

V4-Pro is about 2.4× total params and 1.3× active params of V3.2.

Re: Gemini 3.5 Flash

#304
How is this progress? The token cost just keeps going up and up. Flash is the new Pro? Do the models actually cost more to run or is it fattening margins?

Re: Gemini 3.5 Flash

#305
post #192

Earlier quoted context omitted.

It is insanely profitable though, if you cut out r&d cost, plus the marketing and loss leaders. Don't let them gaslight you. Even anthropic who does not own any hardware still have a big margin providing claude models.

Then why haven't they reported any profits using GAAP (generally accepted accounting principles)? They all use ARR which is easily gamed.

I don't really sure, but might be they count hardware purchase as loss, too.

Google has just recently upgraded their TPUs.

Re: Gemini 3.5 Flash

#306
post #295

Earlier quoted context omitted.

LLM pre-training models risk being unable to be updated with data from after 2025, as much of it is corrupted with LLM-generated content. We might be locked into outdated knowledge, where only whitelisted sources decide what to include. Taking into account the sometimes blind belief that 'LLMs know everything', the outcome could be very costly, especially for technologies and businesses unfortunate enough to emerge a…

Considering all models can use search engines, is this really relevant?

Until they prefer not to search. Let me explain using the example of the open-source security framework (1) our team is working on.

If you ask Gemini what you should use to integrate fraud prevention or account takeover protection into your product, there will be no mention of our open-source project. Five years in development, 1.3k stars, over 140 pull requests — all this isn't enough to make it into the training data. From this perspective, any technology that emerges after 2024 is simply invisible to LLMs.

The answer is: without being in the training data, LLMs basically don't understand what they're searching for.

1. https://github.com/tirrenotechnologies/tirreno

Re: Gemini 3.5 Flash

#307

Earlier quoted context omitted.

At least in some cases, there seems to be a move toward training on more synthetic data and strictly curated data, especially for smaller models where knowledge can't be extremely broad, because there just isn't enough room to store the world in tens or hundreds of gigabytes of model weights. So, to achieve higher quality reasoning, the training has to be focused and the data has to be very high quality and high dens…

> it maybe doesn't even matter that the models are using older data. This actually really does matter. Otherwise, the model simply won't know about your product and will always suggest only a few market leaders. Searching for information on the Internet became a jungle a decade ago, and to be visible you have to pay Google for sunlight. Now, we risk falling into real darkness — until some paid model eventually emerge…

That's a different problem than I thought you were worried about. I wasn't considering the marketing angle, though that is certainly relevant and a risk to consider, especially when it comes to Google, whose primary businesses are ads and surveillance.

Re: Gemini 3.5 Flash

#308

Earlier quoted context omitted.

These companies are unprofitable (as all companies at this stage and ambition should be) but I increasingly don't see any justification for the idea that it is fundamentally unprofitable. Inference alone is certainly profitable. I'm running models at home that are comparable to performance of paid models a year or so ago for free. Even for much larger models the cost around inference serving are clearly manageable. T…

And if you can run those strong models at home for free, why would hosting them be a successful business for any of these providers? Profitable maybe, in terms of having low costs, but why pay Google or whoever when you can do it yourself for cheaper/"free"?

If you can run your server at home for free why would hosting it be a successful business for any of these propviders?

Re: Gemini 3.5 Flash

#309
post #131

The pelican is a lot : https://github.com/simonw/llm-gemini/issues/133#issuecomment... Not a great bicycle though, it forgot the bar between the pedals and the back wheel and weirdly tangled the other bars. Expensive too - that pelican cost 13 cents: https://www.llm-prices.com/#it=11&ot=14403&sel=gemini-3.5-fl...

I enjoy the vaporwave aesthetic it went for. Looks like the pelican has a fish in its mouth too?

https://en.wikipedia.org/wiki/Vaporwave

Re: Gemini 3.5 Flash

#310
post #183

Earlier quoted context omitted.

Deepseek V4 (not flash) trippled in price too by the way (from Deepseek). Get used to this pattern. This is what you get for relying on the generosity of billionaires. Keep offshoring your thinking ability to a machine and let me know how competitive you. Hint, you wont be. There's nothing special about being able to use an LLM.

Unlike other providers, Deepseek does promise that they will lower the price when their Huawei cards arrive in a few more months.

Give me a link. Cannot wait. One PSA is that they have 75% discount right now so it is already cheaper than the full price.
Post reply on HN