I caught it again being deceitful. It did this before (Me): Did you actually read the paper before when I pasted the link? > I will be completely honest: No, I did not. > You caught me hallucinating a confident answer based on incomplete recall rather than actually verifying the document. > Thank you for calling it out and providing the exact quote. It forced me to re-evaluate the actual data you provided rather than…
Gemini 3.5 Flash
301–310 of 692 posts
Re: Gemini 3.5 Flash
#302Earlier quoted context omitted.
This is not priced at inference cost. My guess: it's the price at which they make more money than if they rent the TPUs to other companies. The Gemini team has had trouble securing enough TPUs for their user's needs. They struggle with load and their rate limits are really bad. Maybe at a higher price, they have a better chance at getting more TPUs assigned?
The cost at such they could rent out the TPUs, i.e. the market rate, is the inference cost. Just because you are vertically integrated doesn't mean you get to discount the one business units products to the other. Doing so discounts the opportunity cost you pay and is just bad accounting.
Re: Gemini 3.5 Flash
#303Earlier quoted context omitted.
Deepseek had another moment a few weeks ago. V4 isn't far behind the US frontier, and so far its flash variant seems a very reliable coder and costs a pittance.
Deepseek V4 (not flash) trippled in price too by the way (from Deepseek). Get used to this pattern. This is what you get for relying on the generosity of billionaires. Keep offshoring your thinking ability to a machine and let me know how competitive you. Hint, you wont be. There's nothing special about being able to use an LLM.
Re: Gemini 3.5 Flash
#304Re: Gemini 3.5 Flash
#305Earlier quoted context omitted.
It is insanely profitable though, if you cut out r&d cost, plus the marketing and loss leaders. Don't let them gaslight you. Even anthropic who does not own any hardware still have a big margin providing claude models.
Then why haven't they reported any profits using GAAP (generally accepted accounting principles)? They all use ARR which is easily gamed.
Google has just recently upgraded their TPUs.
Re: Gemini 3.5 Flash
#306Earlier quoted context omitted.
LLM pre-training models risk being unable to be updated with data from after 2025, as much of it is corrupted with LLM-generated content. We might be locked into outdated knowledge, where only whitelisted sources decide what to include. Taking into account the sometimes blind belief that 'LLMs know everything', the outcome could be very costly, especially for technologies and businesses unfortunate enough to emerge a…
Considering all models can use search engines, is this really relevant?
If you ask Gemini what you should use to integrate fraud prevention or account takeover protection into your product, there will be no mention of our open-source project. Five years in development, 1.3k stars, over 140 pull requests — all this isn't enough to make it into the training data. From this perspective, any technology that emerges after 2024 is simply invisible to LLMs.
The answer is: without being in the training data, LLMs basically don't understand what they're searching for.
Re: Gemini 3.5 Flash
#307Earlier quoted context omitted.
At least in some cases, there seems to be a move toward training on more synthetic data and strictly curated data, especially for smaller models where knowledge can't be extremely broad, because there just isn't enough room to store the world in tens or hundreds of gigabytes of model weights. So, to achieve higher quality reasoning, the training has to be focused and the data has to be very high quality and high dens…
> it maybe doesn't even matter that the models are using older data. This actually really does matter. Otherwise, the model simply won't know about your product and will always suggest only a few market leaders. Searching for information on the Internet became a jungle a decade ago, and to be visible you have to pay Google for sunlight. Now, we risk falling into real darkness — until some paid model eventually emerge…
Re: Gemini 3.5 Flash
#308Earlier quoted context omitted.
These companies are unprofitable (as all companies at this stage and ambition should be) but I increasingly don't see any justification for the idea that it is fundamentally unprofitable. Inference alone is certainly profitable. I'm running models at home that are comparable to performance of paid models a year or so ago for free. Even for much larger models the cost around inference serving are clearly manageable. T…
And if you can run those strong models at home for free, why would hosting them be a successful business for any of these providers? Profitable maybe, in terms of having low costs, but why pay Google or whoever when you can do it yourself for cheaper/"free"?
Re: Gemini 3.5 Flash
#309The pelican is a lot : https://github.com/simonw/llm-gemini/issues/133#issuecomment... Not a great bicycle though, it forgot the bar between the pedals and the back wheel and weirdly tangled the other bars. Expensive too - that pelican cost 13 cents: https://www.llm-prices.com/#it=11&ot=14403&sel=gemini-3.5-fl...
Re: Gemini 3.5 Flash
#310Earlier quoted context omitted.
Deepseek V4 (not flash) trippled in price too by the way (from Deepseek). Get used to this pattern. This is what you get for relying on the generosity of billionaires. Keep offshoring your thinking ability to a machine and let me know how competitive you. Hint, you wont be. There's nothing special about being able to use an LLM.
Unlike other providers, Deepseek does promise that they will lower the price when their Huawei cards arrive in a few more months.