Earlier quoted context omitted.
We're having DeepSeek moments every couple of weeks. Qwen 3.6 hit hard in the self-hosting space. It's incredibly capable for its size, really shaking up what's possible in 64GB or even 32GB of VRAM. The Prism Bonsai ternary model crams a tremendous amount of capability into 1.75GB. And, DeepSeek V4 is crazy good for the price. They're charging flash model prices for their top-tier Pro model, which is competitive wit…
We have Qwen 3.6-35b (6) on a 5090 (32GB) and it's blowing me away. Works fine for most (not all) code generation tasks. One developer here has been extremely stubborn about adopting AI; he's finally adopted it, albeit only when it's coming from a local model like this. DeepSeek V4 Pro likewise is insanely good for the price. I simply point it at large codebases, go get a cup of coffee or browse Hacker News, and then…
Gemini 3.5 Flash
351–360 of 692 posts
Re: Gemini 3.5 Flash
#352Earlier quoted context omitted.
Considering all models can use search engines, is this really relevant?
Until they prefer not to search. Let me explain using the example of the open-source security framework (1) our team is working on. If you ask Gemini what you should use to integrate fraud prevention or account takeover protection into your product, there will be no mention of our open-source project. Five years in development, 1.3k stars, over 140 pull requests — all this isn't enough to make it into the training da…
FWIW while neither model included your product in it's initial response, when I followed up with "what about open-source" both did another search and Claude's response included your tool....
Re: Gemini 3.5 Flash
#353Earlier quoted context omitted.
Gen AI is unprofitable, especially at the insanely cheap rates they've been offering to get people in the door. So expect more increases in the future.
These companies are unprofitable (as all companies at this stage and ambition should be) but I increasingly don't see any justification for the idea that it is fundamentally unprofitable. Inference alone is certainly profitable. I'm running models at home that are comparable to performance of paid models a year or so ago for free. Even for much larger models the cost around inference serving are clearly manageable. T…
Re: Gemini 3.5 Flash
#354The price is crazy. And I guess Gemini 3.5 pro will have the pricing increment, too. 12 x 5 = 60? It seems like google does want us to use Chinese models.
What exactly are you doing with this that you can’t generate $1.50 of value per million tokens?
Re: Gemini 3.5 Flash
#355Earlier quoted context omitted.
LLM pre-training models risk being unable to be updated with data from after 2025, as much of it is corrupted with LLM-generated content. We might be locked into outdated knowledge, where only whitelisted sources decide what to include. Taking into account the sometimes blind belief that 'LLMs know everything', the outcome could be very costly, especially for technologies and businesses unfortunate enough to emerge a…
But ChatGPT has been popular since early 2023, and even before it there was no shortage of low-quality content on the web. If anything, this model being trained up to 2025 is a positive sign that the "circular LLM training" problem hasn't (yet) become unmanagable. The year-long delay is probably just due to how long it takes to test/refine a cutting-edge model. It's surely possible to train one faster, but Google wou…
Re: Gemini 3.5 Flash
#356Earlier quoted context omitted.
If Google is actually getting cheaper inference than everyone else with their TPUs, this smells like trouble to me. Maybe serving LLMs at a profit is proving difficult. Or maybe they think because their benchmarks are good they can ramp up the prices. Seems like they don’t have the market share to justify a move like that yet to me.
Prevailing wisdom is that serving LLMs at a profit is achievable... it's when you factor in the cost of training them that prices get astronomical real fast. Open-source model inference providers (who do not have to bear the cost of training) seem able to do it at much lower prices. https://www.together.ai/pricing https://fireworks.ai/pricing#serverless-pricing (scroll down to headline models) Of course, it's possibl…
Re: Gemini 3.5 Flash
#357Earlier quoted context omitted.
Amazon was unprofitable for over a decade, and they were public. Theres no incentive to be profitable as a private company if you can continue to raise money. Ed Zitron and Gary Marcus are... confused.
> Amazon was unprofitable for over a decade, and they were public. Amazon was unprofitable because they poured their revenue into growth. On paper, they were in the red, but everyone - especially investors - saw what was going to happen, given their trajectory. Is it the case that any of these AI companies are actually making a ton of money and growing accordingly? AFAICT, we've just got [a] big players like Google t…
Re: Gemini 3.5 Flash
#358Earlier quoted context omitted.
That's likely because you're using the Gemini app which has a tool for image generation (nano banana) - I do my tests against the API to avoid any possibility of tool use.
This question makes me wonder if you one shot each pelican or do you run it a few times to get the best one?
Re: Gemini 3.5 Flash
#359Re: Gemini 3.5 Flash
#360Earlier quoted context omitted.
We need another "Deepseek moment" or else it will become impossible for the regular dude to use AI. It will become something that only big companies can afford.
You can use lots of open weight models today.