Earlier quoted context omitted.
This is what we do at gertlabs.com - the foundation labs are actually starving for better data. Having quality data is not the same as having a lot of data. Human curated data / RLHF cannot scale to a 5T model and synthetic data pipelines are very much a work in progress in the industry. Some interesting notes: - Training a small model with large model output resulted in LESS improvement than distilling a less smart…
Wouldn’t it be good to start investigating into a micro model architecture? Like first model checks the context and routes to the Java optimized model, etc. would make it also simpler to load/unload models in memory. So extremely small models that are only good for a certain task like programming languages. A little bit of a model at the front that is extremely good in classification of tasks and than a more complex…
Gemini 3.5 Flash
591–600 of 692 posts
Re: Gemini 3.5 Flash
#592Re: Gemini 3.5 Flash
#593Earlier quoted context omitted.
To me this is almost like a tone-deaf naming change. Empty Slot (new Pro as Mythos competitor?) Old Pro -> now Flash Old Flash -> now Flash Lite Old Flash Lite -> now Gemma (and not served by Google) I say "almost" because the situation is more fluid and unstable than a normal naming change. If Apple were to do this with laptops, maybe it'd be like, Air gets better and pricier and becomes Pro-level model, Neo same wa…
> Old Flash Lite -> now Gemma (and not served by Google) > which is now Gemma territory, and I can't get that served by Google anymore Gemma is served by Google. They're serving Gemma 4 26B A4B at $0.15/$0.60. https://console.cloud.google.com/agent-platform/publishers/g... https://cloud.google.com/gemini-enterprise-agent-platform/ge...
Re: Gemini 3.5 Flash
#594Earlier quoted context omitted.
This is what we do at gertlabs.com - the foundation labs are actually starving for better data. Having quality data is not the same as having a lot of data. Human curated data / RLHF cannot scale to a 5T model and synthetic data pipelines are very much a work in progress in the industry. Some interesting notes: - Training a small model with large model output resulted in LESS improvement than distilling a less smart…
Wouldn’t it be good to start investigating into a micro model architecture? Like first model checks the context and routes to the Java optimized model, etc. would make it also simpler to load/unload models in memory. So extremely small models that are only good for a certain task like programming languages. A little bit of a model at the front that is extremely good in classification of tasks and than a more complex…
Obviously, I have no idea but I guess it’s not as simple as “just train only on Java code and reduce size to 1/10th”.
Re: Gemini 3.5 Flash
#595Click on "Listen to article", make sure the voice is "Umbriel" and skip to 4:15 - there's a hallucinated part at the end in Russian (I think). On a blog post about the latest and greatest AI model. Oh the irony.
Re: Gemini 3.5 Flash
#596Earlier quoted context omitted.
Actually, deepseek v4 was 1/3 promotional price for the first month or so. This was pretty clearly communicated. The promotions window just ended is all.
thus proving ops point
There’s a pretty significant difference between saying someone tripled their prices, and a temporary promotion ended. It’s even more so the case if someone is using it as an example for raising prices as a trend.
I’m 100% in the camp that prices are going up and quality is going down; companies are retiring models and requiring you to use more expensive ones. This has happened to me and there are dozens of examples that one can point to.
But a promotion ending is a strawman argument and does the point a disservice.
Re: Gemini 3.5 Flash
#597Earlier quoted context omitted.
This combined with locally runnable models getting pretty good recently (e.g. Qwen 3.6) tells me that it's time to seriously consider local dev setup again
This should become the new Apple's hardware and software play. I am hopeful about the new CEO
Re: Gemini 3.5 Flash
#598Re: Gemini 3.5 Flash
#599Earlier quoted context omitted.
thus proving ops point
If you run out of 50% coupons to your local pizza joint, did they double their prices? Does every company double or triple their prices after Black Friday? There’s a pretty significant difference between saying someone tripled their prices, and a temporary promotion ended. It’s even more so the case if someone is using it as an example for raising prices as a trend. I’m 100% in the camp that prices are going up and q…
Re: Gemini 3.5 Flash
#600Oh and double the cost is assuming you're not using Google cloud for anything else, because data transfer, storage, anything but compute is 10x the going rate outside of GCP at least.
Plus you can run both Kimi K2.6 and MiMo V2.5 locally at marginal cost (ie. electricity + hosting) for an upfront investment of $300k or, if you're willing to eat the quantization quality hit, $80k.