Live data from Hacker News

Grok 4.5

x.ai

191–200 of 1001 posts

Re: Grok 4.5

#191
post #75

With each release from the the other major labs, it becomes harder for Google to tell a compelling story about Gemini 3.5. Edit: Gemini 3.5 Pro . Expectations grow with each day it is not released.

for what it's worth, it's fairly popular among my non-technical coworkers here in Russia. we have unlimited access to all models so it's not about the cost, and they still prefer Gemini over Claude and GPT. I never bothered to ask why, but I assume it's better at communicating in Russian.

This from the country whose entire IT population is still to this day entirely enamored with windows.

Not sure it's a valid data point.

Re: Grok 4.5

#192

Can someone breakdown to me how this makes any sort of economical sense? Spending billions and billions to have the 3rd best model while even the number 1 and 2 players already seem to struggle making a profit. What am I missing here? Not trying to go full Ed Zitron but this doesn’t make sense to me.

With that frame of mind, nothing would be done. Why make another search service if Altavista and Lycos already do it?

Re: Grok 4.5

#193
post #38

Earlier quoted context omitted.

Let’s face it, there is no best model for something because the input is natural language. Some models may fit better some users‘ way of prompting.

Yeah, I think this seems more true than "X is better at iOS than Y", the way you prompt seems a lot more important, and some models react differently to the same prompts.

It's almost like there is no replacement for human expertise when we need to make usable products for other humans.

Re: Grok 4.5

#194
post #68

Is there a reason the AI companies usually announce new products so close to each other. Like not just the same day but literally hours apart. GPT Live then an hour later Grok 4.5. As if they try to one up. I expect something new from Anhtropic as well today.

Maybe it‘s the Nash equilibrium from a timing perspective? Like the reason that close to a McDonals there is usually a Burger King.

reminds me of SBC's (Seattle's Best Coffee) strategy, which was decidedly not Nash: put a store across the street from every Starbucks.

Re: Grok 4.5

#195
post #179

Earlier quoted context omitted.

Ollama is just a local app wrapper/cloud service serving third party apis and models idk why it made it into this list tbh

Claude isn’t a model either.

We can assume (outside of Ollama) that they meant the strongest model from each lab. If you limit yourself to just looking at the literal strings in the list, literally none of these are models. What model is "Deepseek" or "GPT"?

Re: Grok 4.5

#196

Can someone breakdown to me how this makes any sort of economical sense? Spending billions and billions to have the 3rd best model while even the number 1 and 2 players already seem to struggle making a profit. What am I missing here? Not trying to go full Ed Zitron but this doesn’t make sense to me.

You could be typing the same about Google or a number of the other labs right now. A diverse market full of choices keeps it from becoming the browser wars all over again.

> A diverse market full of choices keeps it from becoming the browser wars all over again.

This is a great analogy but I worry you might be implying something I don't agree with but you didn't explicitly say what I'm worried about, so let me call it out:

Microsoft played a dirty game with I.E, but they are in the dirty game business. It wasn't only I.E, it was their OS, Office suite and everything else they do business in.

Google Chrome took advantage of that dirty game and now you have the Chromium engine that powers a lot of browserlike frameworks.

No one born in the LLM age even knows what I.E means or stands for, as it should be - a horribly designed, poorly working product foisted upon users via the Windows distribution system - a dishonorable product from an ethically corrupt company forever lost in history, right alongside Clippy and DCOM.

OTOH, I am glad that Microsoft played a dirty game with I.E and didn't just stop playing dirty there - they jacked up the price of Windows if an OEM even dared to bundle in Netscape Navigator instead - who knows, if they hadn't done that, there wouldn't have been a Google or Apple. We would all be using Windows and Windows Search and Windows Phone.

And without Google, we might not have had the modern LLM as we know it. We would have had some trashy Windows Autocomplete Copilot Clippy. Ugh!

Re: Grok 4.5

#197

Earlier quoted context omitted.

Google is using AI at such scale internally they don't need external customers to recoup their investment.

> Google is using AI at such scale internally they don't need external customers to recoup their investment. That's assuming their flagship product remains relevant in an AI-powered world. Which brings to mind: most of the big shops product (chatgpt, claude, grok, etc...) ALL rely on search, and NONE of them actually have a running search stack. Which means, they must all be calling Google, no? How does Google make m…

> Which means, they must all be calling Google, no?

Incorrect. Alternate search providers exist, such as Bing (used by DuckDuckGo, for example) and Brave.

Re: Grok 4.5

#198

Can someone breakdown to me how this makes any sort of economical sense? Spending billions and billions to have the 3rd best model while even the number 1 and 2 players already seem to struggle making a profit. What am I missing here? Not trying to go full Ed Zitron but this doesn’t make sense to me.

You could be typing the same about Google or a number of the other labs right now. A diverse market full of choices keeps it from becoming the browser wars all over again.

Google is playing a different game. I don't really know what game they're playing, but they're not trying to beat Claude Code. They have coding capabilities and Antigravity, but I'd be surprised if it's much more than an afterthought. They're focusing on efficiency, models at the edge, human interaction, image and video, etc. in ways Anthropic, in particular, is not.

Google wants its AI to be pervasive in everyone's daily life. Merely being the best at coding is not how you get there.

I am more bullish on Google in AI than most folks, I think, as they have been focused on efficiency in a way most US vendors have not. They've published a ton of papers on ways to make LLMs more efficient and capable on smaller devices.. Google wants to own the on-device market for AI, and I don't see many credible competitors in that space.

Re: Grok 4.5

#199
post #102
post #68

Earlier quoted context omitted.

Maybe it‘s the Nash equilibrium from a timing perspective? Like the reason that close to a McDonals there is usually a Burger King.

The joke is that McDonald's spends hundreds of thousands of dollars to identify new locations - traffic studies, visibility, demographics, nearby traffic generators, site characteristics, drive-thru feasibility, etc. They have one of the most rigorous processes in the industry. Burger King's process is to open a location across the street.

I think the story was Starbucks -> Seattle's Best Coffee, not McD's --> BK. but it does work I guess.

Re: Grok 4.5

#200
post #28

Its remarkable how Anthropic is able to maintain their edge against all competition. Anyone have any idea what the secret sauce is that has Anthropic at the top of all leaderboards for the past few years?

I think the "secret sauce" is not juicing the benchmarks. Claude models just feel like they are better than the benchmarks suggest, in terms of smarts and creativity, while models from every other company feel worse relative to what you'd think from the benchmarks. Only company to really internalize Goodhart's Law, IMO.

Yeah every model has great benchmarks. Claude is the only model I want to use when I'm not worried about the marginal cost of tokens (which is most of the time at work.)

I then use cheaper models like GLM for personal projects but they're noticeably much worse despite being similar in benchmarks.

Post reply on HN