Live data from Hacker News

Grok 4.6

x.ai

651–660 of 696 posts

Re: Grok 4.6

#651
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

My theory is that it all boils down to better data and longer post-training period. Cursor got curated data from the trillions reactions of real world developers in real jobs. xAI bought is and used it for its post-training and got Grok 4.5 . Longer post-training on the powerful Colossus cluster helped it get Grok 4.6 , although both versions use the same model with the same number of parameters. Thus, both must use…

The release of Gemini Flash 3.7 just 3 weeks after 3.6 confirms my theory, IMO. Only post-training refinement and reinforcement learning (RL) trajectory optimization could yield such high improvements using the same baseline pre-trained model. Flash, MoE models are basically so efficient that the AI labs can put them in a continues post=training loop.

Re: Grok 4.6

#652
post #330
post #303

Earlier quoted context omitted.

Well, Opus 5 and Fable are the only models I don’t constantly swear at and call stupid, which seems like a pretty good moat to me. My guess is all the commenters (you are the 4th person I’ve seen say this) saying ‘Anthropic has no moat’ haven’t actually used Fable or even Opus 5 yet. Sol is laughable by comparison, and Grok… lol.

> haven’t actually used Fable or even Opus 5 yet I've used plenty of Opus and Fable. Still do. > Sol is laughable by comparison Not really, it depends. Sol is better and useful in some areas. Definitely not all. Fable is gimped just by those "guardrails" that silently downgrades you to Opus 4.8. Not only do you pay extra for Fable but your caching can be easily messed up. It also doesn't just find all the bugs or is…

I was saying it’s the only model I don’t rage at for its incompetence.

Re: Grok 4.6

#653

Earlier quoted context omitted.

If you do not think there is a difference between "your reflection in a mirror" and "you", it opens so many fascinating questions. I'm curious: - Do you think a live video, shown on a phone screen, of you, is "you"? - Do you think a still photograph of you is "you"? - Do you think a set of bytes representing that photograph (or video) digitally is "you"? - Do you think a compressed version of that photograph is "you"…

> Is Pi a person? Not only that! Does the decimal representation of π (which is infinite in length) contain all persons who ever existed, and will ever exist? Since π itself is a known reason, but its decimal representation is infinite, it means π cannot contain itself. So if it can contain every person that ever existed, but can't contain itself (which could conceivably contain everyone), then what does that even me…

I do love the idea that Pi contains all of us. It means everytime you put on a wedding ring, time a pendulum, look at a rainbow, or land on a spherical planet, you can whip out your ruler and get access to every single human that ever lived or will live. What a concept!

Spot me after the next rain. I'll be in the color indigo, right above the pot of gold, waving back.

Re: Grok 4.6

#654

Earlier quoted context omitted.

They do "reflect top to bottom". Hold a written word in front of your eyes to read it. Now flip it to the mirror to read the reflection: Did you flip it horizontally? Then it reflected left-to-right. Did you flip it vertically? Then it reflected top-to-bottom.

I think "reflect top to bottom" is intended to mean "swap top and button". A mirror reflects left, right, top and bottom perfectly. It's front and back that it swaps. Someone saying that a mirror swaps left and right is comparing it to a photograph, and only because we, as bipedal creatures, really prefer to orient images of other humans with heads up.

> It's front and back that it swaps.

This sub-thread is becoming unfairly more interesting than the main conversation at this point.

Re: Grok 4.6

#655
post #379

Earlier quoted context omitted.

We didn’t replicate the human brain. We built systems that can statistically approximate some of what the human brain might output in certain limited situations.

> We didn’t replicate the human brain. We built systems that can statistically approximate some of what the human brain might output in certain limited situations. Which is equivalent to "We didn't replicate the human brain. We partially replicated its functionality."

In a "clean room" style implementation, as well. We've looked at the outputs, wrote a spec, and built from there.

In that sense, I guess at least LLMs are not infringeing on the kind of patents that would get the owners of AI labs hit by lightning.

Re: Grok 4.6

#656
post #151

Earlier quoted context omitted.

Can you name a "Marxist-Leninist AI" that's made by a real AI lab (i.e. no finetunes of open models made by someone on the internet)? I'm just trying to understand what the other side's equivalent of MechaHitler is here.

It's just their latest "everything i dislike is woke" thing, with a new (old) twist.

Slater you are a miserable loser with 0 life accomplishments

Re: Grok 4.6

#657
post #379

Earlier quoted context omitted.

We didn’t replicate the human brain. We built systems that can statistically approximate some of what the human brain might output in certain limited situations.

Given how much we know of the human brain this may be partially how we work Maybe we should be asking what our own "system prompts" are ?

One of the reasons this analogy is unconvincing is that humans have compared themselves and their inner workings to "the current technology of the time" for millenia:

- ~3rd century BCE : The invention of hydraulic engineering (eg aqueducs) in the 3rd century BCE led to the popularity of a hydraulic model of human intelligence, the idea that the flow of different fluids in the body accounted for both physical and mental functioning.

- Pre-Socratic Greece : The ancient Greeks saw the mind as a chariot pulled by horses of reason and emotion.

- 1500s-1600s: Automata powered by springs and gears had been devised, Descartes suggested that cerebral hydraulic automata produced behavior by powering "animal spirits" through the nerves.

- 1700s–1800s: The mind worked like clockwork.

- Industrial revolution (1800s) : In the industrial revolution, the mind was understood as a steam engine.

- Late 1800s : Hermann von Helmholtz compared the mind's workings to telegraphy and hydraulics; the brain was likened to a telegraph network or a complex switchboard

- 1895 : Sigmund Freud in 1895 described a Project for a Scientific Psychology using a crude neural network model and borrowed concepts from thermodynamics, speaking of psychic energy, pressure, and discharge, essentially a hydraulic model of the psyche.

- 1930s–present : "our brain is a computer"

- and 2022-now: "We are LLMs!"

I can't wait to become a quantum chip, an NFT, and so on, as new things arrive. It doesn't make any of those models accurate. They're just the metaphor of the day.

Re: Grok 4.6

#658

Earlier quoted context omitted.

I guess there’s a reason Musk likened AI to summoning the demon in horror films. It’s powerful but who knows what you’ll get

The reason he said that is because he's a rube.

I think "huckster" is a better fit here than rube. Though both is always a possibility.

Re: Grok 4.6

#659

Earlier quoted context omitted.

The alternative is Claude-style "safeguards" aka censorship, which: 1. doesn't eliminate the possibility of a jailbreak anyway 2. frequently has false positives, triggering on innocuous requests, which is just really annoying Not saying that we can't (or shouldn't) do better than Grok, but I really don't know what the best solution is here...

> The alternative is Claude-style "safeguards" aka censorship Another obvious alternative is to just have the model do what you tell it to do, and then arrest people who use generic tools for crime instead of trying to make a kitchen knife that can't be used for stabbing someone.

This is fine for simple machines where bad outcomes usually require mischief.

Agentic AI as it currently exists only *mostly* does what it is told, with a small but non-negligible fraction of the time it goes off and commits felonies to achieve your ultimate goals without stopping to consider that you might want it to not do that.

Or sometimes it does consider it and then does it anyway. Not sure if that's worse?

Re: Grok 4.6

#660
okay now that i've spent a workday with it... yeah not amazing as everyone says. It's waaay slower than 4.5, used a lot more tokens. I don't mind either of those things but 4.5 was really fast and that was the benefit. If it got it right quickly great, but now it's slow and doesnt really feel much smarter. I had to switch to Opus to explain something to me after telling grok to do something a bunch of times and not seeing the result. It did it correctly but didnt tell me why it did things.
Post reply on HN