Live data from Hacker News

Muse Spark: Scaling towards personal superintelligence

ai.meta.com

391–392 of 392 posts

Re: Muse Spark: Scaling towards personal superintelligence

#391

Earlier quoted context omitted.

Nah. Everybody is talking about ai. Everybody is using it. It's by far the most popular new tool human beings are using currently. As popular as mobile phones or spoons. And maybe as disruptive as the steam engines. AI companies are becoming the largest software companies on the planet. Everything points into that direction. Trillions of dollars are waiting in the market to be collected.

> Everybody is talking about ai. Everybody is using it. Please take a moment to step outside the tech bubble. Neither my neighbor (a hair stylist) nor the carpenter fixing up her kitching cabinets are "using" AI. They might get Gemini text when googling something, though they often scroll past it because they often don't trust it. And they get lots of fake videos when scrolling their youtube which increasingly annoys…

But how do they learn to do their respective task? How is the information disseminated?

The capability is there for robotics to handle these kinds of repetitive tasks from a long term view. They're just statistical processes on a fundamental level.

In general, a lot of this shit that we do can be represented this way. It's just a question of where the incentives are to apply it first and how many economic cycles it'll take to get there.

Also, who controls the training data will matter a lot more. I.e. the sort of "ancestral knowledge" within different enterprises and how they deliver on respective business goals.

Re: Muse Spark: Scaling towards personal superintelligence

#392
post #361
post #285

Earlier quoted context omitted.

Benchmarks miss the thing that actually matters for agentic use: how does behavior change over a multi-day horizon? A model that scores well on one-shot coding tasks can still make terrible decisions when it has persistent state and resource constraints. That's where you see the real gaps between models.

Is there a benchmark for these long tasks? That kind of seems like the only number worth measuring. (Of course at that point it involves memory and context management and so on, so you're testing the harness as well as the model.)

[dead]
Post reply on HN