Earlier quoted context omitted.
> Astra is a new step in LLMs I think. I'd be interested in hearing more about your evaluation here. It would be nice if LLMs have gotten past the "tell me" hump of recent Claude/OpenAI verbosity.
So before I got a job this fall, I was working on a side project about compiling a particular language to SQL. To test Astra I pulled it off the shelf and asked it to take the grammar and then create a compiler to SQL. I've done this before with GPT-5.5, 5.6-Sol High. The latter was way better but it was still really verbose and information sparse; it used a lot of words to describe each IR expression but didn't real…
An Alien Mind
81–90 of 488 posts
Re: An Alien Mind
#82Earlier quoted context omitted.
> I want to believe that humanity is trending towards a good outcome here All the trends so far are towards a nightmarish hyper-capitalist end game. None of the AI leadership is trustworthy, and they openly discuss how they are willing to sacrifice everything humans cherish to have a shot at reaching their envisioned utopia (which would be the most obvious dystopia for anyone else)
I'm not really worried about the labs, it's misaligned governments that keep me up at night. ASI landing during the current administration is not ideal. I also would prefer to avoid needing to indoctrinate myself in Xi Jinping Thought.
Re: An Alien Mind
#83> The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI. So the best argument for AI is that it's an arms race. We have to keep pushing every boundary because in any case others will, and we will need to defend against them. If this statement is true, then this particular researchers believes the open source Chines…
Maybe people are rightfully concerned about the capabilities of the models of other (non/less democratic) states. But if we are concentrating power in the hands of few and at the same time allowing the creation of a weapon that thwarts any offense, how do we ensure the health of our democratic societies?
Re: An Alien Mind
#84One of my favorite things to do with these blog posts is to imagine an Alien Museum on the Remains of Humanity, and wonder what the little text flyouts and commentary on the screenshot of this one might say. Some ideas: "Despite a nuanced view of the complexities of what lay ahead, humanity found itself collectively unable to stop the process it had set in motion." "Despite significant progress on the mechanisms of a…
Re: An Alien Mind
#85> Delivering the benefits of scientific progress and economic growth that very intelligent machines enable. I think we're very close to the point where AI-driven breakthroughs outside of pure math and software start to really affect the world. We evaluated GPT-6 Astra in 100 complex, unsaturated multi-agent coding environments, competing and cooperating with other models in open-ended tasks. It's the new frontier mod…
Incredible... software engineers will be joining the breadline soon as managers, executives and PMs take over deliverables. The world will look very different on Jan 1st 2027.
Sorry if I misunderstand the point, just trying to understand.
Re: An Alien Mind
#86> The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI. So the best argument for AI is that it's an arms race. We have to keep pushing every boundary because in any case others will, and we will need to defend against them. If this statement is true, then this particular researchers believes the open source Chines…
Yep. Such a disgusting industry. They created the arm race, push for the arm race, put themselves in position to benefit from the arm race
Re: An Alien Mind
#87Re: An Alien Mind
#88Earlier quoted context omitted.
Playing chess, writing code, finding security bugs, and proving mathematical theorems all seem fairly similar to thinking and don't seem much like being tall.
The crucial distinction here is that "seem" does not at all mean the same thing as "is". Thunder seems like the anger of the gods but it isn't . We've had chess playing programs for a long time now and despite it seeming like thinking is required for them, it isn't. The principle you're using here isn't a scientific one but magical [1]. Abandoning empiricism and rationality is not a good way to make progress. [1] htt…
It makes sense now to say that temperature is what a thermometer measures. However, before there were good thermometers, people often thought that heat and cold were different things. The meanings of the words we use were influenced by scientific progress.
For thinking, we don't have a good thermometer. There are IQ tests, but they aren't aren't necessarily all that useful for comparing what people do to what machines do. And that's why there are a zillion AI benchmarks - none are entirely satisfactory.
So what does "smart" mean to you? How do you define it in practical sense? What definition should scientists settle on?
Without a proper definition, how do you tell the difference between "seems smart" and "is smart?"
Re: An Alien Mind
#89Earlier quoted context omitted.
Incredible... software engineers will be joining the breadline soon as managers, executives and PMs take over deliverables. The world will look very different on Jan 1st 2027.
Maybe I just lack imagination, but I don't really know how jobs are supposed to solidify around the role of giving prompts to agents and then looking at the results. I mean, engineers will be in the breadline because their role was simply to prompt the agents.. only to be superseded by managers or executives who no longer manage engineers but themselves prompt the agents? And, for this previously considered obsolete…
Re: An Alien Mind
#90Earlier quoted context omitted.
Playing chess, writing code, finding security bugs, and proving mathematical theorems all seem fairly similar to thinking and don't seem much like being tall.
The crucial distinction here is that "seem" does not at all mean the same thing as "is". Thunder seems like the anger of the gods but it isn't . We've had chess playing programs for a long time now and despite it seeming like thinking is required for them, it isn't. The principle you're using here isn't a scientific one but magical [1]. Abandoning empiricism and rationality is not a good way to make progress. [1] htt…