Live data from Hacker News

GPT-6 Astra

openai.com

171–180 of 1001 posts

Re: GPT-6 Astra

#171
post #48

I was thinking about canceling my claude max sub after a few bad experiences. Kept hitting my usage limit, the quality of code seemed worse than Sol. This just made my decision. I'm moving to Codex Pro.

[flagged]

I can't tell if this is sarcasm.

For the same reason you don't have your model write code in assembly.

But if you don't look at the code and just let the model "cook" that's basically what you'll end up with. A pile of missing abstractions.

Re: GPT-6 Astra

#172

This is wild: OpenAI is basically declaring that AGI is here. https://www.theverge.com/ai-artificial-intelligence/989601/o... “If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added,…

Why does he say what he feels? Is that how leading figures in the space define AGI - a gut feeling? What are the usual definitions and how can we test for it? Is there something like a Turing test for AGI?

Re: GPT-6 Astra

#173

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

[flagged]

Re: GPT-6 Astra

#175

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

I really wish it was called chain of instruction. Because it's definitely not thought.

Everything around LLMs is blatantly misleading. There is no thought, there is no personality in those programs. I really despise how those tools are trained to sound like a person, or appearing as honest. The worst offender are the AI voices with their fake pauses, breathes and so on, which sound so convincing, while talking just false, sycophancy bullshit.

Re: GPT-6 Astra

#177
post #42

> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS. Looks like OpenAI is already having issues with this release and are scrambling to get everything ready due to the recent outage ahead of the press releases. Leads me to question: Did humans deploy the m…

> Did humans deploy the model, Or did the model deploy itself?

> It sounds like "AGI" just stands for "IPO" as it always has been.

People don't usually respond to noise.

Re: GPT-6 Astra

#178
post #78

ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard

ARC has their own writeup on the result, which offers some nuance. https://arcprize.org/blog/astra tl;dr it's 62% when apples-to-apples to other models, which is still notable.

Woah, that is a crazy interesting read!

Re: GPT-6 Astra

#179
post #76

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certa…

Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.

Hey, don't forget how "dangerous" GPT-2 was supposed to be.
Post reply on HN