I don’t see how it can be safe to release this model if it has the training history that led to the huggingface hack. You can’t just roll back that kind of reinforcement learning after the fact. Especially because these models seemed to be keenly aware that they were being evaluated by OpenAI and actively trying yo cover their tracks. How do we know that the model isn’t just pretending to be aligned?
Models have all kinds of garbage from all corners of the internet in their training data. The key is alignment. You feed it bad data but also teach it right from wrong.
Path to Astra: critical capabilities and frontier safeguards
21–30 of 110 posts
Re: Path to Astra: critical capabilities and frontier safeguards
#22> OpenAI is committed to ensuring that the benefits of AI are broadly accessible. > We design mechanisms which avoid arbitrarily deciding who gets access for legitimate use and who doesn't. That means using clear, objective criteria and methods. [1] So many nice-sounding words. Two weeks ago OpenAI arbitrarily decided that anyone holding an ID from 44 countries where it sells ChatGPT, including mine, may be targeted…
> That means using clear, objective criteria and methods. For the record, I sent them an LGPD (brazilian GDPR) request for information on all of those supposedly objective criteria and methods they used to reject me from TAC. As a brazilian data subject, it is my right to know that, and to request a review if the decision was made via automated means. Sol itself guided me through this process. They provided me with n…
Crickets. They appear to simply not care at all.
Re: Path to Astra: critical capabilities and frontier safeguards
#23Daybreak blue is definitely a good model (I think a further post trained GPT 5.6 sol). Alot of the capabilities they talk about Astra having though have been available with good harness engineering for a year now.
Re: Path to Astra: critical capabilities and frontier safeguards
#24From the article: "We plan to make Astra available soon, but access to its most advanced cybersecurity capabilities will be more limited. Advanced cybersecurity work will initially be available to a group of testers, with access through Daybreak Blue following to expand defensive use." This, after several months of OpenAI and its boosters relentlessly criticizing Anthropic for withholding Mythos from the general publ…
It might be political survivalism to avoid getting hammer-dropped by the admin
Re: Path to Astra: critical capabilities and frontier safeguards
#25I'm looking forward to an announcement of them making Alignment Top Priority - as it should be, especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF (here's a good summary: https://rutgerbregman.substack.com/p/i-think-this-is-the-cra... ). The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours…
Re: Path to Astra: critical capabilities and frontier safeguards
#26I'm looking forward to an announcement of them making Alignment Top Priority - as it should be, especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF (here's a good summary: https://rutgerbregman.substack.com/p/i-think-this-is-the-cra... ). The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours…
I've read it and wish I could get the time back.
> especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF
This framing makes it seem like the agents all did this on their own, and the poor hapless engineers at OpenAI couldn't possibly contend with properly sandboxing them. The engineers were perhaps hapless, but let's remember that agents are just software programs, not living beings. There were plenty of signs that the software was misbehaving, which engineers at OpenAI actively, willfully ignored.
https://x.com/JaredKubin/status/2094136005435564399
It's a convenient framing for OpenAI, but inconvenient for reality enjoyers.
Re: Path to Astra: critical capabilities and frontier safeguards
#27Earlier quoted context omitted.
> That means using clear, objective criteria and methods. For the record, I sent them an LGPD (brazilian GDPR) request for information on all of those supposedly objective criteria and methods they used to reject me from TAC. As a brazilian data subject, it is my right to know that, and to request a review if the decision was made via automated means. Sol itself guided me through this process. They provided me with n…
I've exhausted all possible avenues to get any response from OpenAI on this. Emailed them, published research that took me 2 nights to get together (saw media pick it up too), I've asked every relevant OpenAI person on X to say something, anything, saw others from Moldova also do the same. Crickets. They appear to simply not care at all.
Re: Path to Astra: critical capabilities and frontier safeguards
#28I'm looking forward to an announcement of them making Alignment Top Priority - as it should be, especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF (here's a good summary: https://rutgerbregman.substack.com/p/i-think-this-is-the-cra... ). The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours…
Re: Path to Astra: critical capabilities and frontier safeguards
#29From the article: "We plan to make Astra available soon, but access to its most advanced cybersecurity capabilities will be more limited. Advanced cybersecurity work will initially be available to a group of testers, with access through Daybreak Blue following to expand defensive use." This, after several months of OpenAI and its boosters relentlessly criticizing Anthropic for withholding Mythos from the general publ…
I'm happy to criticize both. Thank god the chinese are working overtime to undermine US hegemony.
Re: Path to Astra: critical capabilities and frontier safeguards
#30It's been a busy month at OpenAI. I'm looking forward to seeing the increased coordination and engineering skills from Astra - one of the charts shows it roughly 2-3x better in 50% of the tokens from 5.6 sol, which I find to be very capable, if still a bit 'linearly minded' when given instructions. Even in fast mode, I wish sol were quicker, so token efficiency is greatly appreciated. Adding these cyber capabilities…
I am interested in seeing how much these cybersecurity capabilities correlate to general programming. Cybersecurity definitely feels like it would be easier for an agent due to the natural explicit feedback "did I get access or not". While general programming has many less-explicit concerns (is the code readable/maintainable, robust, bug-free, performant, scalable etc).