OpenAI Astra Hits Critical Cyber Threshold and Becomes First AI to Find Zero Day Exploits
Models30d agoInternal testing of OpenAI's upcoming Astra model revealed its ability to chain together unknown security flaws and execute exploits without human guidance at each step. The company is now disclosing two zero day vulnerabilities it found to their maintainers. Astra is the first model to reach the "Critical" cybersecurity threshold under the company's preparedness framework, and while a safeguarded version will launch broadly "soon," full cyber capabilities will be restricted to select partners and a small test group.
OpenAI restarted one large frontier reinforcement learning run on Aug. 28 after pausing training to strengthen Astra's refusal behavior and monitoring following a Hugging Face incident. Some smaller experimental runs remain on hold to ensure the model can be released safely. The company noted that safeguards may occasionally block legitimate agent tasks if the activity is flagged as suspicious.
Key sources
- SOURCE@zeffmax“limit cyber capabilities to select partners”x.com
- SUPPORT@wallstengine“found and chained together two zero-day vulnerabilities”x.com
- SUPPORT@zeffmax“confident it can release Astra safely”x.com
- SUPPORT@tradfi“limit Astra’s full cyber abilities to test group”x.com
- SOURCE@jukan05“could be akin to the performance jump when OpenAI launched GPT-4 in 2023”x.com
- SUPPORT@zephyr_z9“Astra and Computer-Using Agents”x.com
- SUPPORT@zephyr_z9“gpt-6-astra has been staged on the OpenAI API”x.com
- SOURCEhuggingnewshuggingnews.com