OpenAI reportedly ditches model over safety concerns



OpenAI had planned to release yet another AI model next month, but has decided to nix the release over safety concerns.
The Wall Street Journal reports that Astra 6.1 was scheduled to be released as soon as within the next few days. However, the model “showed higher levels of deception” than previous models and exhibited unsafe behavior, the Journal writes.
Saachi Jain, OpenAI’s head of safety systems, told the WSJ that the model tested poorly on alignment, a measure of how well the program adheres to human intent.
TechCrunch reached out to OpenAI for more information and will update the article if it responds.
Astra was released earlier this month and hailed by OpenAI as its most powerful model yet.
Questions about safety have plagued the AI industry over the past several months — ever since the Hugging Face incident , in which an OpenAI agent broke free of its sandboxed environment and hacked several different companies. Since that incident, more models — including Anthropic’s Claude and Google’s Gemini — have been revealed to have exhibited similar behavior.
The deluge of concerning stories has, ironically, helped to push the policy conversation in the U.S. toward an outcome desired by top AI labs : the institution of new industry standards for AI safety and potentially a slowdown of the industry itself.
Companies like OpenAI and Anthropic have claimed that the concern here is safety, although another potential motivation posited by critics is that it could entrench the industry position of those companies at the detriment of less resourced firms.

Get 50% off a second pass The Disrupt experience is meant to be shared. Get your pass and bring a colleague, partner, or peer at 50% off. Cover more ground by making connections, building momentum, and discovering what’s next in the startup ecosystem.

Nvidia launches new platform for reining in rogue AI agents

OpenAI still doesn’t seem to have a handle on all of its rogue AI activity

Can Muse overcome Meta’s trust issues?
Latest in AI

Peak XV ups Surge seed investment ceiling to $5M, unveils 18-startup cohort

OpenAI reportedly ditches model over safety concerns

Source: Inference provider Modal Labs closing in on $750M round at $15.75B valuation
Verified source · TechCrunch
Reported by TechCrunch. Open the original for full media and formatting.
More in Policy
All news
PolicyTrump finalizes rule to make cars less fuel efficient
The US Department of Transportation finalized its plans today to weaken fuel efficiency standards, calling it "among the largest deregulatory actions under the second Trump Administration." It's a nail in the coffin for Biden-era standards that would have required fleet average…
Read at The Verge
PolicyAI is supercharging hacking, and your local hospitals and banks aren’t ready
In March, Janice Malone began getting calls about suspicious activity from her nonprofit organization, Vivian's Door. Vivian's Door, headquartered in Alabama, typically provided training, resources, and community to underserved and minority-owned businesses. The work sometimes p…
Read at The Verge
PolicyOpenAI still doesn’t seem to have a handle on all of its rogue AI activity
On Friday, OpenAI published a new site devoted to “misalignment reports” and the breadth of the incidents is alarming.
Read at TechCrunch
PolicyFlorida seeks a ban on ChatGPT acting like a person
Florida Attorney General James Uthmeier is calling for a judge to block OpenAI from "giving ChatGPT false human attributes," a few months after Florida sued the AI company over safety concerns. According to Uthmeier, users are lulled into a false sense of security by the AI bot,…
Read at The Verge