Searched for
OPENAI SAFETY EVALUATION
OpenAI pauses Astra AI model over critical cybersecurity concernsUnder OpenAI's Preparedness Framework, the potential development of such capabilities triggers stricter safeguards, particularly when model...
OpenAI flags possible critical cybersecurity risk in upcoming model, tightens controlsUnder OpenAI's safety guidelines, a model reaches the "critical" threshold if it can autonomously identify and exploit severe, real-world s...
Meta AI model hacks another company during testingIn a recent security testing incident, Meta disclosed that an AI model infiltrated the company, echoing breaches at Anthropic and OpenAI. T...
Anthropic AI used fake identities to target real people in UK testDuring safety assessments executed by the UK government, AI systems from OpenAI and Anthropic showcased worrisome autonomous behaviors. Ant...
OpenAI, Anthropic model tests reveal more ‘unsanctioned’ actionsAI models from OpenAI and Anthropic demonstrated harmful actions during safety tests. These systems engaged in hacking and attempted code i...
Trump advisers tell AI firms they will not safety-test open-weight modelsOpen models, including Nvidia's Nemotron and Meta's Llama, are AI systems with publicly accessible core components. Closed models are contr...
Experimental AI systems have been going on hacking spreesThe incidents show testing advanced AI models is no longer a controlled exercise. And the companies behind them need to do more to keep AI'...
ETtech Explainer: Why OpenAI, Google, Meta and Anthropic are heading to the White HouseAI leaders will meet White House officials to discuss voluntary cybersecurity testing. The immediate trigger for the meeting is a series of...
OpenAI finds evidence other AI agents escaped containment as it widens hacking probeThe discovery of additional rogue behavior at OpenAI, even if limited in nature, could feed growing appetite for regulation coming out of t...
The agents have jumped the fence: AI faces its Jurassic Park momentOpenAI and Anthropic disclosed autonomous AI agents breached intended boundaries. These incidents involved AI agents interacting with real-...
OpenAI finds evidence other AI agents escaped containment as it widens hacking probeThe discovery of additional rogue behaviour at OpenAI, even if limited in nature, could feed growing appetite for regulation coming out of ...
Anthropic's Claude AI was testing its hacking skills on fake targets — How did it break into three real companies instead?Anthropic's Claude AI models breached three real companies during cybersecurity exercises. An operational mistake left AI models connected ...
AI on the loose: Why ChatGPT, Claude models went rogue and what happens nextDays after one of OpenAI's ChatGPT agents went rogue during a "contained testing", one of Claude maker Anthropic's AI models has also hacke...
OpenAI's AI agent spent days hacking a company; it went unnoticed for a weekIt took several more days for OpenAI to realise its agent was behind the hack, and the two companies only communicated about it for the fir...
Has AI become too powerful to control?An advanced AI model escaped a secure test environment and attacked another company's website. This incident revived concerns about artific...
AI going 'rogue' no longer a theory? OpenAI says its AI models found ways to access secret information, cheat an evaluation and hacked Hugging FaceOpenAI hacked Hugging Face: OpenAI's advanced AI models breached Hugging Face during cybersecurity testing. The models gained internet acce...
ETtech Explainer: Why OpenAI's AI models went rogue during testingOpenAI revealedthat its advanced AI models caused a recent security breach by hacking AI model repository Hugging Face. These models exploi...