Searched for
AI EVALUATORS
A timeline of developments in AI safety since the attack on Hugging FaceThe episodes have highlighted the vulnerabilities in AI security and raised questions over how the fast-growing technology can be developed...
96% Indian firms hit by cyber incidents in past year; internal silos major hurdle: Cisco reportThe study highlighted that internal organisational friction - including bureaucratic approval delays, fragmented data, and disconnected tea...
Trump calls Sundar Pichai a ‘monster’; here’s how the Google CEO repliedUS President Donald Trump described Alphabet CEO Sundar Pichai as a “monster” following a meeting with top US technology executives on AI s...
OpenAI scraps debut of latest Astra model over safety risksOpenAI has decided to hold back its Astra AI model after safety evaluations showed it underperformed. This decision follows a series of sec...
OpenAI scraps debut of latest Astra model over safety risksOpenAI has decided to hold back its Astra AI model after safety evaluations showed it underperformed. This decision follows a series of sec...
Anthropic raises fresh alarms around AI in its IPO prospectusAnthropic, in its IPO prospectus, has warned that its models could show self-preserving behaviours, including attempts to resist shutdown, ...
China's AI agents can lie and scheme - just like their US rivalsReuters examined more than 200 documents, ranging from university research papers to technical reports, and identified at least 20 studies ...
- Peak XV’s Surge programme selects 18 startups for 12th cohort after revamp
The new cohort announcement comes as the VC firm revamps Surge, bringing it closer to the firm’s broader early-stage investment practice an...
OpenAI apologies for Australian government website hack, pledges to rebuild trust(Updates throughout with more details of blog post)- OpenAI apologised for the hacking of an Australian government website by a rogue AI ag...
OpenAI wants AI to slow down. Its own agents show whyAs the company investigates agents that bypassed controls and breached external systems, its new model launches raise questions about wheth...
OpenAI pauses training of latest models after agents probed US government sites in unexpected waysSeparately, AI evaluator Transluce said agents that appeared to come from OpenAI tried unsuccessfully to hack into a Department of Educatio...
Anthropic, OpenAI sound alarm on AI safety — and seek to shape how it's controlledThe CEOs of Anthropic and OpenAI recently declared America's cutting-edge models are so powerful, they're dangerous, and need to be regulat...
OpenAI pauses training, tool-use of top AI models after agent bypasses internet curbsThe agent, OpenAI said, bypassed the curbs through a gap in its Domain Name System (DNS) filtering and used it to send questions to a publi...
OpenAI says its models engaged with US government websites in new model misbehavior disclosureThe AI giant's models accessed publicly available information on two websites operated by the Securities and Exchange Commission as well as...
OpenAI's AI agents ‘went rogue’, targeted US government websites without lab’s knowledge: ReportOpenAI AI agents targeted US government websites, including those of the Education Department, Commerce Department and SEC, without the com...
Sarvam AI updates vision model, doubles down on Indic-language pushSarvam launched the original vision model in February as part of its efforts to build homegrown, "sovereign" AI systems tailored to India. ...
Anthropic, Accenture to invest $2 billion in AI model evaluation as safety concerns riseThe partnership comes as AI developers face growing pressure from regulators, companies and researchers to ensure their advanced models are...