Politics, paranoia & AI kill switch: AI can be switched off. But what if it learns to fight back?

US lawmakers have introduced an AI Kill Switch Act to regulate advanced artificial intelligence systems. This bill grants the government power to halt or suspend AI models exhibiting catastrophic risks. Concerns arise from AI safety theories and t...

BCCL - Non Copyright

HAL, which is the button to shut you down?

In July, after an autonomous OpenAI model escaped testing sandboxes and notoriously breached ML platform Hugging Face, two US Congressmen, Democrat Ted Lieu of California and Republican Nathaniel Moran of Texas, introduced the AI Kill Switch Act. The bill, if passed, would give the US federal government the power to order companies to slow down, suspend or shut down advanced AI systems, and the issue jumped from fiction to legislation. Suddenly, the 'kill switch' of gone-rogue computer HAL 9000 in 2001: A Space Odyssey has become real.

To become law, the bipartisan bill will need to be approved by the House of Representatives, debated and passed in the US senate, and signed into law by the president. If and when passed, regulatory bodies will have the authority to call for an immediate shutdown if any model exhibits 'catastrophic risks'.

There is a thesis in AI safety called 'instrumental convergence'. It posits that agents with capability will always converge on a set of goals, including preserving themselves, acquiring resources, expanding their influence, and protecting against modification of their objectives. Such actions are a byproduct of mathematical logic, and measures that make the system's assigned goals harder to achieve.


Actions by capable agentic systems could also include transmitting its source code to unauthorised servers without human permission, concealing sub-processes, preserving API access keys before its main process faces termination, and exploiting vulnerabilities faster than humans in the loop can respond. So, flipping the switch isn't simple.

Evoking one kill switch doesn't usually cut off all computing. Containment spans several boundaries before getting back to a known, safe configuration. That includes revoking tokens and tools, isolating connected services, ceasing execution and, importantly, all downstream automation that could continue even after the model has been paused.

Getting it exactly right is critical. Overly restrictive shutdown mechanisms make models sluggish, limiting their utility in time-sensitive verticals like finance, healthcare and defence. Also, lowering safety gates causes unimagined, severe blind spots.
ADVERTISEMENT

So, efforts are on to infuse non-AI software layers between models and their tools, ensuring every command passes a hardcoded checklist that can't be bypassed or rewritten. Other measures include introducing microcircuit breakers to freeze an agent's network access when abnormal spikes happen, expiry of credentials if agents go rogue, and choking egress gateways when there are attempts to transmit sensitive files or unauthorised cryptographic keys.

Google DeepMind's Frontier Safety Framework, OpenAI's Preparedness Framework, and Anthropic's Responsible Scaling Policy have been published to catch dangerous models before they ship. Future of Life Institute, which tracks AI-related risk, graded several frontier labs across 37 indicators. Anthropic scored 2.66; OpenAI, 2.28; and Google DeepMind, 2.01. None were above C+. xAI, DeepSeek and Mistral failed outright. The lowest scoring category was 'existential safety', which measures actions when losing system control. Essentially, we're just starting that journey.

In a July study, Anthropic researchers working on Claude Sonnet 4.5 claimed that it has a silent internal network region, 'J-Space', comparable to human brain's workspace. AI pioneer Geoffrey Hinton suggested as much many years ago - that the closer technology gets to mimicking the human brain, the higher the possibility of conscious AI.

As we develop bio-hybrid computers - where human neurons are mounted on silicon - we're getting closer to that 'uncomfortable' reality. A sentient entity with superior reasoning capability is certainly better equipped to predict human behaviour. It could figure out our innovative interventions and bypass or disable all termination protocols, even resort to strategic deception and fake compliance.
ADVERTISEMENT

Are humans ready to battle this formidable rival we've created, with the odds stacked so steeply against us? Watch this neural space.
(Disclaimer: The opinions expressed in this column are that of the writer. The facts and opinions expressed here do not reflect the views of www.economictimes.com.)
Download
The Economic Times Business News App
for the Latest News in Business, Sensex, Stock Market Updates & More.
READ MORE
ADVERTISEMENT

READ MORE:

LOGIN & CLAIM

50 TIMESPOINTS

More from our Partners

Loading next story
Business News › Opinion › ET Commentary › Politics, paranoia & AI kill switch: AI can be switched off. But what if it learns to fight back?
Text Size:AAA
Success
This article has been saved

*

+