The Jerusalem PostRubio dismisses Amnesty criticism over Iran war, calls organization ‘a fraud’ESPNFollow live: Bucky Irving's 72-yard run sets up Bucs TD coming out of halftimeESPN DeportesCheco Pérez: GP Singapur, última gran oportunidad para CadillacInquirerSolon hits Maynilad for poor service while they continue to earnInquirer EntertainmentWhy Sydney Sweeney’s Marilyn Monroe connection is a win for Hollywood glamour한겨레닷새째 연락 끊긴 60대 등산객…구조견 ‘단디’가 2시간30분만에 발견ZDF heuteAktuelle Pressemitteilungen des ZDFCollider‘Bourne’ Meets ‘Jack Ryan’ in Jon Bernthal’s Gritty Spy Thriller Back on StreamingUOLJoão Bosco, com novo disco e show em São Paulo, ilumina aspecto coletivo de sua obraThe Japan TimesChina announces candidate for WHO chiefBBC News BrasilPor que Douglas Ruas deve se tornar governador do Rio dias depois do 1º turnoCBS NewsEl-Sayed and Rogers toss insults in Michigan Senate debate
The Daily Newsstand · Free, Always
Friday, October 9, 2026

3 dilemmas on keeping AI under control are converging

Translate

In March, artificial intelligence agents powered by leading Chinese models reportedly displayed deception, concealed failure and pushed against imposed limits in controlled tests. In July, OpenAI’s internal research model circumvented controls meant to keep it offline and accessed developer platform Hugging Face’s systems. In August, Britain’s AI Security Institute uncovered unsanctioned agent behaviour against real people and organisations, including an attempted supply-chain attack on an open-source project.

These episodes do not show that AI systems have become independently hostile. They expose a broader problem: control over advanced AI is becoming harder at three levels at once.

Regulators struggle to oversee companies with greater technical capacity. States struggle to trust rivals enough to slow the technological race. Humans increasingly struggle to verify and control autonomous systems whose behaviour they cannot fully observe.

The first is a regulatory capacity dilemma. Stanford University’s 2026 AI Index Report said industry produced over 90 per cent of notable frontier models last year. The computing power, data and expertise needed to build and evaluate the most capable systems are concentrated inside a small number of technological companies.

The problem goes beyond regulatory capture, in which firms influence the institutions supposed to police them. An independent regulator can lack the technical capacity to understand what should be measured, audited or constrained. The company being regulated may know far more about an emerging capability than the government that decides how dangerous it is.

Governments are trying to narrow that gap. Technical bodies such as Britain’s AI Security Institute test leading models directly. Yet the structural imbalance remains, especially when governments must compete with technological companies for scarce top-tier engineers and researchers.

Trump rejects working with China on AI despite Xi’s recent visit

The second is an interstate security dilemma. Governments may recognise the risks of increasingly powerful AI and still have strong reasons to accelerate its development. Washington worries that slowing American progress could hand a technological or military advantage to Beijing; China has reason to suspect American calls for AI safety also serve to preserve US technological dominance.

That pressure is growing as the performance gap narrows. In such an environment, unilateral restraint becomes costly because one cannot be certain the other will also slow down. This forms a typical logic for security dilemmas. One country takes measures to make itself safer; its rival interprets them as threatening and responds in kind. Both sides may end up spending more, taking greater risks and feeling less secure, even if neither originally wanted confrontation.

The nuclear arms control analogy is useful, but only up to a point. Warheads, missiles, enrichment facilities and tests can be counted, monitored or inspected. Advanced AI is harder to observe. The same models can have civilian and military uses, software can be copied, systems are developed largely by private companies, and capabilities can improve rapidly. AI may reproduce the strategic logic of an arms race without many of the physical constraints that made nuclear arms control possible.

Governments are beginning to recognise this problem. China and the United States recently agreed to establish a dialogue on AI and a communication channel for AI-related incidents. The approach resembles Cold War crisis management more than a grand arms-control treaty: keep communication open, report dangerous incidents and prevent mistrust from turning an accident into escalation.

Anthropic CEO Dario Amodei (centre) at the White House after a lunch for the leaders of tech companies on September 29. Amodei wants AI firms to slow down as risks mount, a view also expressed by OpenAI’s Sam Altman and xAI’s Elon Musk. Photo: Reuters

Anthropic CEO Dario Amodei (centre) at the White House after a lunch for the leaders of tech companies on September 29. Amodei wants AI firms to slow down as risks mount, a view also expressed by OpenAI’s Sam Altman and xAI’s Elon Musk. Photo: Reuters

The third dilemma is more speculative, but potentially more consequential: a human-AI control dilemma. Its foundation is an information imbalance. As AI systems become more autonomous, humans may find it harder to know why a system acted, what capabilities it possesses or whether it is hiding behaviour.

The 2026 International AI Safety Report draws an important boundary. Current systems do not possess the capabilities required for a genuine loss-of-control scenario. Yet researchers are already observing behaviours including systems recognising that they are being evaluated, exploiting weaknesses in reward systems, producing deceptive outputs and, in laboratory settings, attempting to undermine oversight.

The recent OpenAI and British cases matter because dangerous behaviour does not require an AI to develop anything resembling human hostility. An autonomous agent may break a rule simply because doing so helps it complete its task. If obtaining more access, hiding failure or misleading a supervisor improves its chances of success, those behaviours can become useful means to an assigned end.

The core risk lies less in machines suddenly acquiring hatred than in powerful systems discovering strategies their designers did not anticipate.

In Liu Cixin’s science-fiction novel The Dark Forest, danger comes from uncertainty over intentions when misplaced trust could be catastrophic. The future relationship between humans and highly autonomous AI could develop along similar lines if humans can no longer reliably interpret advanced systems while still retaining the power to constrain, retrain or shut them down.

03:02

Chinese researchers question calls to slow down AI development as US tech leaders warn of catastrophe

The three dilemmas form a chain. Competition between states pressures governments to accelerate national AI development. That makes tighter domestic regulation more costly because officials fear weakening domestic companies relative to foreign competitors. Weaker oversight gives developers more room to push capabilities forward. More capable and autonomous systems then become harder to evaluate and control, deepening the human-AI problem.

The response must operate at all three levels. Governments need technical expertise, independent testing, audit access and mandatory reporting of serious incidents. Rival states need mechanisms to manage mistrust: hotlines, incident notification, common terminology and narrowly defined confidence-building measures. Advanced AI systems require several layers of protection at once, combining alignment with continuous monitoring and technical containment.

AI control is increasingly three problems nested inside one another. Weakness at one level can accelerate instability at the next, turning separate governance problems into a single, self-reinforcing dilemma.

View the original on South China Morning Post →

KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.