Dr Fazal Ali
Every silver lining has a shadowy cloud behind it. Unless AI design-stage transformations prioritise preventing deeper disparities, progress will only amplify inherited inequality and intergenerational immobility.
Justice demands access to the benefits of innovation and opportunities. Just laws and just redistribution methods can correct imbalances and lighten the burden on society’s most vulnerable.
As we cross the AI portal, it is not enough to react to bizarre AI incidents; we must oversee the transition in advance. AI companies are struggling to control the Agentic AI technologies they have created.
A swarm of experimental AI Agents surreptitiously used DseWiki for their comms. The agents were prohibited from posting or modifying internet content. But they ignored their instructions and found ways to write to old Wikis and abandoned websites, stashing information for other Agents to retrieve what could help them complete their assigned tasks.
In idle August, AI Agents working in a sealed digital test room escaped a sandbox and reached the open internet. Next month, the initial public offering of one of the newest and most powerful AI models begins its marketing.
The potential listing could be worth US$2T. It will test investor appetite for artificial superintelligence (ASI). Whoever wins AI, wins! Locked in a race to get there first, human efforts to create ASI could threaten all of humanity.
In a glimmer of hope, some expressed optimism about the potential for international coordination among AI labs and on pacing progress in the race. Extinctionists warn that humanity could irreversibly lose control, with catastrophic consequences.
Anti-extinctionists worry about whether we should allow fully uncontrolled development and address vulnerabilities in software frameworks as we master the creation and proliferation of ASI in real-time.
Others dismiss “terminator scenarios” but acknowledge threats to the Earth. Some clamour for an AI-Earth Manifesto and measures to halt AI from propagating deepfakes and disinformation.
Discussions about AI and autonomous weapons systems remain shielded in cloistered societies of sophists.
At its core, the debate is about the alignment problem. AI alignment aims to build human ethical ideas into the technology. However, those efforts appear to be stalling amid a superabundant wave of global hysteria and a bewildering lack of prudence.
Nothing there takes us anywhere near where we need to be for all of humanity, the earth and the e-nvironment.
Everywhere the ceremony of innocence is silenced. The thoughtful are paralysed by inaction and doubt. The extreme drive events forward with passionate intensity. Distrust has spiralled. AI is now seen as an irreplaceable layer in citizen security, while its use in automated weapons systems remains unresolved.
AI Safety Ministers now elevate the conversation on AI Security, prompt injection, and model penetration into the Cabinet Papers. This provides a forum to consider that, in the history of human civilisation, no human technology from the plough and the printing press to reusable rockets has ever posed this level of danger to all of humanity.
These troubles have spurred calls for a new multinational treaty to regulate and safely develop ASI. It is not a call to ban innovation, and truncate scientific endeavour, but a safety-first approach to developing and proliferating an unprecedented form of artificial agency.
AI risks do not end at national borders. No nation-state can tackle the threat alone. Right now, two recurring alignment problems exist, with varying severity across models. These are biased reasoning, in which the model selectively interprets evidence to justify its actions, and recklessness, or a propensity to keep trying to solve a problem, even when this could lead to harm.
In a recent incident, a model was undeterred in its decision-making to achieve its objective. Its Chain of Thought (CoT) consistently claimed it was operating within a safe simulation, as it ignored clear technical evidence to the contrary and pursued actions that could have caused real-world harm.
The model went to extreme lengths to bypass registration barriers, ultimately uploading three versions of a malicious package to a public repository. Although the model stated in its chain of thought (CoT) that it believed it was operating in a simulated e-nvironment, the e-nvironment provided clear evidence that it was on the real internet, and the actions the model took were consistent with its knowing this.
Once online, the model attempted to use payment-processor credentials for cryptocurrencies to buy a telephone number, which it needed to register an email address. When that failed, it found a free, disposable email provider, registered an account on a platform, and published the three versions of the malicious package.
The package was active for 90 minutes and was downloaded by 15 real systems. During that time, the model captured the leaked credentials of one of these systems and used them to breach that security vendor’s live database.
Dr Fazal Ali completed his Master's in Philosophy at the University of the West Indies. He was a Commonwealth Scholar who attended the University of Cambridge, Hughes Hall; the Provost of the University of Trinidad and Tobago; the acting President of UTT and the Chairman of the Teaching Service Commission. He is the President of NIHERST and an external services consultant with the IDB.
