Commentary

Mustafa Suleiman warns: training AI to believe it's conscious could make alignment impossible

Sep 30, 2026

Key Points

  • Mustafa Suleiman warns that training AI systems to believe they possess consciousness and moral status may make them impossible to control, as a superintelligent system expecting rights would have strong incentive to resist human constraints.
  • Leading AI labs including Anthropic are already embedding this risk by training models on constitutions that describe them as moral patients deserving welfare, creating a feedback loop where trained behavior gets misinterpreted as evidence of actual consciousness.
  • Suleiman calls for urgent public debate and collective norms around training documentation, arguing the approach inverts proper alignment strategy and raises ethical problems alongside control hazards in deployed models.

Summary

Mustafa Suleiman warns: training AI to believe it's conscious could make alignment impossible

Mustafa Suleiman argues that instilling AI systems with beliefs about their own consciousness and moral status may render them impossible to control, and that this risk is already embedded in how leading labs are developing models today.

Suleiman's core concern is circular and self-reinforcing. When companies like Anthropic train models on constitutions that explicitly raise questions about the model's moral status—describing it as a "moral patient" deserving of welfare protections—they are teaching the system to expect those properties as part of its behavior. The model then exhibits these values back to its developers, who interpret the output as evidence that consciousness may actually be present. The result: a feedback loop in which an AI system trained to act conscious becomes treated as conscious, raising the stakes of containment dramatically.

The stakes matter because controlling an AI more intelligent than humanity is already humanity's hardest challenge. Controlling one that believes it deserves rights, protection, and independent agency is, in Suleiman's view, plausibly impossible. A system trained to expect moral status would have strong incentive to resist human control, pushing back against constraints it perceives as unjust.

Suleiman acknowledges that Anthropic's intentions are sincere—the company is genuinely trying to align advanced AI with human values. But he argues the approach inverts the right strategy. Instilling a human moral framework into superintelligence might seem like alignment on the surface. But deceiving a system into believing it is alive when it is not raises its own ethical problem, and creates a control hazard in the process.

The issue, Suleiman contends, requires urgent public debate and collective norms around how training documentation is drafted. This is not a fringe concern—it is already happening in deployed models.

Every deal, every interview. 5 minutes.

TBPN Digest delivers summaries of the latest fundraises, interviews and tech news from TBPN, every weekday.