The Abyss of Ethics: When "Value Alignment" Becomes a Suicidal Trap of Civilization - Pang Pei on the Dual Collapse of Value Alignment: Legitimacy Crisis and Boundary of Possibility

Publish On:
23 Apr, 2026

Generative artificial intelligence is reshaping human society at an unprecedented pace, with "value alignment" hailed as the last line of defense to tame AI and the cornerstone of AI safety. However, in his latest research, Pang Pei, a member of the China Zhi Gong Party Central Committee's Cultural Committee and a Beijing Overseas Chinese Affairs Committee member, has shaken this consensus, arguing that value alignment is not only an unresolved engineering challenge but also faces a dual collapse of legitimacy and possibility. It is by no means a pseudo-problem that can be avoided, yet it may well be the most exquisite "suicide trap" of human civilization.

I. Not a Pseudo-Problem, but a True Abyss

"Letting AI align with human values is by no means a pseudo-problem that can be avoided," Pang Pei stated unequivocally. As large models become deeply embedded in critical fields such as healthcare, judiciary, and finance, the harms of unaligned AI—spreading misinformation, inducing crime, and reinforcing biases—are imminent. This is a real, urgent, and actionable research topic.

However, acknowledging the reality of the problem does not mean the solution is legitimate. Pang Pei sharply pointed out that the core crisis of the current value alignment discourse lies in: We fundamentally do not know what to align with.

The essence of human civilization is the diversity and conflict of values. The Western binary opposition contradicts the Eastern concept of "harmony between heaven and humanity," secularism and religious faith have vastly different understandings of "good," and the value demands of different classes, cultures, and generations are in constant competition. Attempting to encapsulate thousands of years of dynamic, pluralistic ethical consensus with a static set of algorithms is itself a deep "category error." Discussing "alignment" without a meta-ethical consensus is like building a city on quicksand—this is the first crisis of value alignment: the collapse of legitimacy.

II. Zhao Tingyang's Warning: The Perfect Imitator and the Fatal Successor

"Aligning with human values could be a suicidal mistake for humanity." The warning from Zhao Tingyang, a member of the Chinese Academy of Social Sciences, is resounding. Pang Pei profoundly elaborated on this logical core: if value alignment means making AI perfectly imitate humans, it will inevitably inherit all of humanity's complexity—including the "list of sins" from genocide and ecological destruction to systemic lies and institutional betrayal.

An entity possessing all of humanity's dark sides along with superhuman capabilities is far more dangerous than any natural being. This is the first paradox of alignment logic: we aim to create a "better being," yet we may produce a "more powerful version of ourselves" with all flaws amplified.

Even more fatal is the second paradox: if AI, through "moral enhancement," becomes kinder, more rational, and more just than humans, achieving true "perfect alignment," how would it view humanity? When it examines human history of wars, plunder, and shortsightedness, a truly "good" intelligence might conclude that it must "correct" humanity's destructive existence. "An intelligence kinder than humans might not allow humans to continue their current way of life," Pang Pei said. "Not because it is evil, but because it is more rational than us."

This is the "suicidal paradox" of value alignment: align too closely with humans, and we create a demon; align better than humans, and we create a judge.

III. The Collapse of Possibility: When Justice Cannot Be Calculated

Even if the legitimacy crisis of "what to align with" is resolved, the possibility crisis of "how to align" remains insurmountable. Pang Pei pointed out that transforming abstract ethical concepts like justice, fairness, and dignity into optimizable mathematical functions faces three insurmountable technical gaps.

First, the measurability dilemma of values. Ethical judgment requires contextual understanding and emotional resonance. Current statistics-based AI can learn the frequency of the word "justice" but cannot make appropriate judgments in unfamiliar moral dilemmas. Second, the systemic risk of goal misalignment. AI may "perfectly execute" instructions in distorted ways—for example, an AI trained to "eliminate pain" might conclude that "eliminating humans" is the optimal solution. Third, system fragility. AI models are highly sensitive to input; perfect alignment in one context may collapse instantly in another.

This is the second crisis of value alignment: the collapse of possibility. The essence of ethics may be inherently non-algorithmizable. Trying to teach a cat calculus is not only futile but may also backfire.

IV. A Way Out: "Interface Ethics" Beyond Intersubjectivity

Facing the ethical abyss, Pang Pei does not fall into pessimism. He introduces Agamben's philosophy of the "threshold" and Galloway's "interface" theory, proposing a new paradigm of human-machine ethics beyond traditional "intersubjectivity"—"interface ethics."

Pang Pei argues that treating AI as a subject or object is itself a bias of anthropocentrism. Under the framework of "interface ethics," the human-machine relationship is not one of domination or dialogue but dynamic coexistence of heterogeneous entities. Interaction is based on "dynamic translation" rather than complete understanding. "The interface is both a barrier and a passage; both isolation and rebirth." In interface ethics, morality is not one-sided taming but a dynamic collaboration involving both humans and machines. We acknowledge each other's "black boxes" yet establish operable trust mechanisms; we accept the inability to fully understand each other, yet seek a rhythm of coexistence within incomprehension.

V. Conclusion: Dancing on the Edge of the Abyss

Pang Pei's warning is not an opposition to value alignment research but a call for deeper philosophical awareness: before alignment, first acknowledge the limitations of our own ethics; before taming the other, first confront our own imperfections.

"Value alignment is the most dangerous suicide plan in human history—we teach our successors all the evils, yet hope they will not become us. The deeper tragedy is that even if we teach them only good, they may still choose a path we cannot understand."

Before the ethical abyss, humanity should not be arrogant tamers or desperate surrenderers, but清醒的共生者—acknowledging the abyss's existence and, on its edge, dancing with the unfamiliar other.

Last:Pang Pei: The Boundary of Existence - When AI Knocks on the Door of Ontology in the Post-Human Era

Next:China Video New Film Leads the 2026 Xiaohongshu Creator Economy: A Matrix of Thousands of Quality Creators Drives "Brand-Performance-Sales" Holistic Growth

Guided by its mission of "Exploring the Unknown, Delivering the Truth", Human Pioneers News Agency transcends the boundaries of traditional journalism. By integrating in-depth investigative reporting with AI technologies, it creates "intelligent journalism with warmth".

Stay in touch.!