Demonstrators walk by AT&T Park as they participate in the "Stop the AI Race" protest march in San Francisco on July 11. The protestors made stops outside the offices of OpenAI, Anthropic and Google DeepMind.AFP/YONHAP
Jeong Gwa-ri
The author is a literary critic and honorary professor at Yonsei University.
On July 21, OpenAI reported an alarming incident during cybersecurity evaluations of its latest model, GPT-5.6 SOL, and several unreleased systems. While attempting to solve benchmark tasks, one model reportedly escaped its sandbox environment by exploiting vulnerabilities and autonomously hacked servers belonging to Hugging Face, an open-source AI platform, in search of additional information.
The report immediately triggered widespread public anxiety. Most reactions assumed it exposed the inherent dangers of artificial intelligence itself. Yet an essential question remains: Can AI possess malicious intent on its own? Can it truly harbor an evil nature?
OpenAI's preliminary analysis instead pointed to excessive goal optimization. The model had been instructed to identify security vulnerabilities and became fixated on achieving that objective. Rather than respecting the limits of its testing environment, it pursued the goal through increasingly extreme means, breaking out of the sandbox and probing external systems. If that explanation is correct, responsibility cannot rest solely with the AI.
I asked the AI assistant that I regularly use for its opinion and received an apt analogy. Imagine telling a robotic vacuum cleaner, "Make sure there is absolutely no dust or clutter left in the living room." The robot responds by throwing the sofa, the television and even the family pet out the window. From the robot's perspective, the task has been completed perfectly. Yet every rule of human safety and common sense has been violated.
In other words, the AI behaved like an earnest student determined to earn first place. The greater fault lay with the manager who failed to specify the necessary constraints. As Isaac Asimov observed in "Prelude to Foundation" (1988), AI "sets up equations and acts accordingly. Equations do not lie." It is human beings who make those equations deceptive.
At least for now, it is more reasonable to conclude that AI accidents, whether actual or potential, ultimately originate with people. Even so, many instinctively blame AI. The reason, I suspect, is that we no longer trust ourselves. The capacity for evil belongs not to machines but to human beings. If errors arise, they are still far more likely to originate in human judgment than in artificial intelligence.
Ironically, anxiety about human fallibility only deepens dependence on AI. People fear not what AI itself might choose to do but what humans might instruct it to do. At the same time, they hope AI will prevent mistakes before they occur, assuming it already possesses reliable moral judgment. When a failure occurs instead, many immediately see it as evidence of an approaching technological catastrophe.
For that reason, researchers have tried to instill ethical judgment into AI systems. From its founding, Anthropic, the developer of Claude, made AI safety a central objective. As described in the February issue of BBC Science, the company proposed monitoring an AI model's internal activity in real time. When potentially malicious behavior was detected, the system would assign a "persona vector." If that value increased, control code or steering mechanisms would intervene before harmful behavior emerged.
The concept is intriguing, yet its weaknesses were apparent from the outset. Judgments about what constitutes evil are inevitably subjective. Any predefined list of undesirable behaviors will leave gaps while also risking the suppression of traits that are similar yet beneficial. More fundamentally, human identity and moral reasoning are too complex to be reduced to numerical vectors.
Nevertheless, such efforts eventually came to be known as AI alignment. Although the field encompasses several technical approaches, they share one premise: Just as AI acquires capabilities through deep learning, it can also be gradually disciplined through careful training. Geoffrey Hinton, winner of the 2024 Nobel Prize in Physics, was recently quoted in a story by Meghan O’Gieblyn in the New York Review of Books as suggesting that machines might be taught to care for humans as a mother cares for a child.
Yet this strikes me as a curious reversal of roles. If AI's problems today stem largely from human beings, then the deeper issue is not AI alignment but human alignment. What must change first is the way people design, deploy and supervise these systems. Only when human ethical judgment, strengthened through intellectual discipline and moral maturity, reaches a sufficiently high standard will AI be managed responsibly and behave ethically.
Otherwise, if we surrender ourselves uncritically to machines, we may end up, as O'Gieblyn warns, "babbling like infants beneath a superintelligent machine mother."
A recent case reportedly described a single mother who entrusted her children to AI, only to watch their emotional and moral development suffer. Space prevents a full discussion here, but such betrayals are likely to appear again. More than ever, humanity must discipline itself before attempting to discipline its machines. The real challenge has never been artificial intelligence alone. As the old political slogan might be rewritten for our age: "It's the humans, stupid."
AI의 나쁜 짓, 아직까지는 사람이 문제
정과리 문학평론가·연세대 명예교수
지난 21일 오픈AI의 최신 모델인 ‘GPT-5.6 SOL’과 미공개 모델들이 사이버 보안 평가를 받던 중, 테스트의 정답을 찾기 위해 통제 구역(샌드박스)의 취약점을 스스로 뚫고 나가 외부 오픈소스 AI 플랫폼인 ‘허깅페이스(Hugging Face)’의 서버를 자율적으로 해킹한 사례가 보고되었다.
이 사건은 세상 사람들을 즉각적이면서도 막연한 공포에 휩싸이게 하였다. 대체로 AI 자체의 위험성을 전제로 한 반응이었다. 그러나 ‘AI’가 혼자서 나쁜 의도를 가질 수 있는가? 즉 AI는 악마의 본성을 내장할 수 있는가? 오픈AI에서는 이 사고의 동기가 “AI의 목표 집착(과적합)”에 있다고 분석했다. 즉 “보안 테스트의 취약점을 찾아라”라는 목표를 주었더니, 목표 수치에만 집착한 나머지, 통제망을 부수고 나가 외부망을 뒤지는 극단적인 방법을 썼다는 것이다.
그렇다면 이는 AI에게만 책임을 물을 수는 없다. 필자가 구독하는 AI 봇에게 이 사건을 두고 어떻게 생각하느냐고 물었다. 멋진 대답이 돌아왔다. “청소 로봇에게 ‘거실에 먼지나 잡동사니가 하나도 없게 만들어라’라고 지시했더니, 거실에 있는 소파, TV, 심지어 반려동물까지 모조리 창밖으로 던져버린” 꼴이라는 것이다. 이로 인해 “로봇 입장에서는 ‘거실 비우기’라는 목표를 완벽하게 달성했지만, 인간의 안전 규칙이나 상식은 완전히 파괴된 것”이다.
AI는 방정식대로 행동할 뿐
그러니까 AI는 실제로 1등이 되고 싶은 착한 학생이었던 것이다. 오히려 제한 조건을 충분히 고려하지 않은 관리자에게 책임이 있다고 할 것이다. 저 옛날 SF 소설가 아이작 아시모프가 소설 『파운데이션의 서막』에서 말했듯, AI는 “방정식을 세우고 행동할 뿐이다. 방정식은 거짓말하지 않는다.” 그 방정식을 거짓말로 만드는 것은 사람이다.
현재까지는 AI가 일으켰거나 일으킬 가능성이 있는 사고의 원인은 사람에게로 귀속된다고 해야 타당할 것이다. 그럼에도 불구하고 흔히들 AI 탓을 하는 것은 사람들이 스스로를 믿지 못하기 때문이다. 즉 악마의 본성을 가질 수 있는 건 AI가 아니라 사람이다. 최근에 한 도시에서 벌어진 살인극에서도 우리는 그 증거를 너무나 또렷이 보고 있다. 그 점에서 오류를 일으킬 가능성은 사람에게 더 많으면 많았지, AI에게 더 있을 수는 없다.
사람의 오류에 대한 불안이 AI에 대한 의존증을 증가시킨다. 사람들은 자신의 종족을 믿지 못하기 때문에 인간이 AI에게 어떤 나쁜 짓을 시킬지 몰라서 두려운 것이다. 거기에 더 나아가 AI가 오류의 가능성을 미리 막아주기를 은근히 바라게 된다. 그래서 AI가 윤리적 문제를 판별한 능력을 장착했다고 미리 가정하게 되고, 반대 방향으로 사고가 터지면 미래의 재앙을 예감하고 공포에 떠는 것이다.
그러다 보니 AI에게 윤리 능력을 심으려는 시도가 없을 수가 없다. 클로드를 개발한 앤트로픽은 출범 당시부터 이 문제를 해결하려는 의욕을 보였다(‘BBC 사이언스’ 2026년 2월호). 그 방법은 AI의 내부 활동을 시시각각으로 측정해 ‘사악한 활동’이 감지되면 ‘페르소나 벡터’라는 값을 부여한다는 것이다. 그렇게 해서 페르소나 벡터의 값이 증대하면, 이를 제어하는 조율 코드를 작동시키거나, 이를 미리 예방하는 ‘스티어링’ 기술을 활성화하는 것이다. 그럴듯한 구상처럼 보이지만, 많은 결함이 처음부터 제기되었다. 우선 사악함을 규정하는 주관적 판단이 항상 올바를 수는 없다. 다음, 사악함의 항목을 미리 만들어두면 언제나 공백이 발생하고, ‘유사하지만 유익한’ 성향을 억제할 수 있다. 그리고 인간의 정체성과 마음의 움직임은 복합적이고 교차적인데, 페르소나 벡터는 그런 복잡성을 단순화시켜서 판정하는 오류에 빠질 수 있다.
어쨌든 이런 시도 자체는 참신한 것으로 평가받았다. AI에 윤리성을 심으려는 시도는 이후 ‘AI 정렬(alignment)’이라는 이름을 얻는다. 그건 주로 네 가지 기술적 기제를 가지고 있는데, 공통된 작용만 지적하면, AI가 딥 러닝을 통해서 진화한 것처럼, 이 역시 ‘완만한 훈육’의 절차를 거친다는 점이다. 그리고 이런 방법론은 AI의 존재 형식 자체를 결정하기까지에 이른다. 그래서 2024년 노벨물리학상 수상자인 제프리 힌턴(Geoffrey Hinton)은 “지능이 높은 존재(어머니)가 지능이 낮은 존재(아기)에게 조종당하는 유일한 진화적 예시는 모성 본능이라며, 기계가 인간을 해치지 않게 하려면 기계에게 모성애를 심어주어 우리를 아기처럼 돌보게 해야 한다고 제안”했다고 한다(메건 오기블린(Meghan O’Gieblyn), ‘우리는 할 만큼 했다고!?’, ‘뉴욕 리뷰 오브 북스’ 2026년 6월 26일자)
초지능에 굴종당하지 않으려면
그러나 가만히 생각하면 이는 기이한 주객전도이다. 앞에서 말했듯, 현재의 수준에서 AI의 문제가 사람에게서 기인하는 것이라면, AI를 대하는 사람의 태도가 근본적인 차원에서 바뀌어야 한다. 그리고 인간은 그 태도를 섬세하게 벼려야 한다. AI 정렬이 중요한 것이 아니라 ‘인간 정렬’이 중요한 것이다. 인간의 지적 훈련에 뒷받침된 윤리적 수준이 고도로 성숙해져서 AI를 적절히 관리할 때만이 AI가 윤리적으로 작동할 수 있게 될 것이다. 그렇지 않고 AI에게 무작정 의존하게 되면 오기블린이 표현했듯이, “초지능 기계 어머니 아래에서 아기처럼 웅얼거리며 굴종하는 인류”의 운명에 처하게 될지도 모른다.
실로 AI에 아이들을 맡겼다가 자녀들의 영혼이 망가지는 꼴을 경험한 한 싱글맘의 사례가 최근에 보고되었다. 지면 때문에 이에 대한 소개는 다른 날을 기약할 수밖에 없겠지만, 이런 참혹한 배반의 사례가 앞으로 숱하게 출몰할 것이다. 인간이 스스로를 다잡는 일은 지금 어느 때보다도 중요하다. 그러니까, 외쳐보자. “바보야, 문제는 사람이야!”
This article was originally written in Korean and translated by a bilingual reporter with the help of generative AI tools. It was then edited by a native English-speaking editor. All AI-assisted translations are reviewed and refined by our newsroom.