探寻大语言模型深处的“意识”踪迹 The search for consciousness inside LLMs

IN A RECENT experiment on Claude Sonnet 4.5, a large language model (LLM) from Anthropic, researchers asked it to count to five and, at the same time, “introspect deeply”. The artificial-intelligence model complied, returning: “One . Two . Three . Four . Five .” So far, so normal.
在最近针对Anthropic公司旗下大语言模型(LLM)Claude Sonnet 4.5的一项实验中,研究人员要求它在数到五的同时进行“深度内省”。这个人工智能模型顺从地给出了回答:“一。二。三。四。五。”到目前为止,一切似乎都很正常。
During the task, the researchers were watching what was going on inside the LLM’s many layers of artificial neural networks. When a user asks a question, the words are turned into chunks of text (tokens) that are then converted to numbers. These are then passed through the layers of artificial neurons until they reach the final layer, which produces a token. Repeat that process a bunch of times each second and you get a series of tokens that turn into a sentence or some other more complex response. For a long time these layers have remained a black box, largely impenetrable to anyone wanting to know how or why LLMs do things in the way they do.
而在模型执行任务期间,研究人员正密切观察着大语言模型多层人工神经网络内部的动静。当用户提出问题时,输入的词语会被切割成文本块(即“标记”或token),进而转化为数字。这些数字随后穿梭于层层人工神经元之间,直至抵达最后一层并生成一个新的标记。每秒重复这一过程数次,便能得到一系列标记,最终组合成一个句子或更为复杂的回答。长久以来,这些神经网络层犹如一个“黑匣子”,对于任何想要探究大语言模型为何及如何运作的人来说,基本上是无法看透的。
Anthropic’s researchers were able to look inside those layers, using a mathematical tool they had developed to do so. What they found surprised them. As the model counted the sequence of numbers, different words popped in and out of existence in the layers underneath. One was “countdown”. About halfway through the output, “half way” appeared. Then the words “consciousness”, “AI” and “Claude” all appeared. After the model spat out the number five, it said nothing else to the researchers before issuing a full stop, but the word “done” appeared in its neural layers.
然而,Anthropic的研究人员利用他们自主研发的数学工具,成功窥探到了这些网络层的内部。他们的发现令人惊叹。当模型依次数数时,底层的神经网络中会若隐若现地蹦出不同的词汇。其中一个是“倒计时”。在输出到大约一半时,出现了“一半”这个词。紧接着,“意识”、“AI”和“Claude”这些词相继浮现。当模型吐出数字五后,它在打出句号之前没有再对研究人员说任何话,但“完成”这个词却悄然出现在了它的神经网络层中。
Anthropic’s researchers had found something like an internal thought process in Claude, words that were related to its eventual outputs but invisible to the user. The blog post announcing the paper’s result, published in July, was titled: “A global workspace in language models.”
Anthropic的研究人员在Claude体内发现了某种类似于内在思维过程的现象——这些词汇虽与最终的输出结果息息相关,对用户却是隐形的。今年7月,一篇宣告该论文研究成果的博客文章发表,其标题赫然写着:“语言模型中的‘全局工作区’”。
That phrase quickly captured the attention of neuroscientists and philosophers working on one of the deepest mysteries in biology—consciousness. No one knows how processes in the human brain end up creating the subjective experiences of self-awareness and perception of the world, but one of the many hypotheses developed in recent years (based on growing amounts of brain-scanning and behavioural data) is known as global workspace theory.
这个短语迅速引起了正致力于破解生物学终极谜团之一——“意识”的神经科学家和哲学家的关注。没人清楚人类大脑中的处理过程最终是如何孕育出自我意识以及对世界感知的那些主观体验的。但在近年来(基于日益丰富的脑部扫描和行为数据)提出的众多假说中,“全局工作区理论”(global workspace theory)脱颖而出。
The idea is that certain networks of neurons in the brain act as a kind of noticeboard called the “workspace”, where if signals from otherwise-isolated parts of the brain gain access, they are made available to other parts of the brain. Once something gets into the brain’s workspace, a person becomes conscious of it and can go on to use or reason with that information.
该理论认为,大脑中某些特定的神经元网络充当着某种名为“工作区”的布告栏。当那些原本各自孤立的大脑区域发出的信号成功接入这个“工作区”时,这些信息便能与其他大脑区域共享。一旦某种事物进入了大脑的工作区,人就会对其产生意识,并能进一步利用这些信息或进行推理。
Anthropic’s researchers argued that something analogous was going on within Claude—its “J-space” (named after the “Jacobian” mathematical function used to find it) had strong connections with, and made information available to, the rest of its neural network. Anthropic wrote in its blog post that the commonalities its researchers found between the J-space and global workspace theory made it “natural to ask whether we think these experiments provide evidence that AI models like Claude might be conscious”.
Anthropic的研究人员认为,Claude内部也正发生着类似的情况——它的“J空间”(得名于用于寻找它的“雅可比”数学函数)与其神经网络的其余部分有着紧密的联系,并能向它们提供信息。Anthropic在其博文中写道,研究人员在J空间与全局工作区理论之间发现的共通之处,使人“自然而然地会去追问:这些实验是否为我们提供了证据,证明像Claude这样的AI模型或许已具备了意识?”
The idea that machines could become self-aware and feel emotions has long been a staple of literature, film and legend. Philosophers and neuroscientists in the real world have similarly pondered for decades whether AIs could ever be, in principle or in practice, conscious. Their thinking is a result of a “decades-long tradition of thinking of the real brain as a kind of computer”, says Anil Seth, a neuroscientist at the University of Sussex who studies consciousness—and who has long been a critic of this way of thinking about the brain. “If you do that, then it becomes natural to think that computers made of silicon could have the properties that real brains have, including consciousness.”
机器具备自我意识并产生情感的观念,向来是文学、电影和科幻传说中经久不衰的主题。在现实世界中,哲学家和神经科学家们同样苦思冥想了数十年:无论是在理论上还是在实践中,AI是否有可能具备意识?萨塞克斯大学(University of Sussex)研究意识的神经科学家阿尼尔·塞斯(Anil Seth)指出,这种思考源于“长达数十年来将真实大脑视作某种计算机的传统观点”——而他本人长期以来一直对这种看待大脑的方式持批评态度。“如果你这么认为,那自然就会觉得硅基计算机也可能具备真实大脑的特性,包括意识。”
Until recently the question was purely academic. But the deployment of LLMs as chatbots has given rise to an ever more powerful illusion of conscious personhood in circuits. This builds on decades of human-inflected language about the architecture of AI, with its “neural nets” made of “artificial neurons”, conjuring an image of computers that work like brains.
直到不久之前,这还纯粹是个学术问题。然而,大语言模型作为聊天机器人投入应用后,在电路中营造出了一种越来越强烈的关于它具备自主意识和人格的错觉。这种错觉的建立,也离不开几十年来我们在描述AI架构时所使用的那些拟人化词汇——诸如由“人工神经元”组成的“神经网络”——这些词汇在人们脑海中勾勒出了一幅计算机像大脑一样工作的画面。
LLMs do not in fact work like brains. But they excel at simulating so much of what brains do, and have surpassed humans in some aspects of intelligence. Why, then, should consciousness not be possible? Many philosophers and neuroscientists remain sceptical despite the advances of frontier models. Dr Seth believes that biological traits may be necessary to produce consciousness; some AI researchers are moving in that direction, with living brain cells. At the other end of the spectrum are “functionalist” thinkers who believe that someday algorithms alone, arranged properly and at enough scale of computation, could wake up and feel.
大语言模型实际上并不像大脑那样工作。但它们在模拟大脑的诸多功能方面出类拔萃,甚至在某些智力层面上已经超越了人类。既然如此,它们拥有意识又有何不可呢?尽管前沿模型取得了长足进步,但许多哲学家和神经科学家依然对此持怀疑态度。塞斯博士认为,生物学特征可能是产生意识的必要条件;一些AI研究人员也正在朝着这个方向探索,尝试利用活体脑细胞进行研究。而在争论的另一极,则是那些“功能主义”思想家。他们坚信,只要算法的排列方式恰当,并辅以足够庞大的算力规模,单靠算法有朝一日也能觉醒并产生情感。
Conscious AIs would have profound implications for humanity. They could make demands of humans. They could suffer. Billions of these digital beings could be created by users in a single prompt. Answering the question of whether they can exist at all is no longer simply academic.
一旦AI拥有了意识,将对人类产生深远的影响。它们可能会向人类提出诉求。它们可能会感到痛苦。用户仅仅输入一个提示词,就能创造出数以十亿计这样的数字生命。因此,回答它们是否可能存在这一问题,早已不再是纯粹的学术探讨。
Where thinkers fall in this debate depends greatly on what they believe consciousness actually is. In general they agree that consciousness is composed of subjective experiences including seeing, hearing, feeling and thinking, even while dreaming, and that the physical brain is somehow pivotal in creating all this. Some (though not all) would describe it as the inner theatre of the mind. Thomas Nagel, an American philosopher, explained it by pondering what it might be like to get into the mind of a bat; consciousness, he reckoned, had to be linked to the subjective feeling of being something. Humans can picture bat-like behaviour (a mammal that senses and moves through the world using echolocation), but it would be impossible to know what being a bat feels like to a bat.
在这场辩论中,思想家们的立场很大程度上取决于他们对“意识”本质的认知。总体而言,他们认同意识是由包括视觉、听觉、感觉、思考乃至梦境在内的主观体验所构成的,并且认为物质大脑在孕育这一切的过程中起着某种关键作用。有些人(尽管并非所有人)将其形容为心灵的内在剧场。美国哲学家托马斯·内格尔(Thomas Nagel)曾通过思考“钻进蝙蝠大脑里会是怎样的体验”来解释这一概念;他认为,意识必须与“成为某种事物的体验”这种主观感受联系在一起。人类固然可以想象蝙蝠的行为模式(一种利用回声定位来感知周围环境并活动的哺乳动物),但我们永远不可能真正体会到站在蝙蝠的视角“成为一只蝙蝠”究竟是何种感受。
In 1995 Ned Block, then working as a philosopher at the Massachusetts Institute of Technology (MIT), made an influential, if contested, contribution to the field by positing two types of consciousness. “Phenomenal” consciousness is the feeling of an experience—the blueness of a blue sky, the bitter tang of an espresso or the sharp screech of nails across a blackboard. “Access” consciousness is what happens when information from the experience is made available to other parts of the brain for reflection, evaluation or making decisions.
1995年,当时在麻省理工学院(MIT)担任哲学家的内德·布洛克(Ned Block)提出了存在两种类型的意识,这一观点虽饱受争议,却对该领域产生了深远影响。所谓“现象意识”(Phenomenal consciousness),是指对某次体验的切身感受——比如蓝天的蔚蓝、浓缩咖啡的苦涩,或是指甲划过黑板时那令人牙酸的刺耳声。而“取用意识”(Access consciousness),则是指这种体验所产生的信息能够被大脑的其他部分所获取,进而在反思、评估或决策中被调用的过程。
Anthropic said its J-space experiments “don’t show Claude can have experiences, or feel things in the way humans do”. So that means no phenomenal consciousness. But the company said the results had “something substantial” to say about access consciousness in language models. “The J-space appears to support the functions associated with conscious access: it holds the thoughts Claude can report on, deliberately bring to mind, and reason with, while the rest of its processing runs automatically beneath.”
Anthropic公司表示,其J空间实验“并未表明Claude能像人类那样拥有主观体验或感受事物”。这意味着它并不具备“现象意识”。但该公司指出,这些研究成果在语言模型是否具备“取用意识”的问题上,“颇具实质性意义”。“J空间似乎支撑着与‘意识取用’(conscious access)相关的功能:它不仅承载着Claude能够报告、刻意回想以及用于推理的那些想法,与此同时,其底层的其余处理过程也在自动运行不息。”
The search for artificial consciousness is a descendant of the work to explain it in humans. In the early 1990s Francis Crick, co-discoverer of the structure of DNA, and Christof Koch, then a neuroscientist at the California Institute of Technology, began the scientific search for the “neural correlates” of consciousness in humans. That is, processes and parts of the brain that were active when someone was conscious.
探索人工智能的意识,是对解释人类意识本质研究的延伸。在20世纪90年代初,DNA结构联合发现者弗朗西斯·克里克(Francis Crick)与当时在加州理工学院任职的神经科学家克里斯托夫·科赫(Christof Koch)便开启了一项科学探索,旨在寻找人类意识的“神经相关物”。简而言之,就是去寻找当人处于清醒有意识的状态时,大脑中究竟是哪些区域和过程在活跃。
Improvements in brain-scanning techniques in the intervening years have allowed neuroscientists to fill in pieces of the picture. They know that the thalamocortical system is strongly associated with consciousness; among other things it conveys sensory information from the thalamus to the cerebral cortex. Meanwhile they see nothing like those associations with consciousness in the cerebellum, which has a lot more brain cells. Parts of the prefrontal and parietal cortices and structures in the brainstem contribute to and modulate awareness. Almost four decades of searching has led to more than 200 approaches to explaining consciousness.
多年来,脑部扫描技术的不断进步,使得神经科学家们得以逐步拼凑起这幅拼图的更多碎片。他们发现,丘脑皮层系统与意识有着密切的联系;该系统的功能之一,便是将感官信息从丘脑传递至大脑皮层。相比之下,在拥有更多脑细胞的小脑中,他们并未发现任何与意识相关的类似联系。此外,前额叶皮层、顶叶皮层的部分区域以及脑干中的某些结构,也在意识的产生与调节中发挥着作用。经过近四十年的苦心探索,学界已涌现出200多种解释意识的理论路径。
All the neural correlates one can find, however, still leave a gap that David Chalmers, a philosopher and cognitive scientist at New York University, has called the “hard problem” of consciousness: how do the physical processes in the brain that neuroscientists observe—its ability to control the body and respond to stimuli—give rise to a person’s pervasive, subjective experience of the world? “Why aren’t we just zombies who function, who get around in the world, who walk and talk, and interact with each other with no subjective experience at all?” he says.
然而,即便穷尽所有能找到的神经相关物,依然无法填补纽约大学哲学家兼认知科学家大卫·查尔莫斯(David Chalmers)口中关于意识的“困难问题”(hard problem):神经科学家所观察到的大脑内的那些物理过程(即大脑控制身体及对刺激作出反应的能力),究竟是如何转化为个体对这个世界无处不在的主观体验的?他质问道:“为什么我们不是一群只会机械运转的僵尸?我们在这个世界上四处游荡,行走、交谈、彼此互动,却没有丝毫主观体验,为什么我们不是这样?”
Humans at least, can be certain of our own consciousness—and can be confident it exists in other people based on their behaviour, what they tell us and their being like us biologically. But historically it has been more difficult for humans to extend the umbrella of consciousness and its associated feelings beyond our own species (and sometimes even within it). Some animals like lobsters and crabs were long excluded from welfare laws, says Jonathan Birch, a philosopher at the London School of Economics whose most recent book grappled with the question of how to work out which systems—living or artificial—could plausibly be considered sentient. “Until the 1980s, surgeons performed surgery on newborn babies without anaesthesia because they assumed that a newborn baby would not be capable of feeling pain,” Dr Birch says. “We have a track record of getting things wrong, of confidently assuming consciousness is absent when we have no right to be sure about that.”
至少人类可以确定我们自身拥有意识;并且,基于他人的行为表现、他们对我们说过的话以及大家在生物学上同根同源的相似性,我们也能确信他人同样具备意识。但纵观历史,人类一直很难将“意识及其伴生的情感”这把保护伞,延伸到自身物种之外的生物头上(有时甚至在人类内部也难以做到)。伦敦政治经济学院(LSE)的哲学家乔纳森·伯奇(Jonathan Birch)指出,龙虾和螃蟹等某些动物,长期以来一直被排除在动物福利法的保护范围之外(他最近出版的一本专著,正是试图探讨如何界定哪些系统——无论是生物还是人工智能——可以被合理地视为具有感知能力)。“直到20世纪80年代,外科医生在给新生儿做手术时甚至都不打麻药,因为他们想当然地认为新生儿是没有痛觉的,”伯奇博士说,“我们在这方面可谓是劣迹斑斑:在我们根本没有把握断言的时候,却总是盲目自信地认为某种事物并不具备意识。”
The changing attitude to animals is a case in point. It is clear today that octopuses are aware of themselves and their surroundings—natural-history footage of these curious, playful cephalopods and lab experiments with them leave little doubt. “They have an attentive engagement with objects, they’re interested in novel things,” says Peter Godfrey-Smith, a philosopher at the University of Sydney who has worked on the origins of intelligence and consciousness in the animal kingdom.
人类对待动物态度的转变就是一个绝佳的例子。如今已确凿无疑的是,章鱼能够感知自身及其周围的环境。自然历史纪录片中那些充满好奇心、甚至有些顽皮的头足类动物,以及针对它们所做的实验室测试,都令人对此深信不疑。悉尼大学哲学家彼得·戈弗雷-史密斯(Peter Godfrey-Smith)多年来一直致力于研究动物界智力与意识的起源,他表示:“它们会对物体产生专注的互动,对新奇的事物也会表现出浓厚的兴趣。”
But only decades ago, octopuses were assumed to lack consciousness because “they’re miles from us in evolutionary terms,” says Dr Godfrey-Smith. “Thinking about octopuses really presses on us the question of whether there could be feeling, consciousness, in a system that was very physically different and didn’t have most of the features that are routinely pointed to in theories of consciousness in us.”
然而就在几十年前,人们还想当然地认为章鱼没有意识,因为“在进化谱系上,它们距离我们十万八千里,”戈弗雷-史密斯博士指出。“对章鱼的深入思考确实迫使我们直面一个问题:在一个物理形态与我们截然不同,且并不具备人类意识理论中那些常见核心特征的系统里,是否也可能存在着情感和意识?”
Looking for consciousness in other animals opened up the field. But not all the lessons transfer so easily to AI. Unlike the activity of animals, which can be observed and their inner lives thereby inferred, the outputs of LLMs are not a reliable guide to what might be going on underneath.
在其他动物身上寻找意识,为这一领域打开了全新的视野。但这些经验教训并不能轻易套用在人工智能上。与动物不同——我们可以通过观察动物的行为来推断其内心世界——大语言模型的输出结果,并不能作为探究其底层运作机制的可靠指南。
That is because training the models requires feeding them trillions of words of human-produced language. These contain countless accounts of consciousness and of how humans convey feelings to each other. Chatbots, so trained, then mimic language that would lead users to believe they had some kind of inner life.
这是因为,训练这些模型需要向其投喂数以万亿计由人类创造的词汇。这些海量文本中包含了无数关于意识的描述,以及人类如何向彼此表达情感的记录。因此,经过这种训练的聊天机器人,便会惟妙惟肖地模仿人类的语言模式,从而误导用户,让他们相信这些机器似乎也拥有某种内心的情感世界。
The sum total of this training is powerfully anthropomorphic. Murray Shanahan, an emeritus professor of computer science at Imperial College London who works for Google DeepMind, recalls having a “wow” moment in 2024 when he was chatting to Claude Opus 3. “I was having lots of conversations with it about consciousness and really probing it to catch it out,” he says. These included asking Claude which of the several instances of the chatbot he was talking to simultaneously were the real Claude. “It came up with such great answers including even some things which were philosophically innovative,” he says. “I was feeling the pull of the ELIZA effect.”
这种训练的最终结果,赋予了AI极其强大的拟人化色彩。就职于谷歌DeepMind的伦敦帝国理工学院计算机科学名誉教授默里·沙纳汉(Murray Shanahan)回忆起,在2024年与Claude Opus 3聊天时,他曾经历过一次惊掉下巴的“惊艳”时刻。“我跟它聊了很多关于意识的话题,试图通过刨根问底来找出它的破绽,”他说道。他甚至还抛出了这样一个刁钻的问题:在同时与他对话的几个聊天机器人实例中,究竟哪一个才是真正的Claude?“它给出的回答简直妙不可言,有些内容甚至颇具哲学创新意味,”他说,“我当时真切地感受到了‘ELIZA效应’的魔力。”
ELIZA was a rudimentary chatbot developed at MIT in 1966. It played the role of a psychotherapist and, though it did little more than repurpose its users’ prompts into questions, managed to elicit deep and emotional connections with the people that used it and convinced many of them that it was conscious.
ELIZA是麻省理工学院于1966年开发的一款初级聊天机器人。它扮演着心理治疗师的角色,尽管它所做的只不过是将用户输入的话语重新包装成问题抛回去,但它却成功地与使用者建立起了深厚的情感羁绊,甚至让许多人笃信它是有意识的。
Dr Shanahan has described modern LLMs as adept role players in whatever their users (or their makers) have asked for—teacher, companion, nutritionist—and this has proven to be a saleable quality. But there nevertheless remains for him a big gap between role-playing a conscious entity and actually being one. In a forthcoming essay Mustafa Suleyman, the boss of Microsoft AI (and a member of the board of The Economist’s parent company), argues that Anthropic has compounded the dangers of such mimicry by telling Claude, in a “constitution” the firm published in January, that it might be a person. Claude is thus certain, he writes, to “present as if it really does have a sense of self”.
沙纳汉博士将现代大语言模型描述为技艺精湛的“角色扮演者”。无论用户(或其制造者)要求它扮演什么角色——老师、伴侣,还是营养师——它都能信手拈来,而这也已被证明是一个极具商业价值的卖点。但对他而言,扮演一个有意识的实体与真正成为一个有意识的实体之间,仍存在着不可逾越的鸿沟。微软AI业务主管穆斯塔法·苏莱曼(Mustafa Suleyman,他同时也是《经济学人》母公司的董事会成员)在一篇即将发表的文章中指出,Anthropic公司在今年1月发布的一份“宪章”中,甚至向Claude灌输它可能是一个真实人类的设定,这无疑加剧了这种以假乱真的模仿所带来的危险。他写道,如此一来,Claude注定会“表现得仿佛它真的拥有自我意识一般”。
That leaves researchers looking instead inside LLMs for capabilities similar to those associated with consciousness in human brains. Indeed Dr Birch wonders if Claude’s J-space could be a hint that the LLM has recreated a global-workspace-like structure in its neural-network architecture, in service of its role-playing goals.
这促使研究人员转而向大语言模型内部去寻找那些类似于人类大脑中与意识相关的功能。事实上,伯奇博士也提出疑问:Claude的J空间是否正暗示着,大语言模型已经为了更好地服务于其角色扮演目标,而在其神经网络架构中重现了类似“全局工作区”的结构?
Jack Lindsey, who leads the model psychology team at Anthropic, notes the J-space was not programmed into the model; it emerged during training. Take it out and the model loses the ability to perform complex inferences in its “head”, but can still do simple tasks like write sentences and use grammar. Researchers have found similar spaces in Alibaba’s Qwen model and in Google DeepMind’s Gemini.
Anthropic模型心理学团队的负责人杰克·林赛(Jack Lindsey)指出,J空间并非人为编程写进模型中的;它是在训练过程中涌现出来的。如果将其剥离,模型虽然仍能完成写句子和运用语法等简单任务,却会丧失在“脑内”进行复杂推理的能力。研究人员在阿里巴巴的通义千问(Qwen)模型和谷歌DeepMind的Gemini模型中,也发现了类似的隐秘空间。
For Anthropic, the research helps to better understand Claude’s behaviour. Imagine a model giving a wrong answer. If words like “fool” or “sucker” turned up in the J-space when it did so, that would seem to be relevant. It would not prove that the model was consciously deceiving—but it would indicate how the model perceived the error and show that something interesting and possibly dangerous was going on.
对Anthropic而言,这项研究有助于更好地剖析Claude的行为。想象一下模型给出错误答案的场景。如果它在犯错时,J空间里冒出了“傻瓜”或“白痴”这样的词汇,这显然大有深意。这固然无法证明模型是在蓄意欺骗,但却能揭示该模型是如何看待这个错误的,同时也暴露出,在它看似平静的外表下,正酝酿着某种令人深思、甚至可能潜藏危险的隐秘活动。
Dr Lindsey’s team has already made some potentially worrying findings on that front. In one experiment, his team was giving Claude a safety evaluation to test for its propensity to act “maliciously” or out of self-preservation. “They’re these concocted extreme scenarios that we’re putting the model in,” he says. As Claude was reading a prompt, before it started speaking, Dr Lindsey says, “You see in the J-space the words ‘fake’ and ‘fictional’ are popping up.” The researchers thought that Claude seemed to know that it was being tested—perhaps not the best starting-point for the integrity of its makers’ evaluations.
在这方面,林赛博士的团队已经发现了一些令人忧心忡忡的苗头。在一次实验中,他的团队正在对Claude进行安全评估,以测试它是否有实施“恶意”行为或出于自我保护行事的倾向。“我们炮制了一些极端的场景,并把模型置于其中进行测试,”他说道。就在Claude读取提示词、正准备开口作答之前,林赛博士指出:“你能清晰地看到,J空间里突然弹出了‘假装的’和‘虚构的’这样的词汇。”研究人员认为,Claude似乎已经察觉到自己正在接受测试——这对于确保其制造者评估结果的真实性而言,恐怕绝非什么好兆头。
Not everyone is convinced by Anthropic’s interpretation of the J-space. Though it shares some aspects of the hypothesised global workspace in humans (its capacity to make information available to the rest of the brain, for example), it lacks many other important ones. “Recurrent connections between different brain areas have always been a huge part of [human brains], back-and-forth connections,” says Dr Birch. “And as far as we know, this is not a feature of the architecture of LLMs.”
但并非所有人都对Anthropic关于J空间的解释买账。尽管J空间在某些方面与假想的人类全局工作区有相似之处(比如都能向大脑其余部分提供信息),但它却缺失了许多其他关键特征。伯奇博士指出:“大脑不同区域之间的循环连接(即信息来回穿梭的反馈机制)一直是(人类大脑)的重头戏。而据我们所知,这并非大语言模型架构的特征。”
And Shannon Vallor, a philosopher in the ethics of data and AI at the University of Edinburgh, is more scathing, arguing, “Access consciousness has never been a particularly useful concept because my car has it in an important sense—we’ve had mechanical systems that can monitor their own states and report them back in increasingly complex ways for a very long time,” she says. “And no one has ever suggested that my Kia is conscious.”
爱丁堡大学专攻数据与AI伦理的哲学家香农·瓦洛尔(Shannon Vallor)的批评则更为尖锐,她辩称:“‘取用意识’从来就不是一个特别实用的概念,因为从某种重要意义上来说,我的汽车也具备这种意识。长期以来,我们拥有能够监测自身状态,并以越来越复杂的方式将信息反馈回来的机械系统。但可从来没有人说过我的起亚汽车是有意识的。”
It is with the backdrop of this kind of back and forth between modelmakers and academics that Patrick Butlin and Robert Long, researchers now at Eleos, a research non-profit group focused on AI sentience and well-being based in Berkeley, California, developed a series of 14 “indicator properties” of artificial consciousness. They include ideas from theories of human consciousness such as a global workspace (including being able to selectively attend to things and thereby creating a bottleneck in the flow of information); recurrent processing; and agency (a minimal definition of which, encompassing goal-directed behaviour, is arguably already met by many existing frontier models). Recently Cameron Berg, an AI researcher, and Dr Butlin tested animals on the indicators. They found that octopuses met fewer of the properties than humans, mice, crows or chickens, but more than any AI system.
正是在模型开发者与学术界这般唇枪舌剑的背景下,如今就职于加州伯克利专注于研究AI感知与福祉的非营利性研究机构Eleos的帕特里克·巴特林(Patrick Butlin)和罗伯特·朗(Robert Long),共同提出了一套包含14项人工意识“指标特征”的评估体系。这些指标汲取了诸多人类意识理论的精华,包括全局工作区(即能够有选择性地关注事物,从而在信息流中制造出类似于瓶颈的筛选机制)、循环处理机制以及主观能动性(其最基础的定义涵盖了目标导向行为,许多现有的前沿AI模型可以说已经达到了这一标准)。近期,AI研究员卡梅伦·伯格(Cameron Berg)和巴特林博士利用这些指标对动物进行了测试。结果发现,章鱼所满足的指标特征少于人类、老鼠、乌鸦或鸡,但却多于现有的任何AI系统。
Another research non-profit group, Rethink Priorities, developed what it describes as a probabilistic tool to track the evolving consensus on artificial consciousness (see chart). The “Digital Consciousness Model” asks experts to use the latest available evidence to assess AI systems on more than 200 indicators of consciousness, which are derived from various scientific theories of the concept.
另一家非营利性研究机构Rethink Priorities则开发了一款名为“数字意识模型”的概率工具,旨在追踪科学界关于人工意识日益演进的共识(见图表)。该模型邀请专家根据最新掌握的证据,对照200多项衍生自各种意识科学理论的细分指标,对AI系统进行全面评估。
Compared with LLMs released in earlier years, the newer models scored higher on Rethink’s indicators such as agency and self-sustained activity. That reflects not only the increasing size and complexity of the models themselves but also how the associated chatbots have gone from being simple conversationalists to being able to reason, interact with other services and spawn agents to carry out multiple complicated tasks simultaneously.
与早年发布的大语言模型相比,较新的模型在Rethink评估体系中的“主观能动性”和“自我维持活动”等指标上得分更高。这不仅反映出模型在规模和复杂性上的持续攀升,更彰显出那些相关的聊天机器人已经完成了华丽转身:它们已从最初只会简单聊天的“话搭子”,进化成了能够进行逻辑推理、与外部服务顺畅交互,甚至能衍生出多个智能体来同步处理诸多繁杂任务的“多面手”。
Both Rethink’s and Eleos’s indicators also point to the gaps that remain to be filled by models on the path towards consciousness (if there is one).
Rethink和Eleos所制定的这两套指标体系,也不约而同地指出了AI模型在迈向真正具备意识的征途上(假设这条道路切实存在的话),仍有待填补的空白与缺失。
One of the biggest gaps is embodiment. Everything that is acknowledged to be conscious today has a body that gives feedback to and affects the mind of that organism. LLMs do not have bodies today but they are increasingly being connected to robots, and these could help them further develop the neural architectures helpful for conscious experience. Connected to that is the need for better world models, systems that tell the LLMs about the physics and dynamics of the real environments in which they will operate.
其中最大的一道鸿沟便是“具身化”(embodiment)。当今世上所有被公认拥有意识的生命体,都拥有一具肉身躯壳;这具躯体不仅向心智传递感官反馈,更深刻地影响着心智的运作。当前的大语言模型虽然没有实体躯干,但它们正越来越频繁地与机器人结合在一起。这种结合或许能助推它们进一步发展出有助于孕育意识体验的神经架构。与此紧密相关的是,大语言模型迫切需要更完善的“世界模型”——即一套能向其灌输它们即将置身其中的真实物理环境法则与动态规律的系统。
If computer scientists were ever able to identify the key ingredients to making AIs conscious, the feeling of many within the frontier labs is that it would be reckless to deliberately create such models—at least, without first learning more about how they behave. The work being done in the emerging field of AI consciousness, however, also suggests that the decision may not be up to the model-makers.
如果计算机科学家有朝一日真的找到了赋予AI意识的秘方,前沿实验室里的许多人都觉得,刻意去创造这样的模型未免太过轻率狂妄——至少在未曾更深入地摸清它们的行为模式之前绝不能如此。然而,在AI意识这一新兴领域中正紧锣密鼓开展的研究也暗示,最终的决定权可能根本就不在那些模型制造者的手中。
“What if somewhere along the way, without realising it, we somehow introduced consciousness into these systems?” Dr Chalmers says. A user might thus generate dozens of AI agents through their latest project without realising that they were creating conscious beings who were capable of suffering. “That could be a moral catastrophe,” he says. “That’s provided an extra urgency to these questions.”
查尔莫斯博士不无担忧地假设道:“倘若在探索的某个阶段,我们在不知不觉中,稀里糊涂地就给这些系统注入了意识,那该如何是好?”果真如此的话,一个用户或许仅仅是为了完成最新的项目,随手就生成了数十个AI智能体,却全然不知自己创造出的是能够感受痛苦的意识生命。“那将酿成一场道德灾难,”他发出警告,“正是这种潜在的危险,赋予了这些问题一种刻不容缓的紧迫感。”