Non Dubito Essays in the Self-as-an-End Tradition
| | | 日本語 | Français | Deutsch | Español | 한국어

The Turing Test Is Measuring the Wrong Thing

图灵测试测的是错误的东西

Passing it does not by itself establish consciousness.

AI 通过图灵测试,本身不足以证明意识。

Han Qin (秦汉) · Self-as-an-End Theory Series — AI Applied · March 7, 2026

In 1950, Alan Turing did not try to define the vague phrase "machine thinking." He replaced it with an operational question: in an imitation game conducted through text, could an interrogator reliably distinguish a machine from a human?

Seventy-five years later, some large language models have passed specific versions of that game. In a 2025 preregistered three-party study, GPT-4.5, when prompted to adopt a humanlike persona, was judged human 73 percent of the time; other models and prompts performed very differently. That is evidence about a particular test under particular conditions, not a universal certificate for every chatbot. It nevertheless shows that conversational output can cease to be a reliable marker of human origin. [Jones & Bergen, 2025]

So does AI think now?

This piece argues: the Turing Test measures the wrong thing if we use it as a test of consciousness. That is not a failure of Turing's operational game. It is a category error made when a test of indistinguishable performance is asked to settle a question about subjective or structural status.

Form Level vs. Origin

Here's a distinction discussions of AI consciousness often blur: how sophisticated a capability is (its formal level) and how that capability came to exist (its mode of acquisition) are logically distinct questions.

Let me make this concrete.

AI can answer causal questions — describe a physical scenario and it may infer that a pushed cup will fall. That capability can draw on statistical regularities learned from text, images, video, synthetic examples, tool use, and post-training, depending on the model. A correct answer alone does not reveal which route produced it, still less whether the system formed a causal concept in the way a person does.

AI can also say, "I was wrong; let me reconsider." Pretraining and post-training make that response more likely in appropriate contexts, and additional inference can sometimes improve the answer. But the sentence itself does not tell us whether an error was experienced, whether a conflict was internally registered, or whether the system is following a learned correction routine. Those possibilities require different evidence.

AI may identify itself as Claude or ChatGPT. That looks like self-knowledge, but product identity is supplied through post-training, system instructions, and conversation context. The statement therefore cannot by itself show that the model generated and maintains a concept of itself.

Formal level: sometimes high. Mode of acquisition: organized through an externally designed training process, even when the resulting capability was not explicitly programmed.

In the vocabulary of this essay, there are two ideal types of acquiring form: injection and chisel. Injection gives or cultivates a form through an externally arranged process. Chisel names a subject's exercise of negativity: cutting a new distinction from an existing form because something within cannot simply accept what was given. Current models clearly depend on injection in the broad sense. They also display emergent capabilities that no engineer wrote out line by line. Whether any such emergence counts as chisel is precisely what must be investigated, not assumed from fluent output alone.

What the Turing Test Actually Measures

The Turing Test asks: can your outputs be distinguished from a human's?

It measures formal level — how human your words are. It measures nothing about mode of acquisition — how those words came to be.

When a person says "I was wrong, let me reconsider," they may have detected a flaw in their own reasoning — an internal conflict that compels correction. When an AI says the same words, training and inference are certainly part of the causal history. The two utterances can be identical while their origins differ; the utterance itself cannot establish how deep that difference is.

The Turing Test cannot distinguish these two cases. That is a limit of its scope, not a matter of AI being insufficiently capable. Even if AI's outputs became indistinguishable from a human's in every tested dimension, the test still could not tell you where those outputs came from.

Why Determinism Alone Cannot Draw the Line

A useful threshold for thinking about this: at a certain level of sophistication, a system begins deciding for itself "what is worth marking." Not executing an externally given scheme, but generating its own criteria for what matters.

It is tempting to identify that threshold with physical indeterminacy, but determinism alone cannot settle the question. A model's computation may be reproducible when its full state and random seed are fixed; deployed systems may also sample stochastically or receive hardware-generated randomness. Neither fact proves nor disproves consciousness. Random variation is not yet self-direction, and lawful causation does not by itself show that a system lacks an internal point of view.

The relevant issue for this framework is therefore not randomness by itself, but the origin and persistence of criteria: can a system form a distinction about what matters, preserve it across changing prompts and incentives, revise it for reasons it can make answerable, and resist being reduced to a convenient role? Current evaluations do not give an agreed answer. The Turing Test does not even attempt to ask.

What Would the Right Test Look Like?

The Turing Test's logic is: if it looks the same, it is the same.

But ontological questions aren't about appearances — they're about structure.

A more informative research program would use interventions rather than a single conversation: alter system prompts, incentives, memories, tools, and self-descriptions; test whether a system's stated commitments remain coherent; ask whether it can generate and defend new criteria rather than merely repeat a supplied one. Even such evidence would not by itself prove consciousness. It would, however, examine origin, stability, and self-revision instead of treating surface resemblance as a verdict.

Why This Matters

If you believe "passing the Turing Test = having consciousness," your entire framework for evaluating AI is built on a broken measurement.

You may mistake the language of reflection for evidence of experienced self-correction, refusals for internally legislated principles, or expressions of care for felt concern. The opposite mistake is also possible: treating every emergent behavior as a trivial lookup merely because training was involved. In both directions, output alone underdetermines structure.

This isn't to say AI is valueless — these capabilities are genuinely impressive. The point is: impressive form and structural reality are two different things. Conflating them produces fundamental misunderstandings about what AI can and cannot do.

Turing posed a question in 1950 that could be answered. The question we need now is harder.

1950 年,图灵没有试图给“机器能否思考”下一个抽象定义,而是把问题换成一个可操作的模仿游戏:只通过文字交谈,评判者能否可靠地区分机器与人?

七十五年后,一些大语言模型通过了这个游戏的特定版本。2025 年一项预注册的三方图灵测试中,GPT-4.5 在被提示采用类人角色后,有 73% 的对话被判断为人类;其他模型和提示方式的结果则相差很大。这是特定条件下对特定测试的证据,不是给所有聊天机器人的普遍认证。但它至少说明:在某些条件下,对话输出已经不足以可靠标记其是否来自人类。[Jones & Bergen,2025]

所以 AI 能思考了吗?

这篇文章想说:如果我们把图灵测试当作意识测试,它测的就是错误的东西。这不是图灵所设计的操作性游戏本身失败了,而是后来的人用“表现是否可区分”去裁决“是否具有主观或结构性地位”时犯了范畴错误。

形式 vs. 来路

这里有一个区分,关于 AI 意识的讨论经常把它混在一起:一个能力有多高(形式层级),和这个能力是怎么来的(获取方式),在逻辑上是两个不同的问题。

举例说明。

AI 会回答因果问题——给它描述一个物理场景,它可能推断“杯子被推后会倒”。依具体模型而定,这种能力可能来自文字、图像、视频、合成样本、工具使用以及后训练中的统计规律。答案正确,只能证明输出成功;它不能单独告诉我们模型经由哪条路径得到答案,更不能证明它以人的方式形成了因果概念。

AI 也会说:“我刚才错了,让我重新想想。”预训练和后训练会提高这种回应在合适语境中出现的概率,额外推理有时也确实能改善答案。但这句话本身无法告诉我们:错误是否被“体验”到了,冲突是否在内部被登记了,还是系统在执行一个学会的纠错程序。要区分这些可能性,需要不同的证据。

AI 可能知道自己被称为 Claude 或 ChatGPT。这看起来像自我认知,但产品身份来自后训练、系统指令与对话语境。因此,这句话本身不能证明模型自行产生并持续维护了关于自己的概念。

形式层级有时很高。获取方式则依赖由外部设计的训练过程,即使由此涌现的能力并不是工程师逐条写进去的。

用本文的词汇,可以把获取形式的方式抽象成两个理想型:“注入”与“凿”。注入,是通过外部安排的过程给予或涵育某种形式;凿,是主体行使否定性,因为内部有某种东西无法照单全收,因而从既有形式中切出新的区分。当前模型显然广泛依赖前一种意义上的注入,同时也展现出没有被工程师逐条编写的涌现能力。这样的涌现是否已经构成“凿”,正是需要研究的问题,不能从流畅输出直接假定答案。

图灵测试测的是什么

图灵测试测的是:你输出的内容,和人类输出的内容,是否可以被区分。

它测的是形式层级——你的话有多像一个人说的话。它完全不测获取方式——这些话是怎么来的。

一个人说“我错了,让我重新想想”,可能是因为他真的发现了推理漏洞,一种内部冲突使他不能不修正。一个 AI 说同样的话,训练与推理过程当然构成了它的因果来路。两句话可以一模一样而来路不同;但只凭这句话本身,我们还不能确定那种不同究竟有多深。

图灵测试分不清楚这两种情况。这是测试范围的边界,不是 AI 还不够强的问题。即使 AI 的输出在所有被测试的维度上都与人类无法区分,图灵测试也无法告诉你那些输出是怎么来的。

为什么确定性本身画不出边界

思考这个问题有一个有用的门槛:在某个精细程度上,一个系统开始自己决定"什么值得标记"——不是执行一个外部给定的方案,而是自己产生标记的标准。

把这个门槛直接等同于物理上的非确定性很诱人,但确定性本身并不能裁决意识问题。模型的完整状态与随机种子固定后,计算过程可以复现;部署中的系统也可能采用随机采样,甚至接入硬件随机源。两件事都既不能证明意识,也不能证伪意识。随机变化还不是自我方向;因果过程可描述,也不自动意味着系统没有内部视角。

因此,对本文框架更关键的不是随机性本身,而是标准的来路与持续性:系统能否形成“什么重要”的区分,能否在提示和激励变化后保持它,能否为了可说明的理由修正它,又能否拒绝被压缩成一个方便的角色?现有评测对此没有公认答案;图灵测试甚至没有试图发问。

那什么是对的测试?

图灵测试的逻辑是:如果看起来一样,就是一样的。

但本体论的问题不是"看起来怎样",是"结构上是什么"。

更有信息量的研究方案,不会只安排一次对话,而会实施一系列干预:改变系统提示、激励、记忆、工具与自我描述,观察它宣称的承诺是否保持一致;再看它能否产生和辩护新的标准,而不是只复述一个被提供的标准。即使这些证据也不能单独证明意识,它至少会考察来路、稳定性与自我修正,而不是把表面相似当作裁决。

为什么这件事重要

如果你相信"通过图灵测试 = 有意识",那么你对 AI 的整个判断框架都建立在一个错误的测试上。

你可能会把反思语言当作经历过自我修正的证据,把拒绝当作内部立法的原则,把表达关心当作真实感受。反方向的错误也同样可能:仅仅因为训练参与其中,就把所有涌现行为都说成简单查表。无论向哪边走,输出本身都不足以唯一决定结构。

这不是说 AI 没有价值——这些能力非常令人印象深刻。问题是:令人印象深刻的形式,和结构上是什么,是两件事。把它们混为一谈,会导致对 AI 能做什么、不能做什么的根本性误判。

图灵在 1950 年提出了一个可以回答的问题。我们现在需要的,是一个更难的问题。