Non Dubito Essays in the Self-as-an-End Tradition
| | | 日本語 | Français | Deutsch | Español | 한국어

The More It Means, the Less It Learns

越有意义,AI 越学不会

Frequency shapes what is easiest to learn without measuring what matters.

频率影响什么最容易学,却不能衡量什么最重要。

Han Qin (秦汉) · Self-as-an-End Theory Series — AI Applied · March 2026

Here is a paradox worth treating as a warning rather than a universal law: the more singular and context-bound a meaning is, the less direct statistical support a model may have for learning it.

Not because AI is dumb, and not because frequency is the opposite of meaning. The tension arises because distributional learning is strongest where many examples constrain a pattern, while some of the meanings that matter most to a person, discipline, or culture appear only rarely and in changing contexts.

Language models do not merely count words. They learn parameters by predicting patterns across enormous corpora and can compose familiar pieces to handle expressions they have rarely or never seen. Still, contrast "the" with "thing-in-itself" — Ding an sich, Kant's term for what exists independently of our mode of appearance. The first supplies abundant evidence about ordinary usage; the second demands a sparse and contested philosophical context. A model can say a great deal about both, but equal fluency does not imply equally grounded understanding.

The uncomfortable implication is narrower and more defensible: dense, load-bearing concepts often receive less direct supervision than ordinary connective language. Statistical frequency does not measure meaning, yet it shapes which distinctions are easiest for a model to stabilize.

Two Roads That Don't Meet

To understand this, you need to know something about how AI reads text.

Before a language model processes text, a tokenizer maps it into units — "tokens" — handled by the network. This is not a trivial step. Tokenization affects efficiency, sequence length, and the representation from which learning begins, although a model can represent concepts across several tokens.

One widely used family of approaches, derived from BPE (Byte Pair Encoding), builds a vocabulary by repeatedly merging frequent byte or character sequences. Related tokenizers use different algorithms, but the trade-off is similar: common sequences tend to receive compact representations, while rare sequences may be split into several units.

This approach has one large advantage: it scales without requiring a human to define every concept in advance. Its limit is not that splitting a word destroys meaning automatically; transformers routinely compose meaning across multiple tokens. The limit is that the vocabulary boundary itself is optimized for statistical utility, not offered as a theory of which concepts matter.

There is another task we can call the semantic path: deciding that "thing-in-itself" should be treated as one concept even if its written form occupies several tokens, locating its disputes, and judging when a use is faithful or superficial. Statistical learning can contribute to this task through composition and context. It does not make the task identical to token compression.

The two roads interact, but they do not collapse into one. Better compression can support learning; it cannot by itself decide which distinctions deserve conceptual weight.

The Inverse Law

This is where a structural tension becomes visible.

Some highly significant meanings are also rare: a private vow, a new scientific distinction, a minority idiom, a philosophical term used precisely rather than decoratively. Rarity gives a learner fewer direct examples. But rare does not always mean meaningful, common does not mean empty, and models can generalize compositionally. Frequency is therefore not a law of meaning; it is one pressure on learnability.

Call the warning the inverse law: as direct statistical support decreases, human judgment about context and significance becomes more important. It is a heuristic, not a theorem assigning a fixed meaning score to every token.

More and better data can help, as can retrieval, tools, curated examples, multilingual training, and improved architectures. Yet some scarcity is not merely a dataset accident. New, local, contested, and deeply personal meanings may remain rare because people actually use them rarely. Scaling changes the boundary; it does not guarantee that every remainder disappears.

In the framework I work with — Self-as-an-End — this is an instance of what we call remainder conservation: every act of structuring makes some distinctions cheap and leaves others costly or obscure. Tokenization does not erase meaning wholesale. It is one of several structures through which meaning must pass, and its optimization target is not identical to meaning.

The Remainder a Human Position Can Expose

Here is the deeper problem: no learner sees all of its own blind spots from inside the same representation.

A model operates downstream from a tokenizer and training pipeline it did not choose. It can often infer a multi-token concept from context, and with tools it can even inspect tokenization. What it cannot obtain from fluent continuation alone is an independent guarantee that the distinctions favored by its pipeline are the distinctions that matter in this case.

Humans do not possess that guarantee either. What they can contribute is another position: lived stakes, disciplinary practice, minority usage, and the ability to say that a statistically plausible reading has missed the point.

A human who genuinely understands a philosophical concept can sometimes recognize when a mechanical representation has flattened it. Not because humans have more processing power, or because human learning is free of frequency, but because it is not built on frequency alone. It can include years of disciplinary practice, lived stakes, confusion, argument, and revision.

Humans can sometimes learn a new distinction from one carefully situated example because they bring bodies, projects, shared practices, and prior concepts to it. Models can also perform striking few-shot and zero-shot generalization, but they bring a different history. The practical question is not which learner uses "statistics" and which does not; it is which background makes the relevant distinction available and which hides it.

Compute and data therefore matter greatly, but they are not the only variables. At the frontier of sparse or contested meaning, the quality and plurality of human judgment can set limits that scale alone does not automatically remove.

What This Means for Chinese

The inverse law has a specific geopolitical consequence that deserves naming.

Tokenizer efficiency differs by model and training mixture. In some widely used vocabularies — especially earlier or English-heavy ones — Chinese text is represented with more tokens than comparable English text. Newer multilingual tokenizers have narrowed the gap, and no single ratio describes every system.

Where the gap persists, Chinese can consume more of a context window and incur greater cost for comparable work. This is not because Chinese is intrinsically more complex. It reflects design choices, vocabulary allocation, and the distribution of training data.

This need not be malicious. A method optimized around one distribution can reproduce center–periphery effects: what is efficient near the statistical center becomes costlier at the edge. Tokenizer design can therefore disadvantage Chinese without anyone deliberately deciding to exclude it. Optimization creates trade-offs; their distribution should be measured rather than naturalized.

The Human in the Loop Is Not a Consolation Prize

A common response to AI anxiety is: "Don't worry, humans will always be needed to supervise the machines." This is usually meant as comfort — a promise that we won't be fully replaced.

What the inverse law suggests is something stronger but more limited: in work where stakes, novelty, or local meaning matter, a responsible human role should not be treated as a temporary inconvenience to be engineered away.

The semantic path — judging what a concept means here and why it matters — is not solved merely by assigning it one token. Automated systems can support that judgment, retrieve its history, and expose competing uses. They can also reproduce a dominant reading while missing the case that gave the question its stakes. Human responsibility begins where plausible continuation is not enough.

The quality of human judgment can therefore set a ceiling on the quality of the whole human–AI process, just as poor data, weak architecture, or inadequate compute can. The point is not that only humans can ever draw a conceptual boundary. It is that someone must remain answerable for why this boundary, in this context, deserves to govern.

Self-as-an-End — humans as ends in themselves, not merely means to AI's performance — remains a moral and institutional claim. The technical architecture does not prove it. What the architecture shows is why abandoning human answerability would also be epistemically costly.

A Practical Consequence

If sparse and situated meanings receive less direct support, then one dangerous failure mode in human-AI collaboration is quiet rather than dramatic: humans gradually outsourcing their judgment.

The quieter risk lies in the workflow rather than in a model secretly retraining itself during the conversation. The less a workflow is exposed to independent human judgment, the more its outputs remain anchored to already represented patterns, majority signals, and whatever the system can make most plausible. Rare or situated meanings can then disappear without anyone noticing the loss.

The cure is not clever. It is: go first.

Before asking the model, write your own judgment. Write it badly, incompletely, hesitantly. The moment your fingers pause over the keyboard and you are not sure — that pause is the remainder knocking on the door. Do not skip it. That is exactly what the model cannot do for you.

Then show the model what you wrote. Ask it to reflect it back, to find the gaps, to push against your framing. Not to replace your judgment — to sharpen it.

Then write again, yourself.

The loop only works in that order. Reverse it — start with the model's output and edit from there — and you are no longer bringing your remainder into the system. You are becoming the model's editor. Which is useful. But it is not the same thing.

The Remainder Cannot Be Dissolved

Every technology that structures human knowledge leaves a remainder. Writing left out the gesture, the silence, the timing that spoken language carries. Print left out the handwriting, the marginalia, the individual voice. Digital text left out — we are still discovering what.

AI tokenization and statistical training leave remainders — not all meaning, not always, and not irreparably. The narrower claim is enough: rare and situated meanings can be harder to stabilize, while fluent language can conceal that weakness.

This is not a reason to mistrust AI or to freeze its present limits into metaphysical law. It is a reason to distinguish gains that scale can plausibly deliver from questions that still require situated judgment, better evidence, and answerability.

The remainder cannot be dissolved. It can only be carried — by humans who are willing to stay in the loop, go first, and bring what only they can bring.

We cannot help not knowing — just for now. But the not-knowing is not the model's. It is ours to work with.

有一个听起来矛盾的判断,适合当作警告,而不是普遍定律:一个意义越独特、越依赖具体语境,模型能够直接用来学习它的统计证据就可能越少。

不是因为 AI 不够聪明,也不是因为频率与意义天然相反。张力来自分布式学习的基本条件:许多样本共同约束一种模式时,模型最容易稳定地学会它;但对一个人、一门学科或一种文化最重要的某些意义,可能只在稀少而不断变化的语境中出现。

语言模型并不只是“数词频”。它通过预测海量语料中的模式来学习参数,也能把熟悉的片段组合起来,处理很少甚至从未直接见过的表达。但可以比较“的”与“物自体”:前者提供了大量日常用法的证据;后者指向康德关于独立于显现方式之物的概念,需要进入稀少而且有争议的哲学语境。模型可以流畅谈论两者,流畅却不等于对两者有同样扎实的把握。

更准确的推论是:密度高、承重大的概念,常常比普通连接语得到更少的直接监督。统计频率不是意义的尺子,却会影响哪些区分最容易被模型稳定下来。

两条不相交的路

要理解这个,你需要知道 AI 是怎么读文字的。

在语言模型处理文本之前,分词器(tokenizer)会把文本映射成模型实际处理的单元——token。这一步不是小事:它影响效率、序列长度以及学习开始时的表示方式,不过模型也可以跨越多个 token 表示一个概念。

一类被广泛使用的方案源自 BPE(字节对编码):反复合并高频的字节或字符序列。其他分词器会采用不同算法,但基本取舍相似——常见序列往往得到更紧凑的表示,少见序列则可能被拆成多个单元。

这套方案有一个巨大的优点:它可以规模化,不需要人类事先定义每一个概念。它的局限并不是“拆词必然毁掉意义”——Transformer 经常能跨多个 token 组合出意义——而是词表边界优化的是统计效用,并不承诺回答“哪些概念真正重要”。

还有另一项任务,可以叫作语义路径:即使“物自体”在字面上占据多个 token,仍把它作为一个概念来处理,定位围绕它的争论,并判断一次使用究竟准确还是表面化。统计学习可以通过组合与语境参与这项任务,但这项任务不等于 token 压缩。

两条路会互相作用,却不能合并成同一件事。更好的压缩能帮助学习;仅凭压缩本身,无法决定哪些区分值得获得概念上的重量。

反比定律

一种结构性张力在这里浮现。

有些意义重大的东西确实很稀少:一项私人誓言、一个新生的科学区分、一种少数语言习惯、一个被准确而非装饰性使用的哲学概念。稀少会让学习者缺乏直接样本。但稀少不总等于有意义,常见也不等于空洞,模型还能够组合泛化。因此,频率不是意义的定律;它只是影响可学性的压力之一。

把“反比定律”保留为一个警告:直接统计支持越少,人对语境与重要性的判断就越关键。它是一条启发式原则,不是给每个 token 指派固定意义分数的定理。

更多、更好的数据能够改善问题,检索、工具、精选样本、多语言训练与新架构也一样可以。但有些稀缺并不只是数据集偶然不全。新生的、地方性的、有争议的、极其私人的意义,本来就可能因为人们很少使用而保持稀少。规模化会移动边界,却不保证所有余项都消失。

用 Self-as-an-End 框架的术语来说:这是余项守恒的一种具体形态。每一次构,都会让某些区分变得廉价,同时让另一些区分昂贵或隐蔽。分词不会整体抹掉意义;它只是意义必须穿过的若干结构之一,而它的优化目标并不等同于意义。

人类位置能够照亮的余项

更深层的问题在于:任何学习者都无法在同一套表示内部看清自己的全部盲区。

模型在它没有选择的分词器与训练流程下游工作。它经常可以从语境推断跨 token 的概念,有工具时甚至能够检查分词。但仅凭流畅续写,它无法获得一个独立保证:训练流程偏爱的区分,恰好就是此刻真正重要的区分。

人也没有这样的保证。人能带来的是另一个位置:具身利害、学科实践、少数用法,以及指出“统计上很像,但还是没有说到点上”的能力。

一个真正理解某个哲学概念的人,有时能察觉机械表示把它压平了。不是因为人的算力更高,也不是人的学习与频率无关,而是人的理解不只建立在频率上。它还可以包含多年的学科实践、切身利害、困惑、争论与修正。

人有时能从一个被妥善放置的样本学会新区分,因为他带着身体、目的、共同实践与既有概念进入现场。模型同样能展现惊人的少样本与零样本泛化,只是带入了另一种历史。实际问题不是谁使用“统计”、谁不使用,而是哪一种背景让相关区分变得可见,又是哪一种背景把它遮住。

因此,算力和数据当然非常重要,却不是唯一变量。在稀少或有争议的意义前沿,人类判断的质量与多样性可能构成规模化无法自动移除的边界。

对中文意味着什么

反比定律有一个具体的地缘后果,值得说清楚。

分词效率会随模型和训练混合而改变。在一些被广泛使用的词表里——尤其较早或明显偏向英语的词表——同等内容的中文会比英文占用更多 token。较新的多语言分词器已经缩小了差距,任何单一比率都不能描述所有系统。

在差距仍然存在的地方,中文完成同类工作可能消耗更多上下文窗口并承担更高成本。不是因为中文本身更复杂,而是词表分配、设计选择与训练数据分布共同造成的。

这不一定出于恶意。围绕一种分布优化的方法,可能形成中心—边缘效应:靠近统计中心的表达更高效,边缘表达则更昂贵。因此,分词器可能在没有人刻意排斥中文的情况下,使中文处于不利位置。优化会产生取舍;关键是测量这些取舍如何分布,而不是把它们自然化。

人在循环里不是安慰奖

对 AI 焦虑常见的回应是:"别担心,总需要人来监督机器。"通常这是用来安慰的——承诺我们不会被完全取代。

反比定律所提示的是一件更强、但也更有限的事:在利害重大、意义新生或高度地方化的工作中,负责任的人类位置不应该被当作迟早要工程化消除的麻烦。

语义路径——判断一个概念在这里是什么意思、又为什么重要——并不会因为它被分成一个还是多个 token 就自动解决。自动系统可以帮助判断,检索概念史,展示互相竞争的用法;它也可能复制占主导的解释,却错过让问题真正具有利害的那一例。仅有“续写得像”还不够的地方,就是人的责任开始的地方。

因此,人的判断质量可以像数据质量、架构与算力一样,决定整个人机过程的天花板。重点不是宣称“只有人永远能够划出概念边界”,而是必须有人回答:为什么是这条边界,在这个语境里,应该获得支配地位?

Self-as-an-End——人是目的本身,不只是 AI 性能的手段——仍然是一项道德与制度主张,技术架构本身不能证明它。技术架构能显示的是:放弃人的可问责位置,同时也会付出认识上的代价。

一个实践上的后果

如果稀少而具体的意义得到的直接支持更少,那么人机协作的一种危险失效模式就不是戏剧性的,而是安静的:人逐渐把自己的判断外包出去。

更安静的风险发生在工作流,而不是模型在对话中偷偷更新自己。一个工作流越少接受独立的人类判断,它的输出就越容易停留在已有的统计模式、多数信号与系统最能生成得像的东西里。稀少而具体的意义于是可能在无人察觉时消失。

解药不复杂:先动手。

在问模型之前,先写下你自己的判断。写得残缺、犹豫、不成熟也没关系。手指停在键盘上那个点不下去的瞬间——那就是余项在敲门。不要跳过那个犹豫。那恰恰是模型替代不了你的地方。

然后把你写的东西给模型看。让它照镜子:你漏了什么?你犹豫的地方有没有道理?你的框架有什么盲区?不是让模型替换你的判断——是让它把你的判断磨得更锋利。

然后再自己写。

这个循环只在这个顺序下有效。倒过来——先拿模型的输出,从那里改——你带进系统的就不再是你的余项。你成了模型的编辑。这有用,但不是同一件事。

余项消灭不了

每一种结构化人类知识的技术都留下余项。书写丢掉了口语里的手势、沉默、时机。印刷丢掉了手写体、批注、个体声音。数字文本丢掉了什么——我们还在发现中。

AI 分词与统计训练会留下余项——不是所有意义,不是每一次,也不是无法修复。更窄的命题已经足够:稀少而具体的意义可能更难稳定下来,而流畅语言会遮住这种薄弱。

这不是不信任 AI 的理由,也不是把当前局限冻结成形而上的定律。它提醒我们区分两类问题:哪些提升可以合理期待规模化解决,哪些仍需要具体判断、更好证据与可问责性。

余项消灭不了。它只能被承担——被那些愿意留在循环里、先动手、带来只有自己才能带来的东西的人承担。

我们不得不不知道——just for now。但这个不知道不属于模型。它属于我们,等我们去工作。