Non Dubito Essays in the Self-as-an-End Tradition
| | 日本語 | Français | Deutsch | Español | 한국어
← 凿构周期律·经济系列 ← Chisel-Construct Cycle: Economics
凿构周期律 · 经济
Chisel-Construct Cycle · Economics
第 21 篇,共 23 篇
Essay 21 of 23

第二十一篇 评分社会:声誉,也想把账合上

Essay 21: The Scored Society — Reputation Wants to Close the Books Too

Han Qin (秦汉)

一 一个品格先你而至

1853年,一份商人杂志上出现了一句几乎带着恐吓口气的话:你无论到哪里赊购货物,一个品格都已经先你而至;它要么帮你,要么毁你。

这句话说的是一件当时刚刚成型的事。

十二年前,1841年,一位纽约商人在纽约开办了一家商业征信机构,通常被视为现代大型征信业的开端。它的做法是在各地布下通讯人,定期把商人的偿付能力和为人写成条陈,寄回纽约总部。

那些通讯人要回答的问题里,有资本多少,有欠账多少,而也有这样一条:他是否品行端正,做生意的习惯是否良好。

也就是说,从第一天起,信用就不只是一个关于还不还得上的概率。它同时是一种关于人的判断。

这两样东西为什么会长在一起,不难理解。在没有别的凭据可查的年代,一个人的品行确实是关于他会不会还钱的最好线索。问的人并不是在道德上审查谁,他是在找一个可用的预测变量。而一旦这个变量被写进表格,它就不再只是线索,它成了一条记录。

这件事的历史意义不在于它消灭了不确定性。它做的是另一件更基础的事:把原本流动的风评,固定成了一份可以存档的你。

有研究者把这个过程概括成从传言到成文记录,并且指出十九世纪的征信业实际上发明了一种金融身份。这个身份不是自然长出来的,它是被信贷需求,远距离贸易和信息中介一起制造出来的。

这个身份和一个人自己认识的自己之间,从一开始就有距离。他知道自己那年为什么周转不开,记录只知道他周转不开;他知道自己后来还清了,记录可能还停在没还清的那一页。距离本身不算问题,任何记录都比它记的那件事简略。要紧的是那份简略会代替他去到很远的地方,而他本人到不了那里。

前一篇的结尾说,被做成市场对象的不是那个人本身,是那个人可预测的外观。

而这个外观并不是二十一世纪才有的。它有一个一百八十年的前身,而且那个前身一开始就带着道德色彩。

1853年那句话最要紧的地方,是它把两件事揉在了一起。

一方面,信用开始离开熟人的圈子,变成可以远距离调用的判断。你到一个陌生城市去赊货,店主不认识你,而他能查到关于你的记录。

另一方面,这个判断又并不纯粹是匿名的。它装满了关于你的习惯,体面,交际圈,私生活甚至品行的叙述。

非人格化的结算,并没有把人格性的材料清除掉。它把这些材料搬进了档案。

搬进档案这个动作值得记住,因为它在后面每一节都会重演一次,只是形式越来越体面。条陈变成表格,表格变成分数,分数变成星级,星级变成分级分类监管的依据。每一次都更整齐,更可比,更容易在别处被读出来;而每一次都离那个被写的人更远一点。

前面二十篇里,这条线索反复出现过:冷的那一侧和热的那一侧从来不是谁取代谁。而接下来这一段历史,会把它再推一步。

它要说的是,声誉自己也想闭合。

这一句会在最后一节被兑现。在此之前要先看清楚,那把叫作声誉的尺子,是怎样一代一代变细的:从写在条陈里的品行,到印在手册里的等级,到算出来的三位数,到平均出来的星级,再到写进监管的分类。每一次都更整齐,而每一次都要有人被留在外面。

二 永久记录

到二十世纪三十年代,这台机器已经从商人之间扩展到了一般人。

1936年,一份周刊报道说,美国一千多个城市的零售信用局,掌握着不少于六千万赊购者的记录;任何逾期一百二十天未付的账,都要进入永久记录。

同一篇报道还举了一个例子:如果一个女人从芝加哥搬到洛杉矶,新的店员很快就可能拿到关于她的一整套叙述,包括她守寡,租房,和什么人来往,以及欠了多少钱。

永久这两个字值得停一下。

一笔账逾期一百二十天,可能有很多原因。可能是生意亏了,可能是病了,可能是丈夫死了,可能是当时正好没钱而后来还上了。这些原因彼此差得很远。

而进入记录的是同一件事:逾期一百二十天。

这不是记录者的疏忽。恰恰相反,这是记录能够被使用的条件。如果每一条都附着一段来龙去脉,洛杉矶那个店员就得先读完一篇故事,再自己判断这故事算不算严重。而他没有时间,也没有立场去做这个判断。要让这条信息可以被远距离调用,它就必须先被剥掉那些只有当地人才能掂量的东西。

第五篇里说过一次相近的动作:复式记账把每一笔都锚住,让账可以闭合。这里的动作是另一种:它把一段有前因后果的经历,压成一个可以跨城市调用的条目。压缩正是它的功能,因为不压缩就没法在洛杉矶被读到。

这也是这一整段历史的基本形状。信用从同行之间的判断,变成了一般人口的社会档案。

但这套档案人格从一开始就有剩余,而且剩余得很明显。

研究这段历史的人反复指出,那些条陈里常常混杂着道听途说,偏见,阶层判断和地方政治。它们不是把主观性清除掉了,而是把主观性保存下来,复制出去,再卖掉。

一份关于某个人是否值得信任的判断,被装进档案的时候并没有变得更客观。它只是变得更容易被传播。

这一条对整个系列都要紧。前面几篇里,构做的通常是把一样东西从模糊变清楚。而这里做的是另一件:把一样本来就不清楚的东西,做成一个看上去清楚的形式。判断没有变得更可靠,它只是获得了一种可以流通的外形。而流通到远处之后,它看上去就像是一件事实。

这里可以对着第三篇看一眼。当年硬币被造出来,是给已经存在的重量一个可以携带的身体;而重量本来就在那里,秤一称就知道。这里被携带的是一个判断,而判断没有可以复称的东西。走到远处以后,没有人能把它拆回去看看当初是怎么得出来的,于是它只能被当成给定的。

也正因如此,1970年美国通过一部规范征信业的联邦法律时,法条里写下的是一句折衷的话:征信机构应当采用合理的程序,在满足商业需要的同时,以对消费者公平公正的方式处理信息。

这不是一句胜利宣言。它是承认冲突仍然在那里。

一部法律如果能一句话解决问题,它不会这样写。它这样写,是因为立法者知道商业需要和对个人的公平之间存在张力,而他们没有办法取消这个张力,只能要求两边都被照顾到。

余项在这里第一次被写进法律。法条没有说清楚公平公正具体指什么,也没有给出一个可以照着算的标准;它做的是留一个位置,标明这里有一样东西不能被商业需要吞掉。第十七篇里那条几乎从不使用的条款是同一类东西:构在自己的文本里承认了有些东西它没能处理好。

三 五十家里只有一家

现代信用评分的分界线,不是有没有信用记录,而是能不能把记录里那些性质各异的信息,换算成一个可以比较的数。

1956年,一位工程师和一位数学家在美国合办了一家公司,想把运筹学和统计推断用到消费信贷上。

1958年,他们向五十家大型放贷机构推销这套评分技术。

只有一家回应。

这一家和四十九家的对比,说明当时的放贷业并不觉得自己缺一把尺子。他们已经有一套办法:让贷款员见人,问话,看铺面,打听同行的评价,再自己拿主意。那套办法慢,不一致,而在他们看来是管用的,因为出了问题有人负责。一个数字给不了这个。

而三十年后,同样这批机构会争着买它。中间变的不是数字本身,是要处理的量。放贷从几百笔变成几百万笔的时候,让人去见每一个申请者这件事在算术上就不成立了。构的每一次扩张几乎都是这样发生的:不是有人被说服了,是原来那套办法先撑不住了。

这个细节值得记住,因为它说明今天看上去理所当然的东西,当时并不显得自然。分数曾经是一项不太被人理解的新算术。

而它后来长成的样子,也不是某个天才一次性发明出来的。

有研究者专门追踪过评分卡是怎么被造出来的。它的结论是:银行,百货公司,放款机构,计算设备,历史违约数据和监管变化,一起把申请人拆成若干特征,再给每个特征分配不同的权重,最后压缩成一张可以执行,可以部署,可以调整阈值的表。

这位研究者强调,评分卡不是一面中立的镜子,它是一种市场装置。它重写了贷款员的工作,重写了前台和后台的分工,重写了审批的速度和可复制性。

这里有一处很容易被略过的后果。贷款员原先要做判断,判断是要担责任的,而担责任的人会去问,会去看,也会在拿不准的时候多问一句。分数上来之后,他的工作变成了录入和执行阈值。责任跟着判断一起转移了,转移到了一个没有人可以质问的地方。

1989年,第一套通用型的分数在一家征信机构上线。1995年,一家政府支持的住房金融机构鼓励按揭贷款使用信用评分,随后住房金融的二级市场把这种判断推向了全国。

而支持这件事的论证必须原样摆足,因为它在很大程度上是对的。

一位联储的研究者在 1997年总结过赞成者的典型说法:评分更快,更便宜,更一致,也更客观。她给了一个具体的对照:传统的小企业贷款审批平均要花大约十二个半小时,而评分可以把时间压到远少于一个小时。

联储在 2010年的一份证词里也概括了同一个立场:信用评分让放贷者能够更快,更便宜地评估风险,并且增加了信贷的供给与可负担性。

这不是空话。

一个没有熟人网络的人,在关系型的放贷世界里几乎借不到钱。他不认识银行经理,没有人替他说话,没有一个本地的名声可以调用。而一套只看记录的评分,恰恰让他有了一个可以被读到的东西。

标准化在这里是一种解放。它把判断从少数几个人的裁量里拿了出来。

这一段必须写足,因为后面几节会讲这套东西的代价,而如果代价被讲得孤零零的,整件事就会变成一个谁都看得出的坏主意。它不是坏主意。它让很多本来借不到钱的人借到了钱,这是可以在数据上看到的,也是当年推动它的人真正想做成的事。

第十篇里说过一次同样形状的事:全市场同一比率这句话,砸掉的是按门第开价那一套。这里是同一个动作的又一次现形。

四 只看那份记录里的信息

那家公司的分数通常在三百到八百五十之间。常见的讲法把它分解成五块:付款历史占三成五,负债水平占三成,信用历史的长度占一成五,新开的信用占一成,信用的组合占一成。

公司自己反复强调一件事:分数只看信用记录里的信息。它不直接看收入,不看职业,不看你在现在这家单位干了多少年。

这句话是一种自我限制,而且这种限制有它的道理:少看一些东西,可以少一些歧视的入口。

而自我限制的另一面是,那些没被看的东西不会因此消失。一个人有没有稳定的收入,做的是什么行当,家里有没有人生病,这些和他还不还得上钱的关系,常常比他上个月刷了多少额度更直接。分数不看它们,不等于它们不起作用;它只是把它们推给了别人去看,推给了那些在分数之外还要做决定的人。

而恰恰是这句话,暴露了溢出。

第一,分数覆盖不了所有人。

美国的消费者金融保护机构在 2015年,按 2010年的基准估算过:大约有两千六百万成年人是信用上看不见的,也就是根本没有信用记录;另有一千九百万人虽然有记录,却因为记录太少或者太旧而无法评分。

这两批人的处境和分数低不一样。

分数低是一个坏消息,而它至少是一句可以被讨论的话。你可以问为什么低,可以指出哪一笔记错了,可以计划怎么把它抬上去。

没有分数不是坏消息,是没有消息。

一个没有记录的人走进放贷机构,机器读不出任何东西。他不是被判为高风险,他是根本没有被判。而在一套按分数分流的系统里,读不出等于过不去。

更麻烦的是,这批人没有办法靠自己走出来。要有记录,就得先借到钱;要借到钱,就得先有记录。这个圈在设计的时候没有人打算把它做成一个圈,它是分数这套办法的必然副产品:一把只读记录的尺子,对没有记录的人只能报出空白,而空白在流程里会被当作风险。

第十六篇里说过,册子外面站着人。这里的册子外面也站着人,只是他们被挡在外面的方式不同:不是被规定不许上册,是那本册子读不出他们。

第二,分数的准确不是自动得到的。

美国的联邦贸易委员会在 2013年向国会说明,有二成六的消费者在三大征信机构的记录之一上,识别出了可能的重大错误并提出了争议;有百分之五的消费者所涉及的错误,可能导致他们为汽车贷款或者保险付出更高的价格。

这两个数字并排放着很说明问题。

分数是一个精确到个位的三位数。它看上去是算出来的,而且确实是算出来的。而它算的那些原始记录,有相当一部分是错的。

数据错了,数学不会帮你变对。

这句话可以再往前推一步。分数越精确,人越容易相信它;而人越相信它,就越少有人回头去核对喂进去的那些东西。精确本身会减少怀疑。第十四篇里那个七位小数,第十八篇里那个每天更新的指数,和这里的三位数,做的是同一件事:用输出端的精细,替被人们省掉了输入端的追问。

第十四篇里说过一种余项:尺子极精确,而被量的对象不可知。这里是它的另一个版本:尺子极精确,而喂进尺子的东西未经核对。

而这两批溢出的人,处境上有一个共同点。

他们都很难自己动手去改。看不见的人无法证明自己可靠,因为证明的方式恰恰是拥有记录;而记录错了的人要去更正,得先知道自己被记了什么,再找到出错的那一家,再走一遍申诉的流程。

一份关于你的判断在你到场之前已经到场,这句话在 1853年是一个比喻。到这里它是一套流程。

而流程是有门槛的。要走完它,得知道自己有权利走,得看得懂那些表格,得有时间打电话,得在被拒绝一次之后还愿意再来一次。这些条件本身就和一个人的处境有关。于是纠错这件事,在最需要它的人那里最难办成。

五 校准与公平打架

公平的争论一直伴随着分数,而且它比通常想象的要硬得多。

1974年,美国通过平等信贷机会法,明文禁止按种族,肤色,宗教,国籍,性别,婚姻状况,年龄,以及是否领取公共救助来歧视。

把这些变量从模型里拿掉,是一件可以做到的事,而且做到了。

但争论并没有因此结束,因为争论的位置换了。

一派研究和监管论证强调:只要评分更能预测违约,它就能降低放贷成本,扩大信贷的可得性,尤其能帮助那些缺少传统关系网络的人进入标准化金融。这一派的证据是实在的。

另一派指出:即使不使用那些受保护的属性,地理位置,职业,消费轨迹,教育经历,账单支付的历史,也都可能成为代理变量,把历史上的不平等重新编码进分数。一个人住在哪里,和他属于哪个群体之间,在很多地方有相当强的对应关系;而模型并不需要知道后者,只要用了前者,结果可能是一样的。

联储的研究者在 2010年讨论的正是这种情形:模型里没有写种族,而效果上有差别待遇。

到这里为止,这还是一场可以想象终点的争论:找出所有的代理变量,把它们也处理掉。

而后来的算法公平研究给出了一个更难对付的结果。

有研究者证明,在一般情形下,几种常见的公平标准无法同时被满足。另有研究者进一步讨论了校准和误差均衡之间的紧张。

翻成日常的话是这样:你可以要求分数在每一个群体里都同样准,也可以要求它在每一个群体里犯错的方式相同,还可以要求不同群体拿到高分的比例相当。这几条听上去都很合理,而在一般情况下,它们数学上不能同时成立。

这一条对这个系列很要紧,所以要说清楚它的性质。

前面二十篇里,竞争性的解释一次次被并排摆出来而不裁决,理由通常是证据还不够,或者各方在回答不同的问题。

这里不是。

这里是一个被证明了的结果。它说的不是我们还没有找到那个最好的定义,而是那个能同时满足几项合理要求的定义不存在。

也就是说,这一次的不可裁决,不是知识的暂时状态。它是这件事情本身的形状。

于是问题从想让分数更公平,变成了想让它对谁更准确,对谁更少伤害。而这是一个需要有人做出选择,并且承担这个选择的问题,不是一个可以靠更好的算法绕开的问题。

构在这里第一次遇到了一堵它自己算出来的墙。

这堵墙的位置很特别。它不在构和外面的世界之间,它在构自己的几项要求之间。前面二十篇里,余项总是在构的对面:量不到的,不肯量的,拿不出凭据的,不在册子上的。而这一次,溢出的东西出现在构自己的内部:它想要的几样东西彼此打架,而它必须放弃其中一样。

它不是被外面的批评者拦住的。是它自己往前推,推到某一步,发现前面几条路互相排斥。

还要补一层。分数并没有真的把人格判断赶出去。

那位联储研究者早就提醒过,评分的扩张会降低传统地方性银行关系贷款的价值。而现实中的放贷者,仍然会在分数之外看收入,看资产,看就业是否稳定,看有没有担保,并且对边缘的申请做人工复核。

那家公司自己也承认,分数只是输入之一。

所以匿名的金融和人格性的判断并不是简单替代,而是不断重新分配位置:分数把大部分人送进流水线,关系和人工裁量则退到例外,复核,申诉和高价值客户那里。

谁能享受到人的判断,本身也成了一件按分数分配的事。

这一句要读慢。分数最初的用处,是让那些没有关系可用的人也能被看见。而它成熟之后,人的判断变成了一种稀缺资源,分配给边缘的申请和高价值的客户。一个分数干净利落的普通人,反而最不可能碰到一个愿意听他讲讲情况的人,因为他的案子根本不需要有人经手。

六 被写进规则的声誉

如果说个人分数是把一个人是否值得借钱压成三位数,那么债券评级就是把一个发行人是否值得借钱压成字母。

1909年,一位分析师出版了一份关于铁路投资的分析,通常被视为现代债券评级的起点。到 1924年,他的评级已经覆盖了美国债券市场的大部分。

这套字母系统的雄心,是把极其复杂的未来偿债风险,压缩成若干个离散的等级,再让无数互不相识的投资者,保险公司,养老金和监管者,在同一张表上行动。

它把声誉做成了标准件。

不是我认识这个银行家,而是这只债券在不在投资级里。

声誉不再系于某个可以被叙述的人的品性,它系于一套可以远距离复制的分类。

这一步和第一节里那家征信机构做的是同一件事,只是对象换成了机构,而且做得更彻底。一家公司不像一个人那样有品性可讲;它有的是报表,是历史,是行业位置。把这些压成一个字母,比把一个人的为人压成一个字母,在直觉上要容易接受得多。而正因为容易接受,它扩散得更快。

这段历史同样不是向纯客观逐步逼近的故事。

二十世纪七十年代,主要的评级机构从投资者付费转向了发行人付费。1975年,美国的证券监管机构以一套资格制度把评级写进了净资本监管,实际上把私人的评级意见变成了合规的门槛。

支持者说,这提高了市场可用的信息密度,便于监管和投资者快速判断。批评者说,这既制造了评级购物和利益冲突,也抬高了行业的进入门槛,让少数几家机构的意见获得了类似公权力的效果。

2008年前后,这套字母声誉的脆弱性暴露得很彻底。后来的调查用了相当严厉的措辞回顾评级机构在危机中的作用,称它们的失败是金融崩溃里的一个关键齿轮;2006年被评为最高一级的按揭证券,到危机中有极大比例被大幅下调。

第十九篇里已经写过这件事的技术一面。这里要补的是另一面。

要紧的不只是评级错了。是评级一旦被嵌进监管和投资约束,错误就会以制度化的方式扩散:基金必须卖,资本占用要重算,流动性迅速枯竭。

评级之所以危险,不只因为它是声誉。

是因为它是被写进规则的声誉。

一个人的名声不好,后果是别人不愿意和他做生意,而这个后果是分散的,可争论的,可以靠时间和行为改变的。一份评级被下调,后果是几百家机构在同一天被迫做同一个动作,因为他们的章程里写着不许持有低于某一级的东西。

声誉一旦成为规则的一部分,它就不再只是判断。它成了触发器。

而这些机构之所以会一起动,靠的是一句谁都没有特别留意过的话:不许持有低于某一级的东西。它写在几百份章程里,平常从不生效。第十八篇里说过,余项长在项目与项目之间的连线上;这里的连线,就是那几百份文件里的同一行字。

于是关于评级的争论一直分成两路。一方强调,没有这种标准化的判断,大规模的资本市场很难运行。另一方强调,一旦全市场都盯着同一套等级,评级就会从帮助判断变成替你判断,并且激励发行人围绕模型做结构设计,围绕门槛做合规表演。

声誉一旦被做成共同语言,就会有人拿它来打表演赛。

七 五颗星

网络平台把这套历史又推进了一步。它们不只给人打分,还把打分这件事嵌进了每一次交易和每一次劳动。

1996年,一家网络拍卖平台上线了评价论坛。创办人写道:通过创建一个鼓励诚实交易的开放市场,我希望让人们更容易在网上与陌生人做生意。

这句话几乎是整个平台声誉经济的纲领,而且它和 1841年那家征信机构要解决的是同一个问题:如何向一个你不认识的人交付信任。

隔着一百五十五年,两次给出的答案在结构上几乎一样:让一个人过去的行为留下痕迹,把痕迹汇总起来,让不认识他的人可以查。差别只在于谁来写。前一次是雇来的通讯人写,后一次是每一个和他打过交道的人写。看上去后者更民主,而它同时意味着,评判的权力被分给了所有人,责任也就没有落在任何人身上。

制度并不神秘。每笔交易之后互相评价,评价公开累积成一份档案,那份档案再反过来影响后续的交易。

它确实管用,而且管用的程度是可以测量的。

有研究者做过一个配对实验:同一位卖家,用一个高声誉的身份和一个全新的身份出售成对的老明信片。老身份能拿到大约百分之八点一的价格溢价。

量化的声誉可以变成钱,这一点被证明了。

而这也意味着它值得被伪造。一个能变成钱的分数,会招来刷分,养号,互评,买评价,以及在差评出现之前先换一个身份。平台此后所有的规则修补,大半都是在追这件事。构造出一样有价的东西,同时也就造出了针对它的产业。

而裂缝也很快出现。

有研究者在分析这家平台的评价机制时指出,互惠性会扭曲声誉信息的生产:人们给好评,可能是在回报别人给的好评;给差评,也可能是在报复别人给的差评。

2008年,这家平台取消了卖家给买家留负面评价的权利,正是对报复性评价的一次制度回应。

这件事说明了一层容易被忽略的东西。

平台试图把声誉做成一种自发的秩序,让它自己长出来。而最后它仍然要靠平台自己去改规则,限阈值,调信息流。

声誉在这里并不比货币更自然。它同样要靠设计,修改和治理才能运转。

这一点对那种回到口碑就好了的想法,是一个很直接的反驳。口碑在小地方之所以管用,靠的是所有人互相认识,而且要长期住在一起;一个说假话的人,明年还得见到这些人。把这套东西搬到几亿陌生人之间,那个约束就没有了,于是必须由平台造一个出来。造出来的那个,已经不是原来那样东西了。

到了共享和零工的平台,评分又往前走了一步:它从交易之后的参考,变成了能不能继续做下去的门槛。

一家短租平台采用十四天的双盲评价窗口,试图减少即时的报复。一家叫车平台则把司机评分做成了持续的账户治理工具:按官方说明,司机的评分是最近五百次评分的平均值,低于所在城市的最低平均值,可能导致失去平台的访问权。

到这一步,星级已经不只是名誉。

它是接单权,是收入权,是继续劳动的资格。

平台公司常把这套系统描述成效率工具:比人工监管更快,比经理的主观印象更广,比传统的认证更便宜。这些说法都不假。

而实证研究显示,偏见会通过评分机制被放大,而不是被数字自动洗净。

一项关于短租平台的田野实验发现,带有非裔美国人特色名字的住客,请求被接受的概率大约低一成六。2023年的一项研究指出,在某些零工平台上,公开显示评级会把少数族裔之间的评分差距放大八成,收入差距放大二成八。

这不是评分失灵的偶然事故。

它是一个可以说清楚的机制:评分把一个一个的个别偏见聚合成了公共的声誉,然后让后面的人把前面的人的偏见当作信息继续使用。

一个人心里的偏见,影响的是他自己那一次决定。一个被平均出来的星级,影响的是此后所有看到它的人。

前一篇说过,构不但奖励可预测,它还会去制造可预测。这里可以补上一句同形的:构不但记录偏见,它还会把偏见变成后来者的依据。

五颗星看上去比三位数更有人情味。

而实际上它更密,更频繁,也更难躲。你持续地被看见,被评分,被平均,被排序;而这个分数既由具体的关系生成,又反过来脱离关系,冻结关系。

它既依赖人,又把人压平。

这两句不矛盾。它依赖人,是因为每一颗星都由一个具体的人在一次具体的经历之后给出;它压平人,是因为几百颗星被平均之后,那些具体的经历一样也留不下。第十二篇里说过,构量的从来不是事情本身,是事情被组织起来的方式。这里被组织起来的,是几百次互不相干的印象。

八 声誉也想闭合

最后一段要处理的是这个题目里最容易被讲坏的部分。

关于中国的社会信用体系,西方流行的想象常常把它画成一个全国统一,人人都有一个总分,乱穿马路要扣分,分低就不能坐高铁的单一系统。

官方文件,后来的法律化整顿,以及近年的研究都表明,这个画法过于漫画。

2014年的规划纲要确实提出到 2020年基本建成社会信用体系,而它的真实结构更像一个由公共信用信息目录,统一代码,部门数据共享,红黑名单,联合奖惩,行业信用监管,司法执行名单和地方试点拼接而成的制度群,而不是一张全国统一的公民总积分表。

把这一点说准是必要的,而理由不只是求实。一个被描述成全民总分的东西,批评起来很省事,而这种省事会让人看不见真正在起作用的部分。前面几篇里已经出现过同样的教训:第十六篇里那三场饥荒,排除人的技术各不相同,只有把每一种的具体形状说清楚,才谈得上知道它做了什么。

2020年国务院办公厅的一份文件尤其能说明问题,因为它表明中央自己也在给早期的泛化实践踩刹车。文件要求公共信用信息的纳入必须有法律,法规或者中央政策文件的依据,并实行目录制管理;严重失信主体名单和惩戒措施也要有明确的定义,依法依规;个人信息公开须有明确的法律依据或者本人同意,还要做必要的脱敏。

2025年的意见继续强调统一社会信用代码,基础目录,惩戒措施基础清单,信用修复,以及分行业的信用建设。

读这些文件会发现,国家层面真正反复建设的,是信用的基础设施和合规的约束机制,远多于一个全民总分。

这并不等于说它温和。

研究者普遍指出,这套体系真正成熟,真正有穿透力的部分,常常不在个人总分,而在企业监管,法院执行,行业的分级分类监管,以及政府部门之间的联动。有研究明确指出,企业是文件里更主要的目标群体之一,而且这套体系仍然高度依赖人工的调查,核查和行政决定。

最容易被误认成社会信用总分的,其实有两类东西。

第一类是法院系统的失信被执行人名单和限制高消费措施。官方的执行信息平台显示,截至 2019年三月六日,全国累计公布失信被执行人名单一千三百二十六万八千八百零九例,限制乘坐飞机一千九百五十八万二千一百七十九人次,限制乘坐火车五百六十三万五千零八十六人次。

这里的逻辑不是某个日常行为的分数太低,而是与特定的司法文书和执行义务挂钩的具体名单。

把它说成全民总分并不准确。而说它不疼,也不准确。它的代价是实时的,具体的,会堵住出行,融资和经营。

第二类是地方的积分试点。有一个县级市的管理办法明确规定:社会成员的信用积分采用千分制,默认一千分,分若干等级;同时又规定可以通过履约,解释,志愿服务,慈善捐助等方式进行信用修复。有报道说,当地已经编制了涉及六百多项经济社会活动的信息征集目录,归集了一千九百万条基础信息和一百三十万条守信失信信息,把自然人,机关,村居组织和法人都纳入了数据库。

还有一件事必须分开:一家电商集团的信用产品和官方的社会信用体系不是一回事。2015年有八家公司获准开展个人征信的试点准备,那家产品是其中之一,后来并没有正式拿到个人征信牌照。多位研究者都强调,把它当成全国的社会信用总分,是最普遍也最顽固的误读之一。

把漫画式的误解纠正过来,并不意味着可以轻描淡写。

比较妥当的说法是:迄今并没有形成流行叙事里那种全国统一的个人总分;而已经形成并且仍在扩展的,是一套跨部门归集,名单化约束,差异化监管,信用修复与统一识别码并行的信用治理基础设施。它对企业,法院执行对象和特定行业主体的治理效果,往往比对普通居民的总分治理更强,也更稳定。

这个分布本身值得留意。最有穿透力的地方,恰恰是那些有明确法律依据,有具体文书,有确定义务的地方。也就是说,这套东西真正硬的部分并不神秘,它靠的是可以被指认的规则;而那些看上去最像科幻的部分,反倒是最松散,最地方性,也最不稳定的。

现在可以回到这一整段历史真正的问题上了。

关于评分社会的批评,很容易滑向一种浪漫:仿佛货币和分数是冷的,而声誉,口碑,关系才是热的;只要回到后者,问题就会少一些。

而材料给出的答案正好相反。

声誉自己也总想闭合。

十九世纪的征信员想把品格装进手册。网络拍卖平台想把陌生人的可信度做成公开的档案。短租平台和叫车平台想把服务体验和劳动纪律做成连续的星级。地方政府会把道德与守法混进同一套千分制。

每一次都是同一个动作:把一种本来依赖具体关系,具体处境和具体裁量的判断,做成可以远距离流通,可以机械调用,可以大规模复制的东西。

声誉并不天然比货币更宽厚。它同样会把人压成一个判断,同样会出错,同样会被操弄,同样会因为地方的偏见和群体的歧视而伤人。

前面二十篇里反复说过,匿名的尺子和人格化的信任从来不是谁替代谁,它们一直缠在一起。

到这里要把话说得更准一点。

它们缠在一起,不只是因为彼此需要。是因为它们本来就是同一个冲动的两种形态。

想把世界压到一把可结算的尺子上的那个冲动,在货币这一侧表现为价格,在声誉这一侧表现为分数。它们看上去一冷一热,而做的是同一件事:把不可比的东西变成可比的,把不可移动的判断变成可移动的,把一段有前因后果的经历压成一个可以在别处被读出的条目。

两侧的分工也是稳定的。价格管的是交换能不能成,分数管的是这个人能不能进来。前者决定一件事值多少,后者决定你有没有资格站在那张桌子边上。前一篇的结尾说,可预测的外观在广告里被换算成价格;到了这里,同一个外观被换算成资格。

也正因如此,溢出的东西在两边是同一样。

一个逾期一百二十天的人,他为什么逾期,不在记录里。一个星级三点九的司机,他那天为什么迟到,不在星级里。一个信用上看不见的人,他这些年怎么过来的,不在任何一个数里。

这三样东西有一个共同点:它们都不是缺失的信息。它们都存在,都可以被讲出来,都有人知道。它们只是没有一个可以被远距离读出的形式。构不是没看见它们,是它按定义不收这一类东西。它要的是可以在别处被机械调用的条目,而理由从来不是这样的条目。

而每一套评分系统,最后都会长出一个叫作修复的机制。

修复这两个字是一份供词。它承认了一件事:一个人可以改变,而分数是过去的压缩,两者之间有一段差。

有了这段差,就要有一条路让人走回来。

而这条路是构自己开的,开在它自己造的墙上,开完之后还要规定走多久,交什么,做什么才算走完。

一个人变了这件事,不会自动进入账本。它要经过申请,审核,公示,期限,然后才被承认。

在这段时间里,他是他自己,而账上是另一个人。

账还没有算平,它仍旧在记。

1. A Character Arrives Before You Do

In 1853, a merchants' magazine printed a sentence that carried something close to a threat. Wherever you went to buy on credit, it said, a character had already arrived ahead of you; that character would either serve you or ruin you.

The sentence described something that had only just taken shape.

Twelve years earlier, in 1841, a New York merchant had opened a mercantile credit-reporting agency in that city, an enterprise usually credited as the start of the modern credit-reporting industry. Its method was to station correspondents in towns across the country, who would periodically write up a merchant's ability to pay and his standing as a person, and mail these reports back to the New York office.

Among the questions those correspondents were expected to answer were how much capital a man held and how much he owed — but also whether he was of good character, and whether his habits of doing business were sound.

Which is to say: from its very first day, credit was never simply a probability about whether a debt would be repaid. It was, at the same time, a judgment about the man himself.

It is not hard to see why these two things grew up together. In an age with no other evidence to consult, a person's character genuinely was the best available clue to whether he would pay what he owed. The correspondent asking these questions was not conducting a moral inquisition; he was hunting for a usable predictor. But the moment that predictor was written into a ledger, it stopped being merely a clue. It became a record.

The historical significance of this development does not lie in its having eliminated uncertainty. What it did was something more basic: it took a reputation that had once drifted loosely through a community and fixed it into an archivable version of you.

Historians of the period describe this as a shift from rumor to written record, and argue that the nineteenth-century credit-reporting trade effectively invented a financial identity. That identity did not grow naturally out of anything. It was manufactured jointly by the demands of credit, the needs of long-distance trade, and the appearance of information intermediaries willing to sell what they gathered.

From the outset, a gap opened between this identity and the self a person knew from the inside. He knew exactly why his cash had run short that year; the record knew only that it had run short. He knew that he had since paid everything back; the record might still be sitting open on the page where he hadn't. The gap itself was not the problem — every record is cruder than the thing it records. What mattered was that this crude version of him could travel to places he himself could never reach.

The previous essay ended by observing that what gets turned into a market object is never the person, only the person's predictable exterior.

That exterior did not originate in the twenty-first century. It has an ancestor a hundred and eighty years old, and that ancestor was moralized from the start.

What matters most about the 1853 sentence is that it welded two things together.

On one side, credit was beginning to leave the circle of people who knew each other and become a judgment that could be summoned across a distance. Ride into a strange city to buy on credit, and the shopkeeper who has never laid eyes on you can nonetheless look up a file about you.

On the other side, that judgment was never purely anonymous. It arrived loaded with narrative about your habits, your respectability, the company you kept, your private life, even your character.

Depersonalized settlement did not clear personal material out of the way. It moved that material into the files.

That move into the files is worth holding onto, because it will recur in every section that follows, each time in a more respectable disguise. The correspondent's dispatch becomes a printed table; the table becomes a score; the score becomes a star rating; the star rating becomes the basis for tiered regulatory classification. Each version is tidier, more comparable, easier to read somewhere else — and each version stands one step further from the person it was written about.

Across the previous twenty essays in this series, one thread keeps resurfacing: the cold side of an arrangement and the hot side never simply replace one another. What follows here pushes that thread one notch further.

It is going to argue that reputation, too, wants to close the books.

That claim will be redeemed only in the final section. Before it can be, we need to watch, generation by generation, how the scale called reputation kept being ground finer: from character written into a dispatch, to a grade printed in a manual, to a three-digit number arrived at by computation, to a star rating produced by averaging, to a classification written into regulatory law. Each version more orderly than the last — and each one leaving somebody standing outside it.

2. Entered Permanently, Explained Never

By the 1930s, this machine had expanded outward from merchants to ordinary people.

In 1936, a weekly magazine reported that retail credit bureaus in more than a thousand American cities held records on no fewer than 60 million people who bought on credit; any account left unpaid for a hundred and twenty days had to be entered into the permanent record.

The same report offered an example: if a woman moved from Chicago to Los Angeles, a new store clerk there could quickly obtain an entire narrative about her — that she was widowed, that she rented rather than owned, whom she associated with, and how much she owed.

The word permanent is worth pausing over.

An account gone unpaid for a hundred and twenty days can have any number of causes. A business might have failed. Someone might have fallen ill. A husband might have died. Or the money might simply have been unavailable for a while and then repaid in full. These causes differ enormously from one another.

What enters the record is always the same thing: a hundred and twenty days overdue.

This is not an oversight on the part of whoever kept the records. It is, in fact, the very condition that makes the record usable at all. If every entry arrived with its full backstory attached, the clerk in Los Angeles would first have to read an entire story and then decide for herself whether it counted as serious — and she has neither the time nor the standing to make that judgment. For a piece of information to be summonable across a distance, it first has to be stripped of everything that only a local could properly weigh.

Essay 5 described a related move: double-entry bookkeeping anchors every entry so that the books can be made to close. What happens here is a different operation: an experience with its own causes and its own context gets compressed into an entry that can be pulled up in another city entirely. Compression is exactly the function being performed, because without it the entry could never be read in Los Angeles at all.

This is also the basic shape of this entire stretch of history. Credit moved from being a judgment exchanged among peers in the same trade to being a social file kept on the general population.

But this archived version of personhood carried a remainder from the very beginning, and an obvious one.

Historians of this period point out again and again that these dispatches were routinely mixed through with hearsay, prejudice, class judgment, and local politics. What the system did was not strip subjectivity away. It preserved that subjectivity, copied it, and sold it onward.

A judgment about whether someone could be trusted did not become more objective by being filed. It only became easier to circulate.

This point matters for the series as a whole. In the earlier essays, what a construct typically did was take something vague and render it clear. What happens here is different: something that was never clear to begin with gets built into a form that merely looks clear. The judgment did not become more reliable. It simply acquired a shape that could travel. And once it had traveled far enough, it came to look like a fact.

It is worth glancing back at Essay 3 here. Coinage, when it was first struck, gave an already-existing weight a body that could be carried; the weight was real before the coin was made, and a scale could always confirm it again. What is being carried here is a judgment, and a judgment has nothing that can be reweighed. Once it has traveled far enough, no one can take it apart to check how it was originally reached, so it can only be accepted as given.

This is precisely why, when the United States passed a federal law regulating the credit-reporting industry in 1970, the statute settled on a deliberately compromised phrase: credit-reporting agencies should adopt reasonable procedures to meet commercial needs while treating consumers fairly and equitably.

This is not a declaration of victory. It is an admission that the conflict is still there.

A law that could resolve a problem in a single sentence would not be written this way. It is written this way because the lawmakers understood that a tension existed between commercial need and fairness to the individual, that they had no way of dissolving that tension, and that all they could do was require both sides to be attended to.

The remainder makes its first appearance in statutory law right here. The text never specifies what fair and equitable concretely means, never supplies a formula anyone could apply. What it does instead is reserve a place, marking that something exists here which commercial need cannot simply absorb. The almost-never-invoked clause discussed in Essay 17 belongs to the same family: a construct admitting, in its own text, that there is something it has not managed to handle.

3. Forty-Nine Said No

The dividing line for modern credit scoring is not whether a credit record exists, but whether the qualitatively different pieces of information inside that record can be converted into a single number that allows comparison.

In 1956, an engineer and a mathematician together founded a company in the United States, wanting to apply operations research and statistical inference to consumer lending.

In 1958, they pitched this scoring technology to fifty large lending institutions.

Only one responded.

The gap between that one and the other forty-nine tells us that the lending industry at the time did not feel it was missing a scale. It already had a method: send a loan officer to meet the applicant, ask questions, look over the storefront, ask around among his peers, and then make a call. That method was slow and inconsistent, but it worked, in their eyes, because when something went wrong, there was a person who could be held responsible for it. A number cannot give you that.

Thirty years later, these same institutions would be competing to buy the technology. What had changed in between was not the number itself but the volume that had to be processed. Once lending grew from hundreds of loans to millions, sending someone to meet every applicant stopped being arithmetically possible. Nearly every expansion a construct undergoes happens this way: not because someone was persuaded, but because the old method could no longer bear the weight.

This detail is worth keeping in mind, because it shows that what looks self-evident today was not self-evident then. The score was once a poorly understood new kind of arithmetic.

And the form it eventually grew into was not invented all at once by any single genius.

Researchers who have traced the construction of the scorecard describe a joint production: banks, department stores, lenders, computing machinery, historical default data, and shifting regulation together broke the applicant down into a set of characteristics, assigned each characteristic a different weight, and compressed the result into a table that could be executed, deployed, and have its threshold adjusted at will.

Such a researcher insists that the scorecard is not a neutral mirror; it is a market device. It rewrote the loan officer's job. It rewrote the division of labor between the people at the counter and the people behind the scenes. It rewrote how fast an approval could be made and how reliably it could be repeated.

One consequence here is easy to miss. The loan officer had once been required to exercise judgment, and judgment carries responsibility with it — someone who bears responsibility asks questions, looks things over, and when in doubt, asks one more. Once the score arrived, his job shrank to data entry and threshold enforcement. Responsibility traveled along with judgment, off to a place where no one could be called to account for it.

In 1989, the first general-purpose score went live at a credit-reporting agency. In 1995, a government-sponsored housing finance institution began encouraging the use of credit scoring in mortgage lending, after which the secondary market for housing finance carried this practice out across the entire country.

The case made in favor of all this has to be laid out in full, because to a considerable extent it was correct.

A Federal Reserve researcher, summarizing the case made by supporters in 1997, offered a concrete comparison: a traditional small-business loan approval took an average of roughly twelve and a half hours, while a scoring model could compress that same decision to well under an hour.

The Federal Reserve, in testimony given in 2010, described the same position: credit scoring lets lenders evaluate risk faster and more cheaply, and it increases both the supply and the affordability of credit.

None of this is empty rhetoric.

A person without a network of personal connections could barely borrow anything at all in a relationship-based lending world. He does not know the bank manager. No one is prepared to vouch for him. There is no local reputation he can call on. A scoring system that consults only the record gives him, for the first time, something that can actually be read.

Standardization, here, functions as a form of liberation. It took judgment out of the hands of a small number of people exercising discretion.

This case has to be made in full, because the sections that follow are going to describe the costs of this system, and if those costs were presented on their own, the whole arrangement would look like an obviously bad idea that anyone could see through. It was not a bad idea. It let a great many people who could not previously borrow money borrow money — a fact that shows up plainly in the data, and one that the people who pushed this technology genuinely meant to achieve.

Essay 10 already described something of the same shape: the phrase one rate for the whole market was precisely what demolished pricing by pedigree. What appears here is another instance of that same move.

4. Everything the File Leaves Out

That company's score typically runs somewhere between 300 and 850. In the commonly cited breakdown, payment history accounts for 35 percent, amounts owed for 30 percent, length of credit history for 15 percent, new credit for 10 percent, and the mix of credit types for the remaining 10 percent.

The company itself insists, over and over, on one point: the score looks only at what is in the credit record. It does not look directly at income. It does not look at occupation. It does not look at how many years you have held your current job.

This is a self-imposed limit, and the limit has its own logic: looking at less can mean fewer openings for discrimination.

But the other side of self-limitation is that whatever goes unseen does not therefore stop existing. Whether a person has a stable income, what trade he works in, whether someone in his household is sick — these facts often bear a more direct relationship to whether he will repay a debt than how much of his available credit he used last month. The score's refusal to look at them does not mean they stop functioning. It only means the work of looking gets handed off to somebody else — to whoever, beyond the score, still has a decision left to make.

And it is exactly this self-description that exposes the overflow.

First: the score cannot cover everyone.

The United States Consumer Financial Protection Bureau, in 2015, estimated using a 2010 baseline that roughly 26 million American adults were credit invisible — meaning they had no credit record whatsoever — while another 19 million had records too thin or too outdated to be scored at all.

The situation of these two groups is not the same as having a low score.

A low score is bad news, but it is at least a sentence that can be argued with. You can ask why it is low. You can point to an entry that was recorded wrong. You can make a plan for raising it.

Having no score at all is not bad news. It is no news whatsoever.

When a person with no record walks into a lending institution, the machine reads nothing back. He has not been judged high-risk; he has not been judged at all. And in a system that sorts people by score, unreadable amounts to the same thing as unable to get through.

Worse still, this group cannot climb out of the hole on its own. To have a record, one must first borrow; to borrow, one must first have a record. Nobody designed this loop on purpose. It is simply the inevitable byproduct of the scoring method itself: a scale that reads only records can report nothing but blankness for someone who has none, and in the process, that blankness gets treated as risk.

Essay 16 observed that outside the register, there are always people standing. Outside this register, too, there are people standing — only they are kept out differently: not by a rule that forbids their entry, but because the register simply cannot read them.

Second: the score's accuracy is not something that arrives automatically.

The United States Federal Trade Commission told Congress in 2013 that 26 percent of consumers had identified a potentially significant error on one of their files at the three major credit bureaus and disputed it, and that errors affecting 5 percent of consumers could lead to their paying more for auto loans or insurance.

Set side by side, these two figures say a great deal.

The score is a three-digit number, exact to the last integer. It looks computed, and it genuinely is computed. But a considerable share of the raw material it is computed from is simply wrong.

Bad data does not become good because the math applied to it is correct.

This point can be pushed a step further. The more precise a score appears, the more people trust it, and the more people trust it, the fewer of them ever go back to check what was fed into it in the first place. Precision itself dampens suspicion. The seven decimal places in Essay 14, the daily-updating index in Essay 18, and the three-digit number here are all doing the same work: using fineness at the output end to spare everyone the trouble of interrogating the input end.

Essay 14 described one version of a remainder: a scale that is extremely precise, measuring an object that cannot be known. Here is another version of the same thing: a scale that is extremely precise, fed by material that has never been checked.

These two overflowing groups share a further trait.

Neither can easily fix its own situation. The invisible cannot prove their reliability, because the very way one proves it is by having a record; and those whose records are wrong must first discover what has been written about them, then track down which bureau made the error, then work their way through an appeals process.

A judgment about you that arrives before you do — in 1853 that was a figure of speech. By now it is a procedure.

And a procedure has thresholds built into it. Getting through it requires knowing you have the right to try, being able to read the forms, having time to make the calls, and being willing to try again after being turned down once. These are all conditions bound up with a person's circumstances. Which means the work of correcting an error turns out to be hardest to accomplish for exactly the people who need it most.

5. A Wall of Its Own Making

The fight over fairness has accompanied the score from the beginning, and it runs deeper than most people assume.

In 1974, the United States passed the Equal Credit Opportunity Act, explicitly forbidding discrimination on the basis of race, color, religion, national origin, sex, marital status, age, or receipt of public assistance.

Removing these variables from a model is something that can be done, and it was done.

But the argument did not end there, because the ground it was fought on shifted.

One line of research and regulatory argument insists that as long as a score predicts default more accurately, it lowers the cost of lending, widens access to credit, and helps precisely those people who lack traditional networks gain entry into standardized finance. The evidence behind this position is real.

Another line points out that even without touching those protected categories directly, variables like geography, occupation, spending patterns, education, and bill-paying history can all serve as proxies, quietly re-encoding historical inequality back into the score. Where a person lives and which group he belongs to correlate strongly enough in many places that a model never needs to know the latter; using the former can produce the identical outcome.

Federal Reserve researchers, writing in 2010, discussed exactly this scenario: race appears nowhere in the model, and yet the effect is disparate treatment all the same.

Up to this point, the argument still has an endpoint one can imagine: track down every proxy variable and deal with it too.

But later research into algorithmic fairness produced a result far harder to work around.

Researchers have proven that, in the general case, several commonly used fairness criteria cannot all be satisfied at once. Others have gone on to show the specific tension between calibration and the balancing of error rates across groups.

Put in plain terms: you can ask that a score be equally accurate for every group, or you can ask that it make its mistakes in the same way across every group, or you can ask that the proportion of people receiving high scores be comparable between groups. Each of these demands sounds entirely reasonable. And in the general case, mathematics will not let all of them hold true together.

This point matters enormously to this series, so its exact nature needs to be stated with care.

Across the previous twenty essays, competing explanations were repeatedly laid side by side without being adjudicated between, usually on the grounds that the evidence was not yet sufficient, or that the different sides were really answering different questions.

Not here.

Here we have a proven result. It does not say that we have simply failed, so far, to find the best definition of fairness. It says that a definition capable of satisfying several reasonable requirements at once does not exist.

Which is to say: this particular impasse is not a temporary condition of our knowledge. It is the shape of the thing itself.

The question therefore shifts, from wanting the score to be fairer, to deciding whom it should be more accurate for, and whom it should harm less. And that is a question that requires someone to make a choice and answer for it — not a question that a cleverer algorithm can be built to route around.

Here, for the first time, the construct runs into a wall of its own calculation.

The location of this wall is unusual. It does not sit between the construct and the world outside it; it sits among the construct's own several demands on itself. In the previous twenty essays, the remainder always stood on the far side from the construct: what could not be measured, what the construct refused to measure, what had no evidence behind it, what never made it into the register. This time the overflow shows up inside the construct itself. Several things it wants are at war with one another, and it has no choice but to give one of them up.

It was not stopped by critics standing outside it. It pushed forward under its own power, arrived at a certain point, and discovered that the paths ahead of it foreclosed one another.

One further layer needs to be added. The score never actually drove personal judgment out of the picture.

That same Federal Reserve researcher warned long ago that the spread of scoring would diminish the value of old-fashioned relationship banking rooted in local ties. And in practice, lenders still look past the score at income, at assets, at whether employment is stable, at whether a guarantor is available, and they still conduct manual review of borderline applications.

The company itself acknowledges that the score is only one input among several.

So anonymous finance and personal judgment are not simply substitutes for one another. They keep being redistributed: the score sends most applicants down a production line, while relationship and personal discretion retreat to the exceptions, the reviews, the appeals, and the highest-value customers.

Who gets access to a human being's judgment has itself become something allocated by score.

That sentence deserves to be read slowly. The score's original purpose was to make visible the people who had no relationships to rely on. And once it matured, human judgment turned into a scarce resource, rationed out to marginal applications and to premium customers. An ordinary person whose score comes back clean is, paradoxically, the person least likely ever to meet someone willing to hear him out — because his case never requires anyone to handle it at all.

6. The Letter That Became a Trigger

If the personal score compresses whether a person deserves to be lent to into three digits, the bond rating compresses whether an issuer deserves to be lent to into a single letter.

In 1909, an analyst published an analysis of railroad investments generally regarded as the starting point of modern bond ratings. By 1924, his ratings covered most of the American bond market.

The ambition behind this letter system was to compress an extraordinarily complicated future risk of nonpayment into a handful of discrete grades, and then let countless investors, insurance companies, pension funds, and regulators who had never met one another act on the strength of the same table.

It turned reputation into a standardized part.

Not I know this banker, but is this bond investment-grade.

Reputation no longer hung on the character of some person who could be described in a story. It hung instead on a classification that could be reproduced at any distance.

This step performs the same operation as the credit-reporting agency described in the opening section, only its object has shifted from persons to institutions, and it carries the operation through more completely. A company, unlike a person, has no character to speak of in the ordinary sense; what it has are financial statements, a history, a position within its industry. Compressing these into a letter is, intuitively, a far easier thing to accept than compressing a person's conduct into a letter — and precisely because it was easier to accept, it spread faster.

This history, too, is not a story of steady progress toward pure objectivity.

In the 1970s, the major rating agencies switched from a model in which investors paid for ratings to one in which issuers did. In 1975, the American securities regulator, through a system of formal designation, wrote these ratings directly into net-capital regulation, effectively converting a private opinion into a compliance requirement.

Supporters argued that this raised the density of information available to the market and let regulators and investors make faster judgments. Critics argued that it produced both rating shopping and conflicts of interest, while also raising the barriers to entry in the ratings business, so that the opinions of a handful of firms took on something close to the force of public authority.

Around 2008, the fragility of this letter-based reputation was exposed without any ambiguity. Investigations that followed used remarkably harsh language to describe the rating agencies' role in the crisis, calling their failure a key cog in the financial collapse; mortgage securities rated at the very top of the scale in 2006 were, by the time the crisis hit, downgraded in overwhelming numbers.

Essay 19 already covered the technical side of this story. What needs to be added here is the other side of it.

What mattered was not simply that the ratings were wrong. It was that once a rating is built into regulation and into investment covenants, its errors propagate in an institutionalized fashion: funds are compelled to sell, capital charges have to be recalculated, and liquidity dries up almost immediately.

A rating is dangerous not merely because it is a form of reputation.

It is dangerous because it is reputation that has been written into the rules.

When an individual's reputation turns bad, the consequence is that others become unwilling to deal with him — and that consequence is diffuse, arguable, and reversible over time through changed behavior. When a rating is downgraded, the consequence is that hundreds of institutions are forced to take the identical action on the identical day, because their own charters state that they may not hold anything below a given grade.

Once reputation becomes part of the rules, it stops being merely a judgment. It becomes a trigger.

And what makes all these institutions move in unison is a sentence almost nobody paid particular attention to: may not hold anything below a given grade. It sits inside hundreds of charters, ordinarily doing nothing at all. Essay 18 observed that the remainder often lives along the connecting line running between one item and the next; here, that connecting line is that identical sentence, repeated across hundreds of separate documents.

So the argument over ratings has always split into two paths. One side insists that without this kind of standardized judgment, capital markets on this scale could hardly function at all. The other insists that once an entire market fixes its gaze on the same set of grades, a rating stops assisting judgment and starts substituting for it, and it gives issuers every incentive to engineer their structures around the model and stage compliance around the threshold.

Once reputation is turned into a common language, someone will always find a way to use it for a performance staged purely to pass.

7. Five Stars, No Place to Hide

Online platforms pushed this history one step further still. They did not merely assign people scores; they built the act of scoring into every transaction and every unit of labor.

In 1996, an online auction platform launched a feedback forum. Its founder wrote that by creating an open marketplace that encourages honest dealing, he hoped to make it easier for people to do business with strangers online.

That line is very nearly the entire program of the platform reputation economy, and it addresses exactly the problem that the credit-reporting agency of 1841 had set out to solve: how do you extend trust to someone you have never met?

Across a gap of a hundred and fifty-five years, the two answers turn out to be almost identical in structure: let a person's past conduct leave a trace, gather the traces together, and let strangers look them up. The only real difference is who does the writing. The first time, hired correspondents did it. This time, everyone who has ever dealt with the person does it. This looks far more democratic — and it also means that the power to judge has been handed out to everyone, which means responsibility for that judgment now belongs to no one in particular.

The mechanism itself holds no mystery. After every transaction, each party rates the other; the ratings accumulate in public into a file; and that file then shapes every transaction that follows.

It works, and the extent to which it works can be measured.

Researchers ran a matched experiment in which the same seller offered paired sets of old postcards for sale, once under a high-reputation identity and once under a brand-new one. The established identity commanded a price premium of roughly 8.1 percent.

That reputation, once quantified, can be converted into money — this has been demonstrated directly.

Which also means it is worth counterfeiting. Any score that can be turned into money will attract review farming, sock-puppet accounts, mutual back-scratching between raters, purchased reviews, and the trick of switching identities before a bad review lands. Nearly every rule this platform has ever patched has been chasing some version of this problem. To build something valuable is, in the same motion, to build an entire industry devoted to gaming it.

Cracks appeared quickly, too.

Researchers who studied this platform's rating mechanism found that reciprocity distorts the whole production of reputation information: a positive review might simply be repayment for a positive review received, and a negative one might just as easily be retaliation for a negative one received.

In 2008, the platform eliminated sellers' ability to leave negative feedback about buyers — a direct institutional response to retaliatory reviews.

This episode reveals something easy to overlook.

The platform had tried to let reputation form as a kind of spontaneous order, something that would grow on its own. In the end, it still had to rely on itself to rewrite the rules, cap the thresholds, and manage the flow of information.

Reputation, here, turns out to be no more natural than money. It, too, requires design, revision, and active governance in order to function at all.

This is a direct answer to the wish that we could simply go back to word of mouth. Word of mouth works in small communities because everyone knows everyone else, and because people are bound to keep living alongside one another for years to come; a person who lies will still have to face these same neighbors next year. Transplant that arrangement among hundreds of millions of strangers, and the constraint that made it work disappears, forcing the platform to manufacture a substitute. And what gets manufactured is no longer the same thing it was trying to replace.

By the time platforms built around sharing and gig work arrived, scoring took one further step: it stopped being a reference consulted after a transaction was over and became a threshold for whether the work could continue at all.

A short-term rental platform adopted a fourteen-day double-blind review window in an attempt to reduce immediate retaliation. A ride-hailing platform turned driver ratings into an ongoing instrument of account governance: according to its own published policy, a driver's rating is the average of his 500 most recent ratings, and falling below the minimum average set for his city can cost him access to the platform altogether.

By this point, the star rating is no longer simply a reputation.

It is the right to accept fares. It is the right to earn an income. It is the very qualification to keep working.

Platform companies routinely describe this system as a tool of efficiency: faster than human supervision, broader in reach than any single manager's impression, cheaper than traditional credentialing. None of these claims is false.

And empirical research shows that bias is amplified by the scoring mechanism rather than being automatically washed clean by numbers.

A field experiment on a short-term rental platform found that guests with names strongly associated with African Americans were roughly 16 percent less likely to have their booking requests accepted. A 2023 study found that on certain gig platforms, publicly displaying ratings widens the rating gap between racial groups by 80 percent and widens the income gap by 28 percent.

This is not an accident of a system malfunctioning.

It is a mechanism that can be stated plainly: scoring gathers up individual, scattered prejudices and consolidates them into a public reputation, and then lets everyone who comes afterward treat the prejudice of everyone who came before as though it were simply information.

A single person's private bias affects only the one decision he makes. An averaged star rating affects everyone who ever looks at it afterward.

The previous essay observed that the construct does not merely reward what is predictable — it actively manufactures predictability. Here one can add a companion line: the construct does not merely record bias — it converts bias into evidence for whoever comes next.

Five stars looks, on its face, more human than a three-digit number.

In practice it is denser, more constant, and far harder to escape. You are watched continuously, rated continuously, averaged, ranked; and the resulting number is at once generated out of concrete relationships and, at the same time, detached from relationship altogether, freezing relationship in place.

It depends on people, and it flattens people.

These two claims do not contradict one another. It depends on people because every single star is awarded by one specific person after one specific encounter. It flattens people because once hundreds of stars have been averaged together, none of those specific encounters survives the process. Essay 12 observed that what a construct measures was never the thing itself, only the way that thing has been organized. What gets organized here is hundreds of impressions that had nothing whatsoever to do with one another.

8. Reputation Wants to Close the Books Too

This final stretch has to handle the part of the topic most easily mishandled.

Where China's social credit system is concerned, the popular imagination in the West often pictures a single nationwide system in which every citizen carries one aggregate score, jaywalking costs points, and falling below some threshold bars a person from boarding the high-speed train.

Official documents, the legal tightening that came afterward, and recent scholarship all indicate that this picture is far too much a caricature.

The 2014 planning outline did call for basically completing a social credit system by 2020, but its real structure looks much more like a cluster of institutions stitched together — a public credit information catalogue, a unified identifying code, data-sharing across departments, red-and-black lists, joint rewards and punishments, industry-specific credit oversight, judicial enforcement lists, and local pilot programs — rather than a single unified table carrying one aggregate score for every citizen.

Getting this exactly right matters, and not simply for the sake of accuracy. Something described as one national score for every citizen is easy to attack, and that very ease of attack can blind people to the parts that are actually doing the work. The earlier essays already carried this same lesson: Essay 16's three famines each excluded people through technically distinct means, and only by stating each one's specific shape precisely could anyone say what it had actually done.

A 2020 document from the General Office of the State Council is especially telling, because it shows the central government itself applying the brakes to earlier, overgeneralized practice. The document requires that any inclusion of public credit information have a basis in law, in regulation, or in central policy, and be managed through a formal catalogue; it requires that lists of seriously untrustworthy actors and the punitive measures attached to them have clear definitions and proceed strictly according to law and regulation; and it requires that public disclosure of personal information rest on a clear legal basis or the individual's consent, along with the necessary anonymization.

A 2025 policy opinion continues to emphasize the unified social credit code, the foundational catalogue, a baseline list of punitive measures, credit repair, and the building of credit systems tailored to individual industries.

Reading these documents, what becomes clear is that the thing the state has actually and repeatedly built, again and again, is credit infrastructure and mechanisms of compliance constraint — far more than any single aggregate score covering every citizen.

None of this means the system is mild.

Researchers generally agree that the part of this system that has truly matured, the part with real teeth, tends not to be the personal aggregate score at all, but rather corporate regulation, court enforcement, tiered and classified oversight by industry, and coordination among government departments. Some research states explicitly that enterprises are among the more central targets named in these documents, and that the system still relies heavily on manual investigation, verification, and administrative decision-making.

What most often gets mistaken for a nationwide personal score actually falls into two categories.

The first is the judiciary's list of persons subject to enforcement for lack of trustworthiness, together with restrictions on high consumption. The official enforcement information platform shows that, as of March 6, 2019, the list of untrustworthy persons subject to enforcement had been published, cumulatively, for 13,268,809 cases nationwide; 19,582,179 person-instances had been barred from flying, and 5,635,086 person-instances had been barred from riding the train.

The logic at work here is not that someone's score for ordinary daily conduct fell too low. It is a concrete list tied to specific judicial documents and specific enforcement obligations.

Calling this a nationwide aggregate score is inaccurate. But calling it painless is equally inaccurate. Its costs are immediate and concrete: it blocks travel, blocks financing, blocks the ability to run a business.

The second category is local pilot point-systems. One county-level city's administrative measures state explicitly that residents' credit points run on a thousand-point scale, with a default of 1,000 points, divided into several tiers, while also providing that credit can be repaired through fulfilling one's obligations, offering an explanation, performing volunteer service, or making a charitable donation. Reports indicate that this locality has compiled an information-collection catalogue spanning more than 600 categories of economic and social activity, gathering 19 million pieces of basic information and 1.3 million pieces of trustworthy-or-untrustworthy information, and folding individuals, government offices, village and neighborhood organizations, and corporate entities alike into a single database.

One further thing needs to be kept separate: an e-commerce conglomerate's own credit product is not the same thing as the official social credit system. In 2015, eight companies were approved to prepare pilot programs for personal credit reporting, and this product was one of them; it never went on to receive a formal license for personal credit reporting. A number of researchers stress that mistaking it for the nation's aggregate social-credit score is one of the most common, and most stubborn, misreadings in circulation.

Correcting a caricature does not mean the underlying thing can be waved away as harmless.

The more accurate statement is this: no nationwide, unified personal aggregate score of the kind imagined in the popular narrative has been built to date; what has been built, and continues to expand, is a credit-governance infrastructure that combines cross-departmental aggregation, list-based constraint, differentiated regulation, credit repair, and a unified identifying code. Its governing power over enterprises, over the subjects of judicial enforcement, and over particular industry actors tends to be considerably stronger and more stable than its governing power, through any aggregate score, over ordinary residents.

This distribution is itself worth noting. The places where this system truly has teeth are exactly the places with a clear legal basis, concrete documentation, and a defined obligation behind them. Which is to say: the genuinely hard part of this apparatus is not mysterious at all — it runs on rules that can be named and pointed to; while the parts that look most like science fiction turn out to be the loosest, the most local, and the least stable.

We can now return to the real question this entire stretch of history has been building toward.

Criticism of the scored society slides very easily into a kind of romanticism: as though money and scores were cold, while reputation, word of mouth, and personal relationships were warm, and returning to the latter would somehow make the problem smaller.

The evidence gives exactly the opposite answer.

Reputation, too, has always wanted to close its own books.

The nineteenth-century credit reporter wanted to pack character into a manual. The online auction platform wanted to turn a stranger's trustworthiness into a public file. Short-term rental and ride-hailing platforms wanted to turn service experience and the discipline of labor into a continuous star rating. Local governments have mixed morality and lawfulness into a single thousand-point scale.

Every single time, it is the same move: take a judgment that once depended on specific relationships, specific circumstances, and specific discretion, and turn it into something that can travel across distance, be summoned mechanically, and be copied at scale.

Reputation is not naturally more generous than money. It, too, compresses a person into a verdict. It, too, makes mistakes. It, too, gets gamed. It, too, wounds people through local prejudice and group discrimination.

The previous twenty essays said, again and again, that the anonymous scale and personalized trust never simply replace one another — they have always been entangled together.

Here, that claim needs to be sharpened.

They are entangled not merely because each needs the other. It is because they were, from the beginning, two forms taken by a single impulse.

The impulse to compress the world onto one settleable scale shows up, on the side of money, as price, and on the side of reputation, as score. The two look like opposites, one cold and one warm, but they perform the identical operation: they turn the incommensurable into the commensurable, they turn an immovable judgment into a portable one, they compress an experience full of its own causes and context into an entry that can be read somewhere else entirely.

The division of labor between the two sides is equally stable. Price governs whether an exchange can go through; score governs whether a person gets to come in at all. The former decides how much a thing is worth; the latter decides whether you even have standing to sit at the table. The previous essay ended by noting that a predictable exterior gets converted into a price in advertising; here, that same exterior gets converted into a qualification.

And precisely because of this, what overflows on both sides is the very same thing.

Why a man was a hundred and twenty days overdue does not appear in his record. Why a driver rated 3.9 stars was late that particular day does not appear in his rating. How a person invisible to the credit system has actually gotten by these past years appears in no number at all.

These three things share one trait. None of them is missing information. Each one exists, each one could be told, and someone, somewhere, knows it. They simply have no form that can be read across a distance. It is not that the construct failed to notice them — by its very definition, it does not take in this category of thing at all. What it wants is an entry that can be mechanically summoned elsewhere, and the reasons behind that entry are never the kind of thing an entry can hold.

And every scoring system, eventually, grows a mechanism called repair.

The word repair is itself a confession. It admits that a person can change, that a score is only a compression of the past, and that a gap opens up between the two.

Once that gap exists, there has to be some path leading back.

And that path is one the construct opens itself, cut into the very wall it built, and having opened it, the construct still gets to specify how long a person must walk it, what he must submit along the way, and what he must do before it counts as finished.

The fact that a person has changed does not enter the ledger on its own. It has to pass through application, review, publication, and a waiting period before it is finally acknowledged.

During all that time, he is himself, while the ledger still holds someone else.

The ledger has not yet balanced. It is still being kept.