戏曲与歌剧
Opera and Chinese Opera
「梅兰芳与卡拉斯的一次跨传统对照」
"A Cross-Traditional Comparison: Mei Lanfang and Maria Callas"
秦汉 Han Qin | 2026
第一篇在听觉经验中提出了"生定展固"模型:听者可能形成期待,期待被确认或改变,再在闭合中留下痕迹。这是SAE的审美假说,不是已经验证的神经定律。
但如果你去看一场京剧,或者看一场歌剧,你会发现一件事:你同时在处理好几条线。唱腔是一条,身体动作是一条,舞台视觉是一条,如果有叙事的话,故事情节还是一条。每一条线都可以独立运行自己的凿构循环。
这就引出了本篇的核心问题:当多条凿构循环同时运行时,发生了什么新的事情?
答案是:跨通道凿。
一、什么是跨通道凿
第一篇为了分析,暂时把纯音乐当作听觉对象;实际音乐会也有身体、空间与视觉。就听觉内部而言,旋律或节奏都可能改变既有期待。
戏曲和歌剧把另一种关系推到前景:一个通道趋于"定"时,另一个通道可能处于"展"。
什么意思?想象一下:你在听一段唱腔,旋律线非常规整,板式完全可预测——你的听觉预测模型是稳定的,你处于"定"的状态。但同时,舞台上的身体动作做了一个你没预料到的事情——一个突然的转身,一个迟疑的手势,一个不合常理的停顿。你的视觉预测模型被打破了。
你的听觉在"定",你的视觉在"展"。两条凿构循环之间出现了相位差。
本文假设,相位差可能制造一种难以归入单一通道的余项:听觉与视觉并置后,整体经验不等于两者分别相加。但结果会因观众熟悉度、具体表演、座位、字幕和录制版本而改变。
本文把这种分析称为跨通道凿。它在戏曲和歌剧中格外清楚,却不为二者独有;电影、舞蹈、仪式与多媒体艺术也能组织类似关系。
二、京剧:板式提供锚点,身体制造凿
京剧高度程式化。西皮、二黄是主要声腔系统,导板、原板、摇板、散板等板式组织节奏与句法;生旦净丑等行当也塑造声音和表演规范。熟悉京剧的观众往往能预期许多落音与收句,但具体唱腔、流派和演员仍保留变化空间。
这看起来像是纯粹的"定"。如果京剧只有唱腔,它就是一个程式化极强、凿空间极小的形式。
但京剧不只有唱腔。它有身段。
唱腔在做什么?南梆子的板式高度规整,旋律线优美但可预测。你的听觉处于"定"——稳定,舒适,你知道下一句会怎样收。
舞剑则可能让视觉期待不断调整。剑花有严格训练和程式依据,并非每个轨迹都不可预测;但演员对速度、幅度、停顿与视线的处理,可能让熟悉程式的观众感到局部偏离。
两条线同时运行。听觉在"定",视觉在"展"。你的认知系统在两个不同的状态之间被拉扯,这个拉扯本身就是余项。
这种余项也许不能由单一通道完整替代:纯录音保留声音细节,却不呈现身体;无声影像保留动作,却改变唱腔提供的时间关系。说它"只存在"于交汇处,是本文的解释,不是可直接测量的事实。
因此,京剧现场与录音具有不同可供性。录音不是被砍残的现场:它可能把嗓音、吐字和历史表演细节带到前景;现场则让身体、空间与观众共同作用。
另一个例子。梅兰芳在这出戏里做的事情更微妙。唱腔本身就有凿(四平调的一些非常规处理),但核心的跨通道凿发生在身段与唱腔的反向运动上:唱腔在往"收"的方向走的时候,身体动作在往"放"的方向走。声音在闭合,身体在打开。两个通道传递的信号是矛盾的。
这种矛盾不是错误,是设计。它制造了一种你用语言很难描述的余项——你知道你感受到了什么,但你说不清楚它到底是"听到的"还是"看到的",因为它既不完全属于任何一个通道。
三、歌剧:不同编码的有限对照
现在把京剧放在一边,看欧洲歌剧。和声、配器、语言与舞台传统不同;本文要检验的,是"相位差"能否作为有限的比较坐标,而不是预设两者做着同一操作。
这是跨通道凿在歌剧中的极致案例。
在第二幕二重唱中,两个声部与乐队不断接续、重叠并延宕闭合。本文可以把某些时刻读成一条线趋于收束、另一条线重新开启,但这是一种听觉分析,不是作品按四步机械排布的事实。
文本关于黑夜、欲望与死亡,音乐则以半音和声、上行推进和延宕不断改写它。把二者听成"叙事趋闭、音乐重开"是一种解释;文字与音乐并不总在说可被简单对立的两件事。
第二幕持续推迟满足,第三幕伊索尔德终场独唱在最后获得著名的和声闭合。把全剧读成长期延宕后带痕收束是本文的模型,但不能说所有余项都在那里"一次性解决",也不能把感受量化为两个多小时的累计值。
比瓦格纳简单得多,但跨通道凿依然在运行。
《今夜无人入睡》的旋律线宽广而明确,卡拉夫在结尾宣告"我将胜利";乐队与和声却同时维持夜色、悬念与逐步扩张的张力。把旋律听成"确定"、乐队听成"尚未确定",是可讨论的批评性读法。
再加一层:如果你懂歌词,你知道卡拉夫在赌命。文本叙事的紧张感和旋律的"豪迈感"之间有一个缝隙。这个缝隙就是跨通道凿制造的余项。
四、同构对照:虞姬舞剑与特里斯坦二重唱
这两个段落的编码系统几乎没有任何共同点。一个是京剧的南梆子加身段,一个是德国浪漫主义歌剧的半音和声加双人声部。语言不同,音乐体系不同,表演传统不同,观众的文化背景不同。
在跨通道凿这一层,本文提出一个可检验的相似性:
一个通道提供锚点(唱腔/其中一条旋律线在"定"),另一个通道在那个锚点之上制造打破(身体动作/另一条旋律线在"展")。两条凿构循环的相位差产生了不可在单一通道内消化的余项。
虞姬舞剑可以从听觉与视觉的错位来读,特里斯坦二重唱则可以从声部、文本与和声的延宕来读。"相位差制造余项"是连接两者的假说,是否成立取决于具体版本和观众所学会的规则。
这里的同构只是一层分析相似,不等于不同传统共享同一机制。
五、异构同效:昆曲游园惊梦与莫扎特唐·乔瓦尼
昆曲《牡丹亭·游园》与莫扎特《唐·乔瓦尼》终场提供另一组比较:不同跨通道方法,是否都可能留下单一通道难以解释的余项?这是一项检验,不是进一步证明。
昆曲是高度程式化的戏曲形式之一。曲牌规定句格、字数、句数、宫调与旋律骨架,并与字音、行腔和身段相互制约;它不是把每个音和每个手指角度都机械固定,演员与流派仍有润腔和表演空间。
但杜丽娘游园这一折恰恰在程式化的极致中制造了一种独特的跨通道凿:唱词在说"春天真美"(construct,确认性的),但昆曲特有的缓慢节奏和水磨腔的处理方式把每一个字都拉成了一个漫长的时间体验。时间本身成了一个通道。你在文字层面得到了确认(这是春天,这是美),但在时间体验层面得到了打破——这个"春天真美"被拉得太长了,长到你开始在那个延展的时间里感觉到别的东西。杜丽娘对春天的感受下面藏着的那个孤独和欲望,不是唱词说出来的,是时间的拉伸暴露出来的。
余项不在文字里,不在旋律里,在时间和文字的相位差里。
完全不同的跨通道凿方式。石像(骑士长的鬼魂)来赴宴的场景。莫扎特做的事情是:叙事通道在construct(鬼魂来审判唐·乔瓦尼,这是一个道德故事的闭合),但音乐通道在chisel。
音乐的处理超出了叙事的逻辑。石像出现时的和声(d小调,长号的阴暗音色)不只是"配合"剧情,它制造了一种超出叙事框架的恐惧。叙事说"坏人受到惩罚",音乐说"这里有一种你无法用道德叙事框定的力量"。你在叙事层面得到了闭合(坏人倒了),但在音乐层面感受到了打开——那个和声的黑暗不是"正义得到伸张"可以解释的。
余项在叙事闭合和音乐打开的缝隙里。这就是为什么这一幕不只是一个道德故事的结尾,而是音乐史上最令人不安的场景之一。
两个作品可以被读成:昆曲让时间和表演扩展文字,莫扎特让音乐超出道德叙事。但它们留下的历史与情感问题并不等价,也不能把经典地位归因于单一"余项"。
异构同效在这里是比较假说:不同方法都可能重开观看,却不承诺同一种不可穷尽。
5a、与Wagner的Gesamtkunstwerk对话
跨通道凿这个概念有一个天然的对话对象:Wagner的Gesamtkunstwerk(总体艺术品)。
1849年,Wagner在《未来的艺术品》里提出:自古希腊以来,音乐、诗歌、戏剧、舞蹈被割裂了,歌剧的使命是把它们重新统一成一个整体。各艺术门类应该消融边界,汇入同一条河,服务于一个共同的目标。
这与本文的"多通道"相关,但不能简化成完全相反的逻辑。
Wagner追求不同艺术在戏剧目的中的协调与综合,不等于所有通道重复同一信息、消除一切张力。本文把注意力放在要素不同步时产生的关系;统一与差异可以同时存在,而非一个只有construct、另一个才有余项。
《特里斯坦》第二幕也不必被说成"违反"Gesamtkunstwerk。文本、声部、和声与舞台行动的张力,可以被看作综合内部的复杂协调。本文用凿构语言重读它,不代表瓦格纳理论被自己的实践推翻。
这其实是Self-as-an-End框架一以贯之的立场:不追求和谐统一,追求在construct内部制造真实的否定,然后看什么活下来。余项不是在融合中产生的,是在矛盾中暴露的。通道之间的相位差就是一种矛盾——听觉告诉你一件事,视觉告诉你另一件事,你的认知系统在两者之间被撕开,撕开的缝隙就是余项。
附带提一句Brecht的间离效果(Verfremdungseffekt)。他跟Wagner恰好站在对面:Wagner要让观众沉浸在统一的幻觉中,Brecht要打破幻觉,让观众意识到自己在看戏。用本文的语言说,Brecht做的是元层面的凿——打破的不是某个通道内部的预测模型,而是"这些通道应该融为一体"这个更高层级的预测模型。他暴露了通道本身。这是另一种跨通道凿,只不过凿的对象不是通道内容之间的相位差,而是"通道存在"这个事实本身。
三个位置由此清晰:Wagner要统一通道,Brecht要暴露通道,本文要利用通道之间的相位差。三者都在处理多通道的问题,但操作方向完全不同。
六、程式化与凿的张力:程式中的微观差异
戏曲和歌剧把声音、身体与舞台程式结合起来;纯器乐同样可能高度程式化,只是规范落在不同材料上。
京剧的板式、行当、身段有严格规范;欧洲歌剧在不同时期也有咏叙调/咏叹调、声部类型和乐队惯例。昆曲曲牌规定框架而非几乎每一个实际音,能乐的面具、步法与扇法也在规范中容纳流派、角色和表演变化。
程式化是一种极强的construct。它让观众的预测模型在开场之前就已经建立了——你知道西皮原板的节奏会怎样走,你知道咏叹调到最后会有一个高音。
这看起来像是"展"的敌人。程式化越强,可凿的空间越小。但事实恰恰相反。
程式越清楚,微小偏离有时越容易被熟悉的观众察觉;但偏离并非必然有力,严格执行也不等于没有创造。许多艺术恰在音色、呼吸、时值与关系处理中体现差异。
因此,"匠人和大师"在这里应当是批评性问题,而不是结构上的客观等级。
一位演员可能以精确执行揭示程式本身的力量,另一位则通过局部偏离重开关系;两者都可能留下余项。
梅兰芳的眼神或卡拉斯的呼吸常被评论者用来说明个体处理,但"迟疑零点几秒"、"统一换气处挪半拍"不是本文实际测量的事实,只能作为说明性设想。
观众是否察觉这些差异,也取决于训练、版本、距离、录音和注意。"说不出哪里不同"不能自动证明真正的展发生了。
反复观看可能让微观处理进入前景,也可能让另一位以克制见长的演员显出新关系;不能规定匠人一遍耗尽、大师十遍不穷尽。
可检验的结构判断应该说明具体差异、比较版本与观众位置,而不是把直觉直接升级为客观品级。
七、反例:当"展"被系统性删除
现在把问题移到制度层面:当国家压制被定义为异端、形式主义或种族污染的艺术时,形式变化会受到怎样的限制?
纳粹德国与斯大林时期苏联都提供了重要案例,但它们不是为本文模型设计的"两个独立实验",历史条件也不可互换。
纳粹德国,1933—1945。 政权以种族主义标准清除犹太及被视为政治、文化上不可靠的音乐家,打击现代主义、爵士与所谓"堕落音乐",同时推崇Bach、Beethoven、Bruckner、Wagner等被纳入"德国性"叙事的作曲家。政策残酷,却并非只允许贝多芬到布鲁克纳一条线;禁令、宣传、市场需要和个人庇护之间也有矛盾。
苏联,1930年代—1950年代。 社会主义现实主义与"形式主义"批判确实限制创作。1936年《真理报》匿名社论猛烈攻击《姆岑斯克县的麦克白夫人》,此前斯大林曾观看演出;但不能写成已证实由斯大林亲自下令封杀,也不能说肖斯塔科维奇此后再无严肃歌剧或芭蕾创作——他仍有未完成、改编和舞台项目。
两种政权都把文化纳入国家控制,却不能被压缩为同一个"删除展、保留生定固"的操作。压制对象、行政结构、审美教条和艺术家的应对方式各不相同。
凿构语言可以帮助提出问题:哪些形式变化被政权解释为失序或威胁?但不协和本身不天然反抗,调性作品也不天然服从。
官方支持的艺术并非一律无聊,也并非无人自愿复听;被批准、妥协、抵抗和暗中编码常在同一作品中纠缠。审美价值不能由作品与政权的距离自动推定。
肖斯塔科维奇作品中的讽刺与双重话语是重要解释传统,但具体段落是否"暗藏颠覆"仍有争论,不能规定第一遍听不出、第十遍才感觉到。
这些历史不能验证"人类认知需要凿"。它们更可靠地说明:国家权力能够伤害艺术家、缩窄制度空间并重塑可公开听见的声音,而作品与听者的反应仍比四步公式复杂。
八、为什么现场和录制是两种体验
跨通道凿也可以比较现场、音频与视频,但三者应被理解为不同可供性,而不是完整版本与残缺版本。
京剧录音不呈现身段,却可能放大吐字、嗓音、伴奏与历史表演细节;现场让身体、空间、观众反应和不可重复事件共同工作。
歌剧录音同样改变注意分配。卡拉斯的舞台表演很重要,但她的传奇性也来自声音、角色塑造、传播史与批评话语,不能断言某个微观偏离在录音中"完全丢失"。
视频同时保留并重新构造视觉:镜头距离、剪辑、收音和屏幕尺度会创造现场没有的关系,也会失去共处空间的重量。
因此,通道较少不必然意味着余项较少。媒介选择改变哪些关系可见、可听、可重复;现场也不天然优于录制。
九、从单通道到跨通道:第一篇到第二篇的进展
回顾一下我们到目前为止建立的东西。
第一篇提出:"生定展固"或许能描述一些听觉期待如何形成、改变与闭合;它不是所有音乐已经得到证明的通用结构。
第二篇提出:多个通道可以分别或相互地组织期待,通道之间的相位差有时会产生单线分析难以说明的余项。这个资源也不为戏曲和歌剧独有。
程式化既可能让微观差异更清楚,也可能让严格执行本身产生力量;"大师/匠人"不能由是否偏离程式客观划分。
这就引出了下一个问题:如果凿构循环可以跨通道运行,那它是不是依赖于某个特定的通道组合?如果我们把听觉通道完全去掉,只留身体动作,凿构循环还能不能成立?
下一篇,我们进入芭蕾和舞蹈——身体作为凿的主通道。当音乐降为背景甚至完全消失时,广播体操和皮娜·鲍什的差别在哪里?
Han Qin | 2026
Essay I proposed Arise-Settle-Unfold-Fix within auditory experience: listeners may form expectations, have them confirmed or altered, and carry traces into closure. This is an SAE aesthetic hypothesis, not an established neural law.
But if you attend a performance of Peking opera, or a performance of Western opera, you notice something: you are processing several lines simultaneously. The singing is one line; bodily movement is another; stage visuals are another; if narrative is present, the story is yet another. Each line can independently run its own chisel-construct cycle.
This leads to the core question of this essay: when multiple chisel-construct cycles run simultaneously, what structurally new phenomenon occurs?
The answer is: cross-channel chisel.
I. What Is Cross-Channel Chisel
For analysis, Essay I temporarily treated pure music as an auditory object; actual concerts also involve bodies, space, and vision. Within audition, melody or rhythm may alter established expectations.
Opera and Chinese opera bring another relation to the foreground: one channel may tend toward Settle while another occupies Unfold.
Consider: you are listening to a vocal passage whose melodic line is entirely regular, its rhythmic-modal framework fully predictable — your auditory predictive model is stable; you are in Settle. But simultaneously, the bodily movement on stage does something you did not anticipate — a sudden turn, a hesitant gesture, an illogical pause. Your visual predictive model is broken.
Your audition is in Settle; your vision is in Unfold. A phase difference has appeared between two chisel-construct cycles.
This essay hypothesizes that phase difference can produce remainder difficult to assign to one channel: after audition and vision are combined, the whole experience is not simply their sum. Results vary with familiarity, performance, seat, surtitles, and recording format.
This is what the essay calls cross-channel chisel. It is especially visible in opera and Chinese opera, but not exclusive to them; film, dance, ritual, and multimedia art can organize comparable relations.
II. Peking Opera: Modal Framework Provides Anchor, Body Produces Chisel
Peking opera is highly formalized. Xipi and erhuang are major tune families, while daoban, yuanban, yaoban, sanban, and other metric patterns organize rhythm and phrasing; sheng, dan, jing, and chou role types shape vocal and performance norms. Familiar viewers can anticipate many cadences, but tunes, schools, and performers still leave room for variation.
This appears to be pure Settle. If Peking opera consisted only of singing, it would be a form with extremely strong formalization and minimal space for chiseling.
But Peking opera has more than singing. It has body work (身段).
What is the singing doing? The nanbanzi modal framework is highly regular, the melodic line beautiful but predictable. Your audition is in Settle — stable, comfortable; you know how the next phrase will cadence.
The sword dance may make visual expectation adjust continuously. Its patterns are trained and conventional rather than wholly unpredictable, but a performer's speed, amplitude, pause, and gaze can create local deviations for viewers familiar with the form.
Two lines run simultaneously. Audition in Settle, vision in Unfold. Your cognitive system is pulled between two different states, and this pulling is itself remainder.
Such remainder may not be fully replaced by either channel alone. Audio preserves vocal detail but not the body; silent images preserve movement while altering the temporal relation supplied by song. To say it "exists only" at the intersection is this essay's interpretation, not a directly measured fact.
Live Peking opera and audio recording therefore offer different affordances. A recording is not an amputated live event: it may foreground voice, diction, accompaniment, and historical detail, while live performance joins body, space, and audience.
Another example. What Mei Lanfang does in this piece is subtler. The singing itself contains chiseling (certain non-standard treatments in the sipingdiao mode), but the core cross-channel chisel occurs in the counter-motion between body work and singing: when the vocal line moves toward "closing," the bodily movement moves toward "opening." Sound is converging; the body is diverging. The two channels transmit contradictory signals.
This contradiction is not error; it is design. It produces a form of remainder that is difficult to describe in language — you know you sensed something, but you cannot determine whether it was "heard" or "seen," because it belongs entirely to neither channel.
III. Western Opera: A Limited Comparison Across Encodings
Set Peking opera aside and consider European opera. Harmony, orchestration, language, and stage tradition differ. The question is whether phase difference can serve as a limited comparative coordinate, not whether the two forms perform an identical operation.
This is the extreme case of cross-channel chisel in opera.
In the Act II duet, the voices and orchestra repeatedly succeed, overlap, and delay one another. Some moments can be heard as one line tending toward closure while another reopens, but this is a listening analysis rather than a fact that the work mechanically follows four steps.
The text concerns night, desire, and death, while chromatic harmony, rising motion, and delay repeatedly revise it. Hearing this as "narrative tending to close while music reopens" is one interpretation; words and music do not consistently state two simply opposed messages.
Act II prolongs deferred satisfaction, while Isolde's final solo in Act III reaches the famous harmonic closure. Reading the opera as extended delay followed by traced closure is this essay's model, but not every remainder is "resolved at once," nor can its weight be quantified as two hours of accumulation.
Far simpler than Wagner, but cross-channel chisel is still operating.
The melody of "Nessun dorma" is broad and direct, and Calaf ends by declaring that he will win. Orchestra and harmony nevertheless sustain night, suspense, and expanding tension. Hearing the melody as "certainty" and the orchestra as "not yet certain" is a debatable critical reading.
Add another layer: if you understand the text, you know Calaf is wagering his life. The tension of textual narrative and the "triumphant feeling" of the melody create a gap. That gap is remainder produced by cross-channel chisel.
IV. Isomorphic Comparison: Consort Yu's Sword Dance and the Tristan Duet
The encoding systems of these two passages share almost no common features. One is Peking opera's nanbanzi mode plus body work; the other is German Romantic opera's chromatic harmony plus dual vocal parts. Language differs, musical system differs, performance tradition differs, audience cultural background differs.
At the level of cross-channel chisel, this essay proposes a testable resemblance:
One channel provides an anchor (the vocal/one melodic line in Settle); another channel creates breaking atop that anchor (body movement/the other melodic line in Unfold). The phase difference between two chisel-construct cycles produces remainder that cannot be absorbed within any single channel.
The sword dance can be read through the offset of audition and vision, and the Tristan duet through delays among voice, text, and harmony. "Phase difference produces remainder" is the hypothesis connecting them; whether it holds depends on the version and on rules the audience has learned.
Isomorphism here means similarity at one analytical level, not a mechanism shared identically by distinct traditions.
V. Heteromorphic Equivalence: Kunqu's "Dream in the Garden" and Mozart's Don Giovanni
Kunqu's The Peony Pavilion: Garden Stroll and the final scene of Mozart's Don Giovanni offer another comparison: might different cross-channel methods both leave something that no single channel explains? This is a test, not further proof.
Kunqu is among the highly formalized forms of Chinese opera. Its fixed tunes prescribe prosody, line and syllable counts, mode, and melodic skeleton in relation to speech tone, vocal elaboration, and body work. They do not mechanically fix every performed note or finger angle; schools and performers retain room for ornamentation and interpretation.
Yet the "Garden Stroll" scene produces a unique form of cross-channel chisel precisely at the extreme of formalization: the lyrics say "spring is beautiful" (construct, confirmatory), but Kunqu's characteristic slow tempo and the水磨腔 (water-polished singing) treatment stretches every syllable into an extended temporal experience. Time itself becomes a channel. At the textual level you receive confirmation (this is spring, this is beauty), but at the level of temporal experience you receive breaking — this "spring is beautiful" is stretched too long, so long that within the stretched time you begin to feel something else. The loneliness and desire beneath Du Liniang's perception of spring is not stated by the lyrics; it is exposed by the stretching of time.
Remainder exists not in the text, not in the melody, but in the phase difference between time and text.
An entirely different cross-channel chisel method. The stone statue (the Commendatore's ghost) comes to dine. What Mozart does is: the narrative channel is in construct (the ghost comes to judge Don Giovanni — this is a moral story's closure), but the music channel is in chisel.
The musical treatment exceeds the logic of the narrative. The harmony at the statue's appearance (D minor, the dark timbre of trombones) does not merely "accompany" the plot; it produces a terror that exceeds the narrative framework. The narrative says "the wicked man is punished"; the music says "there is a force here that your moral narrative cannot frame." At the narrative level you receive closure (the villain falls); at the musical level you feel opening — the harmonic darkness is not something "justice is served" can explain.
Remainder exists in the gap between narrative closure and musical opening. This is why this scene is not merely the conclusion of a moral story but one of the most unsettling scenes in the history of music.
The works can be read as follows: Kunqu lets time and performance expand text, while Mozart lets music exceed a moral narrative. But the historical and emotional questions they leave are not equivalent, and canonical status cannot be attributed to a single remainder.
Heteromorphic equivalence is a comparative hypothesis here: different methods may reopen viewing without promising the same inexhaustibility.
5a. Dialogue with Wagner's Gesamtkunstwerk
Cross-channel chisel has a natural interlocutor: Wagner's concept of Gesamtkunstwerk (total work of art).
In 1849, Wagner argued in The Artwork of the Future that since ancient Greece, music, poetry, drama, and dance had been severed from one another, and that opera's mission was to reunify them into a whole. All art forms should dissolve their boundaries and flow into a single river, serving a common purpose.
This relates to the essay's discussion of multiple channels, but the two positions should not be simplified as exact opposites.
Wagner sought coordination and synthesis of the arts toward a dramatic purpose, not the repetition of one message in every channel or the elimination of all tension. This essay focuses on relations created when elements are out of phase; unity and difference can coexist, rather than one producing only construct and the other alone producing remainder.
Tristan need not be said to "violate" Gesamtkunstwerk. Tensions among text, voices, harmony, and stage action can be understood as complex coordination within synthesis. A chisel-construct reading does not show Wagner's practice refuting his theory.
This is, in fact, the consistent stance of the Self-as-an-End framework: not to pursue harmonious unity, but to produce genuine negation within the construct and observe what survives. Remainder is not produced in fusion; it is exposed in contradiction. The phase difference between channels is a form of contradiction — audition tells you one thing, vision tells you another, your cognitive system is torn between the two, and the tear is remainder.
A brief note on Brecht's Verfremdungseffekt (alienation effect). He stands opposite Wagner: Wagner wants the audience immersed in a unified illusion; Brecht wants to shatter the illusion and make the audience aware they are watching a performance. In the language of this essay, Brecht performs chisel at the meta-level — what he breaks is not the predictive model within any single channel, but the higher-order predictive model that "these channels should fuse into one." He exposes the channels themselves. This is another form of cross-channel chisel, only the object of chiseling is not the phase difference between channel contents, but the fact of "channel existence" itself.
Three positions are thus clear: Wagner unifies channels, Brecht exposes channels, this essay exploits the phase difference between channels. All three address the problem of multiple channels, but in entirely different operational directions.
VI. Formalization and the Tension of Chisel: Micro-Difference Within Convention
Opera and Chinese opera combine vocal, bodily, and stage conventions. Purely instrumental music can be highly formalized too; its norms simply reside in different materials.
Peking opera codifies tune families, roles, and body work; European opera in different periods has conventions of recitative and aria, voice types, and orchestration. Kunqu fixed tunes prescribe frameworks rather than nearly every performed note, while Noh masks, gait, and fan work also permit differences of school, role, and performance.
Formalization is an extremely strong construct. It allows the audience's predictive model to be established before the performance even begins — you know how xipi yuanban's rhythm will proceed; you know the aria will end with a high note.
This appears to be the enemy of Unfold. The stronger the formalization, the smaller the space for chiseling. But the truth is precisely the reverse.
The clearer a convention, the easier a familiar audience may find some small deviations to notice. Yet deviation is not necessarily powerful, and strict execution is not uncreative; timbre, breath, duration, and relation can differentiate performances.
"Artisan and master" should therefore remain a critical question, not an objective structural rank.
One performer may reveal the force of convention through exact execution, another may reopen relations through local deviation, and both may leave remainder.
A glance by Mei Lanfang or breath by Callas is often invoked to illustrate individual handling. But "a fraction of a second" and "half a beat from where everyone breathes" are illustrative scenarios, not measurements made by this essay.
Whether an audience notices such differences depends on training, version, distance, recording, and attention. "I cannot say why it differs" does not automatically prove that genuine Unfold occurred.
Repeated viewing may foreground micro-handling, but it may also reveal new relations in a performer whose art lies in restraint. We cannot decree that an artisan is exhausted once and a master remains inexhaustible after ten encounters.
A testable structural judgment should identify differences, versions, and audience position instead of turning intuition directly into an objective grade.
VII. Counter-Example: When "Unfold" Is Systematically Deleted
Now move the question to institutions: when a state suppresses art defined as deviant, formalist, or racially contaminating, how is formal change constrained?
Nazi Germany and the Stalin-era Soviet Union are important cases, but not two independent experiments designed to test this framework, and their histories are not interchangeable.
Nazi Germany, 1933–1945. The regime used racist criteria to purge Jewish and allegedly unreliable musicians, attacked modernism, jazz, and so-called "degenerate music," and promoted Bach, Beethoven, Bruckner, Wagner, and others folded into a story of Germanness. The policy was brutal, but did not permit only a line from Beethoven to Bruckner; bans, propaganda, practical demand, and personal protection also conflicted.
The Soviet Union, 1930s–1950s. Socialist Realism and campaigns against "formalism" did restrict creation. In 1936 an anonymous Pravda editorial savagely attacked Lady Macbeth of Mtsensk after Stalin had attended a performance. It should not be written as a proven personal ban ordered by Stalin, nor as the end of every serious operatic or ballet project by Shostakovich; unfinished, revised, and stage works followed.
Both regimes subjected culture to state control, but they cannot be compressed into the same operation of "deleting Unfold and retaining Arise-Settle-Fix." Targets, institutions, doctrines, and artists' responses differed.
Chisel-construct language can help pose a question: which formal changes did a regime interpret as disorder or threat? Dissonance is not inherently resistant, and tonality is not inherently compliant.
Officially supported art was not uniformly tedious or devoid of willing listeners. Approval, compromise, resistance, and coded meaning can coexist in one work; aesthetic value does not follow automatically from distance to a regime.
Irony and double speech in Shostakovich are important interpretive traditions, but whether a specific passage conceals subversion remains disputed. We cannot prescribe that it is inaudible once and evident on the tenth hearing.
These histories do not validate a human cognitive need for chisel. More securely, they show that state power can injure artists, narrow institutions, and reshape what may be publicly heard, while works and listeners respond more complexly than a four-step formula.
VIII. Why Live and Recorded Performance Differ
Cross-channel chisel can also compare live, audio, and video performance, but the three are better understood as different affordances than as complete and diminished versions.
A Peking opera recording does not show body work, yet it may foreground diction, vocal timbre, accompaniment, and historical performance detail. Live performance lets bodies, space, audience response, and unrepeatable events interact.
Opera recordings redistribute attention in the same way. Callas's stage acting mattered, but her reputation also arose from voice, characterization, media history, and criticism; one cannot say every micro-deviation is "completely lost" on audio.
Video both preserves and reconstructs vision. Camera distance, editing, recording, and screen scale create relations absent from the theatre while losing the weight of a shared space.
Fewer channels therefore need not mean less remainder. A medium changes which relations are visible, audible, and repeatable; live performance is not inherently superior to recording.
IX. From Single Channel to Cross-Channel: The Progression from Essay I to Essay II
Reviewing what has been established.
Essay I proposed that Arise-Settle-Unfold-Fix may describe how some auditory expectations form, change, and close. It is not a proven universal structure of all music.
Essay II proposes that multiple channels can organize expectation separately and relationally, and that phase difference sometimes leaves remainder a single-line analysis cannot explain. This resource is not exclusive to opera and Chinese opera.
Formalization may clarify micro-difference, while exact execution may itself be powerful. "Master and artisan" cannot be objectively divided by whether a performer deviates from convention.
This raises the next question: if the chisel-construct cycle can run across channels, does it depend on a particular channel combination? If we remove the auditory channel entirely, leaving only bodily movement, can the chisel-construct cycle still hold?
In the next essay, we enter ballet and dance — the body as the primary channel of chisel. When music recedes to the background or vanishes entirely, where lies the difference between military drill and Pina Bausch?