State
Compact semantic feedback tracks what has already been sung.
AAAI 2027 · BUPT × Li Auto · Project page
Change the words. Keep the melody, timing, singer identity, and expression of the reference performance.
01 / Overview
A reference recording contains exactly what we want to preserve—and the obsolete words we need to reject. CLASVS separates these roles instead of compressing the whole performance into one ambiguous history stream.
Compact semantic feedback tracks what has already been sung.
Target lyrics and reference melody remain globally authoritative.
One raw acoustic predecessor preserves short-range continuity.
02 / Method
CLASVS generates normalized 64-D AudioVAE latents in 100-ms patches. A Qwen3-0.6B causal planner retains the immutable controls; repeated Flow-DiT blocks realize the acoustic sequence.
The stop head determines output length, so insertions and deletions can expand or compress the realization.
Chroma-derived control emphasizes pitch and timing instead of source lexical content.
A frozen CAM++ embedding bypasses the planner and modulates each Flow-DiT block.
Progressive grounding
03 / Results
Confirmatory results use one candidate per item/run on 320 Mandarin edits. Bold cells mark the best reported value among systems with the same automatic reference-and-target input policy.
| System | PSub PER↓ | FSub PER↓ | Del PER↓ | Ins PER↓ | Macro-PER↓ | FPC↑ | SIM↑ | Onset ms↓ |
|---|---|---|---|---|---|---|---|---|
| Vevo2 | .0617 | .0563 | .1033 | .0581 | .0699 | .7440 | .892 | 69.58 |
| YingMusic-Singer-Plus | .0171 | .0245 | .0847 | .0382 | .0411 | .9340 | .906 | 18.80 |
| CLASVS (1.74B) | .0505 | .0442 | .0260 | .0298 | .0376 | .9410 | .914 | 7.68 |
Macro-PER 95% CI: Vevo2 [.0617, .0781], YingMusic-Singer-Plus [.0356, .0469], CLASVS [.0326, .0429].
| System | N-MOS | M-MOS | L-MOS |
|---|---|---|---|
| Vevo2 | 3.89 [3.78, 4.00] | 3.74 [3.63, 3.85] | 3.86 [3.75, 3.97] |
| YingMusic-Singer-Plus | 3.65 [3.55, 3.74] | 4.27 [4.19, 4.35] | 4.10 [4.01, 4.19] |
| CLASVS | 4.13 [4.04, 4.22] | 4.28 [4.20, 4.36] | 4.42 [4.34, 4.50] |
N-MOS: naturalness · M-MOS: melody adherence · L-MOS: lyric pronunciation accuracy. The CLA–Ying M-MOS difference is not statistically detectable (adjusted p = .79).
00 / Real-song showcase
These three examples start from complete music recordings—not clean, pre-aligned laboratory clips. We automatically separate the vocals, rewrite consecutive lyric lines with CLASVS, and mix the edited singing back into each original accompaniment.
REAL-WORLD CASE 01 · MANDARIN POP
A continuous 45.2-second excerpt with the opening instrumental preserved and every vocal line from 14.2–45.2 seconds rewritten.
The unedited excerpt, including its 14.2-second instrumental opening and the original eight lyric lines.
New lyrics sung in the original melodic and acoustic context, then returned to the same accompaniment.
The same generated vocal without accompaniment, provided to make pronunciation, timing, and phrase transitions easier to inspect.
Original mixed audio + target lyrics. No MIDI, manual note durations, phoneme alignment, or hand-marked edit boundaries.
Keep the reference melody, timing, singer identity, expression, accompaniment, and untouched instrumental context.
风带着潮湿
→风吹散潮湿
水在讲故事
→谁还在等我
无处停留的我
→无处告别的我
绕着江若无其事
→沿着江看夜色落
三两人经过
→三两盏灯火
路中央有人在唱歌
→桥那边有人轻声唱
我蹲下听着
→我停下听着
她也望向我
→晚风拥抱我
REAL-WORLD CASE 02 · MANDARIN POP
A continuous 69.2-second excerpt that preserves the full opening and rewrites ten consecutive lines into a new, optimistic narrative.
The unedited excerpt, including its opening instrumental and the original ten lyric lines.
Ten selected lyric edits returned to the original accompaniment, including equal-length and syllable-added rewrites.
The same generated vocal without accompaniment, provided to make pronunciation, timing, and phrase transitions easier to inspect.
Original mixed audio + target lyrics. No MIDI, manual note durations, phoneme alignment, or hand-marked edit boundaries.
Eight equal-length edits and two one-character additions are realized while preserving the original musical context.
醉卧于沙场 听呐喊的沙哑
→平凡人登场 也敢笑着出发
笑看人世间 火树银花
→没披风铠甲 心中有光
数风云叱咤 不过道道伤疤
→带着一身伤疤 也要笑着出发
成王败寇 一念之差
→不争输赢 一路到家
生死一霎那 豪气永放光华
→就算偶尔害怕 也把希望高挂
江山如此大 何处是家
→世界那么大 有爱是家
过重重关卡 看盛世的烟花
→走过山与海 仍带笑看晚霞
赢尽了天下 输了她
→平凡的我们 去远方
颠覆了天下 贪一夜浮夸
→不争那天下 只守着牵挂
人生只不过 一场厮杀
→平凡人也能 勇敢出发
REAL-WORLD CASE 03 · MANDARIN POP
A continuous 83.7-second excerpt with eighteen lyric lines rewritten across two vocal passages while the original humming transitions remain untouched.
The original 83.7-second excerpt, including all vocal lines, accompaniment, and the two intervening humming passages.
Eighteen edited lines returned to the original accompaniment, covering equal-length, insertion, and deletion rewrites.
The generated vocal timeline without accompaniment, including the untouched humming transitions for continuity inspection.
Original mixed audio + target lyrics. No MIDI, manual note durations, phoneme alignment, or hand-marked edit boundaries.
The original long and short hums at 36.8–48.7 and 71.5–73.1 seconds pass through unchanged between edited lyric passages.
忘了什么时候开始
→还记得什么时候开始
到清晨才能入睡
→在清晨安然入睡
也忘了什么叫做结尾
→也懂了什么才算结尾
又有谁在乎呢
→总有人在乎呢
凌晨三点的窗前
→城市三点的窗前
播放着那段时光
→回想着那些时光
有一个骄傲的少年
→有一个很勇敢的少年
隐藏他的青春
→珍藏他的青春
不如让我忘了自己
→不如让我找回自己
你觉得怎么样呢
→你说这样好吗
在每个向往的地方
→在努力过的地方
释然一个遗憾
→拥抱这遗憾
躺在我怀里的吉他
→怀里抱着那吉他
好像厌倦了我
→一直陪着我
重复最熟悉的段落
→弹着熟悉的段落
好像无话可说
→好像仍有话说
这生命正值春光
→这生命正有春光
别装作刀枪不入的模样
→别装作不受伤的模样
04 / Controlled audio demos
Every example uses the same reference singing and target lyrics for all systems. Highlighted characters show the requested edit. Headphones are recommended.
Loading the offline audio manifest…
Semantic-state shuffling reduces target-text preference; raw-patch shuffling harms melody continuity more. The frozen progress probe recovers phonetic progress more accurately from semantic state than from the raw acoustic predecessor.
Research demo
Audio excerpts are presented for non-commercial academic comparison. Generated samples may contain synthetic singing voices and altered lyrics. Rights in any underlying music remain with their respective rights holders.