AAAI 2027 · BUPT × Li Auto · Project page

CLASVS
Continuous-Latent
Autoregression for
Melody-Preserving
Lyric Editing

Change the words. Keep the melody, timing, singer identity, and expression of the reference performance.

Yizhong Geng1, Tian-Hao Zhang2, Chunfeng Wang2, Wenxin Fu1, Yingming Gao1, Ruimin Wang2, Zhou Pan2, Kun Zhan2, Liang Li3, Ya Li1

1 Beijing University of Posts and Telecommunications 2 Li Auto Inc. 3 Tsinghua University

.0376Macro-PER ↓ best balanced lyric accuracy
.0279Del / Ins PER ↓ mean on variable-length edits
.9410FPC ↑ reference melody preservation
4.13N-MOS ↑ naturalness, 95% CI [4.04, 4.22]

01 / Overview

Selective preservation is the real editing problem.

A reference recording contains exactly what we want to preserve—and the obsolete words we need to reject. CLASVS separates these roles instead of compressing the whole performance into one ambiguous history stream.

S

State

Compact semantic feedback tracks what has already been sung.

C

Control

Target lyrics and reference melody remain globally authoritative.

T

Transition

One raw acoustic predecessor preserves short-range continuity.

02 / Method

State–Control–Transition routing

CLASVS generates normalized 64-D AudioVAE latents in 100-ms patches. A Qwen3-0.6B causal planner retains the immutable controls; repeated Flow-DiT blocks realize the acoustic sequence.

Figure 1. Target lyrics and melody remain in the causal planner cache; semantic feedback records progress, while the raw predecessor is confined to the local acoustic transition.
10 Hz

Variable-length recurrence

The stop head determines output length, so insertions and deletions can expand or compress the realization.

6.25 Hz

Reference melody tokens

Chroma-derived control emphasizes pitch and timing instead of source lexical content.

192 D

Singer conditioning

A frozen CAM++ embedding bypasses the planner and modulates each Flow-DiT block.

Progressive grounding

PSCG teaches every route to do one job well.

  1. State Grounding Ground semantic feedback in realized phonetic progress.
  2. Control Grounding Follow explicit lyrics and melody under imperfect history.
  3. Task Grounding Transfer broad Mandarin coverage to curated singing.
Figure 2. State, control, and task grounding form a three-stage curriculum for robust variable-length lyric editing.

03 / Results

Balanced accuracy without sacrificing the performance.

Confirmatory results use one candidate per item/run on 320 Mandarin edits. Bold cells mark the best reported value among systems with the same automatic reference-and-target input policy.

Objective evaluation

CLA-LyricEdit-320

↓ lower is better · ↑ higher is better
System PSub PER↓ FSub PER↓ Del PER↓ Ins PER↓ Macro-PER↓ FPC↑ SIM↑ Onset ms↓
Vevo2 .0617.0563.1033.0581 .0699.7440.89269.58
YingMusic-Singer-Plus .0171.0245 .0847.0382.0411.9340 .90618.80
CLASVS (1.74B) .0505.0442.0260 .0298.0376 .9410.914 7.68

Macro-PER 95% CI: Vevo2 [.0617, .0781], YingMusic-Singer-Plus [.0356, .0469], CLASVS [.0326, .0429].

Human evaluation

20 native Mandarin listeners

MOS [95% CI]
SystemN-MOSM-MOSL-MOS
Vevo2 3.89 [3.78, 4.00] 3.74 [3.63, 3.85] 3.86 [3.75, 3.97]
YingMusic-Singer-Plus 3.65 [3.55, 3.74] 4.27 [4.19, 4.35] 4.10 [4.01, 4.19]
CLASVS 4.13 [4.04, 4.22] 4.28 [4.20, 4.36] 4.42 [4.34, 4.50]

N-MOS: naturalness · M-MOS: melody adherence · L-MOS: lyric pronunciation accuracy. The CLA–Ying M-MOS difference is not statistically detectable (adjusted p = .79).

Figure 3. Common-input comparison of discrete AR, continuous NAR, and CLASVS continuous AR systems.

00 / Real-song showcase

Edit the words inside real mixed songs.

These three examples start from complete music recordings—not clean, pre-aligned laboratory clips. We automatically separate the vocals, rewrite consecutive lyric lines with CLASVS, and mix the edited singing back into each original accompaniment.

REAL-WORLD CASE 01 · MANDARIN POP

湘江中路

A continuous 45.2-second excerpt with the opening instrumental preserved and every vocal line from 14.2–45.2 seconds rewritten.

8 / 8lines with PER = 0 .985mean FPC 45.2 sfull-context audio
01 Original song 02 Vocal separation 03 CLASVS lyric edit 04 Accompaniment remix
A · SOURCE Original mixed recording

The unedited excerpt, including its 14.2-second instrumental opening and the original eight lyric lines.

B · EDITED CLASVS edit + accompaniment

New lyrics sung in the original melodic and acoustic context, then returned to the same accompaniment.

C · ISOLATED OUTPUT Edited vocal timeline

The same generated vocal without accompaniment, provided to make pronunciation, timing, and phrase transitions easier to inspect.

INPUT POLICY Fully automatic

Original mixed audio + target lyrics. No MIDI, manual note durations, phoneme alignment, or hand-marked edit boundaries.

PRESERVATION TARGET Performance continuity

Keep the reference melody, timing, singer identity, expression, accompaniment, and untouched instrumental context.

THE EIGHT EDITS

What changed inside the performance

PER ↓ · FPC ↑
01

风带着潮湿

风吹散潮湿

0.000PER.999FPC
02

水在讲故事

谁还在等我

0.000PER.945FPC
03

无处停留的我

无处告别的我

0.000PER.990FPC
04

绕着江若无其事

沿着江看夜色落

0.000PER.990FPC
05

三两人经过

三两盏灯火

0.000PER.998FPC
06

路中央有人在唱歌

桥那边有人轻声唱

0.000PER.971FPC
07

我蹲下听着

我停下听着

0.000PER.996FPC
08

她也望向我

晚风拥抱我

0.000PER.988FPC

REAL-WORLD CASE 02 · MANDARIN POP

真英雄

A continuous 69.2-second excerpt that preserves the full opening and rewrites ten consecutive lines into a new, optimistic narrative.

10 / 10lines with PER = 0 .951mean FPC 69.2 sfull-context audio
01 Original song 02 Vocal separation 03 CLASVS lyric edit 04 Accompaniment remix
A · SOURCE Original mixed recording

The unedited excerpt, including its opening instrumental and the original ten lyric lines.

B · EDITED CLASVS edit + accompaniment

Ten selected lyric edits returned to the original accompaniment, including equal-length and syllable-added rewrites.

C · ISOLATED OUTPUT Edited vocal timeline

The same generated vocal without accompaniment, provided to make pronunciation, timing, and phrase transitions easier to inspect.

INPUT POLICY Fully automatic

Original mixed audio + target lyrics. No MIDI, manual note durations, phoneme alignment, or hand-marked edit boundaries.

REWRITE COVERAGE Flexible lyric length

Eight equal-length edits and two one-character additions are realized while preserving the original musical context.

THE TEN EDITS

What changed inside the performance

PER ↓ · FPC ↑
01

醉卧于沙场 听呐喊的沙哑

平凡人登场 也敢笑着出发

0.000PER.954FPC
02

笑看人世间 火树银花

没披风铠甲 心中有光

0.000PER.966FPC
03

数风云叱咤 不过道道伤疤

带着一身伤疤 也要笑着出发

0.000PER.966FPC
04

成王败寇 一念之差

不争输赢 一路到家

0.000PER.872FPC
05

生死一霎那 豪气永放光华

就算偶尔害怕 也把希望高挂

0.000PER.957FPC
06

江山如此大 何处是家

世界那么大 有爱是家

0.000PER.963FPC
07

过重重关卡 看盛世的烟花

走过山与海 仍带笑看晚霞

0.000PER.947FPC
08

赢尽了天下 输了她

平凡的我们 去远方

0.000PER.973FPC
09

颠覆了天下 贪一夜浮夸

不争那天下 只守着牵挂

0.000PER.987FPC
10

人生只不过 一场厮杀

平凡人也能 勇敢出发

0.000PER.921FPC

REAL-WORLD CASE 03 · MANDARIN POP

认真的老去

A continuous 83.7-second excerpt with eighteen lyric lines rewritten across two vocal passages while the original humming transitions remain untouched.

18 / 18lines with PER = 0 .948mean FPC 83.7 sfull-context audio
01 Original song 02 Vocal separation 03 CLASVS lyric edit 04 Accompaniment remix
A · SOURCE Original mixed recording

The original 83.7-second excerpt, including all vocal lines, accompaniment, and the two intervening humming passages.

B · EDITED CLASVS edit + accompaniment

Eighteen edited lines returned to the original accompaniment, covering equal-length, insertion, and deletion rewrites.

C · ISOLATED OUTPUT Edited vocal timeline

The generated vocal timeline without accompaniment, including the untouched humming transitions for continuity inspection.

INPUT POLICY Fully automatic

Original mixed audio + target lyrics. No MIDI, manual note durations, phoneme alignment, or hand-marked edit boundaries.

CONTINUITY CHECK Untouched vocal gestures

The original long and short hums at 36.8–48.7 and 71.5–73.1 seconds pass through unchanged between edited lyric passages.

THE EIGHTEEN EDITS

What changed inside the performance

PER ↓ · FPC ↑
01

忘了什么时候开始

还记得什么时候开始

0.000PER.973FPC
02

到清晨才能入睡

在清晨安然入睡

0.000PER.990FPC
03

也忘了什么叫做结尾

也懂了什么才算结尾

0.000PER.988FPC
04

又有谁在乎呢

总有人在乎呢

0.000PER.991FPC
05

凌晨三点的窗前

城市三点的窗前

0.000PER.979FPC
06

播放着那段时光

回想着那些时光

0.000PER.994FPC
07

有一个骄傲的少年

有一个很勇敢的少年

0.000PER.991FPC
08

隐藏他的青春

珍藏他的青春

0.000PER.988FPC
09

不如让我忘了自己

不如让我找回自己

0.000PER.972FPC
10

你觉得怎么样呢

你说这样好吗

0.000PER.960FPC
11

在每个向往的地方

在努力过的地方

0.000PER.962FPC
12

释然一个遗憾

拥抱这遗憾

0.000PER.888FPC
13

躺在我怀里的吉他

怀里抱着那吉他

0.000PER.754FPC
14

好像厌倦了我

一直陪着我

0.000PER.975FPC
15

重复最熟悉的段落

弹着熟悉的段落

0.000PER.969FPC
16

好像无话可说

好像仍有话说

0.000PER.802FPC
17

这生命正值春光

这生命正有春光

0.000PER.904FPC
18

别装作刀枪不入的模样

别装作不受伤的模样

0.000PER.987FPC

04 / Controlled audio demos

Listen across four lyric-edit operations.

Every example uses the same reference singing and target lyrics for all systems. Highlighted characters show the requested edit. Headphones are recommended.

Loading the offline audio manifest…

05 / Analysis Open SCT mechanism and rollout diagnostics +
Figure 4. Paired route interventions, a frozen progress probe, and failure rates by output-length quartile.

Semantic-state shuffling reduces target-text preference; raw-patch shuffling harms melody continuity more. The frozen progress probe recovers phonetic progress more accurately from semantic state than from the raw acoustic predecessor.

Research demo

Listen critically and use responsibly.

Audio excerpts are presented for non-commercial academic comparison. Generated samples may contain synthetic singing voices and altered lyrics. Rights in any underlying music remain with their respective rights holders.