Abstract
Text-to-speech systems are increasingly used for drama dubbing, yet evaluation protocols remain sentence-level, leaving critical scene-level behaviors insufficiently measured. We present SceneTTS-Bench, a benchmark that evaluates TTS along three dimensions: timbre consistency across character turns, emotional expressiveness on high-tension utterances, and rhythm coherence under segmented long-form synthesis. The corpus comprises real-world and generated drama scripts totaling 160 bilingual scenes (approximately 10,300 utterances). A backend-agnostic Canonical Intermediate Representation ensures fair cross-system comparison. Three automatic pipelines produce per-utterance diagnostics: Speaker Consistency Score (SCS), Under-Acting Ratio (UAR), and Rate Discontinuity Ratio (RDR). Experiments on four TTS systems show that each system exhibits distinct weaknesses, and that scene-level rankings diverge substantially from sentence-level metrics.
Keywords
TTS benchmark, multi-speaker dubbing, scene-level evaluation, voice stability, automatic evaluation
Key Figures
The figures below summarize the benchmark design, the scene-level diagnostic view, and the cross-system comparison highlighted in the paper.
Failure Cases
Representative audio cases are organized by benchmark task. Task 1 presents ref-voice-guided timbre-drift groups with extra normal references, Task 2 presents insufficient-arousal and insufficient-dominance under-acting cases with prompt conditioning, and Task 3 corresponds to the rhythm-jump examples from the benchmark pipeline.
Loading curated cases…