Abstract

Text-to-speech systems are increasingly used for drama dubbing, yet evaluation protocols remain sentence-level, leaving critical scene-level behaviors insufficiently measured. We present SceneTTS-Bench, a benchmark that evaluates TTS along three dimensions: timbre consistency across character turns, emotional expressiveness on high-tension utterances, and rhythm coherence under segmented long-form synthesis. The corpus comprises real-world and generated drama scripts totaling 160 bilingual scenes (approximately 10,300 utterances). A backend-agnostic Canonical Intermediate Representation ensures fair cross-system comparison. Three automatic pipelines produce per-utterance diagnostics: Speaker Consistency Score (SCS), Under-Acting Ratio (UAR), and Rate Discontinuity Ratio (RDR). Experiments on four TTS systems show that each system exhibits distinct weaknesses, and that scene-level rankings diverge substantially from sentence-level metrics.

Keywords

TTS benchmark, multi-speaker dubbing, scene-level evaluation, voice stability, automatic evaluation

Key Figures

The figures below summarize the benchmark design, the scene-level diagnostic view, and the cross-system comparison highlighted in the paper.

Overview of the SceneTTS-Bench pipeline
Figure 1. Overview of SceneTTS-Bench. The benchmark combines a backend-agnostic execution protocol with three evaluation pipelines for timbre consistency, emotional expressiveness, and rhythm coherence.
Scene-level diagnostic example for a high-tension dialogue
Figure 2. Example scene-level diagnostics for a high-tension dialogue, showing how per-utterance analysis exposes issues that are difficult to capture with sentence-level evaluation alone.
Radar comparison of four TTS systems on SceneTTS-Bench
Figure 3. Radar comparison of four TTS systems across the three scene-level dimensions, highlighting distinct performance profiles and failure modes.

Failure Cases

Representative audio cases are organized by benchmark task. Task 1 presents ref-voice-guided timbre-drift groups with extra normal references, Task 2 presents insufficient-arousal and insufficient-dominance under-acting cases with prompt conditioning, and Task 3 corresponds to the rhythm-jump examples from the benchmark pipeline.

Loading curated cases…