Toward Long-Form and Expressive Speech Synthesis

0:00/0:00

speakers

information

performance location
Ircam, Salle Igor-Stravinsky (Paris)
date
June 29, 2026

Théodor Lemerle's thesis defense

Théodor Lemerle, PhD candidate in the EDITE Doctoral School (ED130) at Sorbonne University, conducted his doctoral research entitled “Toward Long-Form and Expressive Speech Synthesis” within the Sound Analysis and Synthesis team of the STMS laboratory (IRCAM, CNRS, Sorbonne University, French Ministry of Culture), under the supervision of Axel Roebel and Nicolas Obin.

This work was carried out as part of the ANR EXOVOICES project, in collaboration with Lunii, the Laboratoire de Sciences Cognitives et Psycholinguistique (LSCP), and IRCAM.

The examination committee will consist of:

  • Ricard Marxer — Professor, University of Toulon — Reviewer
  • Geoffroy Peeters — Professor, Télécom Paris, Institut Polytechnique de Paris — Reviewer
  • Gérard Biau — Professor, Sorbonne University — Examiner
  • Simon King — Professor, University of Edinburgh — Examiner
  • Berrak Sisman — Assistant Professor, Johns Hopkins University — Examiner
  • Alexandre Défossez — Chief Exploration Officer, Kyutai — Examiner
  • Nicolas Obin — Associate Professor, Sorbonne University — Co-supervisor
  • Axel Roebel — Research Director, IRCAM — PhD Supervisor

Abstract:
This PhD thesis focuses on neural text-to-speech (TTS) synthesis, and more specifically on its adaptation to expressive story telling. It is conducted within the framework of the ANR EXOVOICES project, in collaboration with Lunii, the Laboratoire de Sciences Cognitives et Psycholinguistique (LSCP), and IRCAM. The advent of neural network-based generative models has led to remarkable progress in speech synthesis, making it possible to generate voices that are difficult to distinguish from human speech. However, these advances rely on increasingly large computational infrastructures and ever-growing training datasets, which predominantly consist of short utterances. As a result, current systems are becoming increasingly costly to reproduce and still struggle to generate long-form narration that is coherent, stable, and expressive. First, we propose a speech synthesis system that enables finer conditioning on stylistic and emotional attributes, along with a method for localized control of these attributes in the absence of specifically annotated training data. Second, we introduce a neural speech codec designed to provide a representation well suited for generation while remaining reproducible on commonly available hardware. Finally, we propose a speech synthesis model capable of continuous generation over arbitrarily long durations without loss of stability or coherence. Our approach stems from an empirical analysis of attention mechanisms in conventional neural speech synthesis systems, which suggests an underutilization of the receptive field. This observation motivates a dedicated windowing strategy, which we show enables synthesis to generalize beyond the training horizon while preserving speaker identity and synthesis quality. Taken together, these contributions advance the development of expressive and controllable speech synthesis systems better suited to long-form narration, while maintaining a reproduction cost compatible with the resources typically available in public research laboratories.

IRCAM

1, place Igor-Stravinsky
75004 Paris
+33 1 44 78 48 43

opening times

Monday through Friday 9:30am-7pm
Closed Saturday and Sunday

subway access

Hôtel de Ville, Rambuteau, Châtelet, Les Halles

Institut de Recherche et de Coordination Acoustique/Musique

Copyright © 2022 Ircam. All rights reserved.