Research Summary (Grant-in-Aid for Scientific Research S)

【Grant-in-Aid for Scientific Research (S)】
Rethinking Native Speaker Disfluency: A Cross-linguistic and Interdisciplinary Study and its Applications in Language Teaching and AI Speech Synthesis
SADANOBU Toshiyuki
Principal
Investigator
Kyoto University, Graduate School of Letters, Professor
SADANOBU Toshiyuki Researcher Number:50235305
Project
Information
Project Number:26K21714 Project Period (FY):2026-2030
Keywords: native speaker disfluency (NSD), spoken language, language difference, language teaching, speech synthesis
Purpose and Background of the Research
Outline of the Research

This research builds upon our predecessor project on native speaker disfluency (hereafter NSD), which focused exclusively on Japanese. In that project, we discovered that NSD follows well-ordered patterns and often produces positive outcomes. Based on these findings, we proposed that much of NSD may be understood not as speech trouble or failure, but rather as a form of behavior that ritualizes hesitation. This perspective raises new questions concerning the languages in which NSD occurs, the extent to which it constitutes "ritualized hesitation," the forms it takes, and the reasons for this phenomenon. By addressing these questions, this study aims to clarify the nature of NSD and provide new insights into the mechanisms of speech and human communication.

Predecessor Project (Regarding Japanese)
  • ・ NSD is well-ordered.
  • ・ NSD often produces positive outcomes.
Hypothesis Most NSDs are actually "ritualized hesitation."
Current Project
Investigation
・ When and why does NSD function as "ritualized hesitation"?
・ Cross-linguistic investigation (nine languages)
・ Conversational analysis (three languages)
・ The nature of disfluency, speech, and communication
Applications
・ Second language teaching
・ AI speech synthesis
Feedback to NSD theory
Figure 1. Overview of the research project
Expected Research Achievements
Languages to be Investigated

The languages investigated in this project include nine languages representing a range of language types: Japanese, Korean, Turkish, Hungarian, Sinhala, Finnish, English, French, and Chinese (Table 1).

Among them, in-depth conversation analyses will be conducted for Japanese, English, and Chinese, which have particularly large speaker populations.

Questions to Elucidate, Issues to be examined

We will investigate, from multiple perspectives, the extent to which different types of NSD in these languages resemble speech troubles or failures, as opposed to "ritualized hesitations."

We will examine the circumstances under which NSD should not be regarded as a problem or failure, what it means for NSD to constitute "ritualized hesitation," and more fundamental questions concerning the nature of disfluency, speech, and communication. Using findings on NSD as a starting point, this project will also explore these broader issues.

Language types Language
Head-final Agglutinative Japanese
Korean
Turkish
Hungarian
Non-agglutinative Sinhala
Head-initial Agglutinative Finnish
Non-agglutinative English
French
Chinese
Table 1. Languages to be investigated
News anchors' speech
(completely fluent)
Learners' speech
(disfluent)
Native speakers' speech
(occasionally disfluent)
Figure 2. Image of disfluency teaching
Application 1: Second Language Teaching

To ensure the empirical validity of the research, the findings will be applied to second language teaching. Specifically, we will develop teaching methods that systematically introduce learners to NSD in Japanese, English, and Chinese (Figure 2).

Currently, second language teaching worldwide focuses primarily on achieving fluent speech, aiming to transform learners' disfluent speech into perfectly fluent speech resembling that of professional news anchors (Figure 2). However, speaking with such fluency is difficult even for many native speakers, making this expectation unrealistic for learners.

Rather than eliminating learners' disfluencies, this project seeks to develop techniques for transforming them into NSD. Because this approach is applicable across languages, it has the potential to significantly influence second language education worldwide.

Application 2: Speech Synthesis

The research findings will also be applied to speech synthesis to develop highly realistic synthetic voices that reproduce the disfluency patterns of native speakers. Speech synthesis is widely used throughout society, and incorporating elements of NSD can bring synthetic speech closer to natural human speech and support smoother interaction. As with second language education, this technology has the potential to influence speech synthesis across many of the world's languages.