Filler Words and Discourse Functions in NotebookLM-Generated Podcast-Style Audio Overviews

Henry Crisna Susanto, Agustinus Hary Setyawan

Abstract


This study investigates filler word use in NotebookLM-generated podcast-style Audio Overviews to examine how AI-mediated spoken discourse simulates spontaneous conversation at the discourse level. The study has two objectives: to identify the types of filler words produced in NotebookLM Audio Overviews and to explain their discourse functions. A qualitative descriptive design was employed. Five English-language CNN news articles representing politics, health, entertainment, public relations, and international conflict topics were uploaded to NotebookLM, and the resulting podcast-style Audio Overviews were transcribed manually. The data were coded for filler type, namely lexical and non-lexical fillers, and analyzed functionally using an adapted functional-pragmatic framework based on Ghanmi’s discussion of filler word functions. The analysis identified 255 filler tokens. Lexical fillers were dominant, with 233 occurrences, whereas non-lexical fillers occurred 22 times. Filler use appeared in all topics, with the politics transcript showing the highest total. All 255 filler tokens were coded according to one dominant discourse function. Turn-taking signals formed the largest functional category, followed by cognitive planning aids and discourse connectors. The findings show that filler words in NotebookLM-generated podcasts are patterned rather than random and contribute to conversational flow, listener orientation, and discourse-level naturalness. The study concludes that filler words in AI-generated podcasts should be understood as pragmatic discourse resources and not merely as signs of disfluency.

Keywords


AI-generated speech; Audio Overview; Discourse function; Filler words; NotebookLM.

Full Text:

PDF

References


K. OSUKA, “Mechanical Speech Synthesizer and Soft Material,” Nippon GOMU KYOKAISHI, vol. 78, no. 8, pp. 307–312, 2005, doi: 10.2324/gomu.78.307.

Z. Zhang, “AI-Powered Intelligent Speech Processing: Evolution, Applications and Future Directions,” Int. J. Adv. Comput. Sci. Appl., vol. 16, no. 2, pp. 918–928, 2025, doi: 10.14569/IJACSA.2025.0160291.

S. Latif, A. Shahid, and J. Qadir, “Generative emotional AI for speech emotion recognition: The case for synthetic emotional speech augmentation,” Appl. Acoust., vol. 210, p. 109425, Jul. 2023, doi: 10.1016/j.apacoust.2023.109425.

N. S. K. Anuradha, “Text-To-Speech Conversion,” Int. J. Adv. Trends Comput. Sci. Eng., vol. 2, no. 6, pp. 269–278, 2013.

Y. Tabet and M. Boughazi, “Speech Synthesis Techniques . A Survey,” 7th Int. Work. Syst. Signal Process. their Appl. SPEECH, pp. 67–70, 2011, doi: 10.1109/WOSSPA.2011.5931414.

X. Tan, T. Qin, F. Soong, and T.-Y. Liu, “A Survey on Neural Speech Synthesis,” 2021, [Online]. Available: http://arxiv.org/abs/2106.15561

R. J. Skerry-Ryan et al., “Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,” 35th Int. Conf. Mach. Learn. ICML 2018, vol. 11, pp. 7471–7480, 2018.

L. Bakkouche et al., “Finding the Human Voice in AI : Insights on the Perception of AI-Voice Clones from Naturalness and Similarity Ratings Finding the Human Voice in AI : Insights on the Perception of AI-Voice Clones from Naturalness and Similarity Ratings,” no. June, 2025, doi: 10.17605/OSF.IO/WF3HA.

Š. Tomková, “Text-To-Speech Technologies in Online Media,” Media Mark. Identity, pp. 699–707, 2024, doi: 10.34135/mmidentity-2024-69.

H. H. Clark and J. E. Fox Tree, “Using uh and um in spontaneous speaking,” Cognition, vol. 84, no. 1, pp. 73–111, 2002, doi: 10.1016/S0010-0277(02)00017-3.

M. Corley and O. W. Stewart, “Hesitation disfluencies in spontaneous speech: The meaning of um,” Linguist. Lang. Compass, vol. 2, no. 4, pp. 589–602, 2008, doi: 10.1111/j.1749-818X.2008.00068.x.

S. Ghanmi, “Filler Words : The Definition and Their Communicative Functions,” Arab. Lang. Lit. Cult., vol. 7, no. 2, pp. 6–10, 2022, doi: 10.11648/j.allc.20220702.11.

H. Bortfeld, S. D. Leon, J. E. Bloom, M. F. Schober, and S. E. Brennan, “Disfluency rates in conversation: Effects of age, relationship, topic, role, and gender,” Lang. Speech, vol. 44, no. 2, pp. 123–147, 2001, doi: 10.1177/00238309010440020101.

S. A. Ruschy, “The Utilization of Filler Words in Relation to Age and Gender,” pp. 1–16, 2024.

B. Clancy, “Discourse-pragmatic markers, fillers and filled pauses: Pragmatic, cognitive, multi-modal and sociolinguistic perspectives, edited by Kate Beeching, Grant Howie, Minna Kirjavainen, and Anna Piasecki,” Contrastive Pragmat., vol. 0393, pp. 1–6, 2024, doi: 10.1163/26660393-bja10125.

U. Setyowati and A. H. Setyawan, “Fillers Used by Speakers on DIVE Studios Podcast,” Linguist. ELT J., vol. 11, no. 2, pp. 131–137, 2023, [Online]. Available: http://journal.ummat.ac.id/index.php/JELTL/article/view/20189

Q. Huang, “A Corpus-based Comparative Analysis on the Use of Non-lexical Filler Words in Spoken English Between Chinese EFL Learners and Native English Speakers,” J. Humanit. Arts Soc. Sci., vol. 8, no. 1, pp. 266–272, 2024, doi: 10.26855/jhass.2024.01.046.

S. Khojastehrad, “Hesitation Strategies in an Oral L2 Test among Iranian Students Shifted from EFL Context to EIL,” Int. J. English Linguist., vol. 2, no. 3, pp. 10–21, 2012, doi: 10.5539/ijel.v2n3p10.

L. R. Nurjamin, A. Nurjamin, and R. Melati, “What EFL Learners Say in Managing Their Speech During Academic Presentations,” in Proceedings of the Twelfth Conference on Applied Linguistics (CONAPLIN 2019), Paris, France: Atlantis Press, 2020, pp. 116–120. doi: 10.2991/assehr.k.200406.023.

Google, “Generate Audio Overview in NotebookLM.” Accessed: Jun. 11, 2025. [Online]. Available: https://support.google.com/notebooklm/answer/16212820?ref_topic=16164070&sjid=3352195109469844277-NC

S. Z. Hassan, P. Lison, and P. Halvorsen, “Enhancing Naturalness in LLM-Generated Utterances through Disfluency Insertion,” 2024, [Online]. Available: http://arxiv.org/abs/2412.12710

S. Wang, J. Gustafson, and É. Székely, “Evaluating Sampling-based Filler Insertion with Spontaneous TTS,” 2022 Lang. Resour. Eval. Conf. Lr. 2022, pp. 1960–1969, 2022.

R. Sharma, P. V. Shah, and A. M. Joshi, “Naturalization of Text by the Insertion of Pauses and Filler Words,” no. August, pp. 1–14, 2020, [Online]. Available: http://arxiv.org/abs/2011.03713

L. Doyle, C. McCabe, B. Keogh, A. Brady, and M. McCann, “An overview of the qualitative descriptive design within nursing research,” J. Res. Nurs., vol. 25, no. 5, pp. 443–455, 2020, doi: 10.1177/1744987119880234.

V. A. L. C. E. Lambert, “Editorial: Qualitative Descriptive Research: An Acceptable Design,” Sch. Inq. DNP Capstone, no. 4, pp. 255–256, 2012.

L. F. de Jesus, T. C. S. Andres, and C. B. Tuazon, “Coding Methods for Enhanced Manual Data Transcripts: Contributing to Sustainable Development Goals (SDG),” J. Lifestyle SDGs Rev., vol. 4, no. 1, p. e01864, 2024, doi: 10.47172/2965-730x.sdgsreview.v4.n00.pe01864.

C. Hacking, H. Verbeek, J. P. H. Hamers, and S. Aarts, “Comparing text mining and manual coding methods: Analysing interview data on quality of care in long-term care for older adults,” PLoS One, vol. 18, no. 11 November, pp. 1–13, 2023, doi: 10.1371/journal.pone.0292578.

J. Bóna and M. Bakti, “The effect of cognitive load on temporal and disfluency patterns of speech,” Target. Int. J. Transl. Stud., vol. 32, no. 3, pp. 482–506, 2020, doi: 10.1075/target.19041.bon.

J. Göbel and K. Mccarthy, “The Impact of Cognitive Load on Speech Production in German-English Bilinguals,” 2023.

О. Новікова and І. Суїма, “The Strategies of Argumentation in Political Speeches,” J. “Ukrainian sense,” no. 2, pp. 76–89, Jan. 2024, doi: 10.15421/462320.

X. Qiu, “Functions of oral monologic tasks: Effects of topic familiarity on L2 speaking performance,” Lang. Teach. Res., vol. 24, no. 6, pp. 745–764, 2020, doi: 10.1177/1362168819829021.




DOI: https://doi.org/10.31764/leltj.v14i1.39395

Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Henry Crisna Susanto, Agustinus Hary Setyawan

Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

_____________________________________________________

Linguistics and ELT Journal

p-ISSN 2339-2940 | e-ISSN 2614-8633

Creative Commons License
LELTJ is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

_____________________________________________________

LELTJ is abstracting & indexing in the following databases:

             

_____________________________________________________

LELTJ Editorial Office: