شماره ركورد
25970
شماره راهنما
LIN2 264
عنوان
بازشناسي گوينده در گفتار تلفني زبان فارسي با استفاده از الگوريتمهاي يادگيري عميق
مقطع تحصيلي
كارشناسي ارشد
رشته تحصيلي
زبانشناسي رايانشي
دانشكده
زبانهاي خارجي
تاريخ دفاع
21/11/1404
صفحه شمار
100 ص.
استاد راهنما
دكتر هما اسدي
استاد مشاور
بتول علي نژاد
كليدواژه فارسي
بازشناسي گوينده , گفتار تلفني , گفتار آزمايشگاهي , الگوريتم¬هاي يادگيري عميق و نرمالسازي نمره
چكيده فارسي
چكيده
بازشناسي گوينده در بستر گفتار تلفني، بهدليل محدوديت پهناي باند، نويزهاي واقعي شبكه و اعوجاجهاي كانالي، با چالش «عدم تطابق دامنه» نسبت به گفتار آزمايشگاهي روبهرو است و همين مسئله سبب افت محسوس عملكرد سامانههاي متداول ميشود. هدف اين پژوهش پيادهسازي و ارزيابي يك سامانهي بازشناسي گوينده براي زبان فارسي و مطالعهي رفتار آن در مواجهه با شرايط گفتار تلفني و عدم تطابق دامنه است. در اين راستا، رويكرد مبتني بر يادگيري عميق و استخراج بردارهاي تعبيهشدهي گوينده بهكار گرفته شد و از دو معماري ECAPA-TDNN و ResNet-34 استفاده شده است. مسئله هم در قالب تشخيص گوينده (طبقهبندي چندكلاسه) و هم در قالب تأييد گوينده (تصميمگيري مبتني بر شباهت) مورد ارزيابي قرار گرفت. براي كاهش اثر عدم تطابق دامنه، مدلها با استفاده از دادههاي تميز آزمايشگاهي آموزش ديده و سپس بر روي دادههاي واقعي گفتار تلفني ارزيابي شدند. در بخش تشخيص معيارهاي كسينوسي، منحنيهاي ROC و DET، مقدار AUC و نرخ خطاي برابر به دست آمد. همچنين اثر يك مرحله پس پردازشي نرمالسازي نمرهها در ارزيابي تأييد و بررسي شد. نتايج نشان داد كه هر دو مدل در مواجهه با دادههاي تلفني دچار افت عملكرد ميشوند كه نشاندهندهي نقش پررنگ عدم تطابق دامنه است؛ با اين حال معماري ECAPA-TDNN با بهرهگيري از مكانيزم توجه كانالي و تجميع چند مقياسي، عملكرد پايدارتري نسبت به ResNet-34 از خود نشان داد و به دقت تشخيص بالاتر و نرخ خطاي برابر كمتري دست يافت. علاوهبر اين، ارزيابي تأييد گوينده نشان داد كه اعمال Z-norm به عنوان يك مكمل ارزيابي، ميتواند توزيع نمرات را منظمتر كرده و معيارهاي نرخ خطاي برابر و AUC را در هر دو مدل بهبود بخشد. در مجموع، يافتهها تأكيد ميكنند كه براي كاربردهاي تلفني، انتخاب معماري مناسب و استفاده از روشهاي نرمالسازي نمره، راهكارهاي مؤثري براي كاهش اثرات نامطلوب كانال و دستيابي به سامانهاي مقاوم تر در بازشناسي گوينده فارسي محسوب ميشود.
كليدواژه¬ها: بازشناسي گوينده، گفتار تلفني، گفتار آزمايشگاهي، الگوريتم¬هاي يادگيري عميق و نرمالسازي نمره
كليدواژه لاتين
Speaker recognition , telephone speech , laboratory speech , deep learning algorithms and score normalization
عنوان لاتين
Speaker Recognition in Persian Telephone Speech Using Deep Learning Algorithms
گروه آموزشي
زبان شناسي
چكيده لاتين
Abstract
Speaker recognition in the context of telephone speech faces the challenge of "amplitude mismatch" compared to laboratory speech due to bandwidth limitations, real network noise, and channel distortions. This leads to a noticeable decline in the performance of conventional systems. The aim of this research is to implement and evaluate a speaker recognition system for the Persian language and to study its behavior in the face of telephone speech conditions and domain mismatch. In this regard, a deep learning-based approach and speaker embedding vector extraction were employed, utilizing the ECAPA-TDNN and ResNet-34 architectures. The issue was evaluated both in terms of speaker identification (multi-class classification) and speaker verification (similarity-based decision-making). To reduce the effect of domain mismatch, the models were trained using clean laboratory data and then evaluated on real telephone speech data. In the detection section, the cosine criteria, ROC curves, and DET curves yielded the AUC value and equal error rate. The effect of a post-processing step of normalizing the scores on the evaluation was also confirmed and reviewed. The results showed that both models experienced a decline in performance when faced with telephone data, indicating the significant role of domain mismatch. However, the ECAPA-TDNN architecture, utilizing a channel attention mechanism and multi-scale aggregation, demonstrated more stable performance than ResNet-34, achieving higher recognition accuracy and a lower equal error rate. Additionally, speaker verification evaluation showed that applying Z-norm as an evaluation supplement can regularize the score distribution and improve equal error rate and AUC metrics in both models. Overall, the findings emphasize that for telephone applications, choosing the appropriate architecture and using score normalization methods are effective strategies for reducing the adverse effects of the channel and achieving a more robust system for Persian speaker recognition.
Keywords: Speaker recognition, telephone speech, laboratory speech, deep learning algorithms and score normalization
تعداد فصل ها
5
فهرست مطالب pdf
161857
نويسنده