• شماره ركورد
    25970
  • شماره راهنما
    LIN2 264
  • عنوان

    بازشناسي گوينده در گفتار تلفني زبان فارسي با استفاده از الگوريتم‌هاي يادگيري عميق

  • مقطع تحصيلي
    كارشناسي ارشد
  • رشته تحصيلي
    زبانشناسي رايانشي
  • دانشكده
    زبانهاي خارجي
  • تاريخ دفاع
    21/11/1404
  • صفحه شمار
    100 ص.
  • استاد راهنما
    دكتر هما اسدي
  • استاد مشاور
    بتول علي نژاد
  • كليدواژه فارسي
    بازشناسي گوينده , گفتار تلفني , گفتار آزمايشگاهي , الگوريتم¬هاي يادگيري عميق و نرمال‌سازي نمره
  • چكيده فارسي
    چكيده بازشناسي گوينده در بستر گفتار تلفني، به‌دليل محدوديت پهناي باند، نويزهاي واقعي شبكه و اعوجاج‌هاي كانالي، با چالش «عدم تطابق دامنه» نسبت به گفتار آزمايشگاهي روبه‌رو است و همين مسئله سبب افت محسوس عملكرد سامانه‌هاي متداول مي‌شود. هدف اين پژوهش پياده‌سازي و ارزيابي يك سامانه‌ي بازشناسي گوينده براي زبان فارسي و مطالعه‌ي رفتار آن در مواجهه با شرايط گفتار تلفني و عدم تطابق دامنه است. در اين راستا، رويكرد مبتني بر يادگيري عميق و استخراج بردارهاي تعبيه‌شده‌ي گوينده به‌كار گرفته شد و از دو معماري ECAPA-TDNN و ResNet-34 استفاده شده ‌است. مسئله هم در قالب تشخيص گوينده (طبقه‌بندي چندكلاسه) و هم در قالب تأييد گوينده (تصميم‌گيري مبتني بر شباهت) مورد ارزيابي قرار گرفت. براي كاهش اثر عدم تطابق دامنه، مدل‌ها با استفاده از داده‌هاي تميز آزمايشگاهي آموزش ديده و سپس بر روي داده‌هاي واقعي گفتار تلفني ارزيابي شدند. در بخش تشخيص معيارهاي كسينوسي، منحني‌هاي ROC و DET، مقدار AUC و نرخ خطاي برابر به دست آمد. همچنين اثر يك مرحله پس پردازشي نرمالسازي نمره‌ها در ارزيابي تأييد و بررسي شد. نتايج نشان داد كه هر دو مدل در مواجهه با داده‌هاي تلفني دچار افت عملكرد مي‌شوند كه نشان‌دهنده‌ي نقش پررنگ عدم تطابق دامنه است؛ با اين حال معماري ECAPA-TDNN با بهره‌گيري از مكانيزم توجه كانالي و تجميع چند مقياسي، عملكرد پايدارتري نسبت به ResNet-34 از خود نشان داد و به دقت تشخيص بالاتر و نرخ خطاي برابر كمتري دست يافت. علاوه‌بر اين، ارزيابي تأييد گوينده نشان داد كه اعمال Z-norm به عنوان يك مكمل ارزيابي، مي‌تواند توزيع نمرات را منظم‌تر كرده و معيارهاي نرخ خطاي برابر و AUC را در هر دو مدل بهبود بخشد. در مجموع، يافته‌ها تأكيد مي‌كنند كه براي كاربردهاي تلفني، انتخاب معماري مناسب و استفاده از روش‌هاي نرمال‌سازي نمره، راهكارهاي مؤثري براي كاهش اثرات نامطلوب كانال و دستيابي به سامانه‌اي مقاوم تر در بازشناسي گوينده فارسي محسوب مي‌شود. كليدواژه¬ها: بازشناسي گوينده، گفتار تلفني، گفتار آزمايشگاهي، الگوريتم¬هاي يادگيري عميق و نرمال‌سازي نمره
  • كليدواژه لاتين
    Speaker recognition , telephone speech , laboratory speech , deep learning algorithms an‎d score normalization
  • عنوان لاتين
    Speaker Recognition in Persian Telephone Speech Using Deep Learning Algorithms
  • گروه آموزشي
    زبان شناسي
  • چكيده لاتين
    Abstract Speaker recognition in the context of telephone speech faces the challenge of "amplitude mismatch" compared to laboratory speech due to ban‎dwidth limitations, real network noise, an‎d channel distortions. This leads to a noticeable decline in the performance of conventional systems. The aim of this research is to implement an‎d eva‎luate a speaker recognition system for the Persian language an‎d to study its behavior in the face of telephone speech conditions an‎d domain mismatch. In this regard, a deep learning-based approach an‎d speaker embedding vector extraction were employed, utilizing the ECAPA-TDNN an‎d ResNet-34 architectures. The issue was eva‎luated both in terms of speaker identification (multi-class classification) an‎d speaker verification (similarity-based decision-making). To reduce the effect of domain mismatch, the models were trained using clean laboratory data an‎d then eva‎luated on real telephone speech data. In the detection section, the cosine criteria, ROC curves, an‎d DET curves yielded the AUC value an‎d equal error rate. The effect of a post-processing step of normalizing the scores on the eva‎luation was also confirmed an‎d reviewed. The results showed that both models experienced a decline in performance when faced with telephone data, indicating the significant role of domain mismatch. However, the ECAPA-TDNN architecture, utilizing a channel attention mechanism an‎d multi-scale aggregation, demonstrated more stable performance than ResNet-34, achieving higher recognition accuracy an‎d a lower equal error rate. Additionally, speaker verification eva‎luation showed that applying Z-norm as an eva‎luation supplement can regularize the score distribution an‎d improve equal error rate an‎d AUC metrics in both models. Overall, the findings emphasize that for telephone applications, choosing the appropriate architecture an‎d using score normalization methods are effective strategies for reducing the adverse effects of the channel an‎d achieving a more robust system for Persian speaker recognition. Keywords: Speaker recognition, telephone speech, laboratory speech, deep learning algorithms an‎d score normalization
  • تعداد فصل ها
    5
  • فهرست مطالب pdf
    161857
  • نويسنده

    كوقان، راضيه