شماره ركورد
26110
شماره راهنما
LIN2 267
عنوان
بازشناسي گوينده در گفتار بداهه فارسي با استفاده از يادگيري عميق
مقطع تحصيلي
كارشناسي ارشد
رشته تحصيلي
زبانشناسي رايانشي
دانشكده
زبانهاي خارجي
تاريخ دفاع
1405/02/22
صفحه شمار
126 ص .
استاد راهنما
هما اسدي
استاد مشاور
اسفنديار طاهري
كليدواژه فارسي
بازشناسي گوينده , تشخيص گوينده , گفتار پيوسته فارسي , تغييرپذيري درون گوينده , يادگيري عميق , شبكههاي عصبي , پردازش سيگنال صوتي , زبانشناسي رايانشي
چكيده فارسي
بازشناسي گوينده در شرايط گفتار بداهه يكي از چالشبرانگيزترين مسائل پردازش گفتار است، زيرا تغييرپذيري درونگوينده ناشي از نوسانات لحن، سرعت بيان، حالات عاطفي و شرايط محيطي، استخراج نمايشهاي برداري پايدار از هويت گوينده را بهطور قابلتوجهي دشوار ميسازد. اين چالش در زبان فارسي، به دليل محدوديت منابع آموزشي و تنوع لهجهاي، ابعاد پيچيدهتري پيدا ميكند. پژوهش حاضر با هدف ارزيابي تطبيقي سه خانواده مدل يادگيري عميق شامل «SincNet» و «RawNet» و «Wav2Vec2» در بازشناسي گوينده با تمركز بر تغييرپذيري درون گوينده در گفتار بداهه فارسي انجام شده است. پيكره صوتي مورد استفاده شامل 56 گوينده مذكر فارسيزبان با سبك گفتاري بداهه و كاملاً مستقل از متن است كه در شرايط محيطي كنترلنشده ضبط شدهاند. بهمنظور شبيهسازي واقعگرايانه شرايط كاربردي، از تقسيمبندي زماني دادهها بهجاي تقسيمبندي تصادفي استفاده شد و هر پارهگفتار با طول ثابت دو ثانيه پردازش گرديد. تمامي مدلها با تابع زيان سافتمكس با حاشيه زاويهاي افزوده آموزش ديدند. نتايج نشان ميدهد كه مدل «Wav2Vec2-large-xlsr-53» مجهز به مكانيزم تجميع توجهي آگاه از گوينده با دقت 99٫58٪ بر روي مجموعه آزمون، بهطور قاطع از ساير رويكردها پيشي ميگيرد. «SincNet» با دقت 92٫52٪ در رتبه دوم قرار گرفت و بهعنوان گزينهاي سبكوزن و تفسيرپذير براي محيطهاي با منابع محاسباتي محدود معرفي ميشود. «RawNet2» و «RawNet3» بهترتيب با دقت 86٫34٪ و 83٫32٪ عملكرد قابلقبولي نشان دادند. يافتههاي پژوهش تأييد ميكنند كه بهرهگيري از پيشآموزش خودنظارتي چندزبانه در تركيب با مكانيزم توجه پويا، مؤثرترين رويكرد براي مديريت تغييرپذيري درونگوينده در گفتار بداهه فارسي است. اين پژوهش همچنين نشان ميدهد كه تقسيمبندي زماني دادهها، ارزيابي واقعگرايانهتري نسبت به تقسيمبندي تصادفي ارائه ميدهد و از خوشبيني كاذب ناشي از همپوشاني آماري جلوگيري ميكند.
كليدواژه لاتين
Speaker recognition , speaker identification , Persian spontaneous speech , intra-speaker variability , deep learning , neural networks , audio signal processing , computational linguistics
عنوان لاتين
Speaker Recognition in Persian Spontaneous Speech using Deep Learning
گروه آموزشي
زبان شناسي
چكيده لاتين
Speaker recognition under spontaneous speech conditions represents one of the most challenging problems in speech processing, as intra-speaker variability arising from fluctuations in tone, speech rate, emotional states, and environmental conditions significantly hinders the extraction of stable speaker identity representations. This challenge is further compounded in the Persian language due to limited training resources and dialectal diversity. The present study conducts a comparative evaluation of three families of deep learning models — SincNet, RawNet, and Wav2Vec2 — for speaker recognition from Persian spontaneous speech. The speech corpus comprises recordings from 56 male Persian-speaking participants produced in a fully text-independent, spontaneous style under uncontrolled environmental conditions. To realistically simulate deployment scenarios, a chronological data split was employed in lieu of random partitioning, and all utterances were segmented into fixed two-second segments. All models were trained using Additive Angular Margin Softmax loss to maximize inter-class separation and minimize intra-class variance. Results demonstrate that the Wav2Vec2-large-xlsr-53 model, augmented with a Speaker-Aware Attention Pooling mechanism, substantially outperforms all competing approaches, achieving 99.58% accuracy on the held-out test set. The SincNet variant ranks second at 92.52%, establishing itself as a lightweight and interpretable alternative for resource-constrained deployment environments. RawNet2 and RawNet3 attain 86.34% and 83.32% accuracy, respectively. The findings confirm that multilingual self-supervised pre-training combined with dynamic attention-based pooling constitutes the most effective strategy for managing intra-speaker variability in Persian spontaneous speech. Furthermore, the study demonstrates that chronological data partitioning provides a more realistic performance estimate by preventing statistical overlap between training and test distributions — a common source of inflated results in randomly split evaluations. These results establish a rigorous benchmark for speaker recognition in low-resource Persian speech and offer actionable insights for future research on robust speaker modeling in real-world acoustic conditions.
تعداد فصل ها
5
فهرست مطالب pdf
167551
نويسنده