• شماره ركورد
    26110
  • شماره راهنما
    LIN2 267
  • عنوان

    بازشناسي گوينده در گفتار بداهه فارسي با استفاده از يادگيري عميق

  • مقطع تحصيلي
    كارشناسي ارشد
  • رشته تحصيلي
    زبانشناسي رايانشي
  • دانشكده
    زبانهاي خارجي
  • تاريخ دفاع
    1405/02/22
  • صفحه شمار
    126 ص .
  • استاد راهنما
    هما اسدي
  • استاد مشاور
    اسفنديار طاهري
  • كليدواژه فارسي
    بازشناسي گوينده , تشخيص گوينده , گفتار پيوسته فارسي , تغييرپذيري درون گوينده , يادگيري عميق , شبكه‌هاي عصبي , پردازش سيگنال صوتي , زبانشناسي رايانشي
  • چكيده فارسي
    بازشناسي گوينده در شرايط گفتار بداهه يكي از چالش‌برانگيزترين مسائل پردازش گفتار است، زيرا تغييرپذيري درون‌گوينده ناشي از نوسانات لحن، سرعت بيان، حالات عاطفي و شرايط محيطي، استخراج نمايش‌هاي برداري پايدار از هويت گوينده را به‌طور قابل‌توجهي دشوار مي‌سازد. اين چالش در زبان فارسي، به دليل محدوديت منابع آموزشي و تنوع لهجه‌اي، ابعاد پيچيده‌تري پيدا مي‌كند. پژوهش حاضر با هدف ارزيابي تطبيقي سه خانواده مدل يادگيري عميق شامل «SincNet» و «RawNet» و «Wav2Vec2» در بازشناسي گوينده با تمركز بر تغييرپذيري درون گوينده در گفتار بداهه فارسي انجام شده است. پيكره صوتي مورد استفاده شامل 56 گوينده مذكر فارسي‌زبان با سبك گفتاري بداهه و كاملاً مستقل از متن است كه در شرايط محيطي كنترل‌نشده ضبط شده‌اند. به‌منظور شبيه‌سازي واقع‌گرايانه شرايط كاربردي، از تقسيم‌بندي زماني داده‌ها به‌جاي تقسيم‌بندي تصادفي استفاده شد و هر پاره‌گفتار با طول ثابت دو ثانيه پردازش گرديد. تمامي مدل‌ها با تابع زيان سافت‌مكس با حاشيه زاويه‌اي افزوده آموزش ديدند. نتايج نشان مي‌دهد كه مدل «Wav2Vec2-large-xlsr-53» مجهز به مكانيزم تجميع توجهي آگاه از گوينده با دقت 99٫58٪ بر روي مجموعه آزمون، به‌طور قاطع از ساير رويكردها پيشي مي‌گيرد. «SincNet» با دقت 92٫52٪ در رتبه دوم قرار گرفت و به‌عنوان گزينه‌اي سبك‌وزن و تفسيرپذير براي محيط‌هاي با منابع محاسباتي محدود معرفي مي‌شود. «RawNet2» و «RawNet3» به‌ترتيب با دقت 86٫34٪ و 83٫32٪ عملكرد قابل‌قبولي نشان دادند. يافته‌هاي پژوهش تأييد مي‌كنند كه بهره‌گيري از پيش‌آموزش خودنظارتي چندزبانه در تركيب با مكانيزم توجه پويا، مؤثرترين رويكرد براي مديريت تغييرپذيري درون‌گوينده در گفتار بداهه فارسي است. اين پژوهش همچنين نشان مي‌دهد كه تقسيم‌بندي زماني داده‌ها، ارزيابي واقع‌گرايانه‌تري نسبت به تقسيم‌بندي تصادفي ارائه مي‌دهد و از خوش‌بيني كاذب ناشي از هم‌پوشاني آماري جلوگيري مي‌كند.
  • كليدواژه لاتين
    Speaker recognition , speaker identification , Persian spontaneous speech , intra-speaker variability , deep learning , neural networks , audio signal processing , computational linguistics
  • عنوان لاتين
    Speaker Recognition in Persian Spontaneous Speech using Deep Learning
  • گروه آموزشي
    زبان شناسي
  • چكيده لاتين
    Speaker recognition under spontaneous speech conditions represents one of the most challenging problems in speech processing, as intra-speaker variability arising from fluctuations in tone, speech rate, emotional states, an‎d environmental conditions significantly hinders the extraction of stable speaker identity representations. This challenge is further compounded in the Persian language due to limited training resources an‎d dialectal diversity. The present study conducts a comparative eva‎luation of three families of deep learning models — SincNet, RawNet, an‎d Wav2Vec2 — for speaker recognition from Persian spontaneous speech. The speech corpus comprises recordings from 56 male Persian-speaking participants produced in a fully text-independent, spontaneous style under uncontrolled environmental conditions. To realistically simulate deployment scenarios, a chronological data split was employed in lieu of ran‎dom partitioning, an‎d all utterances were segmented into fixed two-second segments. All models were trained using Additive Angular Margin Softmax loss to maximize inter-class separation an‎d minimize intra-class variance. Results demonstrate that the Wav2Vec2-large-xlsr-53 model, augmented with a Speaker-Aware Attention Pooling mechanism, substantially outperforms all competing approaches, achieving 99.58% accuracy on the held-out test set. The SincNet variant ranks second at 92.52%, establishing itself as a lightweight an‎d interpretable alternative for resource-constrained deployment environments. RawNet2 an‎d RawNet3 attain 86.34% an‎d 83.32% accuracy, respectively. The findings confirm that multilingual self-supervised pre-training combined with dynamic attention-based pooling constitutes the most effective strategy for managing intra-speaker variability in Persian spontaneous speech. Furthermore, the study demonstrates that chronological data partitioning provides a more realistic performance estimate by preventing statistical overlap between training an‎d test distributions — a common source of inflated results in ran‎domly split eva‎luations. These results establish a rigorous benchmark for speaker recognition in low-resource Persian speech an‎d offer actionable insights for future research on robust speaker modeling in real-world acoustic conditions.
  • تعداد فصل ها
    5
  • فهرست مطالب pdf
    167551
  • نويسنده

    غني پور اميرهنده، اميرحسين