• شماره ركورد
    26010
  • شماره راهنما
    LIN2 265
  • عنوان

    طراحي پيكره زبان‌آموز دوزبانه ارمني ـ فارسي داراي برچسب خطاي املائي

  • مقطع تحصيلي
    كارشناسي ارشد
  • رشته تحصيلي
    زبانشناسي رايانشي
  • دانشكده
    زبانهاي خارجي
  • تاريخ دفاع
    1405/02/20
  • صفحه شمار
    102 ص .
  • استاد راهنما
    رضوان متوليان
  • استاد مشاور
    مرجان كائدي
  • كليدواژه فارسي
    پيكره زبان‌آموز , دوزبانه ارمني‌ـ‌فارسي , خطاي املايي , يادگيري ماشين , مدرسه آرمن , ماشين بردار پشتيبان (SVM) , جنگل تصادفي (RF)
  • چكيده فارسي
    پژوهش حاضر با هدف طراحي پيكره‌ زبان‌آموز دوزبانه‌ ارمني‌ـ‌فارسي داراي برچسب‌ خطاي املايي انجام شده است. با توجه به خلاء موجود در منابع پيكره‌اي دوزبانه، اين تحقيق به دنبال شناسايي الگوهاي خطاي نوشتاري و فراهم آوردن زيرساختي داده‌محور براي تحليل‌هاي زبان‌شناختي و محاسباتي است. داده‌هاي پژوهش حاضر از طريق گردآوري توليدات نوشتاري زبان‌آموزان ارمني در مقاطع دبستان و دبيرستان «آرمن» در سطوح مختلف آموزشي تهيه و با استفاده از نرم‌افزار اينسپشن INCEpTION برچسب‌گذاري گرديد. به‌منظور ارزيابي اعتبار پيكره و تحليل هوشمند داده‌ها، الگوريتم‌هاي يادگيري ماشين شامل جنگل تصادفي Random Forest و ماشين بردار پشتيبان SVM جهت طبقه‌بندي خودكار خطاها و سنجش ويژگي‌هاي پيكره به كار گرفته شدند. تحليل نهايي بر روي 1116 برچسب خطاي املايي نشان داد كه مقوله «نشانه‌هاي اصلي» با اختصاص 50 درصد از كل داده‌ها، بيشترين فراواني را دارد، كه بيانگر چالش‌هاي بنيادي اين زبان‌آموزان در سطح واجي و نظام نوشتاري است. همچنين، «چندنويسه‌ها» به‌عنوان دومين مقوله پرتكرار شناسايي شدند كه پيچيدگي تركيب‌هاي حرفي را براي اين گروه از زبان‌آموزان نشان مي‌دهد. نتايج حاصل از مدل‌هاي يادگيري ماشين نيز كارايي و دقت بالاي اين الگوها را در تشخيص و دسته‌بندي خودكار خطاها تأييد كرد. پيكره توليدشده در اين پژوهش، علاوه بر كاربرد در اصلاح برنامه‌هاي درسي مدارس دوزبانه، منبعي ارزشمند براي توسعه ابزارهاي پردازش زبان طبيعي و سيستم‌هاي خطاياب هوشمند فراهم مي‌آورد.
  • كليدواژه لاتين
    learner corpus , Armenian–Persian bilingual , spelling error , machine learning , Armen school , Support Vector Machine (SVM) , Random Forest (RF)
  • عنوان لاتين
    Designing a Bilingual Armenian-Persian Learner Corpus Tagged with Spelling Error
  • گروه آموزشي
    زبان شناسي
  • چكيده لاتين
    The present study aims to design a bilingual Armenian–Persian learner corpus tagged with spelling error. Given the existing gap in bilingual corpus resources, this research seeks to identify patterns of spelling errors an‎d provide a data-driven infrastructure for linguistic an‎d computational analysis. The data for this research were collected from the written productions of Armenian learners in elementary an‎d secondary school levels at the "Armen" school, across various educational stages, an‎d were tagged using the INCEpTION software. To eva‎luate the validity of the corpus an‎d to perform intelligent data analysis, machine learning algorithms, including Ran‎dom Forest an‎d Support Vector Machine (SVM), were employed for the automatic classification of errors an‎d the assessment of corpus features. The final analysis of 1,116 spelling error tags showed that the category "main markers" had the highest frequency, accounting for 50% of the total data, which reflects the fundamental challenges of these learners at the phonological an‎d orthographic levels. Additionally, "multi-letter combinations" were identified as the second most frequent category, indicating the complexity of letter combinations for this group of learners. The results from the machine learning models also confirmed the high efficiency an‎d accuracy of these patterns in detecting an‎d automatically classifying errors. The corpus generated in this study, in addition to its application in refining bilingual school curricula, provides a valuable resource for developing natural language processing tools an‎d intelligent spell-checking systems.
  • تعداد فصل ها
    5
  • فهرست مطالب pdf
    163513
  • نويسنده

    كشيشيان، فريدا