شماره ركورد
26010
شماره راهنما
LIN2 265
عنوان
طراحي پيكره زبانآموز دوزبانه ارمني ـ فارسي داراي برچسب خطاي املائي
مقطع تحصيلي
كارشناسي ارشد
رشته تحصيلي
زبانشناسي رايانشي
دانشكده
زبانهاي خارجي
تاريخ دفاع
1405/02/20
صفحه شمار
102 ص .
استاد راهنما
رضوان متوليان
استاد مشاور
مرجان كائدي
كليدواژه فارسي
پيكره زبانآموز , دوزبانه ارمنيـفارسي , خطاي املايي , يادگيري ماشين , مدرسه آرمن , ماشين بردار پشتيبان (SVM) , جنگل تصادفي (RF)
چكيده فارسي
پژوهش حاضر با هدف طراحي پيكره زبانآموز دوزبانه ارمنيـفارسي داراي برچسب خطاي املايي انجام شده است. با توجه به خلاء موجود در منابع پيكرهاي دوزبانه، اين تحقيق به دنبال شناسايي الگوهاي خطاي نوشتاري و فراهم آوردن زيرساختي دادهمحور براي تحليلهاي زبانشناختي و محاسباتي است. دادههاي پژوهش حاضر از طريق گردآوري توليدات نوشتاري زبانآموزان ارمني در مقاطع دبستان و دبيرستان «آرمن» در سطوح مختلف آموزشي تهيه و با استفاده از نرمافزار اينسپشن INCEpTION برچسبگذاري گرديد. بهمنظور ارزيابي اعتبار پيكره و تحليل هوشمند دادهها، الگوريتمهاي يادگيري ماشين شامل جنگل تصادفي Random Forest و ماشين بردار پشتيبان SVM جهت طبقهبندي خودكار خطاها و سنجش ويژگيهاي پيكره به كار گرفته شدند. تحليل نهايي بر روي 1116 برچسب خطاي املايي نشان داد كه مقوله «نشانههاي اصلي» با اختصاص 50 درصد از كل دادهها، بيشترين فراواني را دارد، كه بيانگر چالشهاي بنيادي اين زبانآموزان در سطح واجي و نظام نوشتاري است. همچنين، «چندنويسهها» بهعنوان دومين مقوله پرتكرار شناسايي شدند كه پيچيدگي تركيبهاي حرفي را براي اين گروه از زبانآموزان نشان ميدهد. نتايج حاصل از مدلهاي يادگيري ماشين نيز كارايي و دقت بالاي اين الگوها را در تشخيص و دستهبندي خودكار خطاها تأييد كرد. پيكره توليدشده در اين پژوهش، علاوه بر كاربرد در اصلاح برنامههاي درسي مدارس دوزبانه، منبعي ارزشمند براي توسعه ابزارهاي پردازش زبان طبيعي و سيستمهاي خطاياب هوشمند فراهم ميآورد.
كليدواژه لاتين
learner corpus , Armenian–Persian bilingual , spelling error , machine learning , Armen school , Support Vector Machine (SVM) , Random Forest (RF)
عنوان لاتين
Designing a Bilingual Armenian-Persian Learner Corpus Tagged with Spelling Error
گروه آموزشي
زبان شناسي
چكيده لاتين
The present study aims to design a bilingual Armenian–Persian learner corpus tagged with spelling error. Given the existing gap in bilingual corpus resources, this research seeks to identify patterns of spelling errors and provide a data-driven infrastructure for linguistic and computational analysis. The data for this research were collected from the written productions of Armenian learners in elementary and secondary school levels at the "Armen" school, across various educational stages, and were tagged using the INCEpTION software. To evaluate the validity of the corpus and to perform intelligent data analysis, machine learning algorithms, including Random Forest and Support Vector Machine (SVM), were employed for the automatic classification of errors and the assessment of corpus features. The final analysis of 1,116 spelling error tags showed that the category "main markers" had the highest frequency, accounting for 50% of the total data, which reflects the fundamental challenges of these learners at the phonological and orthographic levels. Additionally, "multi-letter combinations" were identified as the second most frequent category, indicating the complexity of letter combinations for this group of learners. The results from the machine learning models also confirmed the high efficiency and accuracy of these patterns in detecting and automatically classifying errors. The corpus generated in this study, in addition to its application in refining bilingual school curricula, provides a valuable resource for developing natural language processing tools and intelligent spell-checking systems.
تعداد فصل ها
5
فهرست مطالب pdf
163513
نويسنده