Doğal Dil İşleme Uygulamaları ve Yaklaşımları
Özet
Bu çalışma, bilgisayar bilimi ve dil biliminin ortak alanı olan Doğal Dil İşleme (DDİ) uygulamalarını, literatürdeki çalışmaları ve kullanılan temel yaklaşımları detaylı bir şekilde ele almaktadır. DDİ süreçlerinde verilerin öncelikle bir ön işleme adımından geçmesi gerektiği, ardından kural tabanlı, istatistiksel veya günümüzde büyük veri kümelerinde yüksek başarı sağlayan yapay sinir ağları temelli yaklaşımlarla işlendiği belirtilmektedir. Metinde DDİ'nin popüler uygulama alanlarından morfolojik çözümleme, morfolojik belirsizlik giderme, kelime türü etiketleme (POS), sözdizimsel çözümleme ve makine çevirisi sistemleri incelenmektedir. Morfolojik çözümlemede Türkçe için Zemberek kütüphanesi ve son teknoloji bir model olan MorphNet örnek verilirken; kelime türü etiketlemede Gizli Markov Modelleri ve LSTM mimarileri öne çıkmaktadır. Sözdizimsel çözümlemede İçerik Bağımsız Dil Bilgisi (CFG) ve çözüm ağaçları ele alınmakta, makine çevirisinde ise istatistiksel (İMÇ) ve nöral (NMÇ) sistemlerin veri boyutu, akıcılık ve cümle uzunluğu gibi kriterlere göre karşılaştırmalı bir değerlendirmesi sunulmaktadır. Sonuç olarak, teknolojinin gelişmesiyle birlikte DDİ çalışmalarında kural ve istatistik tabanlı geleneksel yöntemlerin yerini, büyük veri kümelerinde zaman ve başarı avantajı sunan makine öğrenmesi ve derin öğrenme yaklaşımlarının aldığı vurgulanmaktadır.
This study comprehensively addresses Natural Language Processing (NLP) applications, literature studies, and core approaches, which constitute a common field of computer science and linguistics. It states that in NLP processes, data must first undergo a preprocessing step, and then it is processed using rule-based, statistical, or neural network-based approaches, the latter of which provides high success in large datasets today. The text examines popular NLP application areas such as morphological analysis, morphological disambiguation, part-of-speech (POS) tagging, syntactic parsing, and machine translation systems. While the Zemberek library and the state-of-the-art MorphNet model are given as examples for Turkish in morphological analysis; Hidden Markov Models and LSTM architectures stand out in part-of-speech tagging. Context-Free Grammar (CFG) and parse trees are discussed in syntactic parsing, whereas in machine translation, a comparative evaluation of statistical (SMT) and neural (NMT) systems is presented based on criteria such as data size, fluency, and sentence length. In conclusion, it is emphasized that with the advancement of technology, traditional rule-based and statistical methods in NLP studies are being replaced by machine learning and deep learning approaches that offer time and success advantages on large datasets.
Referanslar
R. Feldman and J. Sanger, The text mining handbook: advanced approaches in analyzing unstructured data. Cambridge university press, 2007.
E. Cambria and B. White, "Jumping NLP curves: A review of natural language processing research," IEEE Computational intelligence magazine, vol. 9, no. 2, pp. 48-57, 2014.
E. Adalı, "Doğal Dil İşleme," Türkiye Bilişim Vakfı Bilgisayar Bilimleri ve Mühendisliği Dergisi, vol. 5, no. 2, 2012.
Y. Kaya and Ö. F. Ertuğrul, "Doküman dili tanıma için yeni bir öznitelik çıkarım yaklaşımı: İkili desenler," Gazi Üniversitesi Mühendislik Mimarlık Fakültesi Dergisi, vol. 31, no. 4, 2016.
Y. Kaya, Ö. F. Ertuğrul, and R. Tekin, "Doküman dili tanıma için ikili örüntüler tabanlı yeni bir yaklaşım," Akademik Bilişim, Eskişehir, 2015.
T. Noyan, F. Kuncan, R. Tekin, and K. Yılmaz, "Döküman dili tanıma için içerik bağımsız yeni bir yaklaşım: Açı Örüntüler," Gazi Üniversitesi Mühendislik Mimarlık Fakültesi Dergisi, vol. 37, no. 3, pp. 1277-1292, 2022.
K. Koskenniemi, Two-level morphology: A general computational model for word-form recognition and production. University of Helsinki, Department of General Linguistics Helsinki, Finland, 1983.
K. Oflazer, "Two-level description of Turkish morphology," Literary and linguistic computing, vol. 9, no. 2, pp. 137-148, 1994.
A. A. Akın and M. D. Akın, "Zemberek, an open source NLP framework for Turkic languages," Structure, vol. 10, no. 2007, pp. 1-5, 2007.
"Zemberek Doğal Dil Kütüphanesi." https://github.com/ahmetaa/zemberek-nlp (accessed 13 December 2022.
D. Kılınç, A. Özçift, F. Bozyigit, P. Yıldırım, F. Yücalar, and E. Borandag, "TTC-3600: A new benchmark dataset for Turkish text categorization," Journal of Information Science, vol. 43, no. 2, pp. 174-185, 2017.
R. Dehkharghani, Y. Saygin, B. Yanikoglu, and K. Oflazer, "SentiTurkNet: a Turkish polarity lexicon for sentiment analysis," Language Resources and Evaluation, vol. 50, no. 3, pp. 667-685, 2016.
M. Eminagaoglu, "A new similarity measure for vector space models in text classification and information retrieval," Journal of Information Science, vol. 48, no. 4, pp. 463-476, 2022.
A. Üstün, M. Kurfalı, and B. Can, "Characters or morphemes: How to represent words?," 2018: Association for Computational Linguistics.
D. Yuret and F. Türe, "Learning morphological disambiguation rules for Turkish," in Proceedings of the Human Language Technology Conference of the NAACL, Main Conference, 2006, pp. 328-334.
P. Smit, S. Virpioja, S.-A. Grönroos, and M. Kurimo, "Morfessor 2.0: Toolkit for statistical morphological segmentation," in The 14th Conference of the European Chapter of the Association for Computational Linguistics (EACL), Gothenburg, Sweden, April 26-30, 2014, 2014: Aalto University.
D. Z. Hakkani-Tür, K. Oflazer, and G. Tür, "Statistical morphological disambiguation for agglutinative languages," Computers and the Humanities, vol. 36, no. 4, pp. 381-410, 2002.
D. Yuret and M. d. l. Maza, "The greedy prepend algorithm for decision list induction," in International Symposium on Computer and Information Sciences, 2006: Springer, pp. 37-46.
E. Yildiz, C. Tirkaz, H. Sahin, M. Eren, and O. Sonmez, "A morphology-aware network for morphological disambiguation," in Proceedings of the AAAI Conference on Artificial Intelligence, 2016, vol. 30, no. 1.
Q. Shen, D. Clothiaux, E. Tagtow, P. Littell, and C. Dyer, "The role of context in neural morphological disambiguation," in Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, 2016, pp. 181-191.
E. Dayanık, E. Akyürek, and D. Yuret, "Morphnet: A sequence-to-sequence model that combines morphological analysis and disambiguation," arXiv preprint arXiv:1805.07946, 2018.
Wikipedia. https://en.wikipedia.org/wiki/Hidden_Markov_model (accessed 13 December 2022.
B. Merialdo, "Tagging English text with a probabilistic model," Computational linguistics, vol. 20, no. 2, pp. 155-171, 1994.
C.-H. Chang and C.-D. Chen, "HMM-based part-of-speech tagging for Chinese corpora," in Very Large Corpora: Academic and Industrial Perspectives, 1993.
N. Saharia, D. Das, U. Sharma, and J. Kalita, "Part of speech tagger for Assamese text," in Proceedings of the ACL-IJCNLP 2009 Conference Short Papers, 2009, pp. 33-36.
A. Yajnik, "Part of speech tagging using statistical approach for Nepali text," International Journal of Cognitive and Language Sciences, vol. 11, no. 1, pp. 76-79, 2017.
D. Cutting, J. Kupiec, J. Pedersen, and P. Sibun, "A practical part-of-speech tagger," in Third conference on applied natural language processing, 1992, pp. 133-140.
B. Babüroğlu, A. Tekerek, and M. Tekerek, "Türkçe için derin öğrenme tabanlı doğal dil işleme modeli geliştirilmesi," Web: https://arxiv.org/ftp/arxiv/papers/1905/1905.05699.pdf, vol. 16, p. 2019, 2019.
Z. Huang, W. Xu, and K. Yu, "Bidirectional LSTM-CRF models for sequence tagging," arXiv preprint arXiv:1508.01991, 2015.
D. Jurafsky and J. H. Martin, Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition, Second ed. Pearson Education, 2009.
E. Adalı, "Doğal Dil İşleme," (in scheme="ISO639-1"), 5, Makaleler(Araştırma) 2016, doi: https://dergipark.org.tr/en/pub/tbbmd/issue/22245/238797.
O. Görgün and O. T. Yıldız, "Using morphology in English-Turkish statistical machine translation," in 2012 20th Signal Processing and Communications Applications Conference (SIU), 2012: IEEE, pp. 1-4.
E. Solak, "On relative clause in Turkish," in Proceedings of the 7th International Conference on Computer Processing of Turkic Languages (turklang 2019), Simferopol, Russia, 2019.
E. BARUT, "İstatistiksel Makine Çevirisi İle Nöral Makine Çevirisinin Dilbilimsel Parametrelerle Karşılaştırılması: Google Translate," Akdeniz Havzası ve Afrika Medeniyetleri Dergisi, vol. 4, no. 1, pp. 103-118, 2022.
P. Koehn and R. Knowles, "Six challenges for neural machine translation," arXiv preprint arXiv:1706.03872, 2017.
S. İlhami, Ü. Hüseyin, and D. HANBAY, "Creating a Parallel Corpora for Turkish-English Academic Translations," Computer Science, no. Special, pp. 335-340, 2021.
E. Satir and H. Bulut, "Preventing translation quality deterioration caused by beam search decoding in neural machine translation using statistical machine translation," Information Sciences, vol. 581, pp. 791-807, 2021.
X. Wang, Z. Tu, and M. Zhang, "Incorporating statistical machine translation word knowledge into neural machine translation," IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 26, no. 12, pp. 2255-2266, 2018.
C. Amrhein and S. Clematide, "Supervised OCR error detection and correction using statistical and neural machine translation methods," Journal for Language Technology and Computational Linguistics (JLCL), vol. 33, no. 1, pp. 49-76, 2018.
J. Moorkens, A. Toral, S. Castilho, and A. Way, "Translators’ perceptions of literary post-editing using statistical and neural machine translation," Translation Spaces, vol. 7, no. 2, pp. 240-262, 2018.
F. Stahlberg, "Neural machine translation: A review," Journal of Artificial Intelligence Research, vol. 69, pp. 343-418, 2020.