Öznitelik Seçme Yöntemleri

Yazarlar

Ali Şenol
https://orcid.org/0000-0003-0364-2837

Özet

Bu çalışma, makine öğrenmesi süreçlerinde büyük veri kümelerinin işlenmesini kolaylaştırmak amacıyla kullanılan öznitelik seçme yöntemlerini kapsamlı bir şekilde incelemektedir. Teknolojik gelişmelerle birlikte veri boyutunun devasa boyutlara ulaşması, modellerin zaman karmaşıklığını artırmakta ve performansını düşürmektedir. Bu problemi aşmak için uygulanan boyut indirgeme yaklaşımlarından öznitelik seçimi; filtreleme, sarmalayıcı ve gömülü yöntemler olmak üzere üç ana gruba ayrılır. Filtreleme yöntemleri istatistiksel ve bilgi tabanlı hesaplamalarla en ilişkili nitelikleri belirlerken; sarmalayıcı yöntemler açgözlü arama algoritmaları ve makine öğrenmesi modelleriyle en ideal alt kümeleri seçer. Gömülü yöntemler ise bu seçimi doğrudan model eğitim sürecinde gerçekleştirir. Öznitelik seçimi; metin madenciliği, biyoinformatik, görüntü tanıma, kümeleme, kural oluşturma ve sistem izleme gibi kritik alanlarda yaygın olarak kullanılmaktadır. Çalışmada, meme kanseri veri kümesi üzerinde farklı filtreleme yöntemleri Python ortamında test edilmiş ve her yöntemin farklı sonuçlar verdiği görülmüştür. Bu durum, veri kümesinin yapısına uygun yöntemin belirlenmesi için karşılaştırmalı analizlerin önemini ortaya koymaktadır. Gelecekte de veri odaklı teknolojik dönüşüme paralel olarak bu alandaki çalışmaların etkinliğini sürdüreceği öngörülmektedir.

This study comprehensively examines feature selection methods used to facilitate the processing of large datasets in machine learning workflows. As technological advancements drive data volumes to massive proportions, the time complexity of models increases while their performance degrades. To overcome this challenge, dimensionality reduction techniques are applied, with feature selection categorized into three main groups: filter, wrapper, and embedded methods. Filter methods identify the most relevant features through statistical and information-based calculations such as Information Gain, Chi-square, Fisher Score, Correlation, and Variance Threshold. Wrapper methods evaluate optimal feature subsets using greedy search algorithms combined with machine learning models via forward, backward, or bi-directional elimination. Embedded methods integrate the selection process directly into the model training phase. Feature selection is widely utilized across critical domains including text mining, bioinformatics, image recognition, clustering, rule induction, and system monitoring. In this study, various filtering methods were tested on a breast cancer dataset within a Python environment, demonstrating that each method yields distinct results. This variation highlights the absolute necessity of comparative analyses to determine the most appropriate method for a specific dataset's structure. In conclusion, it is anticipated that research in this field will maintain its vital importance in parallel with ongoing data-driven technological transformations.

Referanslar

S. Kemp.. Digital 2022: October Global Syayshot Report. Erişim Tarihi: 11.15.2022, Web adresi: https://datareportal.com/reports/digital-2022-october-global-statshot

J. Cai, J. Luo, S. Wang, and S. Yang, "Feature selection in machine learning: A new perspective, Neurocomputing, vol. 300, 70-79, 2018.

F. Korn, B. U. Pagel, and C. Faloutsos, "On the "dimensionality curse" and the "self-similarity blessing"," IEEE Transactions on Knowledge and Data Engineering, vol. 13, no. 1, 96-111, 2001.

P. Dhal and C. Azad, "A comprehensive survey on feature selection in the various fields of machine learning," Applied Intelligence, vol. 52, no. 4, 4543-4581, 2022.

J. Tang, S. Alelyani, and H. Liu, "Feature selection for classification: A review," Data Classification: Algorithms and Applications, 37-64, 2014.

Kaya M., Bilge H.S. " A hybrid feature selection approach based on statistical and wrapper methods", 24th Signal Processing and Communication Application Conference, SIU 2016 - Proceedings, , art. no. 7496186 , pp. 2101-2104, 2016.

A. Jović, K. Brkić, and N. Bogunović, "A review of feature selection methods with applications," in 2015 38th International Convention on Information and Communication Technology, Electronics and Microelectronics (MIPRO), 1200-1205, 2015.

N. Hoque, D. K. Bhattacharyya, and J. K. Kalita, "MIFS-ND: A mutual information-based feature selection method," Expert Systems with Applications, vol. 41, no. 14, 6371-6385, 2014.

U. M. Khaire, R. Dhanalakshmi, "Stability of feature selection algorithm: A review," the Journal of King Saud University - Computer and Information Sciences, 1060-1073, 2022.

L. Yu and H. Liu, "Feature Selection for High-Dimensional Data," in Proceedings, Twentieth International Conference on Machine Learning, T. Fawcett and N. Mishra, Eds., 856-863, 2003.

I. H. Witten, E. Frank, and M. A. Hall, Data mining: practical machine learning tools and techniques, 3rd Edition. Morgan Kaufmann, Elsevier, I-XXXIII, 2011.

R. O. Duda, P. E. Hart, and D. G. Stork, Pattern classification, 2nd Edition. Wiley, I-XX, 2001.

M. Dash and Y.-S. Ong, "RELIEF-C: Efficient Feature Selection for Clustering over Noisy Data,", 23th International Conference on Tools for Artificial Intelligence (ICTAI), 2011. Available: https://doi.org/10.1109/ICTAI.2011.135 http://doi.ieeecomputersociety.org/10.1109/ICTAI.2011.135

M. Robnik-Sikonja and I. Kononenko, "Theoretical and Empirical Analysis of ReliefF and RReliefF," Mach. Learn., vol. 53, no. 1-2, 23-69, 2003.

J. Tang, S. Alelyani, and H. Liu, "Feature Selection for Classification: A Review," in Data Classification: Algorithms and Applications, 37-64, 2014.

D. M. Witten and R. Tibshirani, "A Framework for Feature Selection in Clustering," Journal of the American Statistical Association, vol. 105, no. 490, 713-726, 2010.

A. U. Ahmad and A. Starkey, "Application of feature selection methods for automated clustering analysis: a review on synthetic datasets," Neural Comput. Appl., vol. 29, no. 7, 317-328, 2018.

M. A. F. A. Fida, T. Ahmad, and M. Ntahobari, "Variance Threshold as Early Screening to Boruta Feature Selection for Intrusion Detection System," in 2021 13th International Conference on Information & Communication Technology and System (ICTS), 46-50, 2021.

A. J. Ferreira and M. A. T. Figueiredo, "Efficient feature selection filters for high-dimensional data," Pattern Recognition Letters, vol. 33, no. 13, 1794-1804, 2012.

J. M. Peña, "Learning Gaussian Graphical Models of Gene Networks with False Discovery Rate Control," in Evolutionary Computation, Machine Learning and Data Mining in Bioinformatics, Berlin, Heidelberg,165-176: Springer Berlin Heidelberg, 2008.

R. Peck and J. L. Devore, Statistics: The exploration & analysis of data. Cengage Learning, 2011.

G. Forman, "An Extensive Empirical Study of Feature Selection Metrics for Text Classification," Journal of Machine Learning Research, vol. 3, 1289-1305, 2003.

S. Bahassine, A. Madani, M. Al-Sarem, and M. Kissi, "Feature selection using an improved Chi-square for Arabic text classification," Journal of King Saud University - Computer and Information Sciences, vol. 32, no. 2, 225-231, 2020.

M. Dash, K. Choi, P. Scheuermann, and L. Huan, "Feature selection for clustering - a filter solution," in 2002 IEEE International Conference on Data Mining, 2002. Proceedings., 115-122, 2002.

X. Deng, Y. Li, J. Weng, and J. Zhang, "Feature selection for text classification: A review," Multimedia Tools and Applications, vol. 78, no. 3, pp. 3797-3816, 2019.

A. Wang, H. Liu, and G. Chen, "Chaotic Harmony Search based Multi-objective Feature Selection for Classification of Gene Expression Profiles," in 2021 IEEE 9th International Conference on Bioinformatics and Computational Biology (ICBCB), 107-112, 2021.

F. Özyurt, "Efficient deep feature selection for remote sensing image recognition with fused deep learning architectures," The Journal of Supercomputing, vol. 76, no. 11, 8413-8431, 2020.

E. S. M. El-Kenawy, A. Ibrahim, S. Mirjalili, M. M. Eid, and S. E. Hussein, "Novel Feature Selection and Voting Classifier Algorithms for COVID-19 Classification in CT Images," IEEE Access, vol. 8, 179317-179335, 2020.

S. Solorio-Fernández, J. A. Carrasco-Ochoa, and J. F. Martínez-Trinidad, "A review of unsupervised feature selection methods," Artificial Intelligence Review, vol. 53, no. 2, 907-948, 2020.

D. Han and J. Kim, "Unified Simultaneous Clustering and Feature Selection for Unlabeled and Labeled Data," IEEE Transactions on Neural Networks and Learning Systems, vol. 29, no. 12, 6083-6098, 2018.

M. Z. Asghar, A. Khan, F. Khan, and F. M. Kundi, "RIFT: A Rule Induction Framework for Twitter Sentiment Analysis," Arabian Journal for Science and Engineering, vol. 43, no. 2, 857-877, 2018.

R. Jensen, C. Cornelis, and Q. Shen, "Hybrid fuzzy-rough rule induction and feature selection," in 2009 IEEE International Conference on Fuzzy Systems, 1151-1156, 2009.

A. S. N. Huda and S. Taib, "Suitable features selection for monitoring thermal condition of electrical equipment using infrared thermography," Infrared Physics & Technology, vol. 61, 184-191, 2013.

X. Jin, E. W. M. Ma, L. L. Cheng, and M. Pecht, "Health Monitoring of Cooling Fans Based on Mahalanobis Distance With mRMR Feature Selection," IEEE Transactions on Instrumentation and Measurement, vol. 61, no. 8, 2222-2229, 2012.

V. Nasir and J. Cool, "Intelligent wood machining monitoring using vibration signals combined with self-organizing maps for automatic feature selection," The International Journal of Advanced Manufacturing Technology, vol. 108, no. 5, 1811-1825, 2020.

UCI Machine Learning Repository: Breast Cancer Wisconsin (Diagnostic) Data Set. Erişim Tarihi: 1 Aralık 2022, Web adresi: https://archive.ics.uci.edu/ml/datasets/Breast+Cancer+Wisconsin+%28Diagnostic%29

Yayınlanan

5 Aralık 2022

Lisans

Lisans