Görüntü İşleme Tabanlı Üretken Çekişmeli Ağ Modelleri
Özet
Bu çalışma, görüntü işleme tabanlı Üretken Çekişmeli Ağ (ÜÇA) modellerini, bu modellerde sıklıkla karşılaşılan istikrar ve çeşitlilik problemlerini ve bunlara yönelik geliştirilen çözüm önerilerini detaylıca ele almaktadır. ÜÇA'larda öne çıkan en temel sorunlar, üreticinin sürekli benzer görüntüler üretmesine yol açan mod çöküşü ve ayrıştırıcının üreticiye üstünlük kurmasıyla oluşan eğitim istikrarsızlığıdır. Bu engelleri aşmak amacıyla Wasserstein uzaklığı kullanan WGAN, gradyan cezası ekleyen WGAN-GP, spektral normalleştirme uygulayan SNGAN ve iki zaman ölçekli güncelleme kuralından yararlanan SAGAN gibi alternatif amaç fonksiyonları ve düzenlileştirme teknikleri önerilmiştir. Bu yöntemler; ağırlık kırpma veya spektral norm hesaplamalarıyla gradyan azalması ve patlaması problemlerini önleyerek modellerin daha kararlı eğitilmesini sağlar. Mimari açıdan ise evrişimsel sinir ağlarını entegre eden DCGAN, belirli sınıf etiketlerini şart koşan koşullu ÜÇA (cGAN), yapısına sınıflandırıcı ekleyen yarı denetimli SGAN, karşılıklı bilgiyi maksimize eden InfoGAN ve çözünürlüğü kademeli olarak artıran PGGAN modelleri geliştirilmiştir. Ayrıca, gerçek görüntü uzayından gizli unsur uzayına haritalama yaparak görüntülerde öznitelik düzenlemeye imkan tanıyan VAEGAN, BiGAN, AGE ve CRGGAN gibi çıkarım modelleri detaylandırılmıştır. Yapılan genel değerlendirmeler, ÜÇA eğitim metodolojisinin derin öğrenme alanındaki birçok karmaşık imge dönüştürme problemine başarıyla uyarlandığını göstermektedir.
This study comprehensively covers image processing-based Generative Adversarial Network (GAN) models, the stability and diversity challenges frequently encountered in these models, and the solution proposals developed to address them. The most prominent issues in GANs are mode collapse, which causes the generator to continuously produce similar images, and training instability, which occurs when the discriminator dominates the generator. To overcome these hurdles, alternative objective functions and regularization techniques have been proposed, such as WGAN utilizing Wasserstein distance, WGAN-GP adding a gradient penalty, SNGAN applying spectral normalization, and SAGAN leveraging a two time-scale update rule. These methods ensure more stable training of models by preventing vanishing and exploding gradient problems through approaches like weight clipping or spectral norm calculations. Architecturally, models such as DCGAN integrating convolutional neural networks, conditional GAN (cGAN) mandating specific class labels, semi-supervised SGAN adding a classifier to its structure, InfoGAN maximizing mutual information, and PGGAN progressively increasing image resolution have been developed. Furthermore, inference models like VAEGAN, BiGAN, AGE, and CRGGAN, which allow attribute editing in images by mapping from the real image space to the latent feature space, are detailed. General evaluations demonstrate that the GAN training methodology has been successfully adapted to many complex image translation problems in the field of deep learning.
Referanslar
A.A, Efros, W. T. Freeman, Image quilting for texture synthesis and transfer, In the 28th annual conference on Computer graphics and interactive techniques, Los Angeles, CA, USA, 2001.
D. Eigen, R. Fergus, Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture, In the IEEE International Conference on Computer Vision, Santiago, Chile, 2015.
J. Long, E. Shelhamer, T. Darrell, Fully convolutional networks for semantic segmentation, In the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 2015.
S. Xie, Z. Tu, Holistically-nested edge detection, In the IEEE International Conference on Computer Vision, Santiago, Chile, 2015.
R. Zhang, J. Y. Zhu, P. Isola, X. Geng, A.S. Lin, T. Yu, A.A, Efros, Real-time user-guided image colorization with learned deep priors, arXiv preprint arXiv:1705.02999, 2017.
J. Guo, J. Li, H. Fu, M. Gong, K. Zhang, D. Tao, Alleviating Semantics Distortion in Unsupervised Low-Level Image-to-Image Translation via Structure Consistency Constraint, In the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 2022.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, Y. Bengio, Generative adversarial nets, In Advances in neural information processing systems, Montréal, Canada, 2014.
P. Isola, J.Y. Zhu, T. Zhou, A.A. Efros, Image-to-Image Translation with Conditional Adversarial Networks, In the Computer Vision and Pattern Recognition, Honolulu, Hawaii, 2017.
I. Goodfellow, NIPS tutorial: Generative adversarial networks, arXiv preprint, arXiv:1701.00160, 2016.
Y. LeCun, C. Cortes, C. J. Burges, MNIST handwritten digit database, Web Sayfası, Erişim Tarihi: 25.12.2022 Web adresi: http://yann.lecun.com/exdb/mnist/
H. Xiao, K. Rasul, R. ollgraf, Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, arXiv preprint arXiv:1708.07747, 2017.
Y. Dogan, H. Y. Keles, Stability and diversity in generative adversarial networks, 27th signal processing and communications applications conference, Sivas, Türkiye, 2019.
Y. Dogan, Üretken çekişmeli ağlarda gizli unsur kodlayıcı ile çıktı imgesi arasındaki ilişkinin hesaplamalı modellenmesi, Web Sayfası, Erişim Tarihi: 25.12.2022 Web adresi: https://dspace.ankara.edu.tr/xmlui/handle/20.500.12575/72848
J. Zhao, M. Mathieu, Y. LeCun, Energy-based generative adversarial network, arXiv preprint arXiv:1609.03126, 2016.
M. Arjovsky, S. Chintala, L. Bottou, Wasserstein generative adversarial networks, In the International conference on machine learning, Sydney, Australia, 2017.
S. W. Park, J. Kwon, Sphere generative adversarial network based on geometric moment matching, In the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, United States, 2019.
Y. Pantazis, D. Paul, M. Fasoulakis, Y. Stylianou, M. A. Katsoulakis, Cumulant gan, In the IEEE Transactions on Neural Networks and Learning Systems, 1-12, 2022.
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, A. C. Courville, Improved training of wasserstein gans, Advances in neural information processing systems, 30, 5769–5779, 2017.
T. Miyato, T. Kataoka, M. Koyama, Y. Yoshida, Spectral normalization for generative adversarial networks, arXiv preprint arXiv:1802.05957, 2018.
A. Radford, L. Metz, S. Chintala, Unsupervised representation learning with deep convolutional generative adversarial networks, arXiv preprint arXiv:1511.06434, 2015.
Y. Li, N. Xiao, W. Ouyang, Improved boundary equilibrium generative adversarial networks, IEEE access, 6, 11342-11348, 2018.
T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, T. Aila, Analyzing and improving the image quality of stylegan, In the IEEE conference on computer vision and pattern recognition, Seattle, WA, USA, 2020.
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, X. Chen, Improved techniques for training gans, Advances in neural information processing systems, 29, 2234–2242, 2016.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, S. Hochreiter, Gans trained by a two time-scale update rule converge to a local nash equilibrium, Advances in neural information processing systems, 30, 6629–6640, 2017.
H. Zhang, I. Goodfellow, D. Metaxas, A. Odena, Self-attention generative adversarial networks, In the International conference on machine learning, Long Beach, CA, USA, 2019.
J. Susskind, A. Anderson, G. E. Hinton, The Toronto face dataset, Web Sayfası, Erişim Tarihi: 25.12.2022 Web adresi: https://www.kaggle.com/discussions/general/50987
A. Krizhevsky, G. Hinton, Learning multiple layers of features from tiny images, Web Sayfası, Erişim Tarihi: 25.12.2022 Web adresi: https://www.cs.utoronto.ca/~kriz/
I. Atas, C. Ozdemir, M. Atas, Y. Dogan, Forensic dental age estimation using modified deep learning neural network, arXiv preprint arXiv:2208.09799, 2022.
M. Ataş, C. Özdemir, İ. Ataş, B. Ak, E. Özeroğlu, Biometric identification using panoramic dental radiographic images withfew-shot learning, Turkish Journal of Electrical Engineering and Computer Sciences, 30(3), 1115-1126, 2022.
I. Atas, Human Gender Prediction Based on Deep Transfer Learning from Panoramic Radiograph Images, arXiv preprint arXiv:2205.09850, 2022.
S. Ioffe, C. Szegedy, Batch normalization: Accelerating deep network training by reducing internal covariate shift, In the International conference on machine learning, Lille, France, 2015.
V. Nair, G. E. Hinton, Rectified linear units improve restricted boltzmann machines, In 27th International Conference on Machine Learning, Haifa, Israel, 2010.
A. L. Maas, A. Y. Hannun, A. Y. Ng, Rectifier nonlinearities improve neural network acoustic models, In the Proceedings of International Conference on Machine Learning, 30, 3, 2013.
J. Deng, W. Dong, R. Socher, L. J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hierarchical image database, In the IEEE conference on computer vision and pattern recognition, Miami, FL, USA, 2009.
M. Mirza, S. Osindero, Conditional generative adversarial nets, arXiv preprint arXiv:1411.1784, 2014.
A. Odena, Semi-supervised learning with generative adversarial networks, arXiv preprint arXiv:1606.01583, 2016.
X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, P. Abbeel, Infogan: Interpretable representation learning by information maximizing generative adversarial nets, In the Advances in Neural Information Processing Systems, Barcelona, Spain, 2016.
T. Karras, T. Aila S. Laine J. Lehtinen, Progressive growing of gans for improved quality, stability, and variation, International Conference on Learning Representations, Vancouver, Canada, 2018.
T. C. Wang, M. Y. Liu, J. Y. Zhu, A. Tao, J. Kautz, B. Catanzaro, High-resolution image synthesis and semantic manipulation with conditional gans, In the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, Utah, 2018.
A. Ghosh, V. Kulharia, V. P. Namboodiri, P. H. Torr, P. K. Dokania, Multi-agent diverse generative adversarial networks, In the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, Utah, 2018.
T. Karras, S. Laine, T. Aila, A style-based generator architecture for generative adversarial networks, In the IEEE conference on computer vision and pattern recognition, Long Beach, CA, USA, 2019.
T. Karras, M. Aittala, S. Laine, E. Härkönen, J. Hellsten, J. Lehtinen, T. Aila, Alias-free generative adversarial networks, Advances in Neural Information Processing Systems, 34, 852-863, 2021.
Z. C. Lipton, S. Tripathi, Precise recovery of latent vectors from generative adversarial networks, arXiv preprint arXiv:1702.04782, 2017.
A. Creswell, A. A. Bharath, Inverting the generator of a generative adversarial network, In the IEEE transactions on neural networks and learning systems, 30(7), 1967-1974, 2018.
A. B. L. Larsen, S. K. Sønderby, H. Larochelle, O. Winther, Autoencoding beyond pixels using a learned similarity metric, In the International Conference on Machine Learning, New York City, NY, USA, 2016.
D. P. Kingma, T. Salimans, R. Jozefowicz, X. Chen, I. Sutskever, M. Welling, Improved variational inference with inverse autoregressive flow, Advances in neural information processing systems, 29, 4743–4751, 2016.
J. Donahue, P. Krähenbühl, T. Darrell, Adversarial feature learning, arXiv preprint arXiv:1605.09782, 2016.
V. Dumoulin, I. Belghazi, B. Poole, O. Mastropietro, A. Lamb, M. Arjovsky, A. Courville, Adversarially learned inference, arXiv preprint arXiv:1606.00704, 2016.
A. Heljakka, A. Solin, J. Kannala, Pioneer networks: Progressively growing generative autoencoder, In Asian Conference on Computer Vision, Perth, Australia, 2018.
D. Ulyanov, A. Vedaldi, V. Lempitsky, It takes (only) two: Adversarial generator-encoder networks, AAAI Conference on Artificial Intelligence, New Orleans, Louisiana, USA, 2018.
Y. Dogan, H. Y. Keles, Semi-supervised image attribute editing using generative adversarial networks, Neurocomputing, 401, 338-352, 2020.