Bilgi Çıkarımı İçin Belge Düzeni Çözümleme Yöntemleri
Özet
Bu bölümde, fatura, kimlik ve pasaport gibi belirli veya karmaşık yapılara sahip taranmış belgelerden bilgi çıkarımı sağlamak amacıyla kullanılan Belge Düzeni Çözümlemesi (BDÇ) yöntemleri kapsamlı bir şekilde ele alınmıştır. Dijital veya basılı dokümanların kolayca işlenip analiz edilebilmesi için fiziksel düzenlerinin çözülmesi gerekmektedir; bu doğrultuda optik karakter tanıma (OCR), şablon eşleştirme ve makine öğrenimi gibi kural veya öğrenme tabanlı metotlardan yararlanılmaktadır. Kural tabanlı yaklaşımlar dikey/yatay tarama eşikleri gibi sabit parametreleri kullanırken, makine öğrenimi yöntemleri istatistiksel modellerle yerleşim bilgilerini ayrıştırır. Son yıllarda başarı gösteren derin öğrenme yaklaşımları arasında LayoutLM, VTLayout, Layoutparser, DocSegTr ve DiT gibi çok modlu veya transformatör tabanlı gelişmiş modeller yer almaktadır. Nesne tespiti ve bölümlere ayırma süreçlerinde ise R-CNN, Fast R-CNN, Faster R-CNN ve gerçek zamanlı işlem yeteneğine sahip YOLO algoritmaları metin yönlerini, tabloları ve görsel ögeleri çözümlemektedir. Akademik çalışmaların doğruluğunu kıyaslamak adına PubLayNet ve DocLayNet gibi açık kaynaklı karşılaştırma veri kümeleri kullanılırken; Amazon Textract, Google Document AI ve Microsoft Azure gibi bulut tabanlı ticari ürünler kurumsal çözümler sunmaktadır. Sonuç olarak BDÇ, homojen olmayan belgelerin otomatik işlenmesinde kritik bir ön adımdır.
In this section, Document Layout Analysis (DLA) methods used to extract information from scanned documents with specific or complex structures, such as invoices, IDs, and passports, are comprehensively discussed. In order to easily process and analyze digital or printed documents, their physical layouts must be resolved; accordingly, rule-based or learning-based methods such as optical character recognition (OCR), template matching, and machine learning are utilized. While rule-based approaches use fixed parameters such as vertical/horizontal scanning thresholds, machine learning methods parse layout information with statistical models. Prominent deep learning approaches that have shown success in recent years include advanced multi-modal or transformer-based models such as LayoutLM, VTLayout, Layoutparser, DocSegTr, and DiT. In object detection and segmentation processes, R-CNN, Fast R-CNN, Faster R-CNN, and YOLO algorithms with real-time processing capabilities resolve text orientations, tables, and visual elements. To compare the accuracy of academic studies, open-source benchmark datasets like PubLayNet and DocLayNet are used, while cloud-based commercial products such as Amazon Textract, Google Document AI, and Microsoft Azure offer corporate solutions. Consequently, DLA is a critical preliminary step in the automated processing of non-homogeneous documents.
Referanslar
G. M. Binmakhashen ve S. A. Mahmoud, “Document Layout Analysis: A Comprehensive Survey”, ACM Comput. Surv., c. 52, sy 6, s. 109:1-109:36, Eki. 2019, doi: 10.1145/3355610.
K. Y. Wong, R. G. Casey, ve F. M. Wahl, “Document Analysis System”, IBM J. Res. Dev., c. 26, sy 6, ss. 647-656, Kas. 1982, doi: 10.1147/rd.266.0647.
M. Atay, M. Kalayci, H. Api̇k, V. Aybar, F. Seri̇n, ve A. O. Akyüz, “An Approach to Analyzing the Layout of Unstructured Digital Documents”, 30th Signal Processing and Communications Applications Conference (SIU), IEEE, 2022.
L. O’Gorman, “The document spectrum for page layout analysis”, IEEE Trans. Pattern Anal. Mach. Intell., c. 15, sy 11, ss. 1162-1173, Kas. 1993, doi: 10.1109/34.244677.
Jaekyu Ha, R. M. Haralick, ve I. T. Phillips, “Recursive X-Y cut using bounding boxes of connected components”, Proc. 3rd Int. Conf. Doc. Anal. Recognit., c. 2, ss. 952-955, 1995, doi: 10.1109/ICDAR.1995.602059.
S. Marinai, M. Gori, ve G. Soda, “Artificial neural networks for document analysis and recognition”, IEEE Trans. Pattern Anal. Mach. Intell., c. 27, sy 1, ss. 23-35, Oca. 2005, doi: 10.1109/TPAMI.2005.4.
M. Shilman, P. Liang, ve P. Viola, “Learning Non-Generative Grammatical Models for Document Analysis.”, Oca. 2005, c. 2, ss. 962-969. doi: 10.1109/ICCV.2005.140.
H. Wei, M. Baechler, F. Slimane, ve R. Ingold, “Evaluation of SVM, MLP and GMM Classifiers for Layout Analysis of Historical Documents”, içinde 2013 12th International Conference on Document Analysis and Recognition, Ağu. 2013, ss. 1220-1224. doi: 10.1109/ICDAR.2013.247.
Y. Xu, M. Li, L. Cui, S. Huang, F. Wei, ve M. Zhou, “LayoutLM: Pre-training of Text and Layout for Document Image Understanding”, içinde Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Ağu. 2020, ss. 1192-1200. doi: 10.1145/3394486.3403172.
D. A. Borges Oliveira ve M. P. Viana, “Fast CNN-Based Document Layout Analysis”, içinde 2017 IEEE International Conference on Computer Vision Workshops (ICCVW), Eki. 2017, ss. 1173-1180. doi: 10.1109/ICCVW.2017.142.
A. R. Katti vd., “Chargrid: Towards Understanding 2D Documents”, içinde Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, Eki. 2018, ss. 4459-4469. doi: 10.18653/v1/D18-1476.
P. Forczmański, A. Smolinski, A. Nowosielski, ve K. Małecki, “Segmentation of Scanned Documents Using Deep-Learning Approach”, içinde Advances in Intelligent Systems and Computing, 2020, ss. 141-152. doi: 10.1007/978-3-030-19738-4_15.
P. Zhang vd., “VSR: A Unified Framework for Document Layout Analysis combining Vision, Semantics and Relations”. arXiv, 13 Mayıs 2021. doi: 10.48550/arXiv.2105.06220.
X. Wu, Z. Hu, X. Du, J. Yang, ve L. He, “Document Layout Analysis via Dynamic Residual Feature Fusion”, içinde 2021 IEEE International Conference on Multimedia and Expo (ICME), Tem. 2021, ss. 1-6. doi: 10.1109/ICME51207.2021.9428465.
Y. Xu vd., “LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding”. arXiv, 09 Ocak 2022. doi: 10.48550/arXiv.2012.14740.
Y. Huang, T. Lv, L. Cui, Y. Lu, ve F. Wei, “LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking”. arXiv, 19 Temmuz 2022. doi: 10.48550/arXiv.2204.08387.
S. Li, X. Ma, S. Pan, J. Hu, L. Shi, ve Q. Wang, “VTLayout: Fusion of Visual and Text Features for Document Layout Analysis”. arXiv, 12 Ağustos 2021. doi: 10.48550/arXiv.2108.13297.
Z. Shen, R. Zhang, M. Dell, B. C. G. Lee, J. Carlson, ve W. Li, “LayoutParser: A Unified Toolkit for Deep Learning Based Document Image Analysis”. arXiv, 21 Haziran 2021. doi: 10.48550/arXiv.2103.15348.
S. Long, S. Qin, D. Panteleev, A. Bissacco, Y. Fujii, ve M. Raptis, “Towards End-to-End Unified Scene Text Detection and Layout Analysis”, program adı: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, ss. 1049-1059. Erişim: 23 Aralık 2022. [Çevrimiçi]. Erişim adresi: https://openaccess.thecvf.com/content/CVPR2022/html/Long_Towards_End-to-End_Unified_Scene_Text_Detection_and_Layout_Analysis_CVPR_2022_paper.html
S. Biswas, A. Banerjee, J. Lladós, ve U. Pal, “DocSegTr: An Instance-Level End-to-End Document Image Segmentation Transformer”. arXiv, 21 Eylül 2022. doi: 10.48550/arXiv.2201.11438.
J. Li, Y. Xu, T. Lv, L. Cui, C. Zhang, ve F. Wei, “DiT: Self-supervised Pre-training for Document Image Transformer”. arXiv, 19 Temmuz 2022. doi: 10.48550/arXiv.2203.02378.
R. Girshick, J. Donahue, T. Darrell, ve J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation”. arXiv, 22 Ekim 2014. Erişim: 27 Aralık 2022. [Çevrimiçi]. Erişim adresi: http://arxiv.org/abs/1311.2524
R. Girshick, “Fast R-CNN”. arXiv, 27 Eylül 2015. Erişim: 27 Aralık 2022. [Çevrimiçi]. Erişim adresi: http://arxiv.org/abs/1504.08083
S. Ren, K. He, R. Girshick, ve J. Sun, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks”. arXiv, 06 Ocak 2016. Erişim: 27 Aralık 2022. [Çevrimiçi]. Erişim adresi: http://arxiv.org/abs/1506.01497
J. Redmon, S. Divvala, R. Girshick, ve A. Farhadi, “You Only Look Once: Unified, Real-Time Object Detection”. arXiv, 09 Mayıs 2016. Erişim: 27 Aralık 2022. [Çevrimiçi]. Erişim adresi: http://arxiv.org/abs/1506.02640
J. Redmon ve A. Farhadi, “YOLO9000: Better, Faster, Stronger”. arXiv, 25 Aralık 2016. Erişim: 27 Aralık 2022. [Çevrimiçi]. Erişim adresi: http://arxiv.org/abs/1612.08242
J. Redmon ve A. Farhadi, “YOLOv3: An Incremental Improvement”. arXiv, 08 Nisan 2018. Erişim: 27 Aralık 2022. [Çevrimiçi]. Erişim adresi: http://arxiv.org/abs/1804.02767
A. Bochkovskiy, C.-Y. Wang, ve H.-Y. M. Liao, “YOLOv4: Optimal Speed and Accuracy of Object Detection”. arXiv, 22 Nisan 2020. Erişim: 27 Aralık 2022. [Çevrimiçi]. Erişim adresi: http://arxiv.org/abs/2004.10934
“YOLOv5 Documentation”. https://docs.ultralytics.com/ (erişim 27 Aralık 2022).
“YOLOv5”. Ultralytics. [Çevrimiçi]. Erişim adresi: https://github.com/ultralytics/yolov5
C. Li vd., “YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications”. arXiv, 07 Eylül 2022. Erişim: 27 Aralık 2022. [Çevrimiçi]. Erişim adresi: http://arxiv.org/abs/2209.02976
C.-Y. Wang, A. Bochkovskiy, ve H.-Y. M. Liao, “YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors”. arXiv, 06 Temmuz 2022. Erişim: 27 Aralık 2022. [Çevrimiçi]. Erişim adresi: http://arxiv.org/abs/2207.02696
X. Zhong, J. Tang, ve A. J. Yepes, “PubLayNet: largest dataset ever for document layout analysis”. arXiv, 15 Ağustos 2019. doi: 10.48550/arXiv.1908.07836.
B. Pfitzmann, C. Auer, M. Dolfi, A. S. Nassar, ve P. W. J. Staar, “DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis”, içinde Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Ağu. 2022, ss. 3743-3751. doi: 10.1145/3534678.3539043.
A. Jimeno-Yepes, P. Zhong, ve D. Burdick, “ICDAR 2021 Competition on Scientific Literature Parsing”, 2021, ss. 605-617. doi: 10.1007/978-3-030-86337-1_40.
A. Abdallah, A. Berendeyev, I. Nuradin, ve D. Nurseitov, “TNCR: Table Net Detection and Classification Dataset”, Neurocomputing, c. 473, ss. 79-97, Şub. 2022, doi: 10.1016/j.neucom.2021.11.101.
H. Desai, P. Kayal, ve M. Singh, “TabLeX: A Benchmark Dataset for Structure and Content Information Extraction from Scientific Tables”, 2021, ss. 554-569. doi: 10.1007/978-3-030-86331-9_36.
B. Smock, R. Pesala, ve R. Abraham, “PubTables-1M: Towards comprehensive table extraction from unstructured documents”. arXiv, 18 Kasım 2021. doi: 10.48550/arXiv.2110.00061.
Z. Wang, Y. Xu, L. Cui, J. Shang, ve F. Wei, “LayoutReader: Pre-training of Text and Layout for Reading Order Detection”. arXiv, 27 Ağustos 2021. Erişim: 23 Aralık 2022. [Çevrimiçi]. Erişim adresi: http://arxiv.org/abs/2108.11591
M. Li, L. Cui, S. Huang, F. Wei, M. Zhou, ve Z. Li, “TableBank: Table Benchmark for Image-based Table Detection and Recognition”, içinde Proceedings of the Twelfth Language Resources and Evaluation Conference, Marseille, France, May. 2020, ss. 1918-1925. Erişim: 23 Aralık 2022. [Çevrimiçi]. Erişim adresi: https://aclanthology.org/2020.lrec-1.236
M. Li vd., “DocBank: A Benchmark Dataset for Document Layout Analysis”, Oca. 2020, ss. 949-960. doi: 10.18653/v1/2020.coling-main.82.
A. Mondal, P. Lipps, ve C. Jawahar, “IIIT-AR-13K: A New Dataset for Graphical Object Detection in Documents”, Ağu. 2020.
L. Gao vd., “ICDAR 2019 Competition on Table Detection and Recognition (cTDaR)”, 2019 Int. Conf. Doc. Anal. Recognit. ICDAR, ss. 1510-1515, 2019, doi: 10.1109/ICDAR.2019.00243.
L. Cui, Y. Xu, T. Lv, ve F. Wei, “Document AI: Benchmarks, Models and Applications”. arXiv, 16 Kasım 2021. Erişim: 22 Aralık 2022. [Çevrimiçi]. Erişim adresi: http://arxiv.org/abs/2111.08609
“Intelligently Extract Text & Data with OCR - Amazon Textract - Amazon Web Services”, Amazon Web Services, Inc. https://aws.amazon.com/textract/ (erişim 22 Aralık 2022).
“Document AI”, Google Cloud. https://cloud.google.com/document-ai (erişim 22 Aralık 2022).
“Form Recognizer – Automated Data Processing Systems | Microsoft Azure”. https://azure.microsoft.com/en-us/products/form-recognizer/ (erişim 22 Aralık 2022).