ỨNG DỤNG CÁC MÔ HÌNH MẠNG HỌC SÂU NHẬN DẠNG BIỂN SỐ XE Ô TÔ TỰ ĐỘNG

Nguyễn Hữu Tuân1, , Nguyễn Duy Trường Giang1
1 Khoa Công nghệ thông tin, Trường Đại học Hàng hải Việt Nam

Nội dung chính của bài viết

Tóm tắt

Nhận dạng biển số xe ô tô tự động là một bài toán của lĩnh vực thị giác máy tính có tính khoa học và thực tiễn cao. Giải quyết tốt bài toán này là nền tảng cho việc chuyển đổi số trong lĩnh vực giao thông đường bộ và nhiều bài toán về quản lý giao thông đô thị khác như giám sát phương tiện, phạt nguội, truy vết phương tiện. Bài báo này nghiên cứu và đề xuất ứng dụng mô hình mạng YOLOv12 và PaddleOCR v3.3.2 để nhận dạng tự động biển số xe ô tô Việt Nam. Ở bước đầu tiên, mô hình YOLOv12, một phiên bản cải tiến từ các phiên bản trước của họ các mô hình YOLO với cơ chế chú ý sẽ được sử dụng để phát hiện biển số xe ô tô. Sau đó phần biển số xe sẽ được nhận dạng bằng PaddleOCR v3.3.2. Kết quả huấn luyện và kiểm thử trên tập dữ liệu biển số xe ô tô Việt Nam gồm 25.364 bức ảnh cho phần phát hiện biển số và 6.340 bức ảnh cho phần nhận dạng biển số cho thấy phương pháp đề xuất có độ chính xác trong phát hiện và nhận dạng rất tốt khi chỉ số phát hiện biển số mAP@0.5 là 99,3% và độ chính xác của nhận dạng biển số lên đến 89,21%. Kết quả khi đánh giá với các video thử nghiệm thu thập được từ thực tế cũng cho kết quả phát hiện và nhận dạng tốt và hoàn toàn có thể áp dụng vào các tình huống thực tế.

Abstract

Automated License Plate Recognition (ALPR) constitutes a significant problem within the field of computer vision, possessing both high scientific and practical relevance. Successfully addressing this challenge forms a foundation for digital transformation within the road transport sector and enables further advancements in urban traffic management applications such as vehicle monitoring, automated enforcement (e.g., photo-based ticketing), and vehicle tracking. This paper investigates and proposes an application of the YOLOv12 and PaddleOCR v3.3.2 models for automated Vietnamese license plate recognition. Firstly, a YOLOv12 model—an improved version leveraging attention mechanisms over previous YOLO architectures—is employed to detect license plates within images. Then, the detected license plate regions are recognized using PaddleOCR v3.3.2. Training and testing results on a dataset comprising 25,364 images for license plate detection and 6,340 images for license plate recognition demonstrate excellent performance in both detection and recognition tasks. The mean Average Precision (mAP) at IoU threshold of 0.5 for license plate detection reaches 99.3%, while the license plate recognition accuracy achieves a precision of 89.21%. Evaluation using real-world video sequences further validates robust performance, indicating suitability for practical deployment in diverse operational scenarios.

Keywords: car license plate recognition; license plate detection; YOLOv12; OCR.

Chi tiết bài viết

Tài liệu tham khảo

[1] M. Dong, D. He, C. Luo, D. Liu, and W. Zeng, “A CNN-Based Approach for Automatic License Plate Recognition in the Wild,” in Procedings of the British Machine Vision Conference 2017, London, UK: British Machine Vision Association, 2017, p. 175. doi: 10.5244/C.31.175.
[2] R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” Oct. 22, 2014, arXiv: arXiv:1311.2524. doi: 10.48550/arXiv.1311.2524.
[3] R. Laroca et al., “A Robust Real-Time Automatic License Plate Recognition Based on the YOLO Detector,” in 2018 International Joint Conference on Neural Networks (IJCNN), Jul. 2018, pp. 1–10. doi: 10.1109/IJCNN.2018.8489629.
[4] Z. Akhtar and R. Ali, “Automatic Number Plate Recognition Using Random Forest Classifier,” SN Comput. Sci., vol. 1, no. 3, p. 120, May 2020, doi: 10.1007/s42979-020-00145-8.
[5] Y. LeCun, L. Bottou, Y. Bengio, and P. Ha, “Gradient-Based Learning Applied to Document Recognition,” 1998.
[6] tesseract-ocr/tesseract. (Mar. 09, 2026). C++. tesseract-ocr. Accessed: Mar. 09, 2026. [Online]. Available: https://github.com/tesseract-ocr/tesseract
[7] M. D. C. Velarde and G. Velarde, “Benchmarking Algorithms for Automatic License Plate Recognition,” Mar. 27, 2022, arXiv: arXiv:2203.14298. doi: 10.48550/arXiv.2203.14298.
[8] Ultralytics, “YOLOv5: A state-of-the-art real-time object detection system.” Accessed: Apr. 14, 2024. [Online]. Available: https://github.com/ultralytics/yolov5
[9] I. R. Khan et al., “Automatic License Plate Recognition in Real-World Traffic Videos Captured in Unconstrained Environment by a Mobile Camera,” Electronics, vol. 11, no. 9, p. 1408, Apr. 2022, doi: 10.3390/electronics11091408.
[10] Z. E. Vargoorani and C. Y. Suen, “License Plate Detection and Character Recognition Using Deep Learning and Font Evaluation,” vol. 15154, 2024, pp. 231–242. doi: 10.1007/978-3-031-71602-7_20.
[11] S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” Jan. 06, 2016, arXiv: arXiv:1506.01497. doi: 10.48550/arXiv.1506.01497.
[12] A. Howard et al., “Searching for MobileNetV3,” Nov. 20, 2019, arXiv: arXiv:1905.02244. doi: 10.48550/arXiv.1905.02244.
[13] G. Jocher, A. Chaurasia, and J. Qiu, Ultralytics YOLOv8. (Jan. 2023). Python. Accessed: Mar. 27, 2024. [Online]. Available: https://github.com/ultralytics/ultralytics
[14] Z. E. Vargoorani, A. M. Ghoreyshi, and C. Y. Suen, “Efficient License Plate Recognition via Pseudo-Labeled Supervision with Grounding DINO and YOLOv8,” in 2025 IEEE 35th International Workshop on Machine Learning for Signal Processing (MLSP), Aug. 2025, pp. 1–6. doi: 10.1109/MLSP62443.2025.11204315.
[15] J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You Only Look Once: Unified, Real-Time Object Detection,” May 09, 2016, arXiv: arXiv:1506.02640. doi: 10.48550/arXiv.1506.02640.
[16] Y. Tian, Q. Ye, and D. Doermann, “YOLOv12: Attention-Centric Real-Time Object Detectors,” Feb. 19, 2025, arXiv: arXiv:2502.12524. doi: 10.48550/arXiv.2502.12524.
[17] T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré, “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness,” Jun. 23, 2022, arXiv: arXiv:2205.14135. doi: 10.48550/arXiv.2205.14135.
[18] C. Cui et al., “PaddleOCR v3.3.2.0 Technical Report,” Jul. 08, 2025, arXiv: arXiv:2507.05595. doi: 10.48550/arXiv.2507.05595.
[19] F. Bordes et al., “An Introduction to Vision-Language Modeling,” May 27, 2024, arXiv: arXiv:2405.17247. doi: 10.48550/arXiv.2405.17247.