Deep Learning-Based Image Recognition Systems: Current Trends and Future Directions
Keywords:
deep learning, image recognition, convolutional neural networks, vision transformers, self-supervised learning, object detection, semantic segmentationAbstract
The power of deep learning has revolutionized image recognition, setting new records for accuracy on standard test sets, and greatly broadened the scope of tasks where it can be used in domains many considered unapproachable for an automatic system to tackle. In this paper, a thorough survey of deep-learning based image recognition systems is provided, starting from the classic convolutional network and moving towards recent vision transformers and hybrid networks. This course critically reviews the latest developments in areas such as self-supervised learning, multi-modal vision-language models, neural architecture search and computer vision foundation models. Medical imaging, autonomous driving, remote sensing, industrial inspection and agriculture monitoring are used as examples of application areas. We compare the models across 10 representative architectures from 2012 to 2024, showing that we've moved from AlexNet's 63.3% top-1 accuracy on ImageNet to EVA-CLIP's 89.7%. The directions for the future include explainability, data efficiency, robustness to distribution shift, and sustainable AI. The results of this review consist of a synthesis of the results from more than 200 publications (2018–2024), which is structured in a valuable resource for researchers and practitioners.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Author(s)

This work is licensed under a Creative Commons Attribution 4.0 International License.