Data Mining Techniques for Big Data Analytics: A Comprehensive Review

Authors

  • Kanika Khulbe 2Student of B.Tech. (AI & ML), Department of Computer and Technology, Amarapali University, Haldwani, 263139, Uttrakhand, India Author
  • Aanchal Dutt Student of B.Tech. (AI & ML), Department of Computer and Technology, Amarapali University, Haldwani, 263139, Uttrakhand, India Author

Keywords:

Data Mining; Big Data Analytics; Machine Learning; Distributed Computing; Pattern Recognition; Stream Mining; AutoML; Explainable AI; Privacy-Preserving Analytics

Abstract

Big data analytics has emerged as a key talent needed by today's organizations as digital services, cloud platforms, sensor networks, enterprise information systems and social media are rapidly growing. Data mining has laid the analytical groundwork for this capability by making patterns, predictions, explanations and decision support out of very large, diverse and rapidly changing data. In this review, the authors have summarized recent research in data mining techniques for big data analytics, focusing on the research published in the period of 2018 to 2025. Supervised, unsupervised, association-based, regression, deep learning, stream mining, privacy-preserving, automated machine learning and explainable data mining approaches are subject of the paper. It also examines the ways in which these techniques are applied in distributed processing applications such as cloud data lakes, Hadoop-like batch systems, in-memory processing (using Apache Spark, for instance), stream processing engines, and new edge-cloud applications. Applications in healthcare, finance, retail, social media, cybersecurity, smart industry and public services are mentioned. A comparative analysis is used to benchmark techniques based on their scalability, interpretability, data suitability, latency, and governance needs. From the review, several key issues such as data quality, veracity, class imbalance, concept drift, privacy, security, computational cost, interoperability, reproducibility, and explainability are identified. Lastly, a conceptual trustworthy big data mining framework is suggested, which is based on data ownership, flexible infrastructure, model automation, human supervision and ongoing monitoring. Future research directions encompass the development of lightweight explainable models for distributed deployment, privacy preserving analytics across massive data streams, adaptive stream mining, foundation model-driven analyses of responsible use and benchmarking in real world big data pipelines.

Downloads

Published

2026-06-10