TruaceTracing the truth around AIMonday, July 20, 2026
Crime·G Space·Evidence-backed gain·Published 2026-07-20

LLMs for Cybersecurity in the Big Data Era: A Comprehensive Review of Applications, Challenges, and Future Directions

This paper presents a systematic review of research (2020–2025) on the role of Large Language Models (LLMs) in cybersecurity, with emphasis on their integration into Big Data infrastructures. Based on a curated corpus of 235 peer-reviewed studies, this review synthesizes evidence across multiple domains to evaluate how models such as GPT-4, BERT, and domain-specific variants support threat detection, incident response, vulnerability assessment, and cyber threat intelligence. The findings confirm that LLMs, parti…

TRV-2026-0456Peer-reviewedPermanent record — cite & verify
LLMs for Cybersecurity in the Big Data Era: A Comprehensive Review of Applications, Challenges, and Future Directions

Maritime cybersecurity: the future of national security by Hayes, Christopher R. Public domain

The quick read

As of 4 November 2025, this systematic review synthesized 235 peer-reviewed studies from 2020-2025 on LLMs such as GPT-4 and BERT applied to cybersecurity within Big Data infrastructures. It reported that coupling LLMs with scalable pipelines improved detection accuracy and reduced response latency for tasks including threat detection, incident response, and vulnerability assessment.

The findings matter because faster, more accurate detection directly affects prevention and mitigation of cybercrime, but the review also documents unresolved risks that could undermine deployment. Questions remain about how to address adversarial susceptibility, data leakage, computational costs, and limited transparency while advancing domain-specific models and explainability.

Main points
  • Systematic review of 235 peer-reviewed studies from 2020-2025 following PRISMA-2020 with risk of bias assessment and random-effects syntheses.
  • Evaluated models including GPT-4, BERT, and domain-specific variants integrated into Big Data infrastructures.
  • Proposed unified taxonomy and identified future priorities: robustness, bias mitigation, explainability, domain-specific models, and distributed integration.
Gain

LLMs coupled with scalable Big Data pipelines improved detection accuracy and reduced response latency for threat detection and incident response compared with traditional approaches.

The rundown

The review curated 235 peer-reviewed studies from 2020 to 2025, last searched 30 April 2025, using PRISMA-2020 methods with risk of bias assessment and random-effects syntheses.

It examined integration of general and domain-specific LLMs into scalable Big Data pipelines for cyber threat intelligence and vulnerability assessment, consolidating fragmented research into a taxonomy and outlining priorities around robustness, bias, explainability, and distributed integration.

Sources

Reader signal

How should this claim be treated?

The debate