CyberIntel ⬡ News
★ Saved ◆ Cyber Reads
← Back ◬ AI & Machine Learning Aug 04, 2026

Multi-LLM Consensus Framework for Evaluating Banking-Sector NIDS Dataset Coverage of MITRE ATT&CK Techniques

arXiv Security Archived Aug 04, 2026 ✓ Full text saved

arXiv:2608.00895v1 Announce Type: new Abstract: The systemic criticality of global banking networks has ren-dered them high-priority targets for advanced persistent threats, neces-sitating Network Intrusion Detection Systems (NIDS) whose operational effectiveness must extend beyond statistical accuracy. However, a signif-icant validation gap persists between experimental NIDS performance and real-world effectiveness: NIDS models that achieve high accuracy on standard benchmarks often fail in ope

Full text archived locally
✦ AI Summary · Claude Sonnet


    Computer Science > Cryptography and Security [Submitted on 1 Aug 2026] Multi-LLM Consensus Framework for Evaluating Banking-Sector NIDS Dataset Coverage of MITRE ATT&CK Techniques Sanjida Khanom, Sadia Afrin Khan, Adrita Rahman Tory, Md. Ahsan Habib, Khondokar Fida Hasan The systemic criticality of global banking networks has ren-dered them high-priority targets for advanced persistent threats, neces-sitating Network Intrusion Detection Systems (NIDS) whose operational effectiveness must extend beyond statistical accuracy. However, a signif-icant validation gap persists between experimental NIDS performance and real-world effectiveness: NIDS models that achieve high accuracy on standard benchmarks often fail in operational banking environments because generic datasets lack sector-specific patterns, such as SWIFT and ATM-related intrusions, that characterize real financial threats. To address this, the paper investigates a sector-aware evaluation method-ology that systematically assesses how well existing NIDS benchmark datasets cover the attack behaviors most relevant to banking infrastruc-ture. The methodology maps documented adversary behaviors from the MITRE ATT&CK knowledge base to NIDS benchmarks while enforcing the realistic sensor limitations defined by NIST SP 800-94. Leveraging a multi-LLM consensus engine with four state-of-the-art models, we evalu-ated 210 banking-specific adversary techniques to derive a baseline of 68 network-observable behaviors for systematic coverage analysis. Results across five benchmark datasets demonstrate that UNSW-NB15 achieves the highest utility with an 82.2% weighted coverage score (though only 18.4% reflects direct, technique-level evidence), while CIC-DDoS2019 re-veals an 89.9% blind spot for core banking behaviors. These findings es-tablish a reproducible foundation for sector-aware NIDS evaluation and highlight the urgent need for banking-native datasets. Comments: 17 pages Subjects: Cryptography and Security (cs.CR) Cite as: arXiv:2608.00895 [cs.CR]   (or arXiv:2608.00895v1 [cs.CR] for this version)   https://doi.org/10.48550/arXiv.2608.00895 Focus to learn more Submission history From: Khondokar Fida Hasan [view email] [v1] Sat, 1 Aug 2026 23:23:13 UTC (1,518 KB) Access Paper: view license Current browse context: cs.CR < prev   |   next > new | recent | 2026-08 Change to browse by: cs References & Citations NASA ADS Google Scholar Semantic Scholar Export BibTeX Citation Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Demos Related Papers About arXivLabs Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)
    💬 Team Notes
    Article Info
    Source
    arXiv Security
    Category
    ◬ AI & Machine Learning
    Published
    Aug 04, 2026
    Archived
    Aug 04, 2026
    Full Text
    ✓ Saved locally
    Open Original ↗