PaperScope
LIVE · 2026-09-09 05:40 UTC

Evidence-Grounded Retrieval for Investigation Hunt Lead Generation from CTI Reports

Akash Prakash, Boubakr Nour, Makan Pourzandi, Chadi Assi, Mourad Debbabi

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.08790 v1
Category
Submitted
2026-09-08

Abstract

Threat hunting increasingly depends on converting unstructured knowledge (e.g., Cyber Threat Intelligence reports) into actionable hunt leads: concise, investigable hypotheses grounded in observable artifacts and adversary techniques. Producing such leads manually is a tedious and hard-to-scale task. Existing automated approaches stop at the entity layer, ignore the defender's operational environment, and analyze each report in isolation. To address these gaps, we introduce AHLERT, a system that automatically extracts relevant, environment-aware, and hunt leads from threat reports through (i) a hybrid retriever that combines dense vector search with multi-hop traversal over a knowledge graph seeded with MITRE ATT&CK; (ii) an ontology-grounding retrieval-augmented generation method that constrains each lead to the defender's own assets and controls; and (iii) an LLM-agnostic framework that emits structured, directly actionable leads rather than loose indicators of compromise. We evaluate AHLERT on public CTI reports for well-known APTs across multiple proprietary and open-weight models. Hybrid evidence retrieval with ontology grounding raises mean F1 by ~2x (0.44 to 0.85) over a single-route flat-RAG baseline, and AHLERT attains the highest effectiveness score (~86.95%) compared with off-the-shelf LLM models.

Comment: Accepted for presentation and publication at the 2026 IEEE Conference on Dependable and Secure Computing (DSC) - Workshop: Cyber Resilience & Attack Intelligence (CRAI)

arXiv abs page · PDF · same-day batch