Skip to content

Build1 publisher3 min readPublished

Microsoft is generating its detection test logs, and admitting what they do not prove

A Defender research write-up turns MITRE ATT&CK procedures into synthetic process logs so rules can be exercised without range time. It also says synthetic logs are not attack reproduction.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Microsoft is generating its detection test logs, and admitting what they do not prove
Generated illustration

What happened

  • Microsoft Security Blog / Microsoft Defender Security Research Team published "Accelerating detection engineering using AI-assisted synthetic attack logs generation", publication date May 12, 2026.
  • The research feeds MITRE ATT&CK attack techniques and specific attack steps to an AI and creates detection test logs that include process names, parent processes, and command lines.
  • In experiments, a method where multiple AIs share the roles of generation, review, and correction worked best.
  • Synthetic logs are not proof of real-world attack reproduction and are limited to supporting lab tests.
  • The goal of the research is not to reproduce real logs word for word; it is to create logs that keep the meaning, parent-child process relationships, command contents, and event order needed for detection rules to trigger.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Microsoft's Defender Security Research Team has published work on generating detection test logs directly from attack procedures, in a post dated May 12, 2026 titled "Accelerating detection engineering using AI-assisted synthetic attack logs generation" [1]. The practical consequence for detection engineers is scheduling: you can exercise a rule against structured telemetry without booking lab time, and the research itself states that synthetic logs are not proof of real-world attack reproduction and are limited to supporting lab tests [4].

The mechanics are unglamorous, which is the point. The system is fed MITRE ATT&CK tactics and techniques, the specific operations an attacker executes, context such as target OS or scenario, and for multi-step attacks the preceding and following operations plus host relationships [7]. It emits new process names, parent process names, command lines, event ordering, and related records where activity spans multiple hosts [9]. The worked example in the write-up is T1202, Indirect Command Execution, combined with forfiles, environment variables, hex representation, and Python [8]. The stated goal is not word-for-word fidelity to real logs but preserving the meaning, parent-child process relationships, command contents, and event order that a detection rule needs in order to fire [5]. That is a claim about rule syntax and logic, not about adversary behaviour.

Four approaches were tried: prompt-based generation, multi-AI collaboration, LLM-as-a-Judge, and reinforcement learning with verifiable rewards [10]. An expert-guided interactive loop, where a human sets the scenario and a second model checks realism and consistency, held up on simple scenarios but became unstable on complex multi-step attacks [11]. The best result came from splitting the work across models: a Generator drafts, an Evaluator flags missing events and contradictions, an Improver revises, and the cycle repeats [3][12]. According to the write-up, that loop filled in missing events in complex attacks and kept process lineage in order [12]. The reinforcement learning variant scored generations against ground truth with partial credit for matching meaning and deductions for strings that do not match, then used the scores and reasons to improve generation, but it required a large amount of labeled training data [13].

Evaluation ran against three corpora: ten repeatable attack reproductions built by Microsoft researchers, the public OTRF Security Datasets, and ATLASv2, which supplies Windows Security, Sysmon, Firefox, and DNS logs from ten multi-step attacks on two Windows virtual machines [14]. That is twenty scripted multi-step runs across the two purpose-built sets [18]. The headline metric is recall, compared by matching meaning rather than exact strings [16], where recall means how many important ground-truth events the synthetic logs contain [17]. For ATLASv2 the scoring was confined to malicious activity inside the attack time window [15].

Two limits follow from that design. A recall-only metric scored inside the attack window measures whether the generator remembered the events, not whether a rule written against those events survives benign background noise [19]. And the dev.to summary of the research names recall as the main metric without carrying the scores, so the size of the gap between synthetic and real telemetry is not visible from it [20].

Watch whether Microsoft publishes per-technique and per-dataset recall figures, and whether the multi-agent loop's stability on multi-step attacks is quantified rather than described. The labeled-data requirement on the reinforcement learning path [13] is the tell for whether this stays a prompt-engineering practice or becomes a trained pipeline. And the motivation is worth keeping in view: the reason this exists is that real attack logs are rare, expensive to label, and carry sensitive customer data that cannot be shared [6].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories