Skip to content

Build1 publisher3 min readPublished

Cloud security POCs have no control group. A Terraform fixture with 30 known bugs is one attempt.

A dev.to post ships a deliberately misconfigured AWS environment and a CSV scorecard so tools can be graded on found-versus-missed. The comparison it argues for has not been published yet.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • In JavaScript, TodoMVC provides the same app implemented in every framework so that comparisons happen on uniform ground, revealing differences such as bundle size, rendering approach, state management and verbosity.
  • Every vendor demos against its own scenario: Prowler shows its best findings, Wiz shows its graph, AWS Config shows its rules, and nobody runs them all against the same deliberately misconfigured environment and publishes what each one found and what each one missed.
  • When a vendor runs a demo, the vendor chooses the environment, which misconfigurations to show, and the narrative.
  • In a POC run against your own environment you cannot control which misconfigurations exist, so you cannot tell the difference between 'the tool did not find it' and 'it does not exist to find'.
  • If Tool A reports 200 findings and Tool B reports 150, the counts alone do not establish whether Tool A is better or merely noisier.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A post on dev.to has published what cloud security tooling has lacked: a Terraform-deployed AWS environment carrying 30 documented misconfigurations across eight services plus five multi-resource attack paths, scored with a CSV where each ground truth ID is marked FOUND, MISSED, PARTIAL or N/A [6][12]. That matters because the two normal ways to evaluate a scanner, watching a vendor demo or pointing candidates at your own account, both fail the same way: the vendor picks the misconfigurations in the first case, and in the second you cannot distinguish "the tool did not find it" from "it does not exist to find" [3][4].

The author's framing is TodoMVC, the same small app implemented in every JavaScript framework so differences surface on uniform ground [1]. The complaint is that Prowler shows its best findings, Wiz shows its graph, AWS Config shows its rules, and nobody runs all of them against one deliberately broken environment and publishes what each found and each missed [2]. The failure mode is quantitative, not aesthetic: if Tool A reports 200 findings and Tool B reports 150, the count alone does not tell you whether A is better or just noisier [5].

The design choice that makes this usable is refusing arguable findings. The post's standard is binary and objective, "this S3 bucket allows public read access" rather than a role that might be overprivileged depending on risk threshold [9]. Each item carries a unique ID, description, severity and manual verification steps, so you can confirm in the console that the thing you are grading against is actually there [7]. Coverage spans S3, IAM, CloudTrail, KMS, EC2, ELBv2, OpenSearch and Config [13], grouped as public exposure, encryption gaps, logging deficiencies, identity issues and network configuration [14]. Deployment takes about ten minutes and roughly $2/day [8], so a week of evaluation costs about $14 [22]. Nothing to install beyond the tools under test, no vendor contact, no account [19].

The five compound paths are where the fixture earns its keep. A publicly reachable instance with an overbroad IAM role that can read an unencrypted bucket is one attack path, not three findings [10], and the scorecard separates tools that report the resources from tools that connect them [11]. Each path documents entry point, pivot and target, with every step referencing a specific atomic finding [15]. The post cites the CSA Top Threats to Cloud Computing 2026 report saying attackers "exploit the seams between products" [16], with a chain that runs from a compromised third-party API through an overprivileged identity to a misconfigured bucket that no single product sees end to end [17]. Most tools check one resource at a time [20].

Two caveats. This is one author's fixture, and the only filled-in scorecard shipped with it is Stave's own results, offered as a worked example rather than a claim [18]. And the resolution is coarse: with 35 scoreable items, of which five are compound [21], a single missed path moves the compound score by 20 percentage points [23], which is a lot of weight on judgment calls about what counts as PARTIAL.

Watch whether anyone publishes side-by-side scorecards for the three tools the post names, because the fixture is only the control, not the experiment [2][18]. Watch how PARTIAL gets adjudicated when a tool reports every resource in a chain but never links them [11]. And watch whether the 30 atomic items stay binary as they age, since the moment an item becomes debatable, the scorecard turns back into a narrative.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories