Build1 distinct publisher3 min readUpdated
A dev.to post ships a deliberately misconfigured AWS environment and a CSV scorecard so tools can be graded on found-versus-missed. The comparison it argues for has not been published yet.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
A post on dev.to has published what cloud security tooling has lacked: a Terraform-deployed AWS environment carrying 30 documented misconfigurations across eight services plus five multi-resource attack paths, scored with a CSV where each ground truth ID is marked FOUND, MISSED, PARTIAL or N/A [6][12]. That matters because the two normal ways to evaluate a scanner, watching a vendor demo or pointing candidates at your own account, both fail the same way: the vendor picks the misconfigurations in the first case, and in the second you cannot distinguish "the tool did not find it" from "it does not exist to find" [3][4].
The author's framing is TodoMVC, the same small app implemented in every JavaScript framework so differences surface on uniform ground [1]. The complaint is that Prowler shows its best findings, Wiz shows its graph, AWS Config shows its rules, and nobody runs all of them against one deliberately broken environment and publishes what each found and each missed [2]. The failure mode is quantitative, not aesthetic: if Tool A reports 200 findings and Tool B reports 150, the count alone does not tell you whether A is better or just noisier [5].
The design choice that makes this usable is refusing arguable findings. The post's standard is binary and objective, "this S3 bucket allows public read access" rather than a role that might be overprivileged depending on risk threshold [9]. Each item carries a unique ID, description, severity and manual verification steps, so you can confirm in the console that the thing you are grading against is actually there [7]. Coverage spans S3, IAM, CloudTrail, KMS, EC2, ELBv2, OpenSearch and Config [13], grouped as public exposure, encryption gaps, logging deficiencies, identity issues and network configuration [14]. Deployment takes about ten minutes and roughly $2/day [8], so a week of evaluation costs about $14 [22]. Nothing to install beyond the tools under test, no vendor contact, no account [19].
The five compound paths are where the fixture earns its keep. A publicly reachable instance with an overbroad IAM role that can read an unencrypted bucket is one attack path, not three findings [10], and the scorecard separates tools that report the resources from tools that connect them [11]. Each path documents entry point, pivot and target, with every step referencing a specific atomic finding [15]. The post cites the CSA Top Threats to Cloud Computing 2026 report saying attackers "exploit the seams between products" [16], with a chain that runs from a compromised third-party API through an overprivileged identity to a misconfigured bucket that no single product sees end to end [17]. Most tools check one resource at a time [20].
Two caveats. This is one author's fixture, and the only filled-in scorecard shipped with it is Stave's own results, offered as a worked example rather than a claim [18]. And the resolution is coarse: with 35 scoreable items, of which five are compound [21], a single missed path moves the compound score by 20 percentage points [23], which is a lot of weight on judgment calls about what counts as PARTIAL.
Watch whether anyone publishes side-by-side scorecards for the three tools the post names, because the fixture is only the control, not the experiment [2][18]. Watch how PARTIAL gets adjudicated when a tool reports every resource in a chain but never links them [11]. And watch whether the 30 atomic items stay binary as they age, since the moment an item becomes debatable, the scorecard turns back into a narrative.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
In JavaScript, TodoMVC provides the same app implemented in every framework so that comparisons happen on uniform ground, revealing differences such as bundle size, rendering approach, state management and verbosity.
When a vendor runs a demo, the vendor chooses the environment, which misconfigurations to show, and the narrative.
In a POC run against your own environment you cannot control which misconfigurations exist, so you cannot tell the difference between 'the tool did not find it' and 'it does not exist to find'.
If Tool A reports 200 findings and Tool B reports 150, the counts alone do not establish whether Tool A is better or merely noisier.
The published fixture is a Terraform-deployed AWS environment containing 30 documented misconfigurations across 8 services, plus 5 multi-resource attack paths.
Each misconfiguration has a unique ID, a description, a severity rating, and manual verification steps so it can be confirmed in the AWS console.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single self-published source, artifact self-described
Every specification - 30 misconfigurations, 8 services, 5 paths, ~10 minute deploy, ~$2/day - comes from one dev.to post written by the kit's own maker. The design is internally coherent and the misconfigurations are described as console-verifiable, which makes the claims checkable in principle, but nothing in the cluster shows anyone outside the publisher deploying the fixture, and the CSA 2026 report is quoted rather than supplied.
Published artifact, no third-party usage shown
There is a concrete release event - the fixture and pre-filled scorecard exist and are usable without vendor contact - and one completed scorecard, but that scorecard is the publisher's own. No downloads, stars, external evaluations, or submitted community results appear in the supplied material, and the pooled 'coverage map' is described as a future possibility.
Mildly overstated: benchmark framing, no comparison shipped
The headline positions the work as cloud security's TodoMVC, but TodoMVC's value is the published side-by-side; here the fixture and scoring sheet ship while the multi-tool found-versus-missed comparison the post argues for does not. The gap is kept modest by the author's own hedges: the kit is explicitly not a ranking or verdict, the reference scorecard is labelled 'not a claim', and the piece says production evaluation is still required.
Vendor-adjacent: kit maker also supplies the only score
The post defines the evaluation ground truth, defines the scoring scheme, and ships one filled-in scorecard belonging to Stave, while naming competing tools (Prowler, Wiz, AWS Config) as the demo-driven status quo. Whoever sets the fixture chooses which misconfiguration categories count, which is a commercial position regardless of the 'not a claim' disclaimer, and the post does not disclose the author's relationship to Stave.
Concrete artifact, uncorroborated and first-party
Confidence is moderate-low: the described mechanics are specific, falsifiable and cheap for any reader to test, which makes the story unlikely to be fabricated, but the entire cluster is one vendor-adjacent post with no independent replication, no repo inspection, no external adoption, and a key premise stated as an unverified absence.
build
CSA's 2026 threat list is a flat line, so ask which threats a config snapshot can prove1 distinct publisher
build
Two mechanisms, one vCPU floor: why db.t3.micro cannot meet a 1-second RPO on RDS1 distinct publisher
security
ToxicPanda 2.0 Widens From 16 Apps to 140, and From Overlays to ADB Shell3 distinct publishers
build
Amazon Q executed code from any repo you opened, and it is not the only one1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 20, 2026