Build1 publisher3 min readPublished Updated
Flaky CI is a review standards problem, and this checklist names the three gates
A published Playwright and Cucumber review checklist puts flakiness at the merge rather than the runner: state isolation, synchronisation strategy and single-behaviour scoping.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- An author publishing as shefali_qa on dev.to released the complete code review checklist their team uses for Playwright + JavaScript + Cucumber BDD test suites, presented as a senior QA guide to reliable CI test suites.
- The post states that to prevent test rot, flakiness and architectural drift, code reviews need to evaluate more than standard syntax and must enforce design boundaries, state isolation and robust synchronization strategies.
- The checklist is declared applicable to new Playwright test scenarios, updated Cucumber feature files and step definitions, changes to page objects, hooks, utilities, config and test data, and reporting, retry, CI and environment-related automation changes.
- The published excerpt lists these review categories: Test Intent and Coverage; Feature File Quality (Cucumber); Step Definition Quality; Page Object Design; Selector Strategy; Assertions and Validation; Waiting and Synchronization; Test Isolation and State Management; Test Data Handling.
- Nine review categories appear in the published excerpt of the checklist.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A senior QA engineer writing as shefali_qa on dev.to has published the pull request review checklist her team uses on Playwright, JavaScript and Cucumber BDD suites [1]. The diagnosis behind it is the part worth arguing about: according to the post, reviews have to enforce design boundaries, state isolation and robust synchronisation strategies, because syntax review alone lets test rot, flakiness and architectural drift through [2].
That relocates flake. A suite that fails at random under parallel execution usually gets treated as an infrastructure problem and answered with retries and a bigger runner. This checklist treats it as an artefact of what was approved. Note what falls inside its declared review scope: not only new scenarios, feature files and step definitions, but page objects, hooks, utilities, config, test data, and reporting, retry, CI and environment changes [3]. Retry configuration is the standard sedative for a flaky suite, and here it is itself a reviewable change.
Three of the nine categories in the published excerpt [4][5] carry most of the load.
State isolation is the first. The checklist asks whether a scenario can run independently and in any order, whether browser context, page, storage state and created records are isolated per scenario, whether cleanup exists for created users, groups, files or seeded data, and whether the change introduces hidden dependencies on previously executed tests [6]. It also asks that shared scenario state live on the Cucumber World object rather than in globals [7], and that unique identifiers be used to avoid collisions in shared environments [8]. Each of those can be answered from a diff. None requires reproducing the failure first.
Synchronisation is second: rely on Playwright auto-waiting where possible, tie explicit waits to real signals such as element state, navigation or API responses, avoid arbitrary sleeps such as waitForTimeout(), and wait on the relevant response or UI state change for network-dependent behaviour [9]. The item doing the real work is the last one, which asks whether a race condition could appear under CI speed or parallel execution [9]. Asking that out loud, per pull request, is cheaper than triaging it out of a nightly run three weeks later.
Third is scoping. The checklist asks whether a scenario is scoped to a single behaviour instead of combining multiple unrelated validations [10]. Multi-behaviour scenarios are where diagnosis gets expensive, because the failure signal points at a scenario name rather than at one behaviour.
Adjacent items serve the same end without being about timing. Selectors are to be centralised in page objects, with data-qa, data-test, data-testid or semantic Playwright locators preferred, positional XPath, generated classes and deep CSS chains avoided, and any fallback rationale documented [11]. Assertions should verify outcomes that matter to the user or workflow, with weak checks avoided, such as confirming a page loaded without confirming expected data or state [12].
Two honest limits. This is one team's checklist, offered for adoption or adaptation [13], not evidence: it reports no flake rates, defect counts or before-and-after CI figures [15]. And the material available here is cut off inside the test data section [14], so the real list is longer than what is quoted.
What to watch is whether teams that pick this up make isolation and synchronisation blocking review items or advisory ones, and whether retry counts in their CI config fall afterwards. A checklist that never blocks a merge is documentation.