Build1 distinct publisher2 min readUpdated
playwright-score is open source, deterministic and AI-free, and built by a firm that sells managed QA. The corpus it graded, Supabase and Grafana included, still fails on raw selectors and empty tests.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The most consequential design decision in this scorer is a penalty cap: after three occurrences of the same rule in the same file, further hits stop counting toward the score, though the full findings list keeps every one [12]. That is the whole reason cal.com's suite can produce 1,215 findings from 53 files, the highest density in the corpus, and still land on 85 and clear the threshold [10]. Mattermost, with 284 files and 1,157 tests, produces 1,476 findings, roughly five per file against a corpus average of 4.5, and grades 90 [11][1]. cal.com alone accounts for about 22 percent of every finding in the corpus while holding 4 percent of its spec files [3]. What actually drags its grade is the locator ratio: 613 role-based calls against 875 raw ones, 41.2 percent native where the corpus sits at 69.9 percent [13][6].
So the grade measures how widely an anti-pattern is spread, not how often it fires. Two teams can quote the same tool and describe the same suite as a B-grade pass or as the worst offender on the list, and neither is lying.
One number in the writeup does not close. Thirty percent of 17,118 locator calls is roughly 5,150 raw selectors, while the no-raw-locators rule, at 60 percent of 5,467 findings, is roughly 3,280 [4]. That leaves on the order of 1,900 raw calls that produced no finding, and the post does not reconcile the two counts [7]. Since the run is one script over sparse clones of named subdirectories [5], it is checkable, which is more than can be said for most vendor charts.
The zero-assertion headline deserves the same arithmetic in the other direction. A hundred and two cases across 12 of 17 repos sounds endemic; measured against 5,943 tests it is 1.7 percent [9][2]. Not comforting, because every one of them is green and no pass/fail gate will ever surface them [9], but it is a rate rather than a rot.
The conflict of interest is disclosed and remains a conflict. A managed QA vendor built the instrument, wrote the rules, set the weights and the threshold of 80, and chose the 17 repos [3][2][4]. The AI-free framing is not what makes this worth reading [2]; the reproducibility is [5]. Anyone who thinks sqs-v1 flatters the vendor's own house style can fork it and reweight. Meanwhile the locator rules alone are 74 percent of all findings and the raw-selector rule is flagged in 16 of the 17 repos [7], which leaves at most one clean suite in the corpus [5], three years after Playwright's documentation began recommending role-based locators [14].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The corpus covers 17 public Playwright repositories, 1,214 spec files, 5,943 tests and 168,902 lines of test code, scored on 2026-08-19 against each project's live main or master branch.
The run produced 5,467 individual rule violations across the corpus.
Scores come from @qaguardian/playwright-score, scoring version sqs-v1, standard profile, threshold 80. It is deterministic and AI-free, wrapping eslint-plugin-playwright and adding a versioned 0-100 score, a locator-ratio metric and assertion-delegation tracing through local imports, with no LLM in the scoring path.
The authors state that every managed QA vendor, including themselves, claims to write good Playwright tests, and that almost none of that claim is measurable.
The corpus mixes heavily engineered platforms including Supabase, Grafana, Mattermost and n8n with smaller, less mature projects found by searching for real @playwright/test usage, and the authors say repos were not picked to make the tool or the industry look good or bad.
Each repo was scored via a fresh, shallow, sparse git clone of the exact subdirectory holding its Playwright suite, with nothing vendored or cached, and the run is reproducible from scripts/validate-corpus.sh in the GitHub repo.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific and self-reproducible, but single-source and unreplicated
The quantitative core is unusually well specified for a vendor post: exact corpus totals, a locator-call census, per-rule shares, named repos, a pinned scoring version and profile, and a named reproduction script. All of it, however, comes from one self-published article by the tool's author, and no third party in the cluster has re-run or checked it. One internal tension is unaddressed: the reported raw-locator call volume implies far more raw calls than no-raw-locators findings.
Public availability only, no usage signal
The cluster establishes that the scorer is publicly installable and that its author ran it across 17 repos, but supplies no downloads, dependent projects, customer counts or evidence that any team outside the vendor uses it. Availability is not adoption, and there is nothing here to score it from.
Mildly overstated framing over solid underlying numbers
The article deliberately deflates its own grades and publishes findings that complicate them, which is the opposite of hype. The overreach is in generalisation and framing: 17 search-found repos are presented as the state of Playwright quality in 2026, the sample is self-selected by the party that benefits from the conclusion, and the post ends in a sales CTA. Small positive rather than neutral.
Vendor benchmarks its own tool and closes with a demo pitch
The publisher is the managed QA firm that authored @qaguardian/playwright-score; it selected the corpus, defined the rules, set the threshold, and closes by offering to score a prospect's suite live on a sales call. The authors name this conflict in the first paragraph and offset it partially with an open-source tool, third-party repos and a reproduction script, but every methodological degree of freedom rests with an interested party.
Internally detailed, externally unverified
Confidence is capped by structure rather than by sloppiness: one publisher, one author, no independent replication, and no adoption signal, against unusually specific and checkable figures plus a pinned scoring version. The deterministic, AI-free design and named script make the claims verifiable in principle, which is why this sits near the middle rather than low.
build
A coding agent deleted the rule its own rename had broken1 distinct publisher
build
Notion's agent stack is live, not slideware, and it only changes one of your decisions1 distinct publisher
build
Prisma v7 stops seeding for you, and the pooled URL will not finish the job1 distinct publisher
build
CI cannot tell a regression from a stale test because nobody wrote the intent down1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 22, 2026