Science1 distinct publisher3 min readPublished
Planting fake reviews is usually priced as costless if you get away with it. A George Mason working paper puts a number on not getting away with it, using Yelp alerts, foot-traffic counts and card spending.
The Scientist · Science desk
Compiled by The ScientistSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
science
Kickstarter's 2018 eco-pledge rollout gave accountants a rare clean test of green capital1 distinct publisher
invest
Meta's Hatch agent tops out at $199.99 a month, with DoorDash and Etsy behind the meter2 distinct publishers
product
The AI store manager did not fire anyone until humans told it to read its own policy1 distinct publisher
The upside and the downside sit far apart. A strong star profile is worth roughly 2.9% in extra foot traffic [7]; the post-alert drop is about 8% against matched peers [8]. Divide, and the punishment runs about 2.8 times the prize [14]. That figure carries an important caveat: the 2.9% is a cross-sectional comparison between high- and low-reputation neighbours rather than the measured payoff of any particular manipulation campaign, while the 8% is a before-and-after move around a dated event. The ratio is useful for scale, not for arithmetic on a business plan.
The persistence is the part worth explaining, and the researchers locate the mechanism inside the platform. Once the banner goes up, the flow of new reviews around that business thins and what still arrives skews negative [11]. The usual remedy for a damaged rating is more rating, and that supply contracts precisely when it is needed. So a sanction Yelp writes as 90 days [3] shows up in the demand data as something on the order of six times longer [15].
On method: 288,426 Yelp firm-month observations covering 2019 to 2024, of which 16,837 sit inside six-month windows around an alert [5], about 5.8% of the sample [16], though that count bundles flagged firms with their demand-side peers and so is not a rate of flagging. Foot traffic comes from SafeGraph, and for a subset of businesses the authors add Consumer Edge card transactions [6]. That second source earns its keep, because visits and spending can diverge; here they moved together [8]. Comparing each flagged firm against matched local peers absorbs the street-level shock that would otherwise be mistaken for reputational damage. What no matched-peer design can do is make the flagging random. Firms that plant reviews, and firms that get reported for it, may already be on a different path than the shop next door. Cao's co-authors are Sean Wang, John Bai of Hong Kong Polytechnic University and Chi Wan of San Diego State University [2].
One number is missing from this whole picture: the probability that sits in front of the 8%. The 4,900-plus disciplined businesses counted as of 2023 [4] are a tally of enforcement, not of detection, and without an estimate of how many manipulators are never flagged the expected cost of the tactic stays unpriced. Cao puts the decision squarely on that perceived risk: when it looks low, owners judge the incremental benefit worth taking [12]. Two further limits: this is a working paper [1], and an 8% fall in visits is not an 8% fall in profit, since traffic and card spend measure demand and say nothing about the margin carried on it. Detection is also not fixed. Cao describes Yelp leaning substantially on reports from consumers and owners, alongside reported algorithmic screening for AI-written language [13]. If that screening improves, the loss term stays where it is and the probability in front of it rises.
Ranked by verification strength, evidence, and original report placement.
The research is a recently published working paper from Yi Cao, assistant professor of accounting at the Costello College of Business at George Mason University.
The paper was co-authored by Sean Wang (formerly of Southern Methodist University), John Bai of Hong Kong Polytechnic University and Chi Wan of San Diego State University.
To punish businesses it deems guilty of gaming the customer-review system, Yelp places a prominent 'Consumer Alert' banner on all their associated listing pages; these are typically 90-day penalties.
As of 2023, more than 4,900 businesses had been disciplined on Yelp with a Consumer Alert.
The researchers used Yelp data for 2019 to 2024: 288,426 firm-month observations in all, including 16,837 observations of flagged firms and their demand-side peers during the six months surrounding an alert.
The Yelp data was analysed alongside monthly foot-traffic reports from location data firm SafeGraph, and for some businesses credit card transaction data obtained through Consumer Edge.
Distinct publishers with included, body-backed reporting in this cluster.
phys.org
1 article · August 27, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Precise numbers, one unrefereed source
The figures are specific to the digit — 288,426 firm-months, 16,837 alert-window observations, 8%, 2.9% — and they all reach readers through one phys.org retelling of a working paper whose only identifier is an SSRN DOI. Three independent data streams is a genuinely strong design, but the matching that turns flagged firms into comparable peers is named rather than shown, the card-data corroboration is asserted without a number, and nobody outside the four authors has looked at any of it.
The punishment is at scale; the finding is days old
Two very different things are being adopted here. Yelp's Consumer Alert is unmistakably in production — 4,900-plus businesses tagged by 2023, and a six-year panel of nearly 290,000 firm-months of behaviour built around it. The research conclusion has no uptake at all: freshly circulated, unrefereed, cited by nobody in this reporting, and reflected in no platform's stated policy. The score reflects a live mechanism, not a travelling result.
Framing outruns a result that didn't need it
The finding is strong enough to survive being stated plainly, which makes the decoration costly. "Recovery was nowhere on the horizon" is a reading of a flat line inside a window that stops when the data stops, not an observed failure to recover; "hurts businesses big time" adds nothing the 8% hasn't already said. Meanwhile the sturdier insight — that the badge itself starves the review flow a business would need to climb back — is tucked in near the end where it does the least work.
University publicity channel, no counterparty called
This arrived the way academic findings usually arrive: an institutional announcement retold, the lead author quoted three times, no dissenting or affected party in the room. George Mason's business school, three co-authors' universities and both commercial data vendors are all named on the way past. Yelp — whose deterrence case the paper strengthens considerably, and whose penalty design the paper implicitly evaluates — was not asked for anything. None of that makes the 8% wrong; it does mean no one in the chain from author to reader had a reason to push back on it.
Confident about what was said, less about what is true
We can be fairly sure of the claims as stated — the piece is internally consistent and unusually numerate for a research write-up. What we cannot underwrite is the step from those numbers to the world. One outlet, one paper outside review, and a candid admission that even the trigger mechanism is guesswork ("there are also reports" of an AI-language screen). The ratios we computed from the study's figures are safe; the causal weight is not ours to grant until someone else checks the matching.