Build1 publisher3 min readPublished
Swapping an LLM for a scored Places API lookup lifted a sales pipeline's website coverage to 81%
Engineers on a freight-forwarder sales pipeline swapped an LLM web-search step for a Google Places lookup and lifted website completeness from 34% to 81%. The author credits the gain to the normalisation and match-scoring layer built around the API.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Re-running the same lead could return a different domain, so the step had no fixture to test against.
- The Places version scores each candidate on name overlap, distance and business type, and returns nothing when the score falls below a threshold.
- The step's LLM cost fell to zero after the swap.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision The author's assertion test gives a way to sort each model call in a pipeline, keeping the model for open tasks like summarising a thread and moving fixed-answer steps to an index.
- exposure Any pipeline that feeds a generated value straight into customer contact carries the same risk, because a fabricated domain looks exactly like a correct one.
- cost Leads below the match threshold now go to manual review, so some of the saving on model spend comes back as staff time.
- constraint The approach depends on lead records carrying city and country, since without them the index returns a similarly named business in the wrong country.
The author's diagnosis rests on one question: "Does this step need judgement, or does it need an index?" [11] In practice it becomes an assertion. If `expected == actual` is a sensible check for a step, the author wrote, the step is a lookup and a model is the wrong tool for it [12]. A company's website passes. It has one correct answer, it is a fact about the world, and someone already maintains an index of it [13].
The LLM version failed that check in ways that were hard to see early. It took about twenty minutes to wire up and looked fine for the first hundred leads [9]. Nothing else in the pipeline could run without a website [1]. Re-running a lead could change the returned domain, so there was nothing to assert against and no way to tell a regression from ordinary variance [6]. Nondeterminism does have one administrative upside: a step with no fixture never officially regresses. The team had filed the 34% website rate as a data problem, blamed on small forwarders with thin web presence, and the author now says it was not one [10].
According to the author, the Places call was the easy part and the layer around it made the swap work [15]. That layer runs in order:
1. Normalise the name before querying: legal suffixes such as Ltd, GmbH and SARL, punctuation, casing and transliteration [16]. 2. Query on name plus city plus country. Without the geography, the index returns a real business with a similar name in the wrong country [17]. 3. Score each candidate on token overlap with the normalised name, geographic distance and business-type agreement, and return nothing below a threshold [18]. 4. Cache the result on normalised name plus country, so a second campaign against the same company costs nothing [21].
Step three is the part I would copy first. An empty result flags the lead and routes it to manual review [19]. "Returning a confident wrong answer is not a state anything can handle, because nothing downstream knows to doubt it," the author wrote [20]. The LLM path produced exactly that state: plausible domains for companies that did not own them, found only when outreach reached the wrong firms [5].
The reported figures are completeness, measured on the same lead set before and after [23]. Website completeness rose 47 points [1]. The share of leads with no usable website fell from 66% to 19% [2]. The supplied text of the post does not include a wrong-match rate or the threshold value. So the coverage gain is measured, and the claim that answers are also right more often rests on the threshold design. The 91% phone figure has no baseline, because the LLM path often returned no phone at all and the metric went untracked [c3, c22].
For the 81% to transfer, your entities have to sit in an index someone maintains, and your records have to carry enough geography to key the query [c13, c17]. I would also expect the threshold to need tuning against your own names before the numbers look like these.
The zero applies to LLM spend on this step [4]. A Places lookup still costs something, which the author puts at a fraction of a cent [14]. The larger change is in how cost scales. The LLM step's cost grew linearly with lead volume, and the same company across three campaigns meant three paid lookups [c8, c7]. With the cache keyed on name and country, the repeat lookups drop out [21].
What to watch
- A wrong-match rate for the Places path on the same lead set, which would show whether the coverage gain is also a correctness gain.
- How fast the manual-review queue grows as lead volume rises, since every below-threshold lead now needs a person.
- Whether the normalise, geography, score and cache sequence holds up on entity types that an existing index covers less well than local businesses.