Build1 publisher3 min readPublished
Allow-list the closed set, block-list the open one: 193 thin geo pages, one gate
A Moscow new-build listings site cleaned up its programmatic metro and district pages with two lists of opposite shape. The shape choice, not the code, is the interesting part.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- On podbor-minuta.ru, listing pages are grouped by metro station and by district, built ahead of time from the database, each with its own URL, title and a real list of apartments.
- The metro and district values come from a scraper that copies them from other sites as they are.
- The metro field often contained strings like "Tushinskaya (13 min)", with the walking time glued onto the station name instead of a clean station name.
- For grouping, "Tushinskaya", "Tushinskaya (5 min)" and "Tushinskaya (13 min)" count as three different values, so one station turned into several near-identical pages.
- The scraper dumped project names ("ZhK Sezar City"), marketing fragments, and sometimes whole sentences from descriptions with prices, areas and build stages into the district field; real districts were less than half of the values.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer on the Russian new-build listings site podbor-minuta.ru has published how they stopped their own scraper from manufacturing pages: grouping apartments naively by scraped metro and district values produced about 193 thin, near-duplicate pages [6][1]. The fix is not a model or a rewrite, it is two filters of deliberately opposite shape, and the reason they differ is the part worth copying.
The site pre-generates one page per metro station and one per district from the database, each with its own URL, title and real apartment list [1]. Both grouping fields arrive from a scraper that copies them from other sites verbatim [2]. In the metro field that meant strings like "Tushinskaya (13 min)", with walking time glued onto the station name [3]. To a reader that is noise; to a GROUP BY it is fatal, because "Tushinskaya", "Tushinskaya (5 min)" and "Tushinskaya (13 min)" are three values and therefore three near-identical pages [4]. The district field was worse: project names such as "ZhK Sezar City", marketing fragments, and whole sentences lifted from descriptions complete with prices, areas and build stages, with real districts making up less than half the values [5]. The author's stated concern is not the pages themselves but the domain: on a young site, a pile of empty lookalikes can drag the whole thing down in search [7].
Metro is a closed set, so it gets an allow list. The team has a full reference of Moscow stations, so a cleaned value that is not in the reference produces no page at all [8]. Collapsing the duplicates is a separate move: the URL is built from the normalized name rather than the raw string, so every walking-time variant lands on one slug [9]. The exception is instructive. Transit lines are not stations, but MCD, MCK and BKL carry real search demand and thousands of apartments each, so they are kept as their own page type via a small hardcoded set of eight line keys [10][11][12].
District could not use the same medicine, because the reference is incomplete: the districts table is missing okrugs and the towns of the Moscow region, which are real groups, so an allow list would discard roughly half the good data along with the junk [13]. Open sets get block lists instead. Junk is defined by what it looks like: leading "ZhK", "MFK", "residential complex", "tower", "residences", or numbers attached to "million", "thousand", "rub", "sq m", "min", or construction words like "being built", "handed over", "phase" [14]. There is also a blunt length rule, since a district name is never a sentence: anything empty or over 60 characters is rejected [15].
Two implementation notes will save someone an afternoon. JavaScript's "\w" does not cover Cyrillic, so word continuations had to be written out as "[а-яё]*" [16]. And short stems need word boundaries, because a bare search for "dom" matches "Domodedovo" and deletes a legitimate district [17]. Both checks meet in one function that returns a display name and slug or null, and it is the only path by which geo values reach the catalog and the sitemap [18].
Worth watching: the write-up gives the pre-filter count but no post-filter count, so the actual page inventory after the gate is unstated [19]. Block lists also rot faster than allow lists, since developers keep inventing new marketing prefixes, and the missing okrug reference means district pages stay governed by pattern-matching rather than by a known universe [13][14].