Build1 publisher3 min readPublished
The OpenAI bill it displaced was about EUR 10 a month. Paying that back takes a second set of calibration values, and the only comparison actually benchmarked in the writeup is embeddings.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Threshold constants are where an embedding model's geometry leaks into application code. A cutoff that reads as "same story" in one vector space reads as "same broad subject" in another, because the two spaces were trained to different objectives. RSSMonster stacks three decisions on top of those numbers: events group articles covering one real-world story, topics link related events into broader subjects, and interest islands are longer-lived clusters inferred from what a user tends to read [10]. Each layer consumes the layer beneath it. A looser event boundary does not stay at the event layer, which is why a change that looks helpful at the bottom of the stack shows up as noise at the top [11].
Both models passed the existing semantic regression suite while the pipeline behaved differently [7], and that combination tells you what the suite actually tests. It checks that embeddings are well formed and stable enough to pass fixed assertions. Nothing in it asserts that the count of event and topic associations lands in the same place. The diagnostic fixture is 23 hand-labelled article-like inputs against 23 taxonomy candidates [8], so at most 529 input-candidate pairings [18], with close neighbours such as Football, Tech and F1 included deliberately. That is a good confusion harness. Cluster-count drift is a different measurement, and a fixture that size will not surface it. Codex was handed the two generated log files and asked to identify the differences and explain what they meant for the application [14].
The money is worth stating in plain arithmetic. Around EUR 10 a month [3] is about EUR 120 a year [17]. That is the entire budget for the migration, and the migration does not end with one commit: similarity thresholds and fallback behaviour are now model-specific values rather than shared defaults [13], which is a second set of constants to keep calibrated whenever either model moves. At any plausible hourly rate, the first afternoon of tuning has already outspent the year. The reason to do it anyway is the one the author gives, which is whether commercial inference should be the default for a self-hosted reader at all [3].
On the question of where small models are good enough, this experiment measures one thing. The benchmarked comparison is the embedding process [16]. ModernBERT's jobs, information quality, promotional content and emotional tone, arrive with a role description and no reported accuracy against the hosted baseline they replaced [6][16]. The author reports that Qwen performed better than expected [19], and also that no model clearly won the embedding benchmark [9], which is the most useful kind of result and the least quotable.
What transfers here is the location of the work: a per-provider threshold table, and a downstream check that counts events and topics before and after the swap. Any pipeline that reads similarity numbers as fixed application constants is running the same experiment whether or not it is measuring it.
Ranked by verification strength, evidence, and original report placement.
RSSMonster is a self-hosted, Google Reader-inspired RSS application its author built to fetch articles, organise them into feeds and folders, mark them as read and track items to return to later.
The author initially used OpenAI for RSSMonster's semantic features, including embeddings and other AI-related processing.
As the amount of content processed increased, the author's OpenAI API costs rose to around EUR 10 per month, which prompted the question of whether commercial inference needed to be the default for a self-hosted RSS reader.
The author added two local small language models to RSSMonster: Qwen and ModernBERT.
Qwen is a family of compact language and embedding models developed by Alibaba, with smaller variants practical to run locally on CPU; in RSSMonster it is used for embeddings, summaries, tags and semantic labels.
ModernBERT is a modernised BERT-style encoder well suited to classification tasks, used in RSSMonster for information quality, promotional content and emotional tone.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
First-hand, unpublished workings
Everything rests on one developer's account of his own application. The regression suite is his, the 23-item fixture is his, and the log comparison behind the verdict was handed to Codex rather than shown. No threshold values, hardware, timings or classification scores appear anywhere, so a reader can accept the shape of the finding but cannot reproduce or size it.
One install, the author's
Deployment is exactly as broad as the piece claims and no broader: Qwen and ModernBERT are in production in a single self-hosted RSS reader run by the person who wrote both the software and the post. No other install, no user or download count, and no third party repeating the swap appears in this reporting.
Title reaches slightly past the measurements
The headline offers lessons from testing Qwen and ModernBERT, while only the embedding path was actually put against OpenAI; the classification work is described by role. Pulling the other way, the body declines to declare a winner and says input preparation sometimes mattered as much as the model, which is the opposite of overselling. The net overstatement is small and lives in the framing rather than the findings.
Reputational, not commercial
The author is writing about his own project and points readers to his own Medium piece on the same work, so there is a pull toward a tidy local-models conclusion. Alibaba, OpenAI and Anthropic all appear as tools rather than sponsors, and the sum at stake, EUR 10 a month, is too small for anyone to be buying an opinion about it.
Believable on the author's word, unverified beyond it
The account is first-hand, internally consistent and modest enough that it does not strain credulity. What keeps confidence middling is the absence of any second account and of the numbers behind the qualitative verdicts. The per-model threshold lesson is the part most likely to hold up if anyone else tries the same swap; the characterisation of Qwen as the broader grouper is the part resting purely on impression.
product
Chinese banks and telcos are retailing AI tokens in a unit their customers cannot price1 publisher
security
Seven AI coding agents run attacker code named in a repository's own .git config2 publishers
build
Twenty-three security checks, zero coverage: AI coding agents as build-pipeline attack surface1 publisher
product
Baidu's AI line grew 25 percent and still lost the arithmetic1 publisher
Publishers with included, body-backed reporting in this cluster.
1 article · September 7, 2026