Build2 publishersIndependently confirmed3 min readPublished
Musubi's open PolicyLM-1.7B reads a platform's moderation rules at inference time
Musubi released PolicyLM-1.7B under Apache-2.0 to score messages against a written policy in a reported 22 milliseconds on an H100. Teams can rewrite rules without retraining, though every speed and accuracy figure so far comes from Musubi's own tests.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The model takes a written policy and a message and returns a score for each category in the policy, with no generated explanation.
- The custom policies in that benchmark were synthetic, and Musubi says most of its evaluation sets were also used during development.
- According to the model card, PolicyLM has not yet been run against live traffic and can flag harmless material too often. It also does worse on disguised text, code-mixed slang and some lower-resource languages.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Editing a rule skips the retrain, but each team still has to build a labeled set from its own traffic before it can trust a score cutoff.
- cost On Musubi's own numbers, the 16-fold latency cut comes with about 1.7 times the errors, paid for in wrongly flagged posts or in human review of borderline cases.
- capability Apache-2.0 weights and inference code let a platform rerun the vendor's comparison on its own labeled traffic before it commits.
- constraint With images and conversation history out of scope, platforms still need separate tooling for media and for abuse that shows up only across a thread.
Musubi designed PolicyLM to sit between two kinds of moderation tool [7][8]. A conventional classifier is fast, but its categories are set during training. Redefining harassment, fraud or spam can mean new labeling and a retrain [7]. A general-purpose LLM can follow a written policy, but it is slower and costs more when it has to judge every message [8]. In my view PolicyLM takes the right approach for a filter that sees every post: it puts the policy in the input, so a team can edit a rule without retraining the model [2]. Up to six categories are scored in one pass, and Musubi reports a 35-millisecond median for a short message on an NVIDIA L4 [4].
A threshold turns each score into a decision [2]. The model card tells teams to calibrate that cutoff against labeled examples from their own platform [14]. The scores are computed against the policy text, so I'd expect a substantive rewrite of a rule to need its cutoff checked again. A score is also not a rationale [14], and a harassment score is hard to paste into an appeal reply.
Musubi's main comparison is against gpt-oss-safeguard-20B. The larger model's median on an H100 was 349 milliseconds [15], about 16 times PolicyLM's [21]. The same card's accuracy figures put PolicyLM's error rate at 15.8% against 9.1%. On Musubi's own test, it was wrong about 1.7 times as often [22].
For the custom-policy score to carry over, two things would have to be true. A platform's rules would have to resemble Musubi's synthetic ones, and its traffic would have to resemble sets the model met during development [16]. The public safety table has a separate caveat. Each competing model was run on its own taxonomy or on whichever setup scored higher, and PolicyLM's scores vary by test [17]. The latency figures are medians on short messages [4]. A moderation pipeline is sized for its slow tail and its long posts, and a platform would have to measure those itself.
The 19-language figure covers messages only. Every policy in Musubi's evaluations was written in English [19]. A platform that writes its rules in Spanish would be running outside the measured conditions. Musubi pitches the model as one component of a Trust & Safety toolkit, alongside moderation workflows and fraud detection [13].
The founders say they started Musubi after running into the cost of malicious users and the limits of the tools platform teams had [9]. CEO Tom Quisel spent a decade building Trust & Safety systems and was CTO at Grindr and OkCupid [5]. Filip Jankovic, the chief AI officer, led data science at Evidation Health and traced the approach to a 2024 GLiNER project [6]. He told TechCrunch that as the amount of content rises, product teams need a clearer view of what goes on across their platforms [12]. Musubi says its customers include Bluesky, Stocktwits and Bumble [11].
What to watch
- An independent evaluation of PolicyLM-1.7B on live platform traffic, with tail latency and false-positive rates on benign posts.
- Results from Bluesky, Stocktwits or Bumble running PolicyLM against their own labeled traffic.
- Benchmarks that use policies written in languages other than English, or that target code-mixed and disguised text.