Musubi puts its moderation policies inside a 1.7B-parameter model

Co-founder Filip Jankovic says the open-weights release builds on an idea he was exploring before decision models became a talking point; Musubi's own tests show both speed and limits.

By · Published

Primary source: TechCrunch

Why it matters

PolicyLM makes moderation rules adjustable without model retraining, but the release's own benchmarks are company-run and do not establish how well it performs on live platform traffic.

A generic computing board rests in an open rulebook beside two blank cards, one passed through a divider and one held back.

Musubi co-founder and chief AI officer Filip Jankovic is betting that platforms can apply their own moderation rules to every message without waiting for a large language model or retraining a fixed classifier. On October 6th, Musubi released PolicyLM-1.7B, an open-weights model that takes a written policy and a message, then returns category scores rather than a generated explanation. The model weights and inference code are available under the Apache-2.0 license.

Jankovic told TechCrunch that product teams need a better picture of what is happening on their platforms as content volumes grow. PolicyLM is designed to produce that picture quickly: Musubi reports median latency of 35 milliseconds for a short message on an NVIDIA L4 GPU and 22 milliseconds on an H100, with up to six policy categories scored in one pass. Those are Musubi measurements, not an independent evaluation.

A moderator's rules, changed at inference time

Jankovic's interest in this approach predates the current attention around decision models. In the TechCrunch interview, he traced it to a 2024 GLiNER project that used related techniques. His career also puts him close to the problem Musubi is trying to solve: before founding the company, he led data science at Evidation Health, building AI for risk scoring on behavioral data. Musubi CEO and co-founder Tom Quisel spent a decade building Trust & Safety systems and held CTO roles at Grindr and OkCupid. The founders say they started Musubi after encountering the operational cost of malicious users and the limits of tools available to platform teams.

The model's design addresses a familiar trade-off. Conventional moderation classifiers are fast, but their categories are typically set during training; changing a platform's definition of harassment, fraud or spam can mean new labeling and retraining. A general-purpose LLM can follow a written policy but may be slower and more expensive when asked to judge every message. PolicyLM reads the policy alongside the content and outputs a score for each category, so a team can edit the policy without retraining the model. A threshold then converts those scores into a moderation decision.

Diagram showing a written policy and message entering PolicyLM-1.7B, which produces category scores that pass through a threshold to a moderation decision; the policy can be edited without retraining.
PolicyLM scores a message against a written policy; a threshold converts the scores into a moderation decision. AI explanatory diagram, not documentary evidence. RuntimeWire, AI-generated diagram.

Teams can edit the policy, but the scores still need validation. A score is not a rationale, and the model card tells teams to calibrate the cutoff against labeled examples from their own platform. It also says the model has not been tested on live traffic, can over-flag benign material, and is weaker on code-mixed slang, some lower-resource languages and disguised text. Although Musubi says it evaluated messages across 19 languages, the policy rules in its evaluations were written in English. These limitations affect platforms whose moderation decisions carry consequences for users.

Benchmark results and limits

Musubi's model card reports 84.2% accuracy on its custom-policy benchmark, against 90.9% for the larger gpt-oss-safeguard-20B model. PolicyLM was faster in the same comparison: 22 milliseconds median per message on an H100, compared with 349 milliseconds for that larger model. The custom policies were synthetic, however, and Musubi says most evaluation sets were also used during development. In the public safety benchmark table, PolicyLM's scores vary by test, and the comparison uses each competing model's own taxonomy or whichever setup scored higher. The tables support a speed-and-flexibility case; they are not a third-party demonstration that PolicyLM will improve moderation outcomes in production.

A table compares company-reported custom-policy benchmark accuracy and median H100 latency for PolicyLM and gpt-oss-safeguard-20B, with a note that the policies were synthetic and most evaluation sets were used during development.
Musubi reports benchmark accuracy and H100 latency for the two models; its model card notes limits to the evaluation. AI explanatory infographic, not documentary evidence. RuntimeWire, AI-generated infographic.

A faster judgment can make it affordable to score a much larger share of messages, but platforms still have to decide what errors they can tolerate, test policies on their own data and decide when a person should review a borderline case. PolicyLM's scores can support those decisions; they do not settle them. The model card also lists image and conversation-history moderation as outside the model's intended use, making it a text-classification component rather than a complete Trust & Safety system.

Musubi is positioning the model as one part of a broader Trust & Safety toolkit, alongside moderation workflows and fraud detection. Musubi says its customers include platforms such as Bluesky, Stocktwits and Bumble, and that those customers collectively protect more than 850 million users. That reach is a Musubi-reported figure, not a measure of PolicyLM adoption or a breakdown of how many users the model will screen.

Musubi raised a $5 million seed round in February 2025, led by J2 Ventures with participation from Shakti Ventures, Mozilla Ventures and existing investor J Ventures, according to the financing announcement. That earlier round provides context for the team behind the release; it is not new financing tied to PolicyLM.

The open release gives platform teams a way to run the model on their own infrastructure and inspect its behavior rather than route every moderation decision through a hosted API. Jankovic's bet is that the practical unit of AI moderation is often a narrow decision, not a conversational answer. Whether that makes policy changes easier to ship will depend on the work that follows the download: testing the model against each platform's rules, measuring false flags and missed violations, and keeping human reviewers in the loop where the stakes warrant it.

Reader comments

Conversation for this story loads after sign-in.