Build1 publisher3 min readPublished
Rubrica separates harsh judges from weak projects by giving every judge the same anchors
Rubrica, a self-hosted hackathon portal, has every judge in a track review the same anchor projects before its k=3 shrinkage formula adjusts any score. That shared set lets the maths tell a harsh judge from one who drew a weak batch.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Rubrica gives each judge only projects in the tracks that judge covers, and never a project from the judge's own team.
- Projects outside the anchor set go to the least-loaded eligible judges until each project has the configured reviews_per_project count.
- The assignment engine is seeded with 42 and keeps existing judge-project pairs, so regenerating the plan produces the same assignments.
- Identity comes only from the session, so a judge requesting another judge's scores with ?judge= gets a 403 from the HTML pages, /api and /api/v1.
- Rubric weights must sum to exactly 1.0, and each stored score records the rubric version it was produced under, so later rubric edits leave old scores intact.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Organizers have to choose anchor_count and each judge's track coverage before scoring opens, because comparability comes from the assignment plan and cannot be added to private batches afterwards.
- cost Every eligible judge reviews every anchor in their tracks, so each judge added to a track adds anchor_count reviews to the event's total workload.
- capability Because Rubrica always shows the raw ranking beside the normalized one, an organizer can see which placings the correction changed before announcing winners.
The author's example is two judges averaging 3.3 and 3.7 [3]. If each saw a private batch, B might be generous or might have drawn better projects, and the averages cannot say which [3]. Anchors put some of the same projects in front of every judge in a track [5]. A gap on those shared projects is about the judges. "Normalization without overlap is just a guess with decimals," the author wrote [18].
Only after that does the formula run. Each judge's mean and variance are blended with the global values, weighted n_j to k, with k defaulting to 3 [12]. A z-score against the blended values is mapped back onto the global scale and clipped to 1 through 5 [12]. For jdg_01, who has one review, the global prior takes 75% of the weight [1]. For jdg_24, with 11 reviews, it takes about 21% [2], and the mean moves only from 3.318 to 3.372 [16].
The post says k behaves the opposite of what you would guess [19]. The one-review case shows why. A single score has zero variance, and the formula reduces to N = mu_g + sqrt(k/(k+1)) * (S - mu_g) [3]. At k=3 the factor is 0.866. jdg_01's 2.00 sits 1.568 below the global mean of 3.568 and lands 1.358 below it, at 2.210 [3]. That matches the post's worked figure [15]. Raise k and the factor climbs toward 1, so the score stays near its raw value. Push k toward zero and the score collapses onto the global mean [4]. Few hackathon tools publish a worked example you can check with a calculator, and fewer have one that checks out.
Dividing by n is deliberate. According to the author, sample variance is undefined at one review and unstable at two, and the k * v_g term already supplies prior mass [17].
For these numbers to transfer to another event, judges have to differ mostly in level and spread. The formula shifts and rescales each judge's scores [12]. It has no term for a judge who ranks projects in a different order from peers. The docs are candid about this. They call the method "a sample-size-shrunk location-scale normalization heuristic, inspired by empirical-Bayes reasoning" [14]. The author wrote that it resembles empirical Bayes but is a linear, fixed-k heuristic [14].
The access model is the part I would copy. Hiding peer scores by not rendering them "is a UI decision, and the API behind it still answers," the author wrote [21]. In Rubrica the same scope function, assert_judge_scope, checks the assignments table before any score write [10]. Errors map to status codes in exactly one place, from 401 for unauthorized to 429 for rate-limited [20]. The whole portal is Python 3.12, FastAPI and SQLite in one container, under the MIT licence [1].
What to watch
- The default anchor_count, which the post does not state, and how many reviews it adds per judge at an event the size of DOGFOOD 2026.
- Whether the other behaviours the author found in the formulas, beyond the one-review case, change how k should be set.
- Published DOGFOOD 2026 results showing how far the normalized ranking moved from the raw one.