The Unjournal · uj-prioritization · learning lab

How a prioritization tool learns from people

Five interactive simulations of the mechanisms behind preference learning, reward models, calibration, and active learning — the ideas from the companion report, runnable. Every simulation uses the same toy world: papers with a hidden true "value of evaluation," and noisy ways of measuring it.

AI-generated — about this page and the prompt behind it

This page (and its companion PDF report) was generated by an AI assistant, Claude (Fable 5, via Claude Code), on 14 July 2026, at David Reinstein's request, with light human direction and review. The simulations are simple self-contained models written for teaching, not analyses of real Unjournal data.

The prompt, in essence: "Consider how to build this prioritization tool well, leveraging human feedback and maybe some ground truth outcomes. Also consider the AI evaluation part of this project (lightly). Build resources to help me learn the key principles (AI, ML, RLHF involved) — both an interactive and a pdf/static tutorial and report. I want to upskill on this and do it well."

Before writing, the assistant read the uj-prioritization codebase (the scoring pipeline, calibration scripts, and feedback infrastructure) so the recommendations and examples are grounded in the actual system. Errors remain possible — treat quantitative claims as illustrative and check anything important against the cited literature.

Comments & annotations welcome. This page has Hypothes.is enabled — select any passage to highlight or annotate it, or open the sidebar (top-right) to read others' notes and reply. Public annotations are visible to anyone with the sidebar open; use the group selector for a private or group note. Feedback on the substance, the simulations, or the design plan is all useful.

Lab 1 · Elicitation

Pairwise votes → Bradley–Terry ranking

Twelve papers have hidden true values (gray). Each click below simulates team members answering "which of these two should we evaluate first?" — with human error. A Bradley–Terry model (a logit on the difference in latent values, i.e. McFadden with two alternatives) turns those noisy comparisons into estimated values (blue). Watch how few comparisons a decent ranking needs — and how targeted comparisons near the top-4 selection boundary beat random ones.

true value Bradley–Terry estimate ± SE top-4 selection bar
0
comparisons
–
Kendall τ (est vs true)
–
top-4 overlap

Why it matters: comparisons remove each rater's personal scale from the data — and a comparison between two papers adjacent at the selection bar carries far more information than "star vs dud." This is what a "which of these two?" dashboard widget would feed.

Lab 2 · Goodhart / winner's curse

Selecting on a noisy score over-states the winners

400 papers. The LLM's score is a proxy: correlated with true value, plus error. Select the top-k by proxy and compare the selected papers' average proxy score against their average true value. The gap is not a bug in any particular model — it is a theorem about selecting on noise, and it is why the top of your dashboard is over-scored in expectation, and why fresh human checks belong exactly there.

not selected selected (top-k by proxy) selected: mean proxy selected: mean truth
–
mean proxy of selected
–
mean truth of selected
–
optimism gap

Try: drop ρ to 0.4 and k to 3 — the gap explodes. Raise ρ to 0.95 and it nearly closes. Improving the scorer shrinks the curse; only fresh human judgment at the top measures it. Never audit the proxy with the data that trained it.

Lab 3 · Calibration

Fix the mapping, not the model

Simulated data in the shape of your real problem: the LLM's raw scores (x) track team ratings (y) but exaggerate at the top — like a scorer that runs hot on prominent papers. Instead of editing the prompt ("score 7 points lower…"), keep the instrument fixed and learn a monotone correction downstream. Compare a two-parameter linear (Platt-style) fit with isotonic regression (best monotone step function).

paper (raw score, team rating) fitted correction "worth evaluating" bar (65)
–
MAE vs team
–
Spearman ρ
–
boundary agreement

Notice: recalibration barely moves Spearman ρ (monotone maps can't reorder) but substantially improves MAE and boundary agreement — which is what your fixed thresholds (≥65 → Zotero, ≥75 → prioritize) actually consume. Also notice isotonic only beats linear once n is respectable; at n=30 it overfits the steps.

Lab 4 · Active learning

Where should five human labels a week go?

80 papers, noisy model scores, and a label budget. Each label replaces a noisy score with the team's true judgment. Three routing strategies: random, uncertainty (label where the scoring ensemble disagrees most), and boundary (label nearest the top-10 selection bar). The chart shows precision@10 — how much of the true top-10 the ranking gets right — as the budget is spent, averaged over 80 simulated worlds.

random uncertainty (ensemble disagreement) boundary-weighted

The point: the same five labels are worth roughly 2–3× more when routed to boundary and disagreement cases than when spent haphazardly — and a slice of the budget should always stay random, as the only unbiased audit of the whole system.

Lab 5 · The RLHF map

What RLHF is, and which piece you actually need

The pipeline behind instruction-tuned LLMs, step by step. Click each stage for what it means — and what it corresponds to in the Unjournal pipeline. The transferable idea is stage 3: a reward model, a small learned proxy for expensive human judgment. Stages 1 and 4 mostly don't transfer: your "policy" is sort and take top-k, which has no parameters to fine-tune.