FitRank
Typed AI decisions for staffing — the model recommends, a human decides
01 · Executive Summary
Staffing a task is a judgement call made dozens of times a week: who has the skills, the right seniority, the domain history, and the room in their calendar. Most "AI matching" answers it by asking a chatbot for a ranked list — free text you cannot audit, calibrate, or improve.
FitRank treats the same problem as a set of typed decisions. Anything code can compute is computed in code. The semantic questions — is this skill match strong, is the level right, is the domain relevant, is delivery risky, is this an overall fit — each have a fixed option list and come back as a probability for every option. Thresholds, not the model, decide what happens next, and a manager makes every final call.
Those questions are answered by a model I trained myself: a 149M-parameter ModernBERT student distilled from an open-weight teacher, served on CPU. Every accept or reject a manager makes flows back as training data, so the system learns what the people using it actually choose.
02 · The Stack
- Backend
- Python 3.12 + FastAPI — router → service → repository, a Postgres job queue and a background worker for match runs.
- Data
- Postgres + pgvector, SQLAlchemy models and Alembic migrations. Synthetic people and tasks only — no real employee data.
- Decision model
- ModernBERT-base + five linear heads, trained on Kaggle GPUs with a KL loss against the teacher's full distributions.
- Teacher
- Qwen3.5-9B in 4-bit, scored through the same prompts, option shuffles and logprob normalisation as the LLM engine.
- Frontend
- React + TypeScript + Vite, API types generated from the backend's OpenAPI, role-aware screens and an admin area.
- Delivery
- JWT + refresh-cookie auth with roles, rate limits, a budget guard, Docker images, GitHub Actions CI, 93.5% backend test coverage.
03 · System Architecture Flow
- 1Task Intake
A manager describes the task in plain English; an LLM only interprets it into fixed fields (skills, level, dates). It never ranks people.
- 2Hard Rules
Code filters on availability, leave, location, timezone, cost band and clearance. Every exclusion stores a reason code.
- 3Retrieval
pgvector similarity blended with must-have skill coverage keeps the top candidates. Nobody with zero must-haves is ever retrieved.
- 4Typed Decisions
The student model answers five bounded questions per candidate in one batch — a probability for every option, never free text.
- 5Policy Bands
A combiner over code facts + model answers gives P(manager accepts). Thresholds band it: shortlist, review or hidden. Flags cap at review.
- 6Human Call
The manager accepts or rejects with a reason. Nothing is auto-assigned; every decision becomes training data for the next model.
04 · Deep Technical Breakdown
A decision is a letter, scored by logprobs
The first engine asked an LLM each question with the options mapped to single letters and a one-token answer limit. The answer is not the letter it writes — it is the probability mass on every allowed letter, read from top_logprobs and renormalised. If too much mass lands outside the allowed letters, the result is flagged and can never reach the shortlist. Options are shuffled per candidate and each question is asked under two orderings to cancel position bias.
# One decision → a distribution over a fixed option list.
mass = {letter: 0.0 for letter in letter_to_option}
for t in first_token.top_logprobs:
letter = t.token.strip().upper()
if letter in mass:
mass[letter] += math.exp(t.logprob)
unmapped = 1 - sum(mass.values())
total = sum(mass.values())
probs = {letter_to_option[k]: v / total for k, v in mass.items()}
flags = {"low_confidence_format"} if unmapped > 0.2 else set()Distilling the judgement into a model I own
Paying an API per candidate does not scale, and a hosted model can change under you. So an open-weight teacher labelled 3,010 task–person pairs on Kaggle using the exact same decision definitions, and a ModernBERT student learned from them. The loss is KL divergence against the teacher's whole distribution, so the student learns how confident to be, not just the top answer. Pairs are split by task and generated from a different random world than the evaluation set, so nothing leaks.
# Five heads on one encoder; learn the teacher's confidence.
h = encoder(pair_text).last_hidden_state[:, 0]
loss = sum(
kl_div(log_softmax(head(h)), teacher_probs[key], reduction="batchmean")
for key, head in heads.items() # skill, level, domain, risk, overall
)Code checks the model, not the other way round
A schema-valid answer can still be wrong. Deterministic cross-checks compare each decision with the facts — a strong skill-match score for someone with zero must-have skills, or "right level" two grades off, is flagged as a contradiction and capped at review. Thin or inconsistent profiles, and a student that is unsure (40–60% on overall fit), are capped the same way. Explanations are built from numbered facts and are rejected if they cite a fact that does not exist.
Learning from the people who decide
The final score is a small logistic combiner over eight code facts and the model's five answers, trained on accept/reject feedback — so the ranking follows what managers actually choose. Each candidate stores the exact text the model scored, rejections carry a reason (skill gap, level, domain), and that becomes partial training targets for the next student. One task in five is always held out to keep the measurement honest.
05 · Results, Honestly
Every evaluation runs the model side by side with a code-only baseline. If the model cannot beat simple rules, it has not earned its place. Latest run: 35 held-out tasks on synthetic data with a hidden ground truth.
| Metric | FitRank | Code baseline |
|---|---|---|
| hit@5 — best person in the top five | 94% | 91% |
| Pairwise ordering | 0.88 | 0.86 |
| NDCG@10 | 0.89 | 0.89 |
| hit@1 — best person ranked first | 54% | 74% |
FitRank now puts the right person in the top five more often and orders pairs better than the baseline, but it still picks the single best person first less often. That is the open problem, and the reason the system shortlists rather than assigns. The real test comes next: managers' own labels, which will replace the synthetic ground truth for evaluation, calibration and retraining.