Decision Model · Applied ML

FitRank

Typed AI decisions for staffing — the model recommends, a human decides

PythonFastAPIModernBERTPyTorchPostgres + pgvectorReact + TypeScript
rules · codeavailablelocationcost bandmust-havesModernBERT · 5 headsskill_match82%level_fit74%domain61%risk_low68%overall_fit79%policyshortlist≥0.80review≥0.50hidden<0.50manager decides — nothing is auto-assignedacceptrejectfeedback → retrain ↺

01 · Executive Summary

Staffing a task is a judgement call made dozens of times a week: who has the skills, the right seniority, the domain history, and the room in their calendar. Most "AI matching" answers it by asking a chatbot for a ranked list — free text you cannot audit, calibrate, or improve.

FitRank treats the same problem as a set of typed decisions. Anything code can compute is computed in code. The semantic questions — is this skill match strong, is the level right, is the domain relevant, is delivery risky, is this an overall fit — each have a fixed option list and come back as a probability for every option. Thresholds, not the model, decide what happens next, and a manager makes every final call.

Those questions are answered by a model I trained myself: a 149M-parameter ModernBERT student distilled from an open-weight teacher, served on CPU. Every accept or reject a manager makes flows back as training data, so the system learns what the people using it actually choose.

02 · The Stack

Backend
Python 3.12 + FastAPI — router → service → repository, a Postgres job queue and a background worker for match runs.
Data
Postgres + pgvector, SQLAlchemy models and Alembic migrations. Synthetic people and tasks only — no real employee data.
Decision model
ModernBERT-base + five linear heads, trained on Kaggle GPUs with a KL loss against the teacher's full distributions.
Teacher
Qwen3.5-9B in 4-bit, scored through the same prompts, option shuffles and logprob normalisation as the LLM engine.
Frontend
React + TypeScript + Vite, API types generated from the backend's OpenAPI, role-aware screens and an admin area.
Delivery
JWT + refresh-cookie auth with roles, rate limits, a budget guard, Docker images, GitHub Actions CI, 93.5% backend test coverage.

03 · System Architecture Flow

  1. 1Task Intake

    A manager describes the task in plain English; an LLM only interprets it into fixed fields (skills, level, dates). It never ranks people.

  2. 2Hard Rules

    Code filters on availability, leave, location, timezone, cost band and clearance. Every exclusion stores a reason code.

  3. 3Retrieval

    pgvector similarity blended with must-have skill coverage keeps the top candidates. Nobody with zero must-haves is ever retrieved.

  4. 4Typed Decisions

    The student model answers five bounded questions per candidate in one batch — a probability for every option, never free text.

  5. 5Policy Bands

    A combiner over code facts + model answers gives P(manager accepts). Thresholds band it: shortlist, review or hidden. Flags cap at review.

  6. 6Human Call

    The manager accepts or rejects with a reason. Nothing is auto-assigned; every decision becomes training data for the next model.

04 · Deep Technical Breakdown

A decision is a letter, scored by logprobs

The first engine asked an LLM each question with the options mapped to single letters and a one-token answer limit. The answer is not the letter it writes — it is the probability mass on every allowed letter, read from top_logprobs and renormalised. If too much mass lands outside the allowed letters, the result is flagged and can never reach the shortlist. Options are shuffled per candidate and each question is asked under two orderings to cancel position bias.

# One decision → a distribution over a fixed option list.
mass = {letter: 0.0 for letter in letter_to_option}
for t in first_token.top_logprobs:
    letter = t.token.strip().upper()
    if letter in mass:
        mass[letter] += math.exp(t.logprob)

unmapped = 1 - sum(mass.values())
total = sum(mass.values())
probs = {letter_to_option[k]: v / total for k, v in mass.items()}
flags = {"low_confidence_format"} if unmapped > 0.2 else set()

Distilling the judgement into a model I own

Paying an API per candidate does not scale, and a hosted model can change under you. So an open-weight teacher labelled 3,010 task–person pairs on Kaggle using the exact same decision definitions, and a ModernBERT student learned from them. The loss is KL divergence against the teacher's whole distribution, so the student learns how confident to be, not just the top answer. Pairs are split by task and generated from a different random world than the evaluation set, so nothing leaks.

# Five heads on one encoder; learn the teacher's confidence.
h = encoder(pair_text).last_hidden_state[:, 0]
loss = sum(
    kl_div(log_softmax(head(h)), teacher_probs[key], reduction="batchmean")
    for key, head in heads.items()   # skill, level, domain, risk, overall
)

Code checks the model, not the other way round

A schema-valid answer can still be wrong. Deterministic cross-checks compare each decision with the facts — a strong skill-match score for someone with zero must-have skills, or "right level" two grades off, is flagged as a contradiction and capped at review. Thin or inconsistent profiles, and a student that is unsure (40–60% on overall fit), are capped the same way. Explanations are built from numbered facts and are rejected if they cite a fact that does not exist.

Learning from the people who decide

The final score is a small logistic combiner over eight code facts and the model's five answers, trained on accept/reject feedback — so the ranking follows what managers actually choose. Each candidate stores the exact text the model scored, rejections carry a reason (skill gap, level, domain), and that becomes partial training targets for the next student. One task in five is always held out to keep the measurement honest.

05 · Results, Honestly

Every evaluation runs the model side by side with a code-only baseline. If the model cannot beat simple rules, it has not earned its place. Latest run: 35 held-out tasks on synthetic data with a hidden ground truth.

MetricFitRankCode baseline
hit@5 — best person in the top five94%91%
Pairwise ordering0.880.86
NDCG@100.890.89
hit@1 — best person ranked first54%74%

FitRank now puts the right person in the top five more often and orders pairs better than the baseline, but it still picks the single best person first less often. That is the open problem, and the reason the system shortlists rather than assigns. The real test comes next: managers' own labels, which will replace the synthetic ground truth for evaluation, calibration and retraining.

Want a system like this built?

I'm Yaseen Khatib — a Senior Full-Stack AI Engineer (MERN + TypeScript) who ships production AI systems solo: RAG pipelines, agent orchestration, real-time and edge architectures. Open to senior / lead roles and contract work, remote or on-site.