Engineering
LLM as a Judge Rubric: Scoring Answers Reproducibly
Abhishek Bahukhandi

Our interview reports contain a line that reads Overall Score: 7/10. For most of this year the dashboard found that number the only way it could: a regular expression over the report prose. Match it and you get a big number in a coloured card. Miss it — lowercase s, an em dash, a percentage instead of a fraction — and the card renders the string "Score not found".
The regex is not the bug. It is the symptom. The bug is asking a model for a score as a sentence. Which is why an LLM as a judge rubric — a fixed set of criteria, each with a named scale and a machine-readable slot to put the answer in — is the first thing you build when you want a score that means the same thing tomorrow as it did today.
This post is the design we are moving Taqari's coding round onto: what the rubric contains, how the output shape is enforced rather than requested, which parts of the score never go near a model at all, and how we check the thing is stable before a candidate ever sees a number.
Why "score out of 10" produces a different number every run
Ask a model to rate an interview answer out of ten and you have asked it to do four jobs in one token span: decide which qualities matter, decide how to weight them, decide what the scale means, and then compress all of that into a single digit. Nothing in the request pins any of those four. The model re-derives them on every call.
Three specific failures follow, and they are worth separating because they have different fixes.
Three things a single free-form score hides
The scale drifts. There is no definition of 7 anywhere in the system, so 7 means whatever the model's prior says it means given the surrounding text. Two transcripts of comparable quality can land two points apart, and neither score is wrong in any checkable sense, because there was never a standard to be wrong against.
The dimensions collapse. A candidate who writes a correct solution and cannot explain it, and a candidate who reasons beautifully toward a solution that fails on empty input, are different candidates. One number cannot tell you which one you are looking at, and the feedback built on top of that number cannot tell the candidate what to practise.
The result is not addressable. A score living inside a paragraph has to be recovered by pattern matching, which means your product has a failure mode that depends on the model's choice of punctuation. You also cannot store it, chart it over a candidate's history, or compare two sessions, because you do not have a field — you have a substring.
Where our own reports are today
Worth being concrete about the starting point, because it is the shape most teams ship first and it is visible in our own repository.
A finished session produces a report document whose main payload is a single reportContent string. The dashboard renders it with white-space: pre-wrap and, separately, scrapes it for the overall score. Reports are currently generated for the DSA round only; the dashboard says so in a banner, because we would rather admit the gap than show a confident number for a round we are not scoring properly yet.
One field in that same document is not prose, and the contrast is the whole argument of this post. The LeetCode-style round carries an integrityScore out of 100, derived from camera presence signals. It is a real numeric field. The UI can band it into green, amber and red at 80 and 50 without parsing anything, because nobody asked a language model to narrate it — it is computed from signals we already hold.
So the architecture already knows how to carry a score. The model-judged half of the report just never got the same treatment.
The rule we landed on: a model should never be asked to produce an overall score. It should be asked, one criterion at a time, which written level an answer matches and what in the transcript supports that. The overall number is arithmetic, and arithmetic belongs in application code.
Designing the LLM as a judge rubric
An LLM as a judge rubric is not a longer prompt. It is a data structure that happens to be rendered into a prompt, and the discipline is that every criterion in it has to survive four questions: what is being measured, what evidence counts, what the levels are, and what it is deliberately not measuring.
One criterion, one question, one scale
We keep the coding round to four criteria, each scored on the same four-point scale. Four is not a magic number; it is what we could write honest anchors for. A dimension you cannot describe at every level is a dimension you are not really scoring.
Correctness of approach
Does the chosen algorithm solve the stated problem, including the edge cases the problem implies? This is explicitly not "did the tests pass" — the tests are a separate, deterministic signal, and a candidate can reason to a correct approach and then fumble an index.
Complexity reasoning
Can the candidate state the time and space cost of their own solution and defend it? Level 0 is "no claim or a wrong claim left uncorrected". Level 3 is "states it correctly unprompted and identifies what would have to change to improve it".
Communication under questioning
When the interviewer pushes, does the explanation get clearer or vaguer? This is the criterion that most needs a model, because it is about the shape of a conversation and nothing in our pipeline computes it.
Response to a hint
Our interviewer is built to ask follow-ups rather than hand out solutions — we wrote about the rules that make an AI interviewer behave like an interviewer separately. That design gives scoring something genuine to look at: what the candidate did with a nudge is often more predictive than where they started.
Write the anchors, not the adjectives
The single highest-leverage part of a rubric is the text defining each level, and the most common mistake is writing adjectives instead of observations. "Good communication" is an adjective; the model has to invent what counts. "Restates the interviewer's objection in their own words before answering it" is an observation, and two different runs can agree on whether it happened.
So every level in our rubric is written as something you could point at in a transcript. Where we have real sessions that sit exactly on a boundary, the boundary text is derived from them. That is also why the rubric lives in version control as data rather than in a prompt string: when a level definition changes, we want the diff.
Forcing the output shape instead of requesting it
With the rubric defined, the remaining failure mode is the model returning the right judgement in the wrong container. "Reply with JSON" is a request. A schema compiled into the sampler is a constraint.
Both major providers now expose this. Anthropic's strict tool use sets strict: true on a tool definition and constrains token sampling to schema-valid output via grammar-constrained sampling, so the tool input matches your JSON Schema rather than usually matching it. Paired with tool_choice to require the call, the judge has exactly one legal way to answer.
Why minimum and maximum cannot carry your scale
Here is the detail that caught us, and it is documented: the supported JSON Schema subset for strict schemas excludes numeric constraints — minimum, maximum and multipleOf are not enforced. A field typed integer with "minimum": 0, "maximum": 3 gives you an integer and no guarantee it is in range.
The fix is to stop expressing the scale as a range and express it as a closed set. An integer enum of [0, 1, 2, 3] is inside the supported subset, so the constraint is actually enforced during sampling. The same trick appears in the provider's own examples, where a passenger count is an integer enum rather than a bounded number.
The schema we settled on
{
"name": "submit_rubric_scores",
"strict": true,
"input_schema": {
"type": "object",
"properties": {
"criteria": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string",
"enum": ["approach", "complexity", "communication", "hint_response"]
},
"level": { "type": "integer", "enum": [0, 1, 2, 3] },
"evidence": { "type": "string" },
"abstained": { "type": "boolean" }
},
"required": ["id", "level", "evidence", "abstained"],
"additionalProperties": false
}
}
},
"required": ["criteria"],
"additionalProperties": false
}
}
Three things in there are deliberate. Every object sets additionalProperties: false, which strict schemas require anyway. evidence is required, so a level always arrives with something to audit — and a reviewer who disagrees with a level can see which part of the transcript produced it. And abstained exists so "the transcript does not contain enough to judge this" is a first-class answer rather than a quiet 0, which is otherwise the most damaging thing a judge can do to a candidate who simply ran out of time.
There are limits worth knowing before you scale this up: strict schemas cap how many optional parameters and union-typed fields a request may carry, and compilation of very large schemas can time out. A rubric with four criteria is nowhere near those ceilings, but a rubric auto-generated per problem could be.
Separate what needs reading from what you can just compute
The cheapest way to make scoring more reproducible is to give the model less to score. Every fact your pipeline already owns should arrive as context, not as a question.
For the coding round that list is long. Which test cases passed and which failed comes back from the judge pipeline, where we submit a batch and assemble per-case callbacks in order. How long the candidate took to the first run, how many submissions they made, and how the attempts trended are all in the session record. Camera integrity is the integrityScore we already compute. None of that needs a language model, and a language model asked to restate it will occasionally get it wrong.
Aggregation belongs in code, not in the prompt
Once the judge returns four levels, turning them into one number is a weighted sum, and it should live where every other business rule lives. Weights in code are reviewable, diffable, and identical on every run. Weights described in a prompt are a suggestion the model re-interprets each time.
const WEIGHTS = { approach: 0.35, complexity: 0.2, communication: 0.25, hint_response: 0.2 };
const RUBRIC_VERSION = "coding-v2";
function overall(criteria) {
const scored = criteria.filter((c) => !c.abstained);
if (!scored.length) return null; // no score beats a wrong one
const mass = scored.reduce((s, c) => s + WEIGHTS[c.id], 0);
const earned = scored.reduce((s, c) => s + WEIGHTS[c.id] * (c.level / 3), 0);
return { value: Math.round((earned / mass) * 100), version: RUBRIC_VERSION };
}
Two details earn their keep. Abstentions are dropped from both numerator and denominator, so a criterion the transcript could not support does not silently punish the candidate. And every stored score carries the rubric version that produced it — because the day you change a weight, every historical score becomes a different measurement, and you need to know which scores are comparable.
Checking the rubric is stable before anyone sees a number
Reproducibility is a claim, so it needs a test. Ours is a replay harness: a frozen set of real sessions, scored several times at the current rubric version, with the spread recorded per criterion rather than only on the overall score. Per-criterion is the part people skip, and it is where the information is — an overall score that looks steady can hide two criteria swinging in opposite directions.
Lowering temperature is worth doing and is not the answer. It narrows the spread; it does not pin it, and it does nothing about the deeper problem, which is an under-specified scale. When a criterion is noisy across runs, the fix is almost always that its level definitions are still adjectives.
Self-consistency is also only half the bar. A rubric can be perfectly stable and perfectly wrong, so the other half is agreement with a human reviewer on a labelled set — the pattern OpenAI's evals framework calls adding a meta-eval for a model-graded eval, scoring the grader itself against human choices. We treat a rubric change like a code change: it does not ship until the replay numbers and the agreement numbers both hold.
What the judge is never given
Two exclusions are structural rather than tuning, and both exist for the same reason: anything in the context window is something the judgement can quietly become about.
The reference solution never enters the judge's context, exactly as it never enters the interviewer's. A judge that has seen the intended answer scores the distance to that answer, which penalises a different-but-valid approach. The judge sees the transcript, the submitted code and the computed signals, and reasons about them on the rubric's terms.
Candidate identity does not enter either. Name, email, college and anything else that could carry a bias has no bearing on whether a complexity claim was correct, so it is stripped before the call. The judge gets a session, not a person.
Where this leaves us
The honest status: the LLM as a judge rubric, the strict schema and the aggregator are what we are building scoring on now, and the prose-plus-regex report is what is still in production for the DSA round while we migrate. The reason we are writing this before the migration is finished is that the sequencing is the actual lesson. We built the report first and the rubric second, and that is backwards — the regex in our dashboard is what it costs to choose a presentation before you have decided what you are measuring.
If you are adding scoring to anything, the order that works is: write the criteria, write the anchors for every level, compile the shape into the request so it cannot come back wrong, compute everything computable outside the model, and put the arithmetic in code. The number you show a candidate is the last thing you design, not the first.
You can see the end of this pipeline from a candidate's side by running a free mock interview, and the pieces it depends on are written up in how we stream test results back to the browser as they land.
Frequently asked questions
What is an LLM as a judge rubric?
+
It is a fixed list of criteria, each with a named scale and written definitions for every level, that a model fills in one slot at a time. Instead of asking for an overall score, you ask for a level per criterion plus the evidence for it, then aggregate in code.
Why is asking an LLM for a score out of 10 unreliable?
+
Because nothing pins the scale. The model re-invents what 7 means on every call, blends unrelated qualities into one number, and emits it as prose you have to parse. Two identical transcripts can land two points apart with no way to tell which part of the answer moved.
Does temperature 0 make LLM scoring deterministic?
+
No. Lowering temperature narrows the spread but does not pin it, and it is the wrong lever anyway. Reproducibility comes from a scale with written anchors, a schema the output must satisfy, and aggregation done in application code rather than inside the prompt.
How do you force an LLM to return a valid score object?
+
Use strict tool use or structured outputs, which constrain token sampling to your JSON Schema. Numeric bounds like minimum and maximum are not supported, so express a four-point scale as an integer enum of 0 to 3 instead of a range the model may step outside.
Should the model score whether the code passed the tests?
+
Never. Test outcomes, runtime, integrity signals and submission counts are facts your pipeline already owns. Compute them, pass them to the judge as context, and reserve the model for the things that genuinely need reading: reasoning, communication and the quality of the approach.
How do you check a scoring rubric is stable before shipping it?
+
Replay a frozen set of real sessions through the judge several times and look at the spread per criterion, not just the overall score. Then compare against human-labelled sessions, the way a meta-eval does, so you know the rubric agrees with a reviewer and with itself.