Engineering

AI Interviewer System Prompt: Rules That Actually Stick

Abhishek Bahukhandi

Abhishek Bahukhandi

•10 min read
Diagram of Taqari's interviewer agent: the backend minting an ephemeral token carrying the system prompt and tool declarations, the browser opening a live session, and the tool-call loop that relays editor state and test results back to the model
Diagram of Taqari's interviewer agent: the backend minting an ephemeral token carrying the system prompt and tool declarations, the browser opening a live session, and the tool-call loop that relays editor state and test results back to the model
An AI interviewer system prompt alone won't stop a model lecturing or leaking the answer. Here are the config limits and tools that make the rules stick.

Our first interviewer was a very good chatbot. It greeted the candidate, asked a sensible question, and then — about ninety seconds in — explained the optimal approach, unprompted, in a warm and encouraging tone. The prompt said, in capital letters, not to do that.

The lesson took us a while to accept: an AI interviewer system prompt is necessary and nowhere near sufficient. A system prompt is a strong prior, not a constraint. Everything an interviewer must reliably not do has to be enforced by something that isn't prose — a token ceiling, a turn-detection setting, a tool surface, or simply never putting the answer in the context window.

This post is what we actually run in production for Taqari's coding round: what stayed in the prompt, what moved into session config, and where the model's autonomy is deliberately fenced in.

What an AI interviewer system prompt can and cannot enforce

It helps to sort the rules by whether the model can quietly disagree with them.

Prose is good at things the model has no competing instinct about. A role ("you are a senior engineer running a 30-minute coding round"), a rubric, a decision list for each turn, a tone. Give an LLM a short, closed set of actions to pick from after every candidate answer — follow up, ask for a clarification, move to the next problem, end the interview — and it picks sensibly most of the time. That is a genuinely good use of a system prompt.

Prose is bad at fighting the model's own training. "Be concise" loses to a model tuned to be thorough. "Never give hints" loses to a model tuned to be helpful, especially when a candidate sounds stuck and slightly upset. These instructions don't fail loudly; they fail on turn nine, once the conversation has drifted far enough from the system message.

The rule we use: if a candidate could get an unfair advantage when the model ignores an instruction, that instruction does not belong in the prompt. It belongs in config, in a tool boundary, or in what we choose not to send the model at all.

Where the prompt comes from, and why the browser never owns it

Both of our interview rounds follow the same shape. The browser asks our backend to start a session. The backend returns an ephemeral credential for the model provider, plus the system prompt for this round, plus the problem, plus the tool declarations the interviewer is allowed to use. The client opens a live session directly to the provider with that credential and immediately sends the configuration.

So the client passes the prompt along; it never authors it. That matters for two reasons.

First, prompt changes ship without a frontend deploy. Rubric wording is the part of this system we touch most often, and it would be miserable if every adjustment needed a rebuild.

Second, and more important, the tool declarations travel inside the ephemeral credential as well as in the setup message. The provider enforces them against the credential, which means the browser cannot widen what the interviewer is able to do by editing a config object in dev tools. The set of actions an interviewer can take is decided server-side and reviewed in one place.

Taqari's interviewer agent loop: the backend mints an ephemeral token carrying the system prompt and tool declarations, the browser opens a live session with the model provider, and tool calls are executed in the browser and answered, with editor state and test results relayed back as platform messages Taqari backend system prompt + rubric tool declarations Browser session setup: instruction, token cap, VAD, tools Realtime model audio out + tool calls no reference solution token setup Tool handlers write_code, run_code, … Judge pipeline per-case results toolCall toolResponse submit [TEST_RESULTS] / [EDITOR_STATE] The model never receives the reference solution; it learns the outcome only from platform messages.

Three config limits that do the work prose cannot

These are the settings we reach for before we reach for another paragraph of instructions.

Cap the monologue with a token ceiling

The single highest-leverage change we made was setting a hard output cap on the session. Our behavioural round runs with a ceiling of 500 output tokens per turn; the coding round, which sometimes has to read a problem statement aloud, gets 1024.

A cap is blunt and that is exactly the point. When the interviewer starts to drift into teaching, the turn ends. Candidates read a truncated sentence as the interviewer pausing, and the model's next turn is conditioned on a conversation where its own turns are short — so brevity compounds instead of decaying. "Keep your questions short" never produced that.

Tune turn detection so thinking is not treated as an interruption

A coding interview is mostly silence. Someone staring at an array problem goes quiet for eight seconds, and a default voice agent fills that silence with encouragement, which is the single fastest way to make a mock interview feel fake.

Both providers expose voice activity detection as config rather than prompt, and the knobs are worth understanding individually. OpenAI's Agents SDK documents semantic_vad as a mode that "aims for more natural turn boundaries and can wait a little longer when the user sounds like they are not finished yet" — which is precisely the behaviour a coding round needs, so we run it with low eagerness. On the Gemini side we set the fields explicitly.

silence_duration_ms

Google's Generative Language API protos describe this as "the required duration of detected non-speech (e.g. silence) before end-of-speech is committed." We set it to 2000ms. Two seconds of quiet feels long in a chat UI and short in an interview; below about 1.5s the agent starts finishing candidates' sentences for them.

prefix_padding_ms

"The required duration of detected speech before start-of-speech is committed," per the same protos. We keep this low, at 150ms, because the cost of a false start is small and the cost of missing the first word of an answer is a transcript that reads as if the candidate began mid-sentence.

end_of_speech_sensitivity

The GenAI JS SDK types define END_SENSITIVITY_HIGH as "automatic detection ends speech more often." Pairing high sensitivity with a long silence window sounds contradictory, and isn't: the detector is quick to decide a phrase ended, but the session still waits out the full silence window before handing the turn over. We get responsive transcription without an impatient interviewer.

Let barge-in be the model's problem, not the prompt's

We set activityHandling to START_OF_ACTIVITY_INTERRUPTS, documented in the SDK types as "if true, start of activity will interrupt the model's response (also called 'barge in')", and turnCoverage to TURN_INCLUDES_ONLY_ACTIVITY — "the users turn only includes activity since the last turn, excluding inactivity." Candidates cut the interviewer off constantly, and that should be handled by the transport and the session, not by an instruction asking the model to be interruptible. We wrote up the client-side half of this in how we handle barge-in in production, and the transport trade-offs behind it in WebRTC vs WebSocket for voice agents.

Difficulty ramping by withholding, not by instruction

The obvious way to build a two-problem round is to put both problems in the system prompt and tell the model to move on after fifteen minutes. We tried it. The model front-loads: it mentions the second problem early, or compares the two, or — worst case — picks the easier one because the candidate sounded nervous.

What works is giving it one problem at a time. The second problem is injected as a message partway through the session. The model cannot ramp ahead of the clock because the material does not exist in its context yet.

The wrinkle is timing. Injecting a message while the interviewer is mid-sentence produces a jarring topic change, so we arm the injection on a timer and send it on the next turn boundary:

// arm at the threshold, send on the next completed model turn
const questionTimer = setTimeout(() => {
  q2ReadyRef.current = true;
  // fallback: if the model isn't speaking, don't wait forever
  setTimeout(() => {
    if (!q2SentRef.current) sendQuestion2();
  }, 3000);
}, 15 * 60 * 1000);

// ...in the event handler
case 'response.done':
  if (q2ReadyRef.current && !q2SentRef.current) {
    q2SentRef.current = true;
    sendQuestion2();
  }
  break;

Two details earn their keep. The sent flag is separate from the ready flag, because the timer fallback and the turn-boundary handler race and exactly one must win. And the fallback exists at all because a silent model produces no turn-completion event — without it, a candidate who goes quiet at the fifteen-minute mark never receives the second problem.

Tools are the real guardrail against handing out the answer

Here is the part that changed our thinking most. In the coding round the interviewer is agentic: it can write to the editor, edit specific lines, switch language, run the code, open the chat panel, and end the interview. Each of those is a declared tool with a handler in the browser.

Giving the model more power made it leak less. When the only way to affect the editor is to narrate code aloud, a model under pressure narrates the solution. When writing to the editor is a tool call that appears on screen, "help me with the loop" becomes a visible, reviewable, logged action — and the model's spoken turn stays conversational because the code went somewhere else.

Ground the model in what the candidate actually wrote

A model that guesses at the editor's contents will confidently discuss code that isn't there. So after any tool call that changed the editor, we send one message back with the real buffer:

export function formatEditorStateForGemini(code) {
  return `[EDITOR_STATE] The code editor now contains:\n\`\`\`\n${code}\n\`\`\``;
}

One sync per batch of tool calls, and only when something actually changed. Re-sending the buffer on every keystroke would bloat the context and cost real money for no added accuracy.

Test results arrive the same way. The judge is asynchronous — we wrote about the batch-submission and callback design in streaming per-test-case results — so when a run reaches a terminal status we format one compact message: pass count, then up to five failing cases with expected-versus-actual, each truncated to a few hundred characters. Capping the failure detail is not just about tokens. Hand a model forty failing cases and it starts debugging out loud, which is the interviewer doing the candidate's job.

Put the instruction where the model is actually looking

Our favourite line in this whole system is the string a tool returns. The run_code handler replies:

return {
  result: `Code submitted. ${outcome.totalTestCases} test cases are running. ` +
          `Wait for the [TEST_RESULTS] message before judging the solution.`,
};

That second sentence fixed a bug that a paragraph in the system prompt had not. The model used to call run_code and immediately congratulate the candidate, because from its point of view the action had succeeded. The instruction not to do that was thousands of tokens away in the system message; the tool result is the most recent thing in the context at precisely the moment the mistake would happen. Proximity beat emphasis.

The same logic shapes the error paths. Every call in a batch gets a response — including unknown tool names and handlers that threw — because a model waiting on a function response that never arrives doesn't recover, it just stops talking. Google's SDK types are explicit that a FunctionResponse carries "the id of the function call this response is for", so matching every id is not optional bookkeeping; it is what keeps the session alive.

Keep the answer out of the context entirely

The strongest guardrail is the simplest. The interviewer is never given the reference solution. It receives the problem statement, the constraints, the candidate's code, and the test outcomes — the same things a human interviewer who hadn't solved the problem beforehand would have.

A model cannot leak what it was never told. Every other rule on this page is a mitigation; this one is a structural fix, and it is the first thing we would check in anyone else's interviewer.

What we would tell you to copy

If you are building something similar, the ordering matters more than any individual setting:

  1. Withhold before you instruct. No reference solution, one problem at a time. Rules you don't need are rules that can't be ignored.
  2. Cap the output. A token ceiling is the cheapest fix for a model that lectures, and it compounds across turns.
  3. Move pacing into VAD config. Silence windows and sensitivity are interview design, not tuning trivia.
  4. Make actions tools, not narration. Anything the model does to the candidate's workspace should be visible and logged.
  5. Put the just-in-time rule in the tool result. The most recent token in the context wins over the most emphatic one.
  6. Answer every tool call. One unanswered id is a silent interviewer.

What is left in our system prompt after all of that is short: who the interviewer is, the rubric, the four actions it may take after an answer, and a refusal rule about solutions that now functions as a backstop rather than as the only line of defence.

You can hear the result in a free mock interview on Taqari — the interviewer that asks a second follow-up when your first answer was vague, and does not tell you the answer when you go quiet.

Frequently asked questions

What should an AI interviewer system prompt contain?

+

A role, a rubric, an explicit refusal rule about giving away solutions, and a turn-level decision list — follow up, clarify, move on, or end. Keep behaviour you can enforce elsewhere out of it: length limits, pacing and grounding belong in session config and tools, not prose.

Why does an AI interviewer keep giving away the answer?

+

Because the base model is tuned to be helpful, and a sentence telling it not to help competes with that tuning on every turn. The reliable fix is structural: never put the reference solution in the context, and make the model reach for a tool to act instead of narrating code.

How do you stop an AI interviewer from monologuing?

+

Cap output tokens at the session level. We set a hard ceiling on the interviewer's spoken turns, so a drifting answer is truncated by the API rather than by a politely ignored instruction. Prose like 'be concise' is a preference; a token ceiling is a limit.

How do you ramp difficulty during an AI interview?

+

Inject the next problem as a message partway through the session rather than listing every problem up front. The model cannot skip ahead to material it has not been given, so the ramp is controlled by the platform's clock instead of the model's judgement.

Should the model or the platform decide when to run the candidate's code?

+

The model decides, the platform executes. Declaring a run tool lets the interviewer choose the moment, while the actual compile and test run happens in our judge pipeline. The model then waits for a results message instead of guessing whether the code passed.

Where should tool declarations live for a voice interview agent?

+

Server-side, baked into the ephemeral session credential. The browser then cannot widen the model's action surface by editing a config object, and the set of things an interviewer can do stays something you can review and change in one place.

Sources

Did you find this helpful?

Share this guide with your circle.

#ai interviewer system prompt#ai interviewer prompt design#llm interviewer agent#voice agent system prompt#difficulty ramping llm#ai mock interview architecture#agentic tool calling voice agent