Engineering
Code Template Wrappers: How Online Judges Run Your Code
Abhishek Bahukhandi

A candidate opens a problem in a Taqari interview and sees about six lines: a class, a function signature, and a comment telling them to write their solution there. They never write a main, never read a line from stdin, never print anything. A code template wrapper does that part — a prefix and a suffix the harness owns, which turn those six lines into a complete program before it reaches the execution sandbox.
This is the layer between the editor and the runner. The submission plumbing underneath it is covered in batch submission with per-test-case callbacks, and the isolation model beneath that in running untrusted code safely. Here we care about one question: how do you let someone write a function body in four languages and still run it against the same test data?
What a code template wrapper is
Three pieces of text get concatenated before anything compiles.
- Prefix — imports, any type definitions the problem needs, and the code that reads stdin and parses it into native values.
- Body — what the candidate typed. Ideally one function, nothing else.
- Suffix — the call into that function with the parsed arguments, and the code that prints the return value in a canonical format.
The candidate sees only the middle piece, pre-filled with a signature. The runner sees all three joined together. That split is what the competitive programming world calls core code mode, as opposed to ACM mode where you are handed raw stdin and expected to parse it yourself.
Our client never assembles the program. It holds a map of starter templates keyed by language, sent down when the session starts, and posts back only the editor text plus the language name:
// Redux state, populated from the session bootstrap response
leetBoilerplateTemplates: null, // { javascript, cpp, java, python }
// and on submit, the client sends exactly this
{ sourceCode, language, sessionId, questionIndex }
No prefix, no suffix, no test data crosses the wire from the browser. The server owns the wrapper, which is the only arrangement that survives contact with a candidate who opens devtools.
Why the candidate should not parse stdin
You could skip all of this. Give people a blank file, pipe the test case into stdin, let them read it. Competitive programming judges have worked that way for decades and nobody complains.
It is the wrong default for an interview for three reasons.
First, it measures the wrong thing. Twenty minutes of a forty-five minute interview disappearing into Scanner versus BufferedReader, or into why input().split() returned strings, tells you nothing about whether the candidate can reason about a hash map.
Second, it makes failure ambiguous. When a submission returns wrong answer, you cannot tell whether the algorithm is wrong or the parsing is. Our interviewer agent reads the failing cases and asks about them, so that ambiguity would propagate straight into the conversation.
Third, it destroys language parity. Reading a matrix of integers is four lines in Python and fifteen in Java. If the candidate writes the parsing, the Java candidate is handicapped for choosing Java.
The rule the whole design rests on
Test data is stored once, in a language-neutral form. Only the readers differ per language. The moment you find yourself keeping a separate copy of a test case "for the Java version", the wrapper design has already failed.
Writing the prefix: getting typed arguments into a function
The prefix is generated from a typed problem definition — argument names, argument types, return type. Something close to this:
{
"name": "twoSum",
"args": [ { "name": "nums", "type": "int[]" },
{ "name": "target", "type": "int" } ],
"returns": "int[]"
}
From that one object you generate four prefixes. The interesting work is in the type mapping.
The type problem
Every supported language needs a reader for every type in your catalogue. We support four languages, which keeps the matrix small enough to reason about.
Scalars and strings
Easy, and still worth pinning down. An integer read in JavaScript is a double; in C++ it might need to be long long. Decide per problem whether the bound fits in 32 bits and emit the wider type when it does not, rather than discovering it through one failing test case.
Arrays and matrices
This is where most of the generated code lives. One line of JSON becomes std::vector<int>, int[], a Python list, a JavaScript array. Nested arrays double the work. Keep one reader function per type in the prefix and compose them, instead of emitting bespoke parsing per problem.
Linked lists and trees
These need the type definition in the prefix as well as a builder, and the definition must match the one named in the candidate's starter comment exactly. A ListNode with next in the comment and nxt in the prefix produces a compile error the candidate cannot diagnose, because they cannot see the prefix.
The entry-point problem
Languages disagree about where a program starts, and the prefix has to absorb that difference.
Python and JavaScript run top-level statements, so the suffix can simply be more statements. C++ needs int main(). Java needs a class with public static void main(String[] args) — and if the candidate's solution is itself a class, as ours is, the harness has to decide whether to nest it, put the main in a separate class, or make the solution method static. Pick one convention and generate every Java template from it. Mixed conventions across problems is how you get a question that only compiles on Tuesdays.
Writing the suffix: printing an answer a comparator can trust
The suffix calls the function and prints the result. Both halves are less obvious than they look, because the printed string is compared against a stored expected output, usually as text.
That comparison is only as good as the formatting rules.
- Booleans. C++ prints
1, Python printsTrue, Java and JavaScript printtrue. Four languages, three spellings, one expected output file. The suffix must normalise, not the comparator. - Floats. Never compare formatted floats for equality. Fix the precision in the suffix and compare with a tolerance, or the same correct algorithm passes in one language and fails in another.
- Arrays. Decide separators and brackets once and emit the same shape everywhere. Python's default
print(list)puts spaces after commas; nothing else does. - Unordered answers. If the problem accepts any order, the suffix sorts before printing, or the comparator is order-insensitive. Choose one. Doing both is fine; doing neither produces flaky verdicts that look like judge bugs.
- Null and empty. An empty list and a null return are different answers. Make them print differently.
The runner itself is deliberately dumb about all of this. It takes source code, a language id and stdin, runs it, and hands back stdout, stderr, compile output, time and memory — the Judge0 project is the canonical example of that interface, and its language seed definitions are what a language_id actually resolves to. Semantics are entirely your suffix's problem.
What the candidate sees versus what runs
The green band is the only part the candidate ever types. Everything grey around it is generated, which is also why the grey parts have to be right for every problem in the bank, not just the one you tested.
Five ways the wrapper breaks
Each of these is cheap to prevent at design time and expensive to debug live, in the middle of someone's interview.
Line numbers drift
The compiler reports errors against the assembled program. A forty-line prefix means an error on editor line 3 arrives as line 43. Subtract the prefix length before you show the message, and clamp anything that still lands outside the body so it renders as a harness error rather than pointing at a line the candidate does not have.
The signature gets deleted
Candidates select-all and paste, or rename the function to something they prefer. The suffix then calls a function that no longer exists and the whole thing fails to compile. Treat that as a distinct failure class — compile error against the template, not wrong answer — and keep the original starter code recoverable.
Switching language throws away work
Our editor replaces the buffer with the new language's template on switch, and the same path runs when the interviewer agent calls the language tool:
// mirrors the dropdown: switch language, load that language's starter template
dispatch(setLeetSelectedLanguage(language));
const template = leetBoilerplateTemplates?.[language];
if (template !== undefined) setEditorCode(ctx, template);
Deliberate — a half-written Python solution under a C++ template is worse than a clean slate — but it means a mis-click costs real work. Keeping a per-language buffer, so switching back restores what was there, is the version we would build now.
The template never arrives
If the session bootstrap response is missing the template map, the editor must still open. Ours falls back to a generic stub per language, and the language list the agent is allowed to pick from is fixed on the client:
export const LEET_LANGUAGES = ['javascript', 'cpp', 'java', 'python'];
An empty editor with no signature is a much worse failure than a slightly generic one. Equally, the editor's language mode is a separate concern from the template — the CodeMirror language packages we load handle highlighting and indentation, and they neither know nor care what the harness will wrap around the text.
Harness names collide with candidate names
If the prefix defines a helper called parse and the candidate also defines parse, one language will shadow silently and another will refuse to compile. Prefix every generated identifier with something nobody types by accident, and never introduce a bare using namespace std; in a prefix that also defines helpers.
The rules we settled on
- The server owns prefix and suffix. The client posts a language name and a body, and nothing else.
- Templates are generated from one typed problem definition, never hand-written per language.
- Test data is stored once, language-neutral. Only readers differ.
- The suffix normalises output; the comparator stays simple.
- Every failure mode gets its own verdict. A template compile error must never surface as wrong answer.
- Line numbers are translated back into editor space before a human sees them.
None of this is visible when it works, which is the point. The candidate writes a function, presses Run, and sees per-case results stream back. If you want the rest of that path, the SSE streaming layer covers how those results reach the browser, and you can watch the whole thing end to end in a free mock interview.
Frequently asked questions
What is a code template wrapper?
+
It is the prefix and suffix an online judge wraps around a candidate's function before compiling it. The prefix holds imports, type definitions and the code that reads and parses stdin. The suffix calls the function and prints the result in a format the comparator can check.
Why do coding platforms give you a function signature instead of a main?
+
Because the interview is about the algorithm, not about parsing input. Handing over a signature also fixes the contract: the harness knows the argument types and the return type, so it can generate the parsing and printing code for every supported language from one problem definition.
How does the harness parse stdin for different languages?
+
Each test case is stored in a language-neutral form, usually JSON. The prefix for each language contains readers that turn those lines into that language's native types: a Python list, a C++ vector, a Java int array, a JavaScript array. Only the readers differ, never the test data.
What happens if a candidate deletes the function signature?
+
Compilation fails, because the suffix still calls a function that no longer exists. A good harness reports that as a compile error against the template rather than a wrong answer, and keeps the original starter code recoverable so the candidate is not stuck with a broken editor.
Why do compiler error line numbers look wrong on coding platforms?
+
The compiler sees the assembled program, not the editor buffer. If the prefix is forty lines, an error on editor line three is reported on line forty-three. The harness has to subtract the prefix length before showing the error, or candidates chase errors in code they cannot see.
Should the template be generated or hand-written per problem?
+
Generate it from a typed problem definition. Hand-writing a prefix and suffix per problem per language means four times the work for every new question and four places for a type to drift. One definition, four generated templates, is the only version that stays consistent.