How One Bootcamp Screens 400 Take-Homes for AI Code

Last hiring round we ran 387 take-home files through peer, web, and AI checks before a human opened the editor. Thirty-one candidates came back flagged. Nine of those turned out to be real problems, and every one of the nine had a reasonable-sounding explanation ready until we put the evidence side by side.

Detecting AI-generated code in take-home interviews is less about a magic score and more about sequencing: catch the outliers automatically, review the evidence by hand, then let the technical interview do the rest. Everything below is what our 12-week full-stack program actually does now, including the parts that still annoy me.

What the take-home actually asks for

Our take-home is a URL shortener with a twist. Candidates get a starter repo, a Python 3.11 or Node 20 track, and a requirements doc that runs about 900 words. They have five days, and we tell them out loud that the work should take four to six hours. Extra features are not rewarded. Working code, tests, and a README that explains the tradeoffs are.

The one line that matters for everything in this post is this one: "You may use AI tools anywhere in your process, but you will be asked to explain, line by line, why the code looks the way it does." That sentence changed how we review submissions more than any tool we've bought. It also turned a detection problem into a conversation, which is where I'd rather spend my time.

Roughly 40% of candidates now admit to using Copilot or ChatGPT somewhere in the process. In our last two rounds, that number tracked almost perfectly with what the detector reported, which is the sort of thing that makes you trust a tool right up until it doesn't hold.

Why we stopped reading every submission first

Until early 2023, three instructors read all 129 take-homes end to end. That's about 40 hours of reading per round, spread over six days, and by submission ninety all three of us were skimming. Skimming is how you miss the candidate who renamed every variable and left the original comment block sitting right there in the middle of the file.

I used MOSS back in 2019 when I was a TA, and I tried JPlag (v4, the Java-based one) for a semester after that. Both work. Both also assume you're comfortable with a command line, a submission directory structure, and reading a match report in a browser tab that looks like it was designed for another decade. That's fine for the course you teach every year. It's less fine when a hiring coordinator needs to check a batch on a Thursday afternoon.

So we moved the first pass to a scanner and kept the human reading for the flagged 25%. The order matters more than the tool.

The three scores that show up on every submission

Every candidate repo now gets three independent numbers before anyone reads a line. They answer different questions, and conflating them is the most common mistake I see in other programs.

  • Peer similarity: how close this submission is to the other 128 in the same round. This catches the two candidates who compared notes in the recruiting Slack.
  • Web similarity: how much of the file appears in public code, on GitHub, Stack Overflow, or in tutorial repos. This catches the candidate who lifted a rate limiter from a 2019 blog post.
  • AI generation score: a per-file estimate that the code was produced by a language model rather than typed by a person.

We run all three through a code plagiarism checker that reports them in one place, because a candidate can be clean on two and dirty on the third, and nobody wants to reconcile three export formats at 11pm. The web pass is the one that surprises people. Around 18% of submissions in our last round had at least one file matching a public source, and most of those were harmless: a standard parse_qs call, a copied Dockerfile, a README section lifted from the starter repo we ourselves published.

Codequiry submissions table with per-submission peer, web and AI score rings and per-file scan states
The submissions table: peer, web and AI scores per student at a glance, with deeper scans still streaming in.

What a scan looks like at minute fourteen

The last round was 387 files across 129 repos. A full peer-to-peer comparison on that set is roughly 74,000 pairs, which sounds like a lot until you remember the engine hashes tokens first and only builds an AST for pairs that survive the fingerprint filter. Ours finished in under nine minutes on a 4-core box while I was on a call.

Codequiry scan monitor mid-run at 92% with CPU and memory gauges and a live per-file activity log
A check in flight: live progress, resource gauges, and a per-file activity log as each submission is scored.

Here's the thing I didn't expect: token-level normalization is what separates a useful report from a wall of false positives. Watch what happens with a trivial rename.

# Submission A
def calc_tot(items, rate):
    return sum(i["price"] * rate for i in items)

# Submission B
def compute_sum(records, multiplier):
    return sum(r["price"] * multiplier for r in records)

A text diff sees two different functions. A token-based engine sees the same shape. An engine that also builds an AST sees the same tree with different leaves, and reports something like 94% similarity on the body. That's the difference between catching a real copy and catching a coincidence, and it's the main reason I stopped maintaining my own normalization script. Reading the Codequiry comparison with MOSS is worth twenty minutes if you're still on the older toolchain, because the gap between them lives almost entirely in that layer.

Detecting AI-Generated Code in Take-Home Interviews Without Turning It Into an Interrogation

This is the part everyone asks about, so let me be blunt about how it works and how it fails. There is no watermark in ChatGPT output, no hidden token, nothing you can grep for. Detection is statistical, and any tool that claims otherwise is selling you something. A AI code detector looks at how predictable each token is given the ones before it, how evenly that predictability is distributed across a file, and a set of structural habits language models reliably pick up.

Those habits are real, and they're visible to a human reader too. Here's a function from a submission that scored high, trimmed to the part that gave it away.

def calculate_total(items: List[Dict[str, Any]], tax_rate: float = 0.08) -> float:
    """Calculate the total price of items including tax.

    Args:
        items: A list of dictionaries containing item data.
        tax_rate: The tax rate to apply. Defaults to 0.08.

    Returns:
        The total price including tax.
    """
    try:
        total = sum(item.get("price", 0) for item in items)
        return round(total * (1 + tax_rate), 2)
    except Exception as e:
        print(f"Error calculating total: {e}")
        return 0.0

Nothing here is wrong. That's the problem. For a take-home that asked for a URL shortener, a fully typed, Google-style docstringed, defensively wrapped helper for a six-line sum is a costume. Candidates under a five-day deadline write sum(i["price"] for i in items) and move on. They don't wrap except Exception around arithmetic and print the error to stdout.

Other tells we see over and over: comments that restate the function name, an if __name__ == "__main__": block with argparse when nothing asked for a CLI, three paragraphs of README explaining that the code "follows best practices," and identical error message phrasing across two candidates who never met.

An AI score of 78 is a reason to read carefully, not a reason to reject. The rejection comes from the transcript, not the dashboard.
Codequiry AI code detection report with average AI score, highest file score and a risk distribution
AI code detection: probability scores per file, flagging submissions likely written by ChatGPT, Copilot, Claude or Gemini.

We interview every flagged candidate the same way. We open the file, ask them to walk us through the exception handling, and then we ask them to change it. "What happens if this raises on the fourth item rather than the first?" Most people who wrote it can answer. Most people who prompted it cannot, and the good ones say so, which we score as honesty rather than failure.

Where the AI signal is weakest

I want to be careful here, because this is where vendors oversell. Our AI signal is weakest on very short files, under about 40 lines, on languages with heavy boilerplate, and on any assignment where the correct solution is short and standard. Write a take-home that says "reverse a linked list" and your AI score is noise, because a competent person and a language model converge on nearly the same twenty lines.

Rust and Kotlin tripped us up early, mostly because the borrow checker and the type system push human code toward the same shapes the model produces. We haven't tested this past a few thousand files across two cohorts, so treat that as an observation rather than a finding. The practical fix is the same every time: score the file, but read the diff and run the interview. A detector that says "this file looks model-generated" is a lead, not a verdict.

What we changed in the take-home prompt itself

The most effective thing we did last year wasn't a tool. It was a rewrite of the assignment. We added two requirements that are almost impossible to satisfy with a single prompt and trivial for a human who actually thought about the problem.

The first: candidates must connect the shortener to a persistent store of their choice and write a paragraph on why they picked it, referencing something specific about our stack (we ship a short description of it in the brief). The second: they must include one deliberately imperfect decision with a comment explaining when they'd revisit it. That second one is the interesting one. Models write confident code. Humans write code with a note that says "this is O(n) and it's fine for now, don't ship it to production."

I borrowed the idea from a colleague who runs an upper-division course at a state school, and it maps cleanly to assignment design in general. If a prompt can be fully answered by one sentence typed into a chat window, your detection problem is really a design problem. The plagiarism checker for code tells you who to look at. The prompt decides how many people there are to look at.

The bill, the API, and what I'd do differently

Our screening costs us under $200 a round, roughly two instructor hours of review. That math isn't close. The dashboard handles the hiring side, and the REST API lets me push submissions out of our recruiting ATS without a human dragging folders around, which matters more than it should.

Codequiry evidence review with a synced diff of two Java files and a list of GitHub and web matches
Evidence review: a synced diff of the matched lines next to every peer, GitHub and web source for the submission.

Where I'd do it differently: I'd have started with the web-source check before the AI check. We spent a month tuned into AI scores while a candidate quietly shipped a rate limiter copied verbatim from a GitHub gist, 61 tokens long, original author's variable names intact. The web and peer passes would have caught that on day one.

One more thing worth saying out loud. I maintain a small open-source library, and I've watched this same class of tooling get adopted on the other side of the fence. Reviewers running similarity checks on inbound pull requests find copy-pasted GPL code faster than any human reading a 400-line diff, and in that setting the license problem is worse than the plagiarism problem. Same engine, very different stakes.

If you're running take-homes or grading submissions at any scale, start by scanning everything and reading only the outliers. Run your next batch through Codequiry before the interview calendar fills up.