A Triage Framework for AI Code Detection in Student Work

The 94% AI probability stamped on a Python submission did not mean what the instructor thought it meant. Twenty minutes of reading her student's code told a different story: a sophomore who had learned Python almost entirely from a YouTube tutor, whose style had soaked in. Short functions. A docstring on every one. Descriptive variable names. That isn't AI. That's what a modern intro tutorial trains you to write.

The detector wasn't broken. It was answering a different question than the one being asked of it.

What an AI code detection score actually measures

Be specific about what these tools measure. Most AI code detectors, Codequiry's included, sample the token stream and compute two things: perplexity (how surprised a language model is by each next token) and burstiness (how much that surprise varies across the file). Human code swings. We indent things inconsistently. We name one variable temp, another temp2, and a third temp_final_FINAL. We leave a debug print in and comment out three lines we might want back.

An LLM produces a smooth distribution instead. It picks the next most likely token, then the next, then the next, and the variance stays low. That's the signal. It's real. It's also noisy in exactly the student populations you'd worry about most.

The score doesn't tell you whether a person wrote the code. It tells you how closely the token stream resembles what a model would produce. Those are not the same question.

A triage framework borrowed from incident response

I run a DevSecOps team at a mid-size fintech, which means most of my day involves deciding what to do with alerts. A scanner flags 400 dependency vulnerabilities overnight. You cannot act on 400. You build a triage pipeline, and you make the pipeline boring. The same discipline applies to academic integrity alerts, and it's surprising how rarely departments apply it.

Here's the shape of it. Four stages, each one cheap to run and each one allowed to kill the alert. An alert has to survive all four before it becomes a case file.

Codequiry AI code detection report with average AI score, highest file score and a risk distribution
AI code detection: probability scores per file, flagging submissions likely written by ChatGPT, Copilot, Claude or Gemini.

Stage one: read the score as a distribution, not a verdict

The most common mistake I see is treating the AI probability as a percentage of guilt. It isn't. A file scored at 91% is not "91% likely to be AI." It's a file whose token distribution sits 91% of the way toward the model's own distribution on the same prompt space. That number moves when you change the tokenizer, the sample size, or the language.

Small files are the worst offenders. A 12-line submission has almost no token stream to sample from, so the variance estimate collapses and the score saturates. I've watched a two-line print statement score 88%. That's not a detector finding AI. That's a detector running out of data.

Write your thresholds down before you need them:

AI score bandDefault action
0 to 54No review. Log only.
55 to 74Review only if peer or web signals also fire.
75 to 89Human review, corroboration required before any conversation.
90 and aboveHuman review, corroboration required, follow-up conversation regardless.

The bands exist so that the answer to "what do we do with an 85" doesn't depend on who opened the ticket. Which is the point.

One thing I'll flag honestly: the published false positive rates for AI code detection are not great. The best detectors I've tested land somewhere in the 85 to 92% range on mixed-language corpora, and they degrade hard on short submissions, non-English comments, and unusual syntax. If a vendor tells you 99%, ask them on what corpus. Then read the corpus.

Stage two: corroborate with peer and web signals

An AI score alone is the weakest signal in the stack. The strong signal is when AI detection, peer similarity, and web provenance all fire on the same submission. That combination is very hard to produce by accident. It's also easy to produce on purpose, which is why the corroboration is the whole game.

Peer similarity, the traditional domain of code plagiarism checkers, catches the student who copied from the person sitting next to them. It works on token-level comparison, and good engines push through renaming and reformatting because they normalize identifiers and compare structure, not text. A student who renames i to idx and swaps the loop for a for-each doesn't change the underlying token fingerprint.

Web provenance catches the student who lifted from a Stack Overflow answer or a GitHub repo. That's a different crime with a different disposition. Ripping a five-line regex from an accepted answer is often legitimate; ripping an entire assignment is not.

When I want the whole stack in one place, I point departments at Codequiry, which runs all three checks (peer submissions, the open web including GitHub, and AI generation) in a single pass and hands back a per-file breakdown. Most of the free tools do one of these. MOSS does peer similarity and nothing else. Turnitin added AI detection to its text pipeline, but its code review is not built on the same token and AST normalization a dedicated code tool uses, which matters for anything beyond literal copy-paste.

Codequiry web results tracing copied code to GitHub repositories and other web sources with per-domain scores
Web results: every domain a submission matched, scored per source, from GitHub repos to tutorial sites.

Here's the stacked decision I hand to TAs. It's crude, and it works:

def disposition(peer, web, ai, tokens):
    # peer and web are 0-100 similarity scores.
    # ai is 0-100 generation probability.
    # tokens is the submission size, in source tokens.

    if tokens < 150:
        return "insufficient_data"   # small files score noisily, skip

    if peer >= 90 or (peer >= 70 and web >= 40):
        return "plagiarism_review"

    if web >= 70:
        return "provenance_review"   # maybe legitimate reuse, read it

    if ai >= 75 and peer < 45 and web < 20:
        return "ai_review"           # solo AI signal, needs a human read

    if ai >= 90 and peer >= 45:
        return "combined_review"     # AI plus peer, strong evidence

    return "close"
Codequiry insights score breakdown separating peer similarity, web similarity and AI generation, with match sources
The score breakdown: peer similarity, web similarity and AI probability reported separately, with where the matches came from.

That tokens < 150 guard is the one people forget. It was added after we watched a 40-line submission score 96% AI because the file was almost entirely boilerplate: imports, a class header, and one method. Nothing in it was interesting enough for the detector to have an opinion about, and the score was noise.

Stage three: the human read, and what to actually look for

No detector replaces reading the code. What changes with AI detection in the mix is what you're reading for. You're not looking for plagiarism anymore. You're looking for the shape of authorship, and the tells are different.

Here's a function a language model writes when you ask for a running average:

def calculate_average(numbers: list[float]) -> float:
    """
    Calculate the arithmetic mean of a list of numbers.

    Args:
        numbers: A list of numerical values.

    Returns:
        The mean, or raises ValueError if the input is empty.
    """
    if not numbers:
        raise ValueError("Input list cannot be empty")
    return sum(numbers) / len(numbers)

And here's what an actual student writes, roughly:

def calcAvg(nums):
    # don't ask why this works it just does
    total = 0
    for i in range(len(nums)):
        total = total + nums[i]
    return total/len(nums)

The second one is worse code. It's also unmistakably human. The first one has four signatures I look for: type hints on every parameter, a docstring with an Args and Returns section, a proactive ValueError the assignment never asked for, and a one-line sum(...) / len(...) instead of a loop. Any one of those is nothing. Four in a 12-line function is a pattern.

Other tells that hold up across languages:

  • Uniform comment density. Humans comment where they got confused. Models comment everywhere, or nowhere, with eerie consistency.
  • Unused defensive code. Bounds checks, null guards, exception handling for conditions the assignment can't produce.
  • Unused imports. The model added import os because the training distribution thought it might need it.
  • Consistent formatting when the class average is ragged. If a student ran black on one submission and nothing else all semester, ask about it.
  • Correct-by-construction edge cases. Whoever wrote it happened to handle the empty list, the single element, and the negative number on the first try.

The reverse tells matter too. Debug print statements, commented-out code, a comment referencing this week's problem set by its actual name, a variable called temp2, a semicolon at the end of every Python line because the student came from Java. Those are human, and they should pull you back toward closing the ticket even when the detector is loud.

AI detection catches a signal. The code itself tells you the story. You still have to read it.

Stage four: disposition, and the policy question

Remember that alert triage is only as good as what comes after it. This is where most departments quietly fall apart, because the technical question ("is this AI?") is easy next to the institutional one ("what do we do about it?").

Three dispositions, and the department should write all three down before the semester starts:

  1. Close. The signal didn't survive corroboration, or the code reads human. No student conversation, no note in the file. Say this out loud in your policy, because TAs default to "flag it just in case," and a flagged-but-closed alert in a shared dashboard is a permanent accusation with no due process.
  2. Conversation. The code reads AI, single signal, weak corroboration. This is a one-on-one with the student, not a case. Ask them to explain a specific function. Ask them to modify it live. What you want is the story they tell about their own code. Students who wrote it will tell it. Students who didn't will get vague fast.
  3. Case file. AI plus peer plus web. Or AI plus in-class performance that doesn't match. Or AI plus a student who, when asked, cannot explain a single line. Now you have corroborated evidence, and the disposition is whatever your honor code says it is.

The policy exists so that "AI detection is a distraction" doesn't become "we don't look." It also exists so that "we take academic integrity seriously" doesn't become a 94% score and a hearing. Somewhere between those two failure modes is a department that reads the code, checks the stack, and acts proportionally. That's a policy problem, not a tooling problem.

Wiring the stack so it runs without you

The reason this triage works at scale is that it's automatable. If your integrity tool exposes an API or a webhook, you can push submissions through the same pipeline shape you'd use for any other scan. A policy file, a webhook, a report. Codequiry exposes a REST API and a CLI for exactly this. The config below is the pattern, not a copy-paste from their docs:

# course-integrity.yml - example policy, adjust to your course
checks:
  peer_similarity:
    enabled: true
    threshold: 70
  web_provenance:
    sources: [github, stackoverflow, docs]
    threshold: 40
  ai_detection:
    enabled: true
    per_file_threshold: 75
    cohort_threshold: 60
  minimum_tokens: 150
output:
  formats: [json, html]
  webhook: https://integrity.example.edu/hooks/scan-complete
  queue: ./review-queue/
Codequiry API keys page with a masked key, signed webhook configuration and API resources
The API surface: an account key, signed webhooks for finished checks, and docs for wiring scans into CI.

Run it on submission, not on deadline. Students who see their work scanned as they submit it adjust their behavior; students who get scanned two weeks later just get a surprise.

Where this goes next

I don't think the detection problem gets solved. The models are getting better at writing like students, and students are getting better at writing like the models, and the gap in the middle is where we all live. What does get solved is triage. A department that can move 500 submissions through a stack of three signals, read the survivors, and close the rest with a clean audit trail has already won most of the fight, regardless of what the next model does.

Which is the argument for running a dedicated AI code detector alongside your similarity checker rather than picking one. Peer similarity answers one question. Web provenance answers another. AI generation answers a third. A single tool that answers all three, with a per-file breakdown and a webhook, is worth more than the sum of the three separate checks, because the corroboration is where the signal gets strong enough to act on.

If your department is still running MOSS and a spreadsheet, comparing Codequiry to MOSS is a reasonable place to start the conversation.

Frequently Asked Questions

What AI score should trigger a review in a student submission?

We default to 75 for a single-file flag and 60 for a cohort-level flag, but the more important rule is that no AI score triggers action on its own. A score above 90 with no peer or web corroboration gets a human read, not a case file. Below 55, log and move on.

Can AI code detection tell the difference between AI-assisted and fully AI-generated code?

Not reliably, and anyone claiming otherwise is overselling. A student who writes 80% of a submission and asks a model to fill in one function will often not register as AI at all, because the surrounding code keeps the token distribution grounded. Detecting the partially-assisted case is an open research problem.

How often are AI code detectors wrong?

More than the marketing suggests. On short files under 150 tokens, false positive rates can exceed 30% in my experience, mostly because there isn't enough signal to estimate variance. On substantial files in common languages, well-tuned detectors put false positives in the 8 to 15% range. Always corroborate before an accusation.

Do I still need a plagiarism checker if I have AI detection?

Yes. They catch different things. AI detection catches machine-written code; similarity and web provenance catch code copied from a peer or from GitHub. A student who copies a classmate's answers passes every AI detector with flying colors, and a student who generates everything fresh leaves no peer match. You want both signals, which is why a stacked tooling setup beats a single-check workflow.

Start with a small pilot: run the next assignment through a full scan, apply the triage bands above, and see how many alerts survive corroboration. If you want a tool that returns peer, web, and AI scores in one pass, that's what the Codequiry code plagiarism checker is built for, and the free tier will tell you in an afternoon whether the signal is clean enough for your course.