Two days after the Assignment 4 deadline, I had 41 flagged pairs in a 214-student section of CS 210 and an email to the dean half-written. Then I actually read the submissions.
Detecting AI-generated code in student submissions is a different job from detecting copying between students, and the two look different in the data. AI submissions cluster: a dozen or more students, moderate pairwise similarity, no clear outlier pair. Copied submissions pair off: one or two pairs sitting at 92% or higher with the same variable names, the same comments, the same wrong answer in the same place. If you run one check, read the top line, and start writing conduct referrals, you will accuse the wrong students.
Why one prompt makes a whole class look like cheaters
Give 200 students the same problem and you get something close to 200 variations of one algorithm, because that is what an LLM is for. Ask for the kth largest element in an array and in 2024 and 2025 ChatGPT, Claude, and Gemini mostly converge on a min-heap of size k or a quickselect. Same imports. Same helper placement. Same bounds check at the top.
The part that surprised me is how much that inflates peer similarity when nobody copies anybody. Our matching compares token sequences and AST shape, not keystrokes, so renaming every identifier does almost nothing to the score. Two independently generated heaps land between 65% and 85% against each other. On the same assignment in 2021 the median pair similarity in my section was 22%. In Fall 2024 it was 51%. Nothing about my teaching changed in those three years except the tooling my students have on their laptops.
// Student A, submitted 10/14
public static int kthLargest(int[] nums, int k) {
PriorityQueue<Integer> heap = new PriorityQueue<>();
for (int n : nums) {
heap.offer(n);
if (heap.size() > k) heap.poll();
}
return heap.peek();
}
// Student B, submitted 10/15
public static int findKth(int[] arr, int position) {
PriorityQueue<Integer> pq = new PriorityQueue<>();
for (int value : arr) {
pq.offer(value);
if (pq.size() > position) pq.poll();
}
return pq.peek();
}
That pair scores around 78% on a token-and-AST comparison. It is also, almost certainly, two people who opened two different chatbots and typed the same sentence. Neither of them has ever seen the other's file.
Telling AI-generated code apart from a copy pair
The distinction is mostly about shape, not about any single score. One high number tells you almost nothing. A distribution of numbers tells you a lot.
| Signal | AI cluster | Genuine copy pair |
|---|---|---|
| Number flagged | 15 to 60 in a large cohort | Two, occasionally three |
| Distribution | Many pairs between 60% and 85%, no break | One pair at 92% or higher, sharp gap behind it |
| Identifiers | Varied, model-flavored, sometimes oddly formal | Identical, including the same misspelling |
| Comments | Uniform, generic, present on nearly every method | Identical, including a student's own typo |
| Mistakes | Plausible but wrong, like an off-by-one in a boundary case | The same wrong answer in the same line |
| Style drift | Flat. No file looks different from the others | One file suddenly formats like a different person |
The fastest confirmation I know costs about 20 minutes per assignment. I keep a folder called prompts_i_tested and I paste the assignment prompt into two current models. If the output structure matches the flagged cluster, I am looking at shared tooling, not a copying ring. If a pair matches something a model would not produce, I go to evidence review. That folder has settled more cases than any score threshold.

What an AI detection score for code actually measures
Text detectors lean on perplexity and burstiness: how surprising each token is, and how much that surprise varies across a document. Code gives you something extra, because code has structure that text does not. An AI code detector can look at how regular the token stream is, how uniform the comment density is across files, how consistent the naming gets, and whether the defensive checks make sense for the spec.
Here is the pattern I see most often in flagged Java and Python files. A student writes 80 lines with no type hints, no docstrings, no input validation. Then one function in the middle arrives fully dressed.
def validate_positive_integer(value: int) -> bool:
"""Validate that value is a positive integer.
Args:
value: The value to validate.
Returns:
True if value is a positive integer, False otherwise.
"""
if not isinstance(value, int):
return False
if value <= 0:
return False
return True
The docstring is not the evidence. The mismatch is. Their Assignment 2 had for(int i=0;i<n;i++){ with no spaces, and now Google Java Format has clearly visited. Human student code carries scars: a commented-out attempt, a debug print someone forgot, an unused import, a helper that exists only because they started down the wrong path. Generated code arrives without scars, and after three weeks of grading, that absence is louder than any percentage.
Which number do you read? Read the highest per-file score and the distribution, not the average. A submission with nine files averaging 38 and one file at 88 is a different animal from nine files averaging 78. One thing to know about the averages: in an early 2024 build we shipped, a perfectly uniform 12-line license header counted as an AI indicator, and it flagged about a dozen students who had pasted the starter header I gave them. It was patched within a week, but I still open the top five hits by hand every semester before I email anybody.


Stacking peer, web, and AI signals in one pass
Each signal has a different blind spot, which is why I run all three. Peer similarity misses the student who works alone with ChatGPT and never talks to a classmate. Web checks miss the roommate. AI scoring misses the student who copied a friend's file from two semesters ago, because that file was written by a human.
This is also where I stopped running three tools. MOSS is free and it is what most of us learned on, but it returns a URL full of HTML pair reports, has no AI signal, and checks against a corpus you have to maintain. JPlag produces better reports and still no AI signal. Dolos has the nicest UI of the three and the same gap. A code plagiarism checker for teachers that covers peer submissions, the open web, and AI generation in one report saves me a full evening per assignment, and it means the scores I am comparing were normalized the same way. Codequiry runs all three engines against peer submissions, GitHub, and the open web, across 65 languages, and it exposes the same scans through a REST API and CLI, which is how our engineering-adjacent programs check contractor code with the same rules.

A Monday morning triage that takes two hours
Here is the actual workflow, with the times, because the times are what make it survivable at 200 submissions.
- Export the cohort to CSV: submission ID, highest peer match, web match count, highest file AI score. Two minutes.
- Sort by highest file AI score, descending. Triage the top 40 at 30 to 45 seconds each, which is enough to see whether the flagged file is stylistically consistent with the rest of the submission. Roughly 25 minutes.
- Cross-reference against the peer table for anyone who shows up in both lists. A student with a 74% peer match and a 62 AI score is more interesting than either number alone. Five minutes.
- Pull the four to six submissions that survive into evidence review, open the synced diff against the matched file or the GitHub source, and read. Ten minutes each.
Two hours, and I have a defensible list. The only step you cannot automate is the last one. When you detect code plagiarism in a large course, the software ranks your suspects; you still have to read the code and decide what it means.
Handle it as a conversation, not a verdict
My first question in the meeting is always the same. Walk me through line 47. Students who prompted a model and read the output can usually do it, badly. Students who submitted a file they never opened cannot, and neither can students who copied from a friend. What separates the two cases for me is the disclosure, not the typing.
Since Fall 2023 my syllabus has said this in one sentence: undisclosed AI use on individual assignments is an integrity violation, and cited AI use is fine on the two open-book labs where I hand out the prompt log template. That sentence has done more for compliance in my sections than any detector I have run. I am honest that we have only tested it across three semesters, so treat my numbers as mine and not as a universal baseline.
Frequently asked questions
Can a professor tell if a coding assignment was written by ChatGPT?
Sometimes, and the signal is rarely a single score. It is the combination of a high per-file AI score, structural uniformity across the submission, an absence of human mess, and a student who cannot explain their own control flow. The score is a triage tool that decides who gets a conversation, not proof of anything.
Is high peer similarity between two submissions always plagiarism?
No. If 30 pairs in your cohort sit between 60% and 85% with no outliers, you are almost certainly looking at a shared prompt, a shared starter file, or shared tooling. Look for the break in the distribution. Copy pairs separate from the pack; AI clusters sit inside it.
Do AI detectors work on short code files?
Weakly. Under about 50 lines of Java or Python, per-file scores get noisy, and a good student with a terse style can trip one. Read the file-level distribution and the whole submission rather than trusting a single small file.
Can students get around AI code detection?
Renaming variables, rewriting comments, and reordering methods will pull the AI score down. It does not change the structure much, so it will not pull the peer similarity down against every classmate who used the same prompt.
If you want to see how peer similarity, web matches, and AI generation stack up in one report before your next deadline, you can run a free check with Codequiry's code plagiarism checker.