Most computer science departments meet the same question the first time a similarity report lands on a chair's desk: 72% means what, exactly? A code plagiarism score threshold is not a verdict, and it was never designed to be one. It reports how much two submissions overlap once identifiers, whitespace, and comments have been normalized away. Whether that overlap came from copying, a shared starter file, a common tutorial, or two students who asked the same model the same question is a separate judgment, and it's the judgment departments most often skip.
I served as the outside reader on a case review at Cascade State University during the spring 2024 semester. Cascade is a mid-size public institution with roughly 1,900 students a year moving through its introductory sequence. The department had been scanning student code for nearly a decade. It had never once written down what a score was for.
The week the flag queue outran the graders
CS 211 Data Structures enrolled 340 students that spring, with four teaching assistants and one instructor, Dr. Priya Raman, who was in her fourth year with the course. Assignment four asked for a self-balancing binary search tree, an in-order traversal that returned a list rather than printing to the console, and a rebalancing routine with a documented rotation policy. Java, roughly 400 lines per submission once students added their own test harnesses.
Raman ran MOSS the weekend after the deadline, pointing the -b flag at the starter files so the provided interfaces would be treated as baseline rather than submissions. The result page came back with 138 pairs above 30% similarity and 41 pairs above 50%. The honor board already had nine open cases carried over from the fall. Forty-one more would have pushed hearings into the following academic year, because a contested hearing at Cascade averages about ninety minutes of committee time before anyone reads a single line of code.
There's a mundane problem in that workflow that rarely comes up in tool comparisons. MOSS returns an HTML page of match pairs sorted by number of matching lines. It carries no student names, no assignment metadata, and no structured export. A TA spent two evenings copying pair identifiers into a spreadsheet and cross-referencing them against the learning management system, then two more evenings opening files two at a time to see what the matches actually looked like. By the time Raman had a ranked list, the semester was over.
What a code plagiarism score threshold actually measures
Before the department could set a line, it had to understand what the number underneath the line was counting. The three families of comparison engines behave very differently, and most instructors I talk to have never had the distinction explained to them.
Text similarity is the simplest kind. The tool computes an edit distance or a percentage of shared lines. It's fast, it needs no parser, and it collapses the moment a student renames a variable or reformats the file.
Token-based fingerprinting is what MOSS popularized with the winnowing algorithm published by Schleimer, Wilkerson, and Aiken in 2003. The engine strips comments, normalizes identifiers and literals, breaks the token stream into overlapping k-grams, hashes each one, and keeps a subset of those hashes as a fingerprint. Renaming every variable in a submission moves the fingerprint not at all.
Abstract syntax tree comparison parses the submission and compares structure: loop nesting, call relationships, control flow. It survives reordering that token matching misses, and it is slower and fussier about language support. JPlag has leaned on tree comparison for years. Dolos, the open-source tool from a research group at Ghent University, builds on suffix trees and token analysis and has become a common MOSS replacement since its 2022 release.
Here is what the difference looks like in practice. These two Java methods are the same algorithm, and neither is a copy of the other in any textual sense.
public static int maxSubarraySum(int[] nums) {
int best = Integer.MIN_VALUE;
for (int i = 0; i < nums.length; i++) {
int running = 0;
for (int j = i; j < nums.length; j++) {
running += nums[j];
best = Math.max(best, running);
}
}
return best;
}
static int bestWindowTotal(int[] a) {
int bestTotal = -1_000_000_000;
for (int start = 0; start < a.length; ++start) {
int acc = 0;
for (int end = start; end < a.length; ++end) {
acc = acc + a[end];
if (acc > bestTotal) bestTotal = acc;
}
}
return bestTotal;
}
A line-based comparison lands somewhere around a third. A token-based engine, with identifiers normalized to a single symbol, reports a match in the high eighties, because structurally there is almost nothing left to differ. An AST engine reports the same and can also tell you the loop nesting is identical in depth and order. If your department's threshold was set by someone who only ever looked at a text diff, you are grading the wrong quantity.
The calibration exercise that set the line
Raman's solution was to stop guessing. She built a validation set of twenty-four submission pairs from prior semesters: twelve pairs with confirmed, documented copying, and twelve pairs from students who had worked independently, verified through office-hour notes and version history. Then she ran the whole set at five thresholds.
| Threshold | Copied pairs flagged (of 12) | Independent pairs flagged (of 12) |
|---|---|---|
| 40% | 12 | 5 |
| 50% | 12 | 4 |
| 60% | 11 | 2 |
| 70% | 10 | 1 |
| 80% | 7 | 0 |
Those numbers come from one department, one language, and a fairly small sample, so I would not treat them as universal. The shape is typical, though. There is no threshold that is both sensitive and clean, because in a well-specified assignment, correct solutions converge. When the spec requires a recursive helper with a particular signature, every student who follows the spec earns fifteen to twenty-five points of similarity for free.
A threshold is not really a line about student behavior. It is a staffing decision wearing a percentage sign.
The department settled on two numbers rather than one. Sixty percent opened a review. Eighty percent, combined with a web match or a pattern across more than two submissions, produced an evidence packet. Nothing automatic happened at either number. That distinction needs to be in the syllabus in plain language, because a policy that exists only inside a tool's settings page will not survive an appeal.
Where AI-written submissions sit in the distribution
Then there was the second problem, the one that arrived the same semester. Raman noticed a cluster of matches in the 55% to 75% band that made no sense as collusion. Eleven submissions each matched six or more others in that band. When she looked at the files, they shared a house style: an explanatory comment block above every method, input validation the spec never requested, an Optional return where the assignment called for a plain value, and a consistent habit of marking every local variable final even when it was never reassigned.
That is the signature of a language model, and it matters for how you read the score. AI-generated submissions do not usually produce one dramatic 95% match. They produce a wide mesh of moderate matches across dozens of students, because dozens of students typed a similar prompt into a similar tool. Peer similarity alone cannot distinguish thirty students who each used an assistant from thirty students who copied from one another.
Statistical signals borrowed from prose detection are weaker in code than people assume. Perplexity measures how surprised a model is by the next token. Burstiness measures how much sentence length varies. Both were designed for English essays, and they degrade on programming languages, where token entropy is low across the board and the grammar is far more constrained. What holds up better in source code is structural and stylistic: comment density relative to method length, defensive patterns nobody asked for, and a uniformity across files that a human writing under deadline rarely achieves. The publishing on this is thinner than the vendor pages suggest, and I'd treat any single AI score above 0.9 as a reason to look rather than a finding.
In one review that semester, a student scored 0.87 for AI generation and had used the same idiom in a public repository two years earlier. The department learned to check a student's earlier coursework and public version history before scheduling a conversation. That step appears in no vendor's documentation, and it is the one that prevents the worst outcome available to a department, which is accusing a strong student of something they did not do.

The starter file that produced sixty-one false matches
The first scan of the semester produced a spectacular and entirely wrong result: sixty-one submissions matched each other at 100%. The cause was mundane. A TA had uploaded the assignment's starter repository into the peer corpus, so every student who kept the provided TreeMap wrapper intact matched every other student on that file, perfectly.
Nothing about the engine was broken. The lesson is that anything you distribute to the whole class belongs in an exclusion list, and that you should read per-file similarity rather than trusting the aggregate score at the top of a report. The same mistake shows up in industry audits, where a shared internal library produces a uniform match across every service in a repository.

The triage policy the department kept
By fall 2024, Cascade had a written process, and it has changed very little since. The mechanics matter less than the order of operations.
A flag opens a review, never a charge. A reviewer reads the two files side by side with the diff in view and records a short note: what matched, what the match sources were, and whether the similarity survives the exclusion of anything the instructor provided. Web matches get traced to their origin, because a 40-line overlap with a 2016 Stack Overflow answer that the student cited in a comment is a different event from a 40-line overlap with a classmate's private repository.

Only after that reading does anyone talk to the student, and the conversation is a code walkthrough, not an interrogation. Raman's version runs about fifteen minutes and starts with a specific line. "Talk me through what happens here if the key already exists." Students who wrote the code can explain why the loop terminates early on a duplicate key. Students who did not usually cannot. The ones who can are the false positives you were about to punish.

Of the 41 original flags, deduplication and the starter-file correction left 26 unique cases. Eleven went to the honor board. Fifteen resolved in a documented conversation with the instructor, four of which ended with a grade penalty for misrepresenting a source and eleven with no penalty at all. The department also wrote down why each non-flagged submission got cleared, which turned out to be the part that made the process defensible when a parent called the dean.
What the department changed about tooling
The process changes were the important part. The tooling changes made the process survivable. Cascade moved off MOSS for routine course use in the fall of 2024, mainly because the HTML result page could not produce the per-file evidence packet the new policy required, and because the department wanted peer similarity, web sources, and AI signals in one report instead of three unrelated spreadsheets.
The replacement was Codequiry, and the reason was mostly operational. It runs peer comparison, web and GitHub source matching, and AI generation analysis in a single check, so a reviewer opens one workspace rather than reconciling three tools. The code plagiarism checker reports per file, which is what surfaced the starter-file problem in the first place, and the review workspace shows a synced diff beside the list of external matches. Assembling that from MOSS, JPlag, and a separate AI code detector is possible, but it costs a graduate assistant and a great deal of patience.
For departments weighing options, the honest summary is that MOSS remains free and fast, JPlag remains excellent for pairwise academic comparison, and Dolos is a solid open-source choice with broad language support. What none of them provide is a dashboard a non-programmer administrator can read, a student-facing portal, or an API that lets an engineering team run the same checks inside a CI pipeline. That last point mattered more than anyone expected. The chair circulated a Codequiry vs MOSS comparison before the vote, largely because cost was the first question raised in the meeting and the second was who would maintain the scripts.
What I would tell another department
Set the threshold in writing before you have a case you care about. Calibrate it against your own assignments, because a first-semester Python course and a senior compilers course produce completely different baseline similarity. Treat AI scores as a reason to look, not a reason to act. Keep the evidence packet short enough that a faculty member can read it in ten minutes.
The alternative is what Cascade had in April: a queue nobody had time to work, a spreadsheet maintained by hand, and a policy that existed only in a tool's default settings. If your department is starting from a MOSS script and a shared drive, the practical first step is to run a full peer, web, and AI check on one assignment and see how many of your 50% matches survive a per-file reading.