How Code Plagiarism Detection Went From Hashes to LLMs

Karl Ottenstein published the first working approach to code plagiarism detection in 1976. It ran on a mainframe, compared student Fortran assignments by hashing their token streams, and got the field most of the way to where it sits today. The next fifty years were mostly spent closing the last twenty percent.

I've sat on both sides of that table. I ran engineering at two startups, hired contractors under deadline pressure, and once shipped a payments module that turned out to contain a stripped-header copy of an AGPL PDF library. The lawyer's letter arrived eight months after launch. Now I consult, and a decent chunk of that work is telling a CTO whether the code in their repo is actually theirs.

The history matters, because it explains a failure mode you'll hit this week. A tool that reports 0% similarity is not telling you the code is original. It's telling you the tool only knows how to ask one question. Here's how we got from one question to three.

How the first code plagiarism detectors worked

Ottenstein's approach was structural. Strip comments and whitespace, normalize identifiers, hash the resulting token stream, compare hash sets. Two programs with the same logic but different variable names produce overlapping fingerprints. That was 1976, and it's still the core of most tools on the market.

The competing approach at the time was simpler and much worse. Run Unix diff, count matching lines, divide by total lines. A lot of honor boards used exactly this. Some still do.

Here's why it collapses:

# Submission A
total = 0
for value in readings:
    total += value
return total

# Submission B
sum_result = 0
for entry in measurements:
    sum_result = sum_result + entry
return sum_result

A line diff on those two produces noise. Different variable names, different line breaks, a different operator form. The score comes back around 20%, and the student who copied gets a pass. Rename three variables and you've beaten a line-based checker.

Token fingerprinting and the winnowing algorithm

Alex Aiken's MOSS landed at Stanford in 1994 and changed the default. It's still free, still used, and still the benchmark everyone compares against. The algorithm underneath came later, in a 2003 SIGMOD paper by Schleimer, Wilkerson, and Aiken: winnowing.

The mechanics are worth understanding, because they explain what your tool can and cannot see. Convert source to a token stream. Strip comments and whitespace. Normalize identifiers. Take every sequence of k consecutive tokens, hash each one, then slide a window of size w across the hash list and keep only the minimum hash in each window.

That last step is the clever part. It thins thousands of hashes down to a sparse fingerprint set, and it comes with a guarantee: any matching run of at least k + w - 1 tokens will produce at least one shared fingerprint. Not "probably." Will.

Winnowing trades completeness at the edges for a mathematical guarantee in the middle. That's why it survived thirty years of refactoring and still catches renamed variables, reformatted braces, and reordered helper functions.

Two things to know before you trust the output. First, the parameters matter. Set k too low and you flag every loop in the class. Set it too high and a copied function that's been lightly edited slips under the match-length threshold. Second, MOSS has a flag almost nobody reads: -m caps the number of reported matches per file pair, and the default is 10. If two submissions share forty matching blocks, you'll see ten and assume you've seen everything. I once watched an instructor build an entire misconduct case on a truncated report.

Token matching is not the end of the story either. It's blind to structure. Students who rename everything and rearrange control flow pass the fingerprint step while leaving the underlying logic untouched.

AST comparison and refactoring-resistant detection

JPlag arrived in 1996 out of the University of Karlsruhe, built by Guido Malpohl, and it took a different route. Parse the source into an abstract syntax tree and compare tree structure instead of token order. Renaming variables changes tokens. Reformatting changes tokens. Wrapping code in a new function changes tokens. None of those reliably change the tree's shape.

// Submission A
for (int i = 0; i < n; i++) sum += values[i];

// Submission B
for (int j = 0; j < count; j++) total += data[j];

Both parse to the same tree: a for-loop with an initialization, a comparison against a length expression, an increment, and a compound assignment into an array element. An AST matcher flags that pair at high similarity even when the token fingerprints have gone thin.

Tree comparison is expensive, so tools mostly use greedy matching rather than true tree edit distance. GumTree, published by Falleri and colleagues in 2014, is the reference implementation a lot of projects borrow from. JPlag 4.x, re-released in 2023 on Java 17 with a Docker image, now covers fifteen-plus languages and reports similarity as a percentage of matched tokens, with a default minimum match of around nine tokens.

Dolos came out of TU Delft around 2021 and sits in a similar place. Open source, tree-based fingerprinting, reasonable language coverage. Nothing against either tool. Both are good at the question they answer.

The problem is the question.

Side-by-side code comparison in Codequiry showing a 91% match between two student submissions
Side-by-side comparison: Codequiry lines up matching code between two submissions, with confirmed and false-positive review labels.

What Stack Overflow and GitHub did to the problem

Cohort comparison made sense when the only copies of an assignment were the ones your students made. That assumption died in 2008, the year Stack Overflow and GitHub both launched.

Ask a room of 200 CS professors where their students copy from and almost none will say "each other." They'll say GeeksforGeeks, a YouTube tutorial, a blog post from 2013, a GitHub repo with an MIT header. A tool that only compares submissions against each other looks at a copied assignment, finds zero internal matches, and hands you a clean report.

MOSS added a web crawl years ago, but it's shallow, and coverage isn't something you can inspect or audit. For academic work that's tolerable. For enterprise code, it's a liability question.

Think about the contractor scenario, because it's the same problem wearing a tie. You pay an agency for a service layer. They deliver working code on time. Six months later your legal team discovers three files came from a GitHub repo under GPLv3 with the headers stripped and the license file deleted. The vendor's "original work" clause is worth exactly what your detection process is worth.

Codequiry web results tracing a submission to a Stack Overflow question with line and token counts
Tracing code to its source: a submission matched to a Stack Overflow answer, down to lines and tokens.

Why AI-generated code broke the old model

GitHub Copilot hit technical preview in June 2021 and went generally available in June 2022. ChatGPT launched on November 30, 2022. Between those dates, the plagiarism problem inverted. Stack Overflow's 2024 developer survey put active AI tool use at 62% of developers, up from 44% a year earlier.

Every technique I've described assumes at least two copies exist. Winnowing needs a fingerprint to match. AST comparison needs a second tree. Web crawling needs a source page. An LLM generates code that exists in exactly one place, never before printed, in your student's exact variable naming style because that's how it was prompted.

Run MOSS on a cohort where half the submissions came from ChatGPT and you get a beautiful report. Low similarity across the board. No matches. Nothing to review.

That's the trap. A 0% plagiarism score is a statement about copying, not about authorship. There's nothing to copy.

So what does the code actually look like? Once you've read a few hundred submissions, the patterns stop being subtle. Uniform docstrings on every function, including trivial ones. Comments that restate the function name in complete sentences. Defensive try/except blocks around operations that cannot fail. Type hints and error handling in a first-week assignment that has covered neither. Even indentation that's too even.

None of that is proof. Statistical detectors score machine-written code on signals like perplexity, which is how predictable each token is given the ones before it, and burstiness, which is how much that predictability varies. Human code is bursty. We write a clever line, then a boring one, then stare at the wall. LLM output is smooth in a way that shows up numerically, and a newer generation of tools reads those numbers.

Where it breaks down matters as much as where it works. Strong students write clean code. Anyone running Black or ESLint in a pre-commit hook gets flagged more often than they should. Boilerplate, generated config, and copy-pasted test scaffolding all look machine-authored. I'd be skeptical of any tool claiming accuracy above the low 90s on real classroom data, and I'd want its false positive rate on the top quartile of the class before letting it make a decision on its own. We haven't stress-tested this past a few hundred submissions per course, and I'd want that number before running it across a 2,000-seat intro sequence.

Codequiry AI code detection report with average AI score, highest file score and a risk distribution
AI code detection: probability scores per file, flagging submissions likely written by ChatGPT, Copilot, Claude or Gemini.

Layered detection, the only thing that holds up

Three questions, three different failure modes, three separate signals:

  1. Does this submission match another submission in the cohort?
  2. Does it match something published on the web or in a public repository?
  3. Does it look machine-generated?

A single tool answering one of those will miss the case covered by the other two. The professor who only runs MOSS misses the Stack Overflow copy and the ChatGPT rewrite. The professor who only runs an AI detector misses the student who copied a classmate and lightly paraphrased. The CTO who only scans dependencies misses the copied function a contractor pasted into the service layer.

This is where the Codequiry vs MOSS comparison gets practical rather than theoretical. MOSS is free and excellent at token matching inside a cohort. It has no AI layer, no persistent API, and it deletes submissions after roughly two weeks, which means you can't re-run last semester's cohort when someone finally talks. JPlag is strong on structural matching and has no web or AI layer either. Turnitin brought AI detection to text in April 2023, but its code support is thin and its corpus is document-shaped, not repository-shaped.

A source code plagiarism checker that runs token, AST, web, and AI analysis over the same submission, and keeps the results, gives you something none of those do individually: a single integrity record per assignment you can open again in March when a pattern shows up. Codequiry runs peer similarity, web sources, and AI generation detection across 65 languages, with a dashboard for courses and a REST API plus CLI so the same pipeline can sit in a CI job for an engineering team. Its AI code detector is worth looking at separately if machine authorship is the specific problem you're chasing.

Codequiry insights score breakdown separating peer similarity, web similarity and AI generation, with match sources
The score breakdown: peer similarity, web similarity and AI probability reported separately, with where the matches came from.

Fifty years in, the technology hasn't gotten smarter about the question. It has just learned to ask more of them at once, which is the only thing that ever worked.

Frequently asked questions

Can MOSS detect AI-generated code?

No. MOSS compares submissions against each other and against a web crawl, looking for shared fingerprints. LLM output is generated fresh, so there's nothing to match. A clean MOSS report on ChatGPT-written code is expected, not reassuring.

Why does a plagiarism checker report 0% for copied code?

Usually one of three reasons. The source they copied from isn't in the tool's comparison corpus, whether that's a private repo, a Discord study group, or an older cohort. Or the tool only does text and line matching and the student renamed enough to break it. Or the code was generated rather than copied, in which case there was never an original to find.

What's the difference between token-based and AST-based code similarity?

Token-based tools such as MOSS flatten source into a token stream and look for shared subsequences. They're fast and they survive renaming and reformatting. AST-based tools such as JPlag and Dolos parse the code and compare tree structure, which catches rearranged control flow and code moved into helper functions. Serious detection runs both.

Do AI code detectors produce false positives?

Yes, and the rate depends heavily on your population. Heavily linted code, boilerplate, and work by strong students all score higher than they should. Treat the score as a routing signal that decides which submissions a human reads closely, never as a verdict by itself.

If you're running this across a course or a codebase and want the three signals in one report, start with a code plagiarism checker that covers peer, web, and AI detection together.