Two submissions landed in the same week of my Spring 2024 backend cohort, both implementing a token bucket rate limiter in Python. They scored 4% against each other. Both scored in the 70s against a Java gist from 2017 that I found by pasting one line of a docstring into GitHub search.
That's the shape of the problem. Cross-language plagiarism detection is not a harder version of regular similarity scoring. It's a different problem, and most of the tooling sitting on a university server right now was never built for it.
The short answer: yes, a Python submission can be traced back to a Java repository, but not by comparing the two files textually. You trace it through the things translation doesn't change, string literals, numeric constants, control-flow shape, and the order of API calls, and by running structural comparison rather than string comparison. If a student used an LLM to do the translating, the AI detector is usually the fastest signal you have.
Why cross-language plagiarism detection breaks ordinary similarity scoring
Every mainstream code similarity tool reduces a program to some representation, then compares representations. The differences between tools are almost entirely about which representation they pick, and what that representation throws away.
MOSS, Alex Aiken's Stanford tool that's been running since the mid-1990s, uses k-gram winnowing. The 2003 paper by Schleimer, Wilkerson, and Aiken gives the guarantee precisely: if two documents share a substring of length at least t, the algorithm will report a match. Below t, all bets are off. That guarantee is beautiful and it is also per-language, because the k-grams come from the raw token stream.
JPlag, originally Guido Malpohl's work at Karlsruhe in the late 1990s and now maintained by the Plagiarism Analysis group at KIT, takes a different route. It parses each submission with a language front-end (the modern Java rewrite uses ANTLR grammars), emits a token sequence, and runs Greedy String Tiling, Michael Wise's 1993 algorithm, to find the longest non-overlapping common subsequences above a minimum match length. The 2002 benchmark by Prechelt, Malpohl, and Philippsen showed JPlag handling renamed identifiers and reformatted code far better than the string-based tools of the era, which is exactly why it's still around.
Both designs assume a shared alphabet. Feed MOSS a directory of .py files and a directory of .java files in one run with -l python and it will compare the Python against the Python and the Java against the Java. There's a well-known hack where you set -l text or -l ascii and let MOSS chew on raw characters, which I have watched people do, and which produces a wall of noise I've never once seen turn into a defensible case at a hearing.
What survives a hand translation
Here's a real pair. The Java is from a public gist, the Python is a student's week-3 submission, lightly anonymized.
public class Luhn {
private static final int RADIX = 10;
public static boolean validate(String cardNumber) {
if (cardNumber == null || cardNumber.length() < 12) {
throw new IllegalArgumentException(
"card number too short: " + cardNumber.length());
}
int sum = 0;
boolean alternate = false;
for (int i = cardNumber.length() - 1; i >= 0; i--) {
int digit = Character.digit(cardNumber.charAt(i), RADIX);
if (alternate) {
digit *= 2;
if (digit > 9) digit -= 9;
}
sum += digit;
alternate = !alternate;
}
return sum % 10 == 0;
}
}
RADIX = 10
def validate(card_number):
if card_number is None or len(card_number) < 12:
raise ValueError(f"card number too short: {len(card_number)}")
total = 0
alternate = False
for i in range(len(card_number) - 1, -1, -1):
digit = int(card_number[i], RADIX)
if alternate:
digit *= 2
if digit > 9:
digit -= 9
total += digit
alternate = not alternate
return total % 10 == 0
Run these through a token-level comparator and you get a mess. Java keywords become Python whitespace conventions. int sum = 0 and total = 0 are two tokens versus one. boolean alternate = false has no Python equivalent at all. My tokenizer sees maybe 30% overlap, most of it coming from punctuation that means different things in each language.
Now look at what actually carried over.
The threshold is 12. Most Luhn implementations use 13, or 16, because that's the ISO/IEC 7812 length for a primary account number. Twelve is unusual, and it appears in both files. The named constant RADIX = 10 survives as RADIX = 10, uppercase, same value, same position above the function. The error message is "card number too short: " in both, down to the colon and the trailing space. The flag variable is called alternate in both, which is not the name I'd have picked out of the air. And the loop walks backward from the last character with an inverted boolean toggled at the bottom, which is a legitimate way to write Luhn but not the only one.
None of those are formatting. They're authorial decisions, and translation preserves them because a translator is thinking about logic, not about what the code looks like.
String literals and numeric constants are the closest thing to a fingerprint that a program has. A human translating code will rewrite every variable name and leave the error message untouched, because the error message isn't code to them.
This is why the useful cross-language comparison operates on a normalized structure, not on tokens. Strip identifiers to positional placeholders, keep literals, canonicalize control flow into a language-neutral skeleton, then align. The skeleton of the Java above and the skeleton of the Python are the same graph: null-check, length-check, throw, reverse loop, conditional doubling, conditional subtraction, accumulation, toggle, modulo.

The three ways students translate, and what each leaves behind
I've now looked at enough of these cases to sort them into three buckets, and the buckets matter because each one needs a different detector.
Hand translation
A student reads a Java implementation, understands it well enough to port it, and writes the Python themselves. This is the regime where structure beats text every time. Identifiers get renamed, comments get rewritten or dropped, and the string literals survive. Constants survive. Unusual loop shapes survive. If the source was a public repository, you can often confirm it by searching for the literal.
The tell I look for is a mismatch between the code's fluency and the student's demonstrated fluency. A week-3 student who writes alternate = not alternate with an inverted boolean on a reverse loop, and who cannot explain at office hours why they started from the end of the string instead of the beginning, has told you something.
LLM translation
This is the fastest-growing bucket and the most interesting one technically. Paste Java into Claude or GPT-4 class models and ask for Python, and you get faithful, idiomatic output. Faithful enough that it preserves the constant, the literal, and often the structure of the original comments. It's genuinely a good translation.
Which is the problem for the student. A hand translation has a human's stylistic noise baked in: inconsistent naming, a stray debug print, a slightly different loop construct because they got bored. An LLM translation is uniformly fluent. Every function has the same voice. Comments are grammatical and evenly distributed, which is not how students write comments at 2 a.m.
That uniformity is a statistical signature, and it's what the AI code detector side of a scan is looking for. Machine-translated code carries the same low-perplexity, low-variance markers as directly generated code, sometimes more strongly, because the model had a rigid source to follow and produced an even smoother output than it would have from a blank prompt.
This is the case where stacking matters. A structural comparator tells you the Python is suspiciously similar in shape to something. The AI detector tells you a model probably wrote it. Around 2019 you needed the first signal and couldn't get the second. In 2025, a lot of the time, the second one arrives first.

Mechanical transpilation
There are real tools for this. py2many translates Python into C++, Rust, Go, Julia, and Java. c2rust moves C to Rust via LLVM IR. For JVM work, you can compile to bytecode and decompile with something like CFR or Procyon and get a plausible Java file out the other side. LLVM's own IR is a shared target that C, C++, Rust, Swift, and Julia can all reach.
Transpiled output is ugly in a very specific way. Type annotations everywhere that don't belong, helper functions with generated names, casts that a human would never write. In my experience it gets caught by eye before any tool sees it. I don't have numbers on how often this route shows up in student work, because of that, and I'd be suspicious of anyone who claims to.
Detector classes ranked by what they actually cost you
If you're evaluating tools for a department, this is the comparison that matters. Everything below is a real technique with real implementations, ranked roughly by how much infrastructure it takes to run at semester scale.
| Approach | Catches across languages | Breaks down on | Cost at 500 submissions |
|---|---|---|---|
| Winnowing fingerprints (MOSS) | No, per-language token streams | Any translation, wholesale renaming | Seconds; MOSS is fast |
| Token + Greedy String Tiling (JPlag) | Limited; needs a shared token vocabulary | Languages with different keyword sets | Seconds to a minute per language set |
| AST hashing (Deckard-style characteristic vectors) | With a hand-built bridge | Idiomatic differences in how languages express loops | Minutes; tree construction dominates |
| Intermediate representation comparison (LLVM IR, JVM bytecode, CPython bytecode) | Yes, within an IR family | Cross-family pairs like C++ vs Python | Minutes, plus a build step |
| Program dependence graphs (control + data flow) | Yes, in principle, best semantic fidelity | Dynamic languages where you can't resolve types | Hours; points-to analysis is expensive |
| Code embeddings (CodeBERT, GraphCodeBERT, UniXcoder, jina-code) | Yes, no bridge needed | Precision; every bubble sort is near every other bubble sort | GPU inference, one forward pass per file |
The embedding row is the one people get excited about and the one I'd caution against relying on alone. CodeBERT arrived in 2020, GraphCodeBERT in 2021, UniXcoder in 2022, and the retrieval-oriented code embedding models from 2024 are noticeably better. They will absolutely cluster a translated solution near its source. They will also cluster a translated solution near every other correct solution to the same assignment, because the assignment is the same and the model is measuring semantics. Cosine similarity of 0.94 across a cohort is not evidence of anything. We haven't pushed embedding comparison past a few hundred submissions in anything I'd call a production setting, so treat my numbers there as soft.
The academic work on this is real and worth reading if you're building something. Bui and colleagues published CLCDSA in 2019, which trains a classifier on AST-derived features plus API call sequences to detect clones across languages. Microsoft's MISIM work in 2020 went after semantic similarity of snippets with a learned model. Neither is packaged as a tool you can point at a Canvas export, which is the gap.
The false positive floor is higher than you think
Here's the part nobody wants to hear. Cross-language comparison has a structurally higher false positive rate than same-language comparison, and no amount of engineering makes it go away entirely.
Consider a counting sort in C++ and a counting sort in Python written independently by two students who have never met. The algorithm has one shape. You allocate an array of size k, count occurrences, compute prefix sums, build the output. There is no room for stylistic variation in a correct counting sort, and the maximum stack depth is three loops. A structural comparator looking at a normalized control-flow graph will see the same graph twice.
Same thing for the standard quicksort with Lomuto partitioning, or a breadth-first search over an adjacency list, or a Dijkstra with a binary heap. These are canonical. If your assignment is "implement Dijkstra," you have set an assignment that is structurally identical across every correct solution in every language, and cross-language comparison will produce noise on it.
Two things follow from that. First, run cross-language comparison as a filter that surfaces candidates, never as a verdict. Second, design assignments with structural slack: give students a choice of input format, require a particular error-handling policy, ask for a specific output ordering that isn't the obvious one. The variation you introduce at the assignment level becomes the signal you look for at the detection level.
A workflow I actually run
This is what the week looks like for the 90-student cohort, and it runs on one afternoon if I don't get distracted.
- Compile a cross-reference corpus. I keep every prior-semester submission in the same language, plus one translated baseline per assignment. The baseline is a reference solution I've ported into the three languages students commonly arrive with: Java, C#, and JavaScript. It anchors the "this is what the algorithm looks like when nobody copied anything" end of the scale.
- Run the peer check first. Same-language, within the current cohort. Fast, and it catches the majority of cases, because most copying is still within a class.
- Run the web and repository check. This is where cross-language cases usually break open, and it's counterintuitive. The Python and the Java gist don't match textually, so a text comparison reports nothing. But the gist is indexed, and the constants and literals inside it are searchable. When a student pulls the docstring or the error string along for the ride, the web pass finds the source even though the similarity pass found nothing.
- Run the AI pass last. It's the slowest to interpret and the one most likely to need a conversation. Anything above a high score gets a manual read, and I look for the uniformity markers: consistent comment density, no debug leftovers, no stylistic drift between functions.
- Then talk to the student before writing anything down. Always. Every process here produces candidates, and I've been wrong.

Running those as three separate tools with three separate output formats is how this becomes a job nobody wants. Which is the practical case for a source code plagiarism checker that reports peer similarity, web matches, and AI scores as three columns on one row instead of three PDFs. Codequiry does that, and it also does it in a dashboard that a TA can read without a training session, which for a department of sixty adjuncts is not a small thing. If you've been running MOSS through a shell script since 2011, the side-by-side Codequiry vs MOSS comparison is the honest version of that conversation, including where MOSS still wins on cost.
What to actually change about your assignments
A few things I've landed on after four years of this, in rough order of how much they helped.
Require a git history. Not as an integrity check, as a pedagogical one. A student who hand-translated a Java gist into Python and then committed it in one shot at 11:47 p.m. has no draft commits, no failing tests, no rename. A student who wrote it has forty commits with names like "oops" and "why is this off by one." The history is more informative than the diff, and it costs you nothing to require.
Run a short oral check on any submission that scores high. Not an interrogation, just "walk me through the loop." The gap between a student who wrote the code and a student who translated it is enormous and immediately visible in the first thirty seconds.
Publish the policy in the syllabus, in plain language, and include translation explicitly. Most academic integrity policies I've read say nothing about porting code between languages, which leaves a student who genuinely doesn't know whether translating a Stack Overflow answer is allowed with no answer. Say it. "Translating code from another language is the same act as copying it" is one sentence.
And keep the reference corpus. Every semester you run a check, you add to next semester's baseline. That's the compounding advantage that commercial detectors have over a homegrown script, and it's the reason I stopped maintaining my own winnowing implementation around 2021.
Frequently Asked Questions
Can a student avoid detection just by translating code into another language?
Not reliably. Translation removes the text similarity that winnowing and token comparison depend on, but it preserves string literals, numeric constants, control-flow shape, and identifier patterns. A structural comparison or a web-source check against the original repository typically surfaces the case. If the translation was done by an LLM, the output also carries machine-generation signatures that AI detectors pick up.
Does MOSS support cross-language comparison?
No. MOSS compares submissions within a single language specified by the -l flag. There's a workaround where you set it to text or ascii mode and feed it raw files from multiple languages, but in practice this generates enormous amounts of noise and I've never seen it produce a usable case. JPlag is the same: one language per run.
What's the highest-signal indicator that a submission was translated from another language?
Preserved string literals and unusual numeric constants. A human translator rewrites every identifier because identifiers are code to them, but leaves the error message and the magic number alone because those feel like content. If a Python submission contains an error string that matches a Java repository, you have something concrete to work with, and it's the kind of evidence that holds up when a student appeals.
How do I check a whole cohort without running three separate tools?
The practical answer is a single platform that runs peer, web, and AI passes in one job and reports them together, so a TA can look at one row per student and sort by risk. That's what Codequiry was built to do, and it handles the 65-language spread you get when half your cohort arrives from a Java sequence and half from JavaScript.
If you want to see what the cross-language picture looks like on your own submissions, upload a batch and look at the web-match column before you look at anything else. It's usually where the real story is.