Cross-language code plagiarism detection is the practice of comparing two submissions written in different programming languages and asking whether one is a translation of the other. It is not a text comparison. Two files, one in Java and one in Python, may share zero identical lines and still carry the same algorithm, the same helper functions, the same wrong edge case. The detection problem is to see the shared structure through the syntax.
I want to describe how that works, because most of the tools we reach for in a grading workflow were built for a narrower job.
In the fall of 2019, my department let students in CS 1 submit one assignment in either Java or Python. The intent was good. We had a handful of transfer students who had learned Python first and we wanted to be kind to them. Two submissions came in that looked nothing alike. One was Java, one was Python. I ran both through MOSS in separate runs, which is the only way MOSS works, and got silence. I ran my own line diff. Zero percent. I then spent a Friday evening with a legal pad, hand-normalizing both files onto a single page of token categories. They agreed at roughly 94 percent. That evening is why I now have opinions about cross-language detection.
What counts as cross-language code plagiarism
The research literature calls these semantic clones. In 2007, Chanchal Roy and James Cordy at Queen's University published the taxonomy that most of the field still uses, and it is worth having in your head before you look at any tool's output.
| Clone type | What differs | Detectable by |
|---|---|---|
| Type 1 | Whitespace, comments, formatting only | Line diff, any text tool |
| Type 2 | Identifiers, literals, types renamed | Token normalization |
| Type 3 | Statements added, removed, or reordered | Token or AST with tolerance for gaps |
| Type 4 | Same behavior, different syntactic form | Structural or semantic comparison |
A translated submission is a Type 4 clone by construction. It is not a renaming exercise, because the target language does not have the same constructs. A Java for (int i = 0; i < n; i++) becomes for i in range(n). An array of arrays becomes a list of lists. The loop is the same loop, and no token survives the trip unchanged except the variable name, which the student usually keeps.
That last detail matters more than it sounds. Students who translate rarely rename everything. They rename some things, forget others, and leave a residue of identifiers that do not belong in the target language's idiom. That residue is often the best evidence you will get.
Why winnowing and token matching cannot cross a language boundary
The dominant technique in academic code comparison is winnowing, described by Saul Schleimer, Daniel Wilkerson, and Alex Aiken in their 2003 paper on fingerprinting. The idea is compact. Convert the source into a stream of tokens, hash every window of k consecutive tokens, then slide a larger window over those hashes and keep only the minimum hash in each window. The surviving hashes are the document's fingerprints, and they are robust to reordering, insertion, and whitespace. MOSS, the Stanford tool that has been running as a web service since the mid-1990s, is a winnowing implementation at its core.
Winnowing solves Type 1 and Type 2 detection beautifully. It has one structural limit: the token stream has to come from a lexer that understands exactly one language. When you invoke moss -l java -m 10 *.java, you have committed to Java for that run. There is no way to add a Python file to a Java run and have anything meaningful come out, because the hash inputs live in different token spaces.
JPlag, developed at the University of Karlsruhe (now KIT) since 1996, made an important refinement. Its front ends parse each language and emit a generic token type, so a Java loop and a C loop produce the same token category. Version 5, released in early 2023, rewrote the core in Rust and moved to tree-sitter grammars, which made adding languages much cheaper. Dolos, from the programming languages group at KU Leuven, takes the same tree-sitter approach with a TypeScript core and a polished web front end. All three are excellent tools and all three are, by default, per-language.
The limitation is not a bug in any of these tools. It is a design boundary. A winnowing index is a dictionary of one language's token vocabulary, and cross-language comparison requires a dictionary that belongs to no language in particular.
JPlag does advertise some cross-language capability, and it can produce output when the two languages map onto compatible token categories, though the maintainers are careful about what the numbers mean and I would treat a cross-language JPlag score as a hint rather than a finding. The mapping is coarse, and the further apart the languages sit, the less of the signal survives.
One practical flag worth knowing if you use JPlag directly: -t sets the minimum number of matched tokens. The default is 9. On a short exercise, that threshold is low enough that two students who both wrote the same three-line input validation routine will register a match. Raise it before you read the report, not after.

What cross-language detection actually compares
There are three families of technique in use, and mature platforms tend to run more than one.
Normalized token streams
Replace every identifier with a single placeholder, replace every numeric literal with another, keep keywords and operators. A Java method and its Python translation then produce token sequences that can be diffed directly. A stripped-down version of the normalization looks like this:
KEYWORDS = {"if", "else", "for", "while", "return", "break", "continue"}
def normalize(tokens):
out = []
for kind, text in tokens:
if kind == "ID":
out.append("V") # any identifier
elif kind == "NUM":
out.append("N") # any numeric literal
elif kind == "STR":
out.append("S")
elif kind == "KW" and text in KEYWORDS:
out.append(text.upper())
else:
out.append(text)
return out
This is honest about its own weakness. It catches renaming, and it catches a fair amount of translation, but it throws away so much structure that unrelated code can start to look similar. Every loop in every language normalizes to the same four tokens.
Language-agnostic intermediate representations
The stronger approach builds a parse tree for each language and then maps language-specific node types onto a small shared schema. A Java EnhancedForStatement, a Python for_statement, and a C++ range-based for all collapse to the same abstract node. Control flow, call structure, and nesting depth survive the mapping. Identifier names do not, unless you deliberately keep them as an additional signal.
This is where cross-language detection starts to earn its keep. The research has followed the same path. CLCDSA, published in 2020, used syntactic and semantic features across languages for clone detection. C4, from Microsoft Research and presented at ICPC in 2022, trained a contrastive model to detect cross-language clones directly. Both are answering the question a professor asks every term: is this the same program wearing different clothes?
Learned embeddings
The newest family embeds a whole function into a vector using a code language model (CodeBERT in 2020, GraphCodeBERT in 2021, CodeT5+ in 2023, and the StarCoder2 family since 2024) and compares the vectors. This is genuinely language-agnostic in a way the other two are not, because the pretraining corpus is multilingual. It also has a calibration problem I will come back to. Everything in a tightly specified assignment is semantically similar to everything else. Cosine distance of 0.86 sounds alarming until you notice the cohort median is 0.81.
A worked example with two languages
Consider a pair of submissions from a data structures exercise. Here is the Java version.
public static int[][] build(int[] data, int n) {
int[][] grid = new int[n][3];
int t2 = 0;
for (int i = 0; i < n; i++) {
if (i >= 0 && i < n) {
int v = data[i % data.length];
grid[i][0] = v;
grid[i][1] = (v * 37) % 1000;
grid[i][2] = t2;
t2 = t2 + v;
}
}
return grid;
}
And here is what came in from a second student, in Python.
def build(data, n):
grid = [None] * n
t2 = 0
for i in range(n):
if i >= 0 and i < n:
v = data[i % len(data)]
grid[i] = [v, (v * 37) % 1000, t2]
t2 = t2 + v
return grid
A line diff reports nothing. MOSS, run per language, reports nothing, because each file is alone in its run. Now look at what the two files share.
The accumulator is named t2 in both. That is not a name anyone reaches for naturally in either language. The magic constant 37 appears in both, and no specification mentioned 37. Both apply a redundant guard i >= 0 && i < n inside a loop that is already bounded by n, which is a defensive habit, not a requirement, and an unusual one to share. Both use modular indexing against the input length. Both return a three-slot row structure that collapses the language's natural nested type in the same way.
Normalize both token streams and the leading sequence is identical for the first thirty tokens. That is the signal. It is not proof on its own, and I would not take it to an honor board alone, but it is a specific, inspectable set of coincidences, and that is what an integrity case needs.

The translation problem after 2023
Before large language models, cross-language copying was rare, because it was expensive. Translating 120 lines of Java into Python by hand takes a competent student the better part of an hour, and the result is usually awkward in ways that give it away. Now it takes one prompt. A student pastes a peer's submission into ChatGPT, Claude, or Gemini and asks for a translation. The output arrives in seconds.
There is a detail here that cuts in an unexpected direction. Language models tend to normalize as they translate. They replace a hand-rolled loop with a comprehension, swap a manual accumulator for sum(), rename awkward variables. That normalization can lower the raw structural similarity between the two submissions while preserving the semantics entirely, which pushes the case further into Type 4 territory, exactly where the weakest detectors operate.
It also means the AI signal and the plagiarism signal are not the same signal. A translated submission may score low on peer similarity against every other student in the cohort, because it is textually unique, and moderate to high on AI generation. Neither number tells you what happened. If you are running one check and not the other, you are reading half a report.
This is one place where a platform that reports peer similarity, web source matches, and AI generation in a single view saves real work, because the three numbers have to be read together. Codequiry is built around that combined view rather than treating AI code detection as a separate product bolted onto a similarity engine.

False positives, base rates, and why the cohort matters
The hardest part of cross-language comparison is not producing a score. It is knowing what score means something.
Consider an intro assignment that says "implement an iterative Fibonacci function that returns the first n values." Every correct submission in the room is going to be highly similar under any normalized representation, because there is essentially one shape the solution can take. Now add a requirement that submissions may be in Java or Python. The cross-language comparison will flag pair after pair, and every one of them will be a false positive.
The fix is to stop reading absolute scores and start reading relative ones. Compute the pairwise similarity for the entire cohort, look at the distribution, and pull out the outliers. If the median pair in a 60-student section agrees at 71 percent and one pair agrees at 96 percent, that pair is worth an hour of your time. If the whole section sits between 88 and 97 percent, your assignment is over-specified and no threshold will help you.
Cross-language similarity scores are rankings, not verdicts. A score is only meaningful next to the distribution it came from.
This is a design decision you can make in the assignment itself, and I would argue you should make it before you make it in the tooling. Open-ended requirements produce spread-out distributions, and spread-out distributions are easy to review. Assignments with a single natural solution produce tight distributions, and tight distributions produce arguments in honor board hearings. I have sat through two of those.

A review workflow for instructors
- Run the per-language pass first. It is cheap, it is precise, and it catches the ordinary cases so you can spend your attention on the hard ones.
- Then run everything in one cohort. Cross-language comparison only works when both files are in the same comparison set, which is the mistake I made in 2019. Tools that organize by course and assignment rather than by individual file make this the default rather than a step you have to remember.
- Sort by cohort outlier, not by raw similarity. The top of the list is where the evidence is. The middle of the list is where the base rate lives.
- Read the evidence in the source language. Open the pair side by side and look for the residue: identifiers that do not belong, constants the assignment never specified, helper functions with the same decomposition as the partner's submission.
- Ask about a specific line. Not "did you write this" but "walk me through why the guard on line 6 is there." A student who wrote the code will tell you. A student who translated it will stall, because the line does not have a reason in the target language. It only has a reason in the source language.
That last step is the one that converts a similarity score into a finding you can defend. Detection narrows the field. The conversation closes it.
Where cross-language detection genuinely breaks down
I want to be honest about the limits, because the limits are also where your judgment has to take over.
Cross-paradigm comparison is close to hopeless in the current generation of tools. Comparing a Haskell submission against a Python submission, or Prolog against C, means the two programs do not share control flow, do not share data representation, and often do not share decomposition. A recursive Haskell solution and an iterative Python solution to the same problem may be semantically identical and structurally unrecognizable.
Very short submissions are also unreliable. Under roughly forty tokens there is not enough structure to fingerprint, and boilerplate from a starter file will dominate whatever comparison you run.
Heavily edited translations defeat most structural detectors, and this is where the field is still moving. If a student takes a translated solution and spends twenty minutes restructuring it, the syntactic evidence thins out considerably. What tends to survive is semantics and behavior, which is why the embedding-based approaches and paired AI detection matter more each year than they did in 2019.
And there is one more case that tools cannot resolve, because it is not a technical question. A student who learned an algorithm from a lecture in Java, then went home and wrote it in Python from memory, has produced a Type 4 clone of the lecture. That is called learning. The line between that and translation of a peer's file is not a similarity score. It is a set of shared idiosyncrasies that have no reason to exist in two independently written programs. Your job is to find those, and the tool's job is to put them in front of you.
What to put in the syllabus before next term
Most academic honesty policies were written before machine translation was free and instant, and many of them say nothing about translating another student's work into a different language. That gap is worth closing in writing, because it comes up in hearings.
A clause that has worked in my department reads roughly like this: submitting a translation of another person's program into a different programming language is plagiarism, whether the translation was performed by hand or with the assistance of a tool. Writing your own solution in a second language, from your own understanding of the problem, is not.
The second sentence is as important as the first. Without it, you have written a policy that punishes students for practicing, which is a poor trade and one that a decent honor board will push back on.
If your department runs a shared submission portal, consider whether it collects a single language per assignment or accepts multiple. A portal that keeps everything in one comparison set makes cross-language review a routine step. A portal that sorts by language file extension, and most do, quietly makes it impossible. For departments running several courses with different language requirements, a code plagiarism checker for teachers that handles mixed-language cohorts in one pass removes an entire category of manual work.
The larger point is not that translation is a new crime. It is that the cost of translating a program dropped from an hour to a few seconds, and any detection method that assumes the old cost is now miscalibrated. Cross-language comparison is not a specialty feature anymore. It belongs in the default workflow, next to the same-language pass and the AI check.
Frequently asked questions
Can MOSS detect code copied between two different programming languages?
Not directly. MOSS compares submissions within a single language run. If you have Java submissions and Python submissions that you suspect are related, you would need a tool that normalizes both into a shared representation before comparison, or a workflow that pairs them manually.
Is translating a classmate's code into another language considered plagiarism?
At most institutions, yes, and it is increasingly written into policy explicitly. The submission is a derived work of someone else's program, and the fact that the surface text differs does not change that. Writing your own solution in a language you are still learning is a different act and should be treated differently.
How accurate is cross-language code plagiarism detection?
Within a family of related languages (Java, C#, C++) it is strong, and normalized token or AST comparison catches most translations. Across distant paradigms (Haskell versus Python) it degrades badly. The reliable signal is always the cohort outlier plus the shared idiosyncrasies you can point to in the source.
Does AI detection catch translated code?
Sometimes, and the two signals are complementary rather than redundant. A machine-translated submission often shows AI-generation indicators even when it is textually unique against the rest of the cohort. Running a source-level AI check alongside the similarity pass is the practical answer.
The short version
Cross-language plagiarism is a semantic clone problem, and semantic clones are the hardest class in the taxonomy. Winnowing and token matching, which handle the other three classes well, cannot see across a language boundary because their indexes are built for one language. What works is a normalized representation: places, node types, and structural fingerprints that do not care which language produced them. What confirms it is a person reading the two files side by side and finding the identifiers, constants, and odd guard clauses that could not have arisen twice.
Run the comparison on the whole cohort, sort by outlier, and read the evidence in the source. If you want a workflow where per-language, cross-language, web source, and AI generation checks all land in one report, that is what a code plagiarism checker built for mixed-language courses is for.