How a Lecturer Catches Code Translated Between Languages

Cross-language code plagiarism detection is the practice of comparing programs that solve the same problem in different programming languages. Most of the tooling in a CS department's stack was never built for it. MOSS, JPlag, and Dolos compare submissions within a single language, so a Python file gets scored against other Python files and a Java file against other Java files. When a student submits a translation of somebody else's Java solution, the similarity score can land in single digits while the underlying logic is identical.

Marta Oyelaran has run the second course in the programming sequence at a mid-sized public university for eleven years, and she keeps a spreadsheet of what her tooling misses. In the spring of 2024, her 180-student cohort submitted in two languages: Java for the morning lab sections that had access to the campus build server, Python for the evening sections that did not. The first pass through a code plagiarism checker came back clean. Then she ran the Python directory against the Java directory by hand, and two submissions matched once she stripped the syntax away.

// Java, submitted by student A
public static int depth(Node root) {
    if (root == null) return 0;
    return 1 + Math.max(depth(root.left), depth(root.right));
}
# Python, submitted by student B three weeks later
def depth(root):
    if root is None:
        return 0
    return 1 + max(depth(root.left), depth(root.right))

That pair is the easy case, and it is also the rarest one. Matching identifiers and matching structure together are the fingerprint of a hand translation, and hand translation is close to extinct. Oyelaran's spreadsheet holds forty-one entries across four years. Only nine of them looked anything like the example above.

Why ported code slides past MOSS and JPlag

MOSS has been the default since 1994, when Alex Aiken's group began distributing it. The fingerprinting method underneath it, described in Schleimer, Wilkerson, and Aiken's "Winnowing: Local Algorithms for Document Fingerprinting" at SIGMOD 2003, hashes k-grams of the token stream and keeps a sampled subset. It is fast, free, and indifferent to whitespace, comments, and most identifier renaming.

It is not indifferent to language. MOSS partitions a mixed directory before it compares anything, and there is no flag that turns that off.

JPlag, the Karlsruhe tool that many departments adopted as a friendlier alternative, works from a token sequence produced by a language-specific parser. Version 4, released in 2022, rewrote the architecture and broadened coverage to Java, C, C++, Python, C#, Go, Kotlin, Rust, Scala, Scheme, Swift, JavaScript, and TypeScript. Every one of those is a separate comparison space. Dolos, from TU Delft's 2021 work, follows the same shape with better visualizations and an interactive pair view; its documentation is direct about the constraint.

The practical consequence is arithmetic. A careful port keeps somewhere between 15 and 25 percent of its token overlap with the original, and most of what survives is standard library calls and keywords. That is the same band where unrelated submissions to the same assignment live. A merge sort assignment produces a lot of genuinely similar merge sorts. Translation pushes a copied submission down into the noise, and the noise is where the threshold sits.

Identifier conventions help the disguise along. Java's camelCase becomes Python's snake_case. Type annotations evaporate. A for loop with an index becomes a comprehension. A student who ports without thinking about it produces a file that no token matcher will flag, and they may not even know they did something the honor code covers.

Side-by-side code comparison in Codequiry showing a 91% match between two student submissions
Side-by-side comparison: Codequiry lines up matching code between two submissions, with confirmed and false-positive review labels.

What the translation step looks like in 2025

Before 2023, a student who translated someone else's solution did it by hand and left traces: a variable named numItems sitting in a Python file, a comment reading // initialize the array above a list comprehension. Lecturers learned to read those traces the way copy editors learn to read a sudden change in voice.

Those traces are gone. A prompt along the lines of "convert this Java class to Python, keep the logic identical" is enough. The output is idiomatic, docstringed, type-hinted, and structurally nothing like the original at the surface level. It also arrives in seconds, which removes the main practical deterrent, the fact that hand-porting a whole assignment is more work than writing it.

What the model leaves behind is a different kind of signature, and it is the same signature that shows up in code the model wrote from scratch. Uniform formatting. Defensive error handling the assignment never asked for. A helper function extracted for no reason the problem statement implies. A docstring on a twelve-line file. A good AI code detector reads those patterns as a statistical profile across a whole submission rather than as any single smoking gun, which is the right way to read them.

"The tell used to be the mess," said Priya Raghunathan, a TA who has graded Oyelaran's course since 2019. "Now the tell is the tidiness."

The two integrity questions have merged. A submission can be a translation of a classmate's work and fully machine-written at the same time, and a check that answers only one of them is half a check. This is where most departments still have a gap: the plagiarism system knows nothing about generation, and the AI detector knows nothing about the code two seats over.

Codequiry AI code detection report with average AI score, highest file score and a risk distribution
AI code detection: probability scores per file, flagging submissions likely written by ChatGPT, Copilot, Claude or Gemini.

Three detection techniques that survive a language change

Intermediate representations

Parse every submission with a multi-language parser, map node types onto a shared vocabulary (loop, branch, call, return, assignment, comparison), and compare shapes instead of text. tree-sitter, the incremental parsing library released in 2018 and now embedded in most editor tooling, ships grammars for well over a hundred languages, which makes it the practical starting point. The mapping layer is where the difficulty lives, because a Python comprehension and a Java stream pipeline have no shared node type.

Pipeline maintainers learn the failure mode the hard way. The tree-sitter Python grammar needed revisions after Python 3.12 changed f-string tokenization under PEP 701, and a pipeline pinned to the older grammar will quietly drop those nodes instead of raising an error. A dropped subtree looks exactly like a clean submission.

Program dependence graphs

Two research systems predate the LLM era and handle translation better than token methods on published benchmarks. Sherlock, built at the University of Warwick and described in 2018, and XLPlag, published by Oscar Karnalim in 2021, both construct graphs of data and control dependencies and compare graph structure rather than syntax. Neither is a product. You build them, you maintain them, and you own the false positives.

Behavioral fingerprinting

Run both submissions against the same input corpus and compare the outputs, along with line counts written to stdout and, where the runtime allows instrumentation, the sequence of functions called. Two programs that fail on the same edge case at the same input size are related in a way no static comparison shows. The cost is a sandbox with no network access and hard CPU limits, which is why most departments reserve it for the handful of pairs that are already suspicious.

Only one of those three is a weekend project. That is the honest calculus for a lecturer deciding how far to go, and most stop after the first.

A workflow for a course that submits in more than one language

The version Oyelaran settled on after two semesters of iteration runs in four passes and takes about four hours for a 180-student cohort.

  1. Within language first. Run structural comparison for each language group against itself and against the previous semester's submissions in that language. This catches the majority of real cases. Use a source code plagiarism checker that treats prior cohorts as a baseline rather than only comparing current peers, because cross-language reuse often travels through a group chat with older members in it.
  2. Pivot. Translate the smaller group into the larger group's language with a deterministic translator, not a model. A model will improve the code on the way through, which changes the structure you're trying to compare. Freeze the pivot version and write it down, because next year's cohort gets compared against this year's results.
  3. Compare with a lower threshold. A pivot loses signal. If your within-language review threshold is 65 percent, expect to start reviewing around 45 percent on the pivoted set, and expect to discard most of what it flags.
  4. Read the evidence before writing an email. This part doesn't automate. Raghunathan spends four to six minutes per flagged pair, looking at the aligned regions rather than the score.

In the spring 2024 run, the pipeline flagged twenty-one pairs. Six survived manual review. Two of those were translations, and in one case the student had also used a model to do the translation, which meant the final report had to make an argument about both.

Codequiry match review workspace with a Java code viewer, match explorer and per-submission analytics
The match review workspace: matched code, every peer and web source, and per-submission analytics on one screen.

What to do with a confirmed translation

Honor codes written before 2010 tend to use words like "copy" and "unauthorized collaboration" and say nothing about translation. A student who translated a classmate's Java into Python will argue, sometimes in writing, that they wrote a different program. Technically they did.

"We had a case in 2022 where the student's defense was that the two submissions were different programs because they were in different languages," said the department chair who reviewed it. "That's true. The code was the same."

Translation of someone else's work is the straightforward half of this. Translation of your own prior work is not, and departments handle it inconsistently. Some treat a Python rewrite of last year's Java lab as an integrity violation. Others treat it as a syllabus question, which is closer to right, because a student who reuses their own solution is guessing at a rule nobody wrote down.

The industry version of the same dispute is not hypothetical either. In 2021 the US Supreme Court ruled that Google's reimplementation of Java API declarations was fair use, and it did so while assuming for the sake of argument that the declaring code was copyrightable. Structure and organization of code carry legal weight. The case simply left the underlying question for another day.

What to look for in tooling for a multi-language course

The honest comparison is short, and it starts with what each tool refuses to do.

ToolLanguagesCross-language pairsAI-generated code scoreWhere it runs
MOSSRoughly 30NoNoCommand line, emailed HTML report
JPlag 4.x (2022)15+NoNoSelf-hosted, web report
Dolos (2021)About 15NoNoSelf-hosted, interactive pair views
Codequiry65 approved languagesPivot workflow, structural engineYes, per file and per submissionHosted dashboard plus REST API and CLI

The comparison that matters for a course submitting in two languages is not about who has the cleverest graph algorithm. It's about whether the peer pass, the web pass, and the generation pass live in one place, because a translation usually arrives with a companion question. Codequiry compares at the token, AST, and fingerprint level, which is what lets it hold up when a student renames everything and reformats the file, and it checks against peer submissions, prior cohorts, GitHub, and the open web in the same run. When a chunk of a submission traces back to a specific Stack Overflow answer or repository, the report shows the lines and token counts behind that match, which is the difference between a number and evidence you can put in front of a student.

It isn't free, and MOSS is, which matters in a department with no budget line for integrity software. The case for paying is roughly the case for any paid tool: you stop maintaining a pipeline, you get support when a semester is on fire, and you get an AI score that no free option currently provides. Departments that have run both tend to describe the tradeoff in exactly those terms. A side-by-side of the engines is worth reading before committing, and Codequiry vs MOSS covers it without pretending the free tool is bad.

Codequiry web results tracing a submission to a Stack Overflow question with line and token counts
Tracing code to its source: a submission matched to a Stack Overflow answer, down to lines and tokens.

Oyelaran's spreadsheet is down to three entries this year, from eleven. The pipeline is not the reason. What changed is that the department added a line to the syllabus in August 2024 stating that translating a submission into another language, by hand or with a tool, is treated the same as copying it. Clear rules reduce the number of cases that need detecting, which is the least glamorous finding in four years of data and probably the most useful one.

For departments running multi-language sections, the practical starting point is a structural code plagiarism checker that reports peer, web, and AI signals together, so the pivot pass has somewhere to land.