Running a Mid-Cohort AI and Plagiarism Sweep on 300 Submissions

A mid-cohort sweep is three scans over the same batch of submissions: peer similarity inside the cohort, similarity against GitHub and the open web, and AI code detection in student submissions, scored per file. Run all three, then only look closely at the ones that trip more than one. On 312 submissions that's roughly 20 minutes of machine time and 40 minutes of my attention.

It's Thursday of week 7 in our 12-week backend program. Four sections, 312 zips, and I want to be finished before 10pm. I've been running this sweep for three cohorts and it has settled into something I can do half-asleep, which is exactly the property you want in a process you'll repeat every term.

Why one signal is never enough

Peer-only comparison is the default and it's the weakest of the three in a bootcamp setting. Everybody starts from the same starter repo. Everybody solves the same three warmup functions. If you don't trim the provided files out, you get a wall of 100% matches that mean nothing.

We learned this the hard way. First cohort we ran through a peer scan, we forgot to add Main.java and the JUnit harness to the ignore list. Every submission matched every other submission at 100%, because roughly 40% of every submission was code we wrote and handed out. Two evenings of "reviewing" 240 pairs before someone noticed the match viewer was highlighting the same provided lines over and over. That was a wasted week.

Web checks catch a different thing. The student who found last year's version of the assignment on a public GitHub repo. The Recursion helper lifted from GeeksforGeeks. The whole function pasted out of a Stack Overflow answer with the variable names intact. Peer comparison never sees any of this, because it isn't shared with anyone in the room.

AI checks catch a third category, and it overlaps with both of the others in ways that turn out to be useful. A student who pastes ChatGPT output into a file usually hasn't copied a classmate and hasn't copied a specific web page, so peer and web both come back clean.

Run one of the three and you'll feel confident about two-thirds of the problem. That confidence is worse than ignorance, because it stops you looking.

Setting up a batch that doesn't drown in starter code

Before the first scan of the term I build the ignore list: starter files, the test harness, generated code, anything under docs/. That list lives in the course repo next to the assignment and every check imports it. Least glamorous part of the workflow and the part that decides whether the reports are readable.

Then we submit the whole section as one batch. Not one check per student, one check per assignment. Codequiry's source code plagiarism checker takes a zip per language and runs peer, web, and AI in a single pass, which matters more than it sounds. In 2022 we were stitching together separate tools for peer and AI, two dashboards, two sets of IDs, and the only join key was a filename. That workflow produced exactly one insight per semester and a lot of spreadsheet.

Codequiry new check dialog with name, course, language and detection engine selection
Starting a check: name it, pick a language and a detection engine, then upload submissions.

Language selection is worth a minute of care. Java and Python are easy. If your cohort writes templated C++, token-based scanners can get noisy on the boilerplate alone, and it's worth running a 30-submission pilot before you trust the numbers on the full batch.

Reading 312 submissions in 40 minutes

The reports come back with a risk distribution, and my first move is to ignore everything below the upper band. Last sweep: 27 submissions above the peer threshold, 14 above the web threshold, 41 above the AI threshold, and 9 submissions that showed up on more than one list. Those 9 are the actual work. Everything else is background noise produced by a cohort with a shared starting point.

I work the review queue rather than the alphabetical list. Sorting by cohort outlier score puts the genuinely strange submissions on top, which means if I get pulled into a meeting at minute 25, I've already seen the ones that matter.

Codequiry peer similarity report with a risk distribution and a smart review queue ranking cohort outliers
The peer report: a class-wide risk distribution and a smart review queue that surfaces the strongest outliers first.

For each of the 9 I open the match view and read the actual lines. This is where token and AST comparison earns its keep over surface text matching. Two students who rename computeTotal to calculateSum, reorder their functions, and convert a for loop to a while loop will still match, and the viewer shows you the matched block. I had a pair last term sitting at 78% after exactly that kind of cosmetic rework. The percentage alone wouldn't have convinced me. The side-by-side block did.

Codequiry web results tracing a submission to a Stack Overflow question with line and token counts
Tracing code to its source: a submission matched to a Stack Overflow answer, down to lines and tokens.

Web results behave differently. They're close to binary. The line either exists on a page or it doesn't, and the useful output is the URL. A submission that traces to a 2019 Stack Overflow answer about reading CSV files, with the same three variable names and the same unused import, is not a judgment call.

What AI-generated code actually looks like in week 3

Not one thing. A cluster of things, and the cluster is what moves the score. Here's a function from a week-3 submission, lightly cleaned.

from typing import Iterable

def summarize_orders(orders: Iterable[dict]) -> dict[str, float]:
    """
    Aggregate order totals by region.

    Args:
        orders: An iterable of order dictionaries.

    Returns:
        A mapping of region name to total order value.
    """
    totals: dict[str, float] = {}
    try:
        for order in orders:
            region = order.get("region", "unknown")
            totals[region] = totals.get(region, 0.0) + order["total"]
    except KeyError as exc:
        logging.error("Malformed order: %s", exc)
    return totals

Nothing in that function is wrong. The problem is the context. The student's repo has no other type annotations, no logging import, no docstrings anywhere else. Three weeks in, they haven't been taught Iterable or exception handling. The code isn't suspicious by itself. The mismatch is suspicious.

Here's what a week-3 student actually writes for the same problem.

def summarize(orders):
    totals = {}
    for o in orders:
        r = o["region"]
        totals[r] = totals.get(r, 0) + o["total"]
    return totals

Other things that show up repeatedly in the AI band: if __name__ == "__main__": in a file that's imported by the test harness and never run directly. try/except wrapped around dictionary access that can't raise. itertools.groupby in a cohort that hasn't been taught itertools. Comments that explain what a line does instead of why it's there, written in complete sentences with consistent capitalization while the rest of the submission has no comments at all.

None of these is proof. Strong students annotate early. What the AI code detector hands you is a ranking, not a verdict, and treating the top of the list as "ask about this one" rather than "this one is guilty" is the difference between a working process and a grievance.

Codequiry per-file AI analysis showing AI versus human probability for each file with written indicators
Drilling into one submission: per-file AI and human probabilities, each with the stylistic indicators behind the score.

The 1:1 that decides it

Every flagged submission gets a 15-minute conversation, not an email. I open the file and say some version of: walk me through line 14. That single question sorts almost everything.

Students who wrote their own code start explaining the problem they were solving, wander into a tangent about an earlier bug, and occasionally tell me the code is bad. Students who didn't will describe what the line does syntactically and stop. That gap is much more reliable than any score, and it's the thing that has kept our false-positive rate from turning into a false-accusation rate.

One case last term is worth sitting with. A student had used Copilot to scaffold the function, then rewritten about half of it, added their own error handling, and could explain all of it. That's a policy question, not a cheating question. We updated the syllabus language the following week instead of writing anyone up. Our honor code at the time said nothing about code completion tools, which was an oversight on our side, not theirs.

Why week 7 and not week 12

Detection at the end of term is a punishment. Detection at week 7 is teaching. Same scan, same data, entirely different outcome, because there are five weeks left to fix the underlying habit.

We started running the sweep mid-term after a cohort where we caught a cluster of copied submissions during final grading. The students involved lost the assignment and learned nothing, and four of them dropped the program. Not a single one of those outcomes required the scan to happen in December instead of October.

Running the same sweep on hiring take-homes

We quietly started running the same three-signal check on take-home exercises for our own instructor hiring, because a 90-minute take-home that was written by Claude tells you nothing about a candidate. It also changed how we structure the interview. Rather than rejecting on a score, we use it to pick which part of the submission to ask about, and a candidate who can explain an AI-scaffolded function in detail is a candidate who knows how to use the tools well. That's a hire in 2025, not a rejection.

Frequently Asked Questions

Can you actually tell if a bootcamp submission was written by ChatGPT?

Sometimes, and never from the code alone. The useful signal is the combination: a high per-file AI score, a mismatch with everything else the student has submitted, and an inability to explain the code in a short conversation. Any one of those on its own is not enough to act on.

How many false positives should I expect from an AI code detector?

More than the marketing suggests. Last sweep, 41 submissions scored above our AI threshold and I would call about six of them genuinely flagged. Tune the threshold against your own cohort and your own assignment, and never make a decision on one number in isolation.

Should you tell students you're running these checks?

Yes, and we put it in the syllabus and the first-day slides. We also tell them what happens on a hit: a conversation first, a policy decision second, a grade penalty only if the conversation goes badly.

If you want to run the same three-signal pass on your own cohort, Codequiry is the code plagiarism checker for teachers we've used since 2023, mostly because peer, web, and AI come back in one report instead of three. See how a Codequiry check works on your own submission set.