What a CI gate for AI code can and cannot decide
Yes, you can wire AI-generated code detection into a CI pipeline and get a usable signal on a pull request. No, it will not give you a clean binary answer, and it should not be the job that fails the build. The pattern that has held up for us is a score per file, a warning threshold on the PR, and a narrower policy that blocks only the highest-confidence cases on a release branch.
I run DevSecOps for a mid-size fintech. Around 140 engineers, mostly Java and Kotlin services, with a Python data platform bolted on the side. Last March, a payments service merged a 380-line retry-with-backoff wrapper. It passed review. It passed tests. Two weeks later nobody on the team could explain the jitter formula. The author, a good mid-level engineer, said Copilot drafted most of it and they "cleaned it up." Nobody did anything wrong, exactly. That's what made it interesting. We had no policy for code that no human fully reasoned through at the time of merge, and we were about to need one.
Why detecting AI-generated code in pull requests is not a plagiarism scan
Plagiarism detection needs a corpus. You take 200 submissions of the same assignment and compare them to each other, then to GitHub, then to the open web. That's an extrinsic problem: the signal comes from what the work matches. A CI gate has no corpus. One engineer, one branch, one diff. Nothing to compare against, so the detector falls back on intrinsic signals inside the code itself, and that changes everything about how you use it.
This is why we bolt a separate AI probe onto the pipeline instead of hoping our SAST tool grows the feature. Semgrep finds the SQL injection. It has no opinion about whether a human typed the surrounding function. The two questions are orthogonal, and conflating them in one dashboard is how you end up with a "Code Quality" score that nobody trusts.
Codequiry is interesting here because it sits on both sides. Peer similarity, web and GitHub matching, and AI generation scoring in one report, which is what a university wants when it grades a cohort. For a repo, you mostly care about the AI and web columns and ignore the peer column entirely. The AI code detector output is per file, not per repo, and that granularity matters more than the headline number.
What a detector is actually measuring under the hood
Most of these systems lean on token-level predictability. Run the code through a language model and ask how surprised it is by each token. Machine-written code tends to score low perplexity: the next token is exactly what the model would have guessed. Human code wanders. People name a variable x when they mean rate, leave a TODO from 2021, and comment out a debug print before pushing.
The second signal is burstiness, the variance in structure across lines or statements. Human code is lumpy. You write three tight lines, then a ten-line block with a weird edge case, then a one-liner. Generated code is flat. Uniform docstrings, uniform error handling, every function the same length and shape.
Here's the compound interest function that showed up in a config service last fall:
def calculate_compound_interest(principal: float, rate: float, time: float, n: int = 1) -> float:
"""Calculate compound interest.
Args:
principal: The initial principal amount.
rate: The annual interest rate (as a decimal).
time: The time in years.
n: The number of times interest is compounded per year.
Returns:
The total amount after compound interest.
"""
if principal < 0 or rate < 0 or time < 0:
raise ValueError("Values must be non-negative.")
amount = principal * (1 + rate / n) ** (n * time)
return round(amount, 2)
And here's the version of the same function that had been in the repo for three years, which a human wrote at some point under deadline pressure:
def calc_ci(p, r, t, n=1):
# round at the end, finance complained about penny drift in Q2
amt = p * (1 + r / n) ** (n * t)
if n == 0:
return p # shouldn't happen, guard from the sim swap
return round(amt, 2)
Both work. One of them carries fingerprints. The dead guard clause, the lowercase abbreviation, the comment that references a real incident: those are the residue of a person who has been annoyed by production. Models don't leave that residue unless a human adds it afterward.

What none of this tells you is intent. A tightly written, heavily documented, beautifully type-hinted function from a senior engineer can score high. A messy one from Claude can score low if a human restructured it before committing. Treat the score as a prior, not a finding.
Where false positives come from, and why generated code is the worst offender
We shipped the gate in March with the generated-code exclusion flag off by default. That was a mistake I'd like back. Our protobuf output lives in gen/, and the first week of scans threw 240 findings across 30 pull requests, nearly all of them in .pb.go files. OpenAPI client stubs did the same thing. So did a 4,000-line i18n table and a Terraform provider directory we vendor.
Generated code is machine-written by definition. It scores like LLM output because structurally, it is LLM-adjacent: uniform, repetitive, no human variance. Any honest detector will flag it. Your job is to keep it out of scope before someone sees a red badge and loses faith in the tool.
# .ci/ai-scan.yml
scope: diff
exclude:
- "gen/**"
- "**/*.pb.go"
- "**/*.g.dart"
- "**/migrations/**"
- "vendor/**"
- "node_modules/**"
- "testdata/**"
thresholds:
warn: 0.55
block: 0.90
The flag everyone forgets: these exclusions are not the same list your SAST tool uses. We keep two ignore files and they drift every few sprints. Someone adds a new codegen directory, the SAST tool gets updated, the AI scanner does not, and suddently Monday's review queue is 60 items deep. If you run both, put the shared paths in one file and have each tool's config include it.
Placing the gate: warn on pull requests, block only on release
A hard merge block is the fastest way to teach a team to route around a tool. We ran one for three weeks in the spring. What we got was engineers regenerating submissions until the score dropped, which is exactly the behavior you don't want, because it optimizes for the detector instead of for readable code. We turned it off and moved to warnings on every PR plus a blocking check on the release branch for services in our tier-1 group.
- name: AI code signal
if: github.event_name == 'pull_request'
env:
CODEQUIRY_TOKEN: ${{ secrets.CODEQUIRY_TOKEN }}
run: |
./scripts/ai-code-check.sh \
--base "origin/${{ github.base_ref }}" \
--mode warn
The wrapper does one thing: turn a report into an exit code and an annotation. Diff scope, not whole repo, because scanning 400k lines on every push is a waste of runners and produces findings in code nobody touched.
score=$(jq -r '.summary.ai_score' out/report.json)
status=$(jq -r '.summary.review_status' out/report.json)
if [ "$(echo "$score > 0.9" | bc)" -eq 1 ]; then
echo "::warning title=AI signal::score $score in changed files ($status)"
fi
# PRs always exit 0. The release job runs the same script with --mode block.
exit 0
Keep the two modes in one script so the thresholds can't diverge. The release job sets --mode block, and a failure there stops the tag, which is a conversation with a release engineer instead of a dead PR at 6 p.m. on a Friday.
What to do when something scores high
You ask a question, you don't open a case. The one that has worked: "Walk me through the backoff math and why we picked 1.7 for the multiplier." If the author can explain it, the score was noise, and you log it as noise. If they can't, you've found something real, and it was never about whether a model wrote it. It's about whether the person signing off understands what ships.
A detection score is a triage tag. The moment you treat it as a verdict, your engineers start writing code for the detector instead of for each other.
We track outcomes now. Of the 22 high-score PRs we've reviewed since April, 19 were legitimate: generated scaffolding, a copied utility, or a well-commented function that just read clean. Three were cases where the author could not explain the logic, and in all three we asked for a rewrite with a test that pinned the behavior. Two of the three rewrote it. The third one was a contractor, and that's a different conversation.
Evidence, retention, and the audit trail
If you're in a regulated shop, the score itself is not the artifact. The artifact is the record: commit SHA, file path, score, threshold at the time, the decision, and who made it. Store the report JSON as a CI artifact next to the diff, keep it for at least a year, and get it out of the vendor's UI and into your own object storage the same day it's produced. Tools change pricing and UI. Your audit trail shouldn't live inside someone else's product.
Signed webhooks are worth the setup. We push a copy of each high-confidence result into our SIEM alongside the SBOM and license scan output, so a single query answers "what changed in this release and who approved it." That's the part leadership actually cares about when a model or a vendor is anywhere near the pipeline.

Where Codequiry fits, and what it doesn't replace
Our SAST stack stays where it is. CodeQL for the deep dataflow work, Semgrep pinned in a container for the fast policy rules, Gitleaks for secrets. None of those answer the AI question and none of them pretend to. What we added is a detector that scores AI generation per file, matches code against GitHub and the open web, and exposes the whole thing through a REST API and CLI so it can sit in a pipeline rather than only in a browser tab.
We looked at MOSS first because it's free and everyone in academia has used it. It also means running a Perl-era server yourself, with no API, no report storage, and no AI column. That tradeoff is fine for a one-off class comparison. It's less fine when you need the same evidence format every quarter for an auditor. The same reasoning applies on the education side, where Codequiry vs MOSS comes down to whether you want to maintain infrastructure or buy reports.
If you're managing a codebase rather than a course, the pieces you'll actually touch are the API key, the CLI, and the webhook. If you're running a department, the Codequiry solutions side is the one that matters: peer comparison, web sources, and AI scores in a single review queue for a TA who has 180 submissions and an hour.

What still doesn't work
Heavily edited model output is close to undetectable, and I'd be suspicious of anyone who says otherwise. The signal lives in the parts a human never touched. Once someone rewrites the function names, adds their own comments, and restructures the error handling, you're left with a weak prior and a lot of noise.
Language coverage is uneven too. Python and JavaScript give you rich token statistics. A DSL, a legacy COBOL module, or a 40-line YAML Helm template gives you almost nothing to work with. We haven't validated this past a few thousand pull requests in one org, so treat our thresholds as a starting point to tune, not a default to copy.
Frequently asked questions
Can you detect GitHub Copilot code in a pull request?
Yes, at the file level, with a confidence score rather than a yes/no. Copilot output tends to score high on predictability and low on burstiness, especially when a human hasn't restructured it afterward. Pair the score with a review question and you get something usable. Treat the raw number as proof and you'll burn trust fast.
What AI score should block a merge?
We block at 0.90, and only on release branches for tier-1 services. Everything else warns. Your mileage depends on how much generated code lives in your repo legitimately. Start with warnings, watch the distribution for a month, then set the block threshold above the noise floor you observe.
Does AI detection work on protobuf or OpenAPI generated files?
It fires on them constantly, which is why exclusion paths are the first thing to configure. Generated stubs are machine-written and look it. If you skip this step, your first week of scans will produce more findings in .pb.go files than in hand-written code, and your team will stop reading the output.
If you want to see what this looks like before you wire it into a pipeline, run a batch of your own repo through a code plagiarism checker that also scores AI generation and read the per-file breakdown. It takes an afternoon and tells you more about your real threshold than any vendor's documentation will.