When Does Copied Code Become an Open Source License Violation?

Two years ago a former student of mine emailed me a screenshot of a code review. He had joined a twenty-person payments company the previous summer, and a contracting firm had delivered a Java module containing a well-known compression routine. The routine worked. It also arrived with no copyright header, no SPDX identifier, and a variable naming style that matched nothing else in the module.

He wanted to know whether this was a legal problem or just a sloppy pull request. Copied code becomes an open source license violation the moment you redistribute it in a way the original license forbids, and internal-only use does not change that. Finding it is the harder half of the problem, because the license scanners most companies already own were never built to look for pasted code.

What Actually Counts as an Open Source License Violation

Start with the legal shape of the thing. An open source license is a copyright license, not a contract. The Federal Circuit made that explicit in Jacobsen v. Katzer in 2008, holding that exceeding an open source license condition is copyright infringement rather than simple breach of contract. That distinction matters, because copyright infringement carries statutory damages and a rights holder does not have to prove they lost money.

Licenses then split into two families that behave very differently, and the difference is the single most important thing to understand before you look at any scan report.

Permissive licenses still have conditions

MIT, BSD-2-Clause, BSD-3-Clause, ISC, and Apache-2.0 let you ship closed-source binaries. Their obligations are mostly about attribution. MIT asks for exactly one thing: the copyright notice and the permission notice must appear in all copies or substantial portions of the software. Apache-2.0 adds a NOTICE file requirement, a patent grant, and a patent termination clause that fires if you sue a contributor over the patent.

This is where people get sloppy. A developer reasons that MIT is "free," strips the four-line header off a 300-line utility class, and pastes it into a proprietary file. That is still a violation. It is quieter than a GPL violation, and it is much less likely to produce a lawsuit, but the license condition was not met and the copyright notice was removed.

Copyleft licenses attach conditions to distribution

GPL-2.0, GPL-3.0, LGPL-2.1, LGPL-3.0, MPL-2.0 (at file granularity), and AGPL-3.0 impose obligations on the distribution of derivative works. GPL-2.0 section 3 requires that you accompany the binary with the complete corresponding source under the same terms. LGPL-2.1 section 6 requires that recipients be able to relink a modified version of the library against your application, which is why static linking an LGPL library without providing object files is a common finding in audits.

AGPL-3.0 section 13 is the one that catches SaaS companies flat-footed. If users interact with a modified version over a network, you must offer them the corresponding source. No binary ever ships. The obligation still triggers.

The three failure modes that produce most findings

  1. Header stripping. Someone finds a function on GitHub, copies two hundred lines, deletes the license block because it looked like noise, and moves on. Nothing enters the build through a package manager. No SBOM will ever record it.
  2. Static linking without meeting the relinking obligation. The library was declared, the license was noticed, and the compliance step was skipped anyway. This is the failure mode a good dependency scanner will flag for you.
  3. Network distribution under AGPL. The code is embedded in a service, never shipped, and the team assumes copyleft does not apply. It does.

Enforcement history is worth a paragraph, because it calibrates how seriously to take any of this. The Linksys WRT54G case in 2003 established that GPL obligations reached consumer routers. The Software Freedom Law Center filed a series of BusyBox suits between 2007 and 2009 against Monsoon Multimedia, Xterasys, High-Gain Antennas, and Verizon, and most settled with an undisclosed payment and a compliance program. The Software Freedom Conservancy sued Vizio in Orange County, California, in October 2021 over source disclosure for SmartCast televisions. The pattern is consistent: rights holders want disclosure and a fixed process, not a scorched-earth judgment.

The license attached to a line of code is a moving target

HashiCorp moved Terraform from MPL-2.0 to BUSL-1.1 in August 2023, and OpenTofu forked. Redis moved from BSD-3-Clause to RSALv2 and SSPLv1 in March 2024, then back to AGPLv3 in May 2024 for Redis 8. Elastic returned to AGPL in 2024 after three years on SSPL. If your compliance record is a snapshot of license names taken at onboarding, it was stale within a year.

What you actually need is a record of where each line of code came from, not which license the vendor was using the quarter you ran the scan.

Why Dependency Scanners Do Not See Copied Code

Dependency and license scanners are good at one job: reading a manifest, a lockfile, a container layer, or a compiled binary and reporting what came in through a declared channel. Syft and Trivy generate SBOMs in SPDX or CycloneDX format. FOSSA, Snyk, Mend, and Synopsys Black Duck assign license conclusions and policy verdicts. ScanCode Toolkit, now in the 32.x series, runs license and copyright detection across a filesystem and is the engine behind a lot of what other tools report.

Synopsys has published its Open Source Security and Risk Analysis for over a decade. The 2023 edition reported that 96% of audited codebases contained open source components, and roughly half of the audits turned up at least one license conflict. That is with commercial tooling deployed. It is also only counting what arrived as a package.

Here is the blind spot in plain terms. If a contractor pastes 400 lines of GPL-licensed Java into a file inside your repository, three things are true at once: there is no new entry in pom.xml, there is no new artifact in the container image, and there is no new dependency in the SBOM diff. Your scanner will report a clean build. The repository root LICENSE file, which is what GitHub's licensee and Google's License Classifier read, says nothing about it either, because that tool answers a different question: which license governs this repository, not which fragments inside it came from somewhere else.

ScanCode can detect license text embedded inside source files, which sounds like it covers this. It almost does, with one practical caveat that trips up nearly every team the first time. The --license-score option defaults to 0, meaning every match is reported regardless of confidence. New users see hundreds of low-confidence hits, conclude the tool is noisy, and stop reading the output. Setting the threshold to 80 or 90 is the difference between a report someone acts on and a file nobody opens. Our own department's rollout slipped a full month in spring 2023 when the JSON schema changed between 31.2 and 32.0 and our parser silently dropped the license array.

Three different questions, three different tools

Question Dependency license scanner Snippet and provenance check Peer similarity check
What does it read? Manifests, lockfiles, binaries, container layers Source text, tokens, identifiers, web and forge indexes A closed corpus of submissions from one cohort
What does it catch? Declared packages and their license terms Pasted code regardless of whether a manifest changed Copying between submissions, including translation and refactoring
What does it miss? Anything that never arrived as a package Copying that never touched a public source Anything from outside the corpus, including GitHub
Typical output SBOM plus policy pass or fail Per-file match score plus source URLs Ranked pairs with side-by-side diff

Most organizations run the first column and believe they have covered the second. They have not. A team that wants to know whether a deliverable is original is asking the question in the middle column, whether the context is a vendor handoff or a student submission.

License Header Stripping Is the Most Common Failure Mode

The reason header stripping dominates is mechanical. Deleting a comment block is the easiest possible edit, and it removes the only artifact that most tooling knows how to look for. Consider what the original looks like.

/*
 * Copyright (C) 2016 Markus Reinhardt
 * SPDX-License-Identifier: GPL-2.0-or-later
 *
 * This library is free software; you can redistribute it and/or
 * modify it under the terms of the GNU General Public License as
 * published by the Free Software Foundation; either version 2 of
 * the License, or (at your option) any later version.
 */
package org.reinhardt.webstats;

public final class Entropy {

    public static double shannonEntropy(byte[] buffer, int offset, int length) {
        int[] counts = new int[256];
        for (int i = offset; i < offset + length; i++) {
            counts[buffer[i] & 0xFF]++;
        }
        double h = 0.0;
        for (int c : counts) {
            if (c == 0) continue;
            double p = (double) c / length;
            h -= p * (Math.log(p) / Math.log(2));
        }
        return h;
    }
}

And here is the same function six months later, in a different repository, with the header gone and the locals renamed.

package com.payments.risk;

public final class PayloadStats {

    // adapted from an internal utility
    public static double entropyOf(byte[] data, int start, int len) {
        int[] freq = new int[256];
        for (int i = start; i < start + len; i++) {
            freq[data[i] & 0xFF]++;
        }
        double result = 0.0;
        for (int n : freq) {
            if (n == 0) continue;
            double ratio = (double) n / len;
            result -= ratio * (Math.log(ratio) / Math.log(2));
        }
        return result;
    }
}

Renaming shannonEntropy to entropyOf and counts to freq changes nothing that matters to a token-based or AST-based comparison. The control flow is identical, the magic constant 256 survives, the bitmask & 0xFF survives, and the entropy formula is reproduced term for term. Structural comparison does not care about the names.

You can see the shape of it in about fifteen lines of Python. This is the winnowing idea that MOSS popularized in the 2003 paper by Schleimer, Wilkerson, and Aiken, reduced to its essential step.

import hashlib, io, tokenize

def fingerprints(src, k=5):
    toks = [t.string for t in tokenize.generate_tokens(io.StringIO(src).readline)
            if t.type not in (tokenize.COMMENT, tokenize.NL, tokenize.NEWLINE)]
    return {hashlib.blake2b(" ".join(toks[i:i + k]).encode(),
                            digest_size=8).hexdigest()
            for i in range(len(toks) - k + 1)}

def jaccard(a, b):
    return len(a & b) / len(a | b)

Strip comments, tokenize, hash overlapping five-token windows, compare the sets. JPlag 5.x and Dolos both operate on this principle, with Dolos defaulting to k-grams of length 23 inside a window of 17, a choice tuned to reduce noise on source code rather than prose. MOSS is still hosted at Stanford and still accepts submissions the same way it did a decade ago.

I have run this on a few hundred files, not a few hundred thousand, so treat any specific threshold as a starting point rather than a published number. On the pair above, the Jaccard index lands around 0.87 with the comments stripped. Anything above roughly 0.6 on a file of real length deserves a human look.

Codequiry web results tracing copied code to GitHub repositories and other web sources with per-domain scores
Web results: every domain a submission matched, scored per source, from GitHub repos to tutorial sites.

How to Detect Copied Open Source Code in a Proprietary Codebase

The process that works is unglamorous and has four steps. First, remove the things that make files look different without being different: comments, whitespace, formatting. Second, produce a structural signature for each file, using tokens and, where the language allows it, an AST so that a renamed method in a different class still maps to the same node types. Third, compare those signatures against two corpora: your own codebase history, so you catch internal copy-paste between teams, and the outside world, meaning GitHub, GitLab, Bitbucket, package registries, and Q&A sites. Fourth, rank the results so a human reads twelve files instead of twelve thousand.

Step three is where most tooling stops short. Commercial snippet matchers like Black Duck's snippet matching and FossID (acquired by Snyk in 2023) compare against curated corpora of open source code and are genuinely useful for finding fragments of known projects. What they do not do well is trace a pasted block back to a specific Stack Overflow answer from 2014, which is where a surprising amount of production code began.

A source code plagiarism checker built for web-scale matching handles that differently. Codequiry compares submissions against peer submissions in a cohort, against the open web, and against public repositories, and it reports the specific source URL for each match rather than a generic "match found." For the payments company in my opening story, that is the difference between knowing a file is suspicious and being able to hand counsel a link with a line number.

The platform runs token, AST, and fingerprint comparison rather than surface text matching, which matters precisely because header stripping and renaming are the two most common evasion patterns. It also scans for AI-generated code from ChatGPT, Copilot, Claude, and Gemini in the same pass, which is relevant here for a reason that is not obvious: code a developer asked an LLM to write often reproduces training data it saw, headers included or not.

Codequiry evidence review with a synced diff of two Java files and a list of GitHub and web matches
Evidence review: a synced diff of the matched lines next to every peer, GitHub and web source for the submission.

What Students and Contractors Have in Common

The honor board hearings I have sat through and the vendor code review my former student described are the same event with different consequences. In both cases a person faced a deadline, found working code, removed the evidence of where it came from, and handed it in. In CS 210 the penalty is a zero and a conversation. In a payments company it is a legal review, a rewrite, and sometimes a disclosure letter.

The detection method is identical, which is why academic integrity tooling keeps turning up in industry procurement. A professor running a code plagiarism check across a cohort and an engineering manager verifying a contractor deliverable are both asking: is this original, and if not, where did it come from? The only real difference is the corpus size and the tolerance for false positives.

I will make one pedagogical argument while I am here. If we teach students to write an SPDX identifier at the top of every file, we get two things. We get compliance habits that carry into their first job, and we get a class of files where a missing header is itself a signal worth investigating. Our software engineering course added a two-week unit on licensing in 2021. The number of students who could tell me the difference between GPL-2.0-only and GPL-2.0-or-later went from about four in thirty to twenty-eight in thirty.

Wiring Provenance Checks Into a CI Pipeline

For engineering teams the goal is to make the provenance check boring. It should run on every pull request, fail loudly only when the score is high, and produce an artifact a non-engineer can read during an audit.

# illustrative request shape, not a literal SDK binding
curl -s -X POST https://codequiry.com/api/v1/check \
  -H "apikey: $CODEQUIRY_KEY" \
  -F "file=@build/source.zip" \
  -F "engine=web,peer,ai"

The response is JSON keyed by file path, with a match score and a list of source URLs for anything that cleared the threshold. Wire that into a build step, gate at whatever score your counsel is comfortable with, and keep the JSON as audit evidence. Codequiry solutions cover both the dashboard workflow for educators and the API path for CI, which is not a combination MOSS, JPlag, or Dolos offer. Those three are excellent research tools with no user interface to speak of and no commercial support behind them. That is fine for a graduate seminar and awkward for a compliance program that needs an audit trail someone can defend.

Codequiry API keys page with a masked key, signed webhook configuration and API resources
The API surface: an account key, signed webhooks for finished checks, and docs for wiring scans into CI.

Run the dependency scanner and the provenance check side by side. They answer different questions, and only one of them will notice the pasted function.

Frequently Asked Questions

Is using GPL code internally an open source license violation?

Generally no, as long as you never distribute the resulting work outside your organization. GPL obligations trigger on distribution, and AGPL-3.0 adds network interaction as a trigger. Internal-only use that stays inside the company is normally fine, though "internal" gets complicated the moment a contractor or a customer gets a copy.

Can a dependency scanner detect copied source code?

Only if the code arrived as a declared package. A pasted function with the header removed leaves no trace in a manifest, a lockfile, or a container image, so the SBOM is unchanged. Snippet matching and provenance checks are a separate step, and most organizations skip it.

How much copied code is enough to be a violation?

There is no percentage threshold, and anyone who gives you one is oversimplifying. Copyright protects expression, so a copied thirty-line function can infringe while a copied three-line idiom probably cannot. The de minimis analysis is fact-specific and belongs with counsel.

Do I need a lawyer or a tool first?

A tool first. The tool produces candidates with evidence attached, which is exactly what counsel needs to give you an opinion. Paying a lawyer to review 12,000 files is not a workable plan.

If you are reviewing contractor deliverables, grading a cohort, or preparing for an acquisition due diligence pass, run a code plagiarism check that compares against both peer submissions and the open web, because the pasted function with the header deleted is the one your dependency scanner will never see.