Plagiarism vs. Paraphrasing: Where Indian Researchers Draw the Wrong Line
Every PhD scholar in India submits their thesis through an anti-plagiarism check. Most of them spend weeks worrying about their similarity percentage. Far fewer understand what that percentage actually measures — and almost none have been taught the forms of plagiarism the software cannot detect at all.
This guide addresses both problems. It explains what plagiarism detection software does and does not do, what the UGC’s official thresholds mean in practice, and where Indian researchers most commonly go wrong — not because they intend to cheat, but because the line between acceptable paraphrasing and academic misconduct is genuinely blurry and rarely taught clearly.
What the UGC Regulations Actually Say
The governing framework is the UGC (Promotion of Academic Integrity and Prevention of Plagiarism in Higher Educational Institutions) Regulations, 2018, published in the Gazette of India on 23 July 2018. These are binding on every UGC-recognised Higher Education Institution in India — no exceptions, no opt-out.
The regulations establish four similarity levels with different consequences for thesis and dissertation submissions:
- Level 0 — below 10% similarity: Treated as minor or acceptable overlap. No formal penalty. Thesis cleared for evaluation.
- Level 1 — above 10% up to 40%: The student must submit a revised script within a stipulated period not exceeding six months.
- Level 2 — above 40% up to 60%: The student is debarred from submitting a revised script for one year.
- Level 3 — above 60%: The student’s registration for the programme is cancelled.
Two points that most guides do not mention clearly. First, UGC sets the national floor — your university may apply stricter internal thresholds. Some universities require below 7% or 5%, or apply separate chapter-level limits for the literature review, which is by nature the most source-dense section of any thesis. Check your specific university’s PhD ordinance, not only the UGC framework. Second, a raw similarity score and the score after standard exclusions are different numbers. Bibliographies, references, quoted material within quotation marks, and standard methodological phrases are routinely excluded from the similarity calculation by the software or by the institution before assessing which Level applies. A raw score of 22% can become 8% after these exclusions — placing it comfortably within Level 0. Always ask for the excluded report before comparing your result against the thresholds.
What Turnitin and iThenticate Actually Detect
This is the part most researchers misunderstand, and Turnitin itself is direct about it. In its own published guidance, Turnitin states plainly: Turnitin does not detect plagiarism. What it detects is similarity — and it calls its output a Similarity Report, not a Plagiarism Report, precisely for this reason.
Turnitin works by using a text-fingerprinting system that breaks a submitted document into small overlapping chunks — called n-grams — and matches them against its repository. That repository includes over 1.5 billion student papers, 99 billion web pages, and content from major academic publishers including Elsevier, Springer, and Wiley. When a match is found, it is flagged and counted toward the similarity percentage. The algorithm assigns that percentage based on the proportion of matched text in the document.
Originally, this was purely a word-matching exercise. More recently, Turnitin has incorporated semantic similarity analysis that can catch closely paraphrased passages even when most individual words have been changed. This is an important development: the common assumption that swapping synonyms defeats the software is increasingly incorrect.
What the software structurally cannot do is determine intent, assess context, or identify ideas that have been genuinely paraphrased with full attribution. A passage that is 100% original in phrasing but completely unattributed in idea generates zero similarity score and zero software alert — and is still plagiarism. The software’s blind spots are where most of the genuine misconduct in Indian academic writing actually lives.
The Four Types of Plagiarism Checkers Cannot Catch
1. Idea plagiarism (thorough paraphrasing without attribution)
This is the most common and most serious form of plagiarism that detection software misses entirely. If a researcher reads a source, thoroughly rewrites every sentence in their own words and sentence structure, but does not cite the source — the similarity score is zero. The software sees no matched text. A human reviewer reading the original source and your thesis, however, would immediately recognise that the argument, sequence, structure, and intellectual contribution belong to someone else.
Plagiarism is the representation of another person’s ideas, processes, results, or words without giving appropriate credit. Words are only one of the four things listed. Appropriating someone’s argument — their way of framing a problem, their analytical sequence, their conclusions — without attribution is plagiarism regardless of whether a single sentence matches.
2. Mosaic or patchwork plagiarism
Mosaic plagiarism — also called patchwork plagiarism — occurs when a researcher copies phrases, passages, and ideas from different sources and combines them, slightly rephrasing passages while keeping the same basic structure and meaning as the original, and inserting their own words to stitch the material together. It borrows the original’s language in fragments, uses synonyms for individual words, but preserves the framework, sequence, and intellectual structure of the source.
This is particularly insidious because it is more effort than direct copying and appears more original to a casual reader. Modern detection software increasingly catches it through semantic analysis, but not always — and not consistently across all disciplines. More importantly, even where the software flags it, mosaic plagiarism is academically dishonest even if you cite the source. Footnoting a source does not license reproducing its structure sentence by sentence with synonym substitution. That is not paraphrasing — it is word substitution.
3. Self-plagiarism
Self-plagiarism means reusing work you have previously submitted or published — presenting it as new without disclosure. The most common form in Indian doctoral research is submitting work from a previously published journal article (or conference paper) into a thesis chapter without clearly marking those sections as drawn from prior work.
Self-plagiarism is a genuine form of academic misconduct because it misrepresents the novelty of the work. Turnitin will often catch it if your prior paper is in its database — and institutional repositories increasingly include previously submitted work. The solution is disclosure, not concealment: clearly indicate in your thesis which sections reproduce or adapt previously published work, and obtain the necessary permissions from the journal or publisher where required.
4. Source-based plagiarism (citing a secondary source as a primary)
This form is almost never discussed in Indian research writing guides. It occurs when a researcher cites a primary source they have not actually read, having encountered it only through a secondary source that discussed it. If you cite Rawls’s A Theory of Justice based on a paragraph about it in a review article you read — without ever opening the primary text — you are misrepresenting your scholarship. No similarity checker detects this because the citation appears correct. Only a reader expert enough in the field to know that the specific framing or emphasis comes from the secondary source, not the primary, would catch it.
The simple rule: cite what you have read. If you have not read the primary source, cite the secondary source you actually read, and make clear you are citing it as quoted or discussed in that secondary source.
Where Indian Researchers Most Commonly Go Wrong
Three patterns recur across disciplines and institutions.
Word substitution mistaken for paraphrasing. The most widespread error is replacing individual words with synonyms while keeping the original sentence structure, word order, and argumentative sequence intact. This is not paraphrasing. It is word substitution — a form of mosaic plagiarism — and it remains misconduct even when the source is cited. Modern tools are increasingly capable of detecting it, and experienced supervisors and examiners recognise it on sight. Genuine paraphrasing requires reading the source, setting it aside, understanding the idea, and writing it from your own comprehension in your own sentence structure. The test is not whether your version looks different word by word — it is whether you could explain the idea in a conversation without looking at the source.
Literature review chapter treated as a compilation. Many Indian PhD theses treat the literature review as a summary of what various authors said, assembled paragraph by paragraph. Each paragraph reproduces the argument of one paper, closely following its structure, with light rephrasing and a citation at the end. This is not a literature review — it is an annotated list. A genuine literature review synthesises sources: it identifies agreements, contradictions, gaps, and themes across the literature and develops an argument about the state of knowledge. The intellectual contribution is yours; the sources are evidence for your analysis, not a script to be paraphrased sequentially.
Methodology boilerplate copied from prior theses. Standard methodological descriptions — definitions of qualitative research, descriptions of interview protocol, explanations of doctrinal analysis — are frequently copied from earlier theses deposited on Shodhganga. Because these are common academic phrases describing standard practices, they often register as similarity in detection software. More importantly, if they are copied without attribution, they are plagiarism regardless of how standard the content is. Write your own methodology section. Describe your specific choices, your specific tools, your specific protocol — not a generic description borrowed from someone else’s thesis.
How to Paraphrase Correctly: A Practical Method
Correct paraphrasing is a skill, not a technique. It cannot be reduced to a word-substitution formula. The following method is reliable:
- Read the source passage fully until you understand the idea — not just the words.
- Set the source aside. Do not look at it while you write.
- Write the idea from your understanding in your own sentence structure and voice. If you need to glance back at the source to write the sentence, you have not understood it well enough to paraphrase it honestly.
- Compare your version with the original. If the structure, word order, or phrasing is similar, rewrite.
- Cite the source. Every time you use an idea that originated elsewhere — regardless of how thoroughly you have paraphrased the language — the citation is mandatory. Paraphrasing correctly eliminates the textual debt. It does not eliminate the intellectual debt.
When in doubt between paraphrasing and quoting: quote. A properly formatted direct quotation with attribution is never plagiarism. An imperfect paraphrase without attribution always is.
The Distinction That Matters Most
The UGC’s 10% threshold, Turnitin’s similarity report, and institutional clearance certificates all measure the same thing: textual overlap with known sources. They do not measure originality, intellectual honesty, or scholarly integrity. A thesis can pass every plagiarism check with a Level 0 score and still be built substantially on unattributed ideas, misrepresented sources, and borrowed arguments written around in sufficiently clever prose.
Academic integrity is not a software compliance problem. The purpose of citation is not to satisfy a detection algorithm — it is to give readers the information they need to trace your sources, verify your claims, and build on your work. The purpose of genuine paraphrasing is not to reduce your similarity score — it is to demonstrate that you have understood an idea well enough to articulate it in your own intellectual voice.
These are the standards that matter. The software measures a proxy. Understanding the difference between the proxy and the thing it is meant to measure is the beginning of honest research writing.
This article is intended as a practical writing guide for researchers and students. The UGC threshold figures cited are drawn from the UGC (Promotion of Academic Integrity and Prevention of Plagiarism in Higher Educational Institutions) Regulations, 2018, as published in the Gazette of India. Descriptions of Turnitin’s functionality are based on Turnitin’s own published guidance and peer-reviewed analyses of similarity detection software. Researchers with institution-specific queries should consult their university’s Research Cell or PhD Coordinator.