How to Read a Plagiarism Report: Similarity Scores Explained

A plagiarism or similarity report is a map of text overlap, not an automatic finding of misconduct. Its overall percentage can help a reviewer decide where to look, but it cannot explain why words match, whether a source was cited correctly, who wrote the passage or which policy applies. Those conclusions require a person to examine the underlying matches in context.

This distinction prevents two opposite mistakes. A high similarity score is not necessarily plagiarism: references, quotations, standard wording and a supplied template can create substantial overlap. A low score is not proof of originality: copied ideas can be paraphrased, a relevant source may not be in the comparison collection, or a short unattributed passage may be serious despite contributing little to the total. The report is evidence for review, not the decision itself.

Similarity, plagiarism and copyright are different questions

Similarity means that a system found matching or closely related text in material available to it. Plagiarism is presenting another person's words, ideas or work without appropriate acknowledgement in a context where acknowledgement is required. Copyright infringement is a legal question about protected expression, permission and exceptions. One passage can raise more than one question, but the terms are not interchangeable.

A correctly quoted paragraph may be a legitimate match yet still be too long for a publication's editorial needs. An uncited paraphrase may raise a plagiarism concern even if it produces little verbatim overlap. A work can be properly attributed but used without the necessary copyright permission. Review each issue under the relevant institutional, publisher or legal process.

What is usually inside a similarity report

Interfaces vary, but most reports contain several familiar elements. Understanding each one makes the review faster and more consistent:

  • Overall similarity score: the portion of submitted text the system associates with material in its searchable collection, calculated according to that provider's rules;
  • Highlighted passages: sections of the submission connected to one or more possible sources;
  • Source list: web pages, publications or other records that may contain matching language;
  • Match percentage or word count: the amount attributed to a particular source, which may overlap with matches assigned elsewhere;
  • Filters: options that may exclude quotations, bibliographies or small matches from the displayed calculation;
  • Source view: a side-by-side or linked comparison that lets the reviewer inspect wording and context.

The exact denominator matters. A provider may exclude some text before calculating the result, and the display can change when filters change. Read the report legend and provider documentation before comparing two percentages. Results from different systems—or from different settings in the same system—should not be treated as equivalent measurements.

How to read a plagiarism report step by step

1. Confirm the document and report settings

Check that the correct, complete version was analysed. Record the submission date, file name and any processing choices. Determine whether the system included the title page, quotations, references, footnotes, appendices and supplied template. If a report was regenerated with exclusions, preserve the original and note exactly what changed.

Do not remove inconvenient matches without a consistent reason. Filters can reduce noise, but they can also hide relevant evidence. A useful protocol states in advance whether bibliographies and properly formatted quotations are excluded and whether a minimum match length is used.

2. Treat the overall percentage as triage

Use the total score to estimate review workload, not guilt. There is no universal safe or unacceptable threshold. A 30% result could consist almost entirely of a required form and cited quotations, while a 4% result could contain a short, uncited paragraph central to the submission. Genre also matters: a technical method, legal clause or standard product specification may legitimately use fixed language.

A threshold can support consistent triage—for example, deciding which reports receive a manual check—but it should not replace that check. If an institution uses thresholds, it should document their purpose, validate them for the relevant work and provide a route for human review.

3. Open the largest source matches

Start with matches that account for substantial or distinctive passages, but do not stop there. Open the source rather than relying only on a coloured label. Compare the submitted text with the source's full context. Ask whether the source predates the submission, whether it is the original publication and whether another page has copied the same wording.

Source attribution can be messy. A report may point to an aggregator, archived copy or later page instead of the original author. Several source records may cover the same sentence, so adding their percentages can double-count text. Trace the material to the best available original source before deciding what citation was needed.

4. Classify each meaningful match

Use a small set of categories so reviewers describe similar evidence consistently:

  • Properly quoted and cited: the wording is clearly marked and the source can be identified;
  • Acceptable common or required language: the match is a standard phrase, title, question, template or unavoidable technical wording;
  • Close paraphrase needing revision: words have changed but the source's structure or distinctive expression remains too close;
  • Missing or unclear attribution: borrowed words or ideas are not acknowledged adequately;
  • Reference or metadata noise: bibliographic entries, author affiliations or navigation elements create the match;
  • Needs more information: the visible source, date or authorship trail is insufficient for a fair conclusion.

The label should describe what the evidence shows, not presume intent. Intent may matter under a policy, but a report alone rarely establishes it.

5. Inspect how the source is used

A citation at the end of a paragraph does not automatically make all copied wording acceptable. Direct language normally needs quotation marks or block formatting as well as a citation, subject to the required style. A paraphrase should express the writer's understanding in a genuinely new structure and still cite the underlying idea when appropriate.

Check whether the submission represents the source accurately. Selective wording can reverse a conclusion or conceal a limitation even when attribution is present. Verify page numbers, links and publication details. If the passage uses a secondary source, consider whether the original should be consulted and cited.

6. Review unmatched material and the argument as a whole

A similarity tool can only compare against material it can search. Private documents, paywalled collections, images, audio, unpublished work and recently changed pages may be absent. Coverage differs between providers and should not be inferred from a score. A no-match result therefore means “no relevant match was returned under these conditions,” not “this work is original.”

Read the whole submission for source use. Look for unsupported shifts in terminology, detailed ideas without attribution, references that do not appear in the text and paraphrases that follow a source's sequence unusually closely. Check the reasoning, not just strings of words.

7. Compare the evidence with the applicable policy

Academic departments, journals, employers and clients may define attribution and permitted collaboration differently. Identify the exact rule in effect when the work was created. Consider the type and extent of the material, its importance to the work, prior guidance and whether correction is possible. Avoid applying a rule retroactively.

For a high-stakes decision, give the author the relevant passages and an opportunity to explain. Notes, version history and research records may clarify how a problem arose. Follow the organisation's formal review and appeal process rather than allowing software to impose an outcome.

8. Document the conclusion match by match

A good review record includes the original file, report date, settings, sources opened, classifications, policy provisions, author response and final decision. State whether the concern involved unattributed wording, inadequate paraphrase, citation formatting, duplicate publication or something else. “The score was high” is not a sufficient rationale.

If the work is revised, record the correction. A revision might add quotation marks and a citation, rewrite a close paraphrase after returning to the source, remove unnecessary borrowed wording, or seek permission. Re-run a report only when that supports the documented review; do not keep editing until an arbitrary percentage is reached.

How to interpret common match types

Reference lists and citations

Bibliographic entries are expected to match other documents citing the same work. They can inflate the total without indicating copied argument. Excluding a reference list may make the display easier to review, provided the setting is recorded and the main text is still checked for accurate citation.

Direct quotations

A visible match may be entirely appropriate when quotation marks or block formatting identify the borrowed words and a complete citation points to the source. Review whether the quotation is accurate, proportionate and permitted. A report cannot determine whether the citation style satisfies the assignment or publication.

Common phrases and technical language

Short conventional phrases often have no meaningful original author, while a technical name may have no sensible substitute. Judge distinctiveness and context. Changing precise terminology merely to lower a score can make the work less accurate.

Templates and previously submitted work

A supplied cover sheet, standard declaration or recurring method can create legitimate overlap. A match to the author's earlier work raises a separate question often called text recycling or self-plagiarism. Whether reuse is acceptable depends on the context, disclosure and publisher or institution policy. Do not assume that ownership removes every obligation to cite or disclose reuse.

Patchwriting and close paraphrase

Patchwriting changes isolated words while retaining a source's syntax or sequence. It may reflect a developing writer, incomplete note-taking or deliberate copying; the report does not reveal which. The remedy is not a thesaurus pass. The writer should set the source aside, explain the idea from understanding, compare for accuracy and cite it. Educational support may be appropriate alongside any formal process.

Matches to an unexpected or later source

The named source is not always the source used. Syndicated content and duplicated web pages can make a later copy appear prominent. Compare publication dates, author information and canonical publication details. If the original cannot be established, describe the uncertainty instead of asserting a direct copying path.

Worked examples

Example A: a high score with mostly legitimate overlap

A laboratory report returns 38% similarity. On inspection, 18 percentage points come from a departmental template, nine from the reference list and six from properly marked quotations in the literature review. The remaining matches are short technical phrases. The reviewer may still assess whether the report demonstrates independent analysis, but the total percentage by itself does not support a plagiarism finding.

Example B: a low score with a material problem

An article returns 6% similarity. One match is a distinctive 90-word explanation copied without quotation or citation, and the article relies on it for its main recommendation. Its small contribution to the document-wide score does not make it trivial. The reviewer should record the passage, verify the source and apply the relevant attribution policy.

Example C: a correctly cited but overly close paraphrase

A paragraph cites the correct paper but follows the source sentence by sentence, changing only a few nouns and verbs. Citation acknowledges the idea, yet the wording and structure may still be too close. A better revision returns to the evidence, summarises only what the new argument needs and makes clear which interpretation belongs to the source and which belongs to the author.

Plagiarism reports and AI-detection reports are not interchangeable

A similarity checker looks for overlap with material in a comparison collection. An AI-text classifier estimates whether patterns resemble categories learned by that system. AI-generated wording may have no source match, and copied human-written text may receive a low AI score. Running both tools does not turn either one into proof.

If AI assistance is relevant, review it under the applicable disclosure rule and examine process evidence. The guide to reviewing AI-assisted content responsibly explains that workflow. EraseGPT also publishes an AI detection methodology and limitations page so readers can distinguish classification from authorship.

A compact reviewer checklist

  • Have you confirmed the correct document, date and report settings?
  • Did you inspect the underlying sources rather than rely on the total score?
  • Have duplicate source matches and legitimate boilerplate been identified?
  • Are quotations both marked and cited according to the required style?
  • Do paraphrases use genuinely independent wording and structure?
  • Have you read unmatched sections for unattributed ideas or unavailable sources?
  • Is the relevant policy clear, current and applied consistently?
  • Has the author had a fair opportunity to explain in a consequential case?
  • Does the written conclusion identify passages and evidence instead of treating a percentage as a verdict?

You can use the plagiarism checker to identify possible overlap, then apply this checklist to every meaningful match. The purpose is not to chase a target score. It is to distinguish legitimate source use from problems that need citation, revision or formal review.

Sources and further reading