What Does an AI Percentage Mean? How to Read the Score

An AI detector percentage looks precise, but it is easy to give that number more meaning than it has. A result such as 72% does not normally prove that 72% of the words came from an AI system. It does not identify who wrote the text, reconstruct how it was produced or establish that a policy was broken. It is an output from one classifier, under one set of conditions, about patterns in the submitted text.

That distinction matters for students, teachers, editors, hiring teams and anyone reviewing important writing. Independent research has found meaningful limitations in AI-text detection, including false positives, false negatives, uneven performance across tools and potential bias against writers whose first language is not English. A percentage can begin an inquiry. It should not end one.

The short answer: an AI percentage is a tool-specific score

The safest interpretation is: this detector found the text more or less similar to examples it has learned to classify as AI-generated. The exact meaning depends on the product. One service may present an estimated probability, another may estimate a share of qualifying text and another may map an internal score onto a 0–100 scale. Those outputs can look interchangeable even when they are not.

Before interpreting any result, read the label beside the score and the provider's methodology. Look for an explanation of the unit, the minimum supported text length, languages covered, evaluation data and known limitations. If those details are missing, avoid converting the displayed number into a stronger claim. EraseGPT's own limitations and intended use are set out in the AI detection methodology.

How an AI-text classifier reaches a score

AI-text detectors are classification systems. In broad terms, a classifier is trained or configured to distinguish examples associated with different categories. When it receives new text, it evaluates signals that may separate its reference examples. The actual features and calculation vary by system, and commercial implementations do not always publish enough information to reproduce their results.

Possible signals include statistical regularities in word choice, repeated structures, predictable transitions or combinations of features learned by a model. None of these signals belongs exclusively to AI writing. A person can write highly regular prose, especially in a constrained format. An AI-assisted draft can also be substantially reorganised, fact-checked and rewritten by a person. The detector sees the submitted text, not the history behind it.

Scores, thresholds and labels are different things

A detector may first calculate an internal score and then compare it with a threshold. The product can convert the outcome into labels such as lower or higher likelihood. Moving a threshold changes the balance between false positives and false negatives: a stricter threshold may reduce one kind of error while increasing another. That is why a polished percentage should not be mistaken for a direct physical measurement.

Calibration is another issue. A score is calibrated when, across suitable test data, values correspond consistently with observed outcomes. Users cannot assume that every percentage is calibrated for every language, genre, subject, model or date. A tool can also change over time as its models and thresholds are updated.

Three common interpretations to avoid

1. “The percentage is the share of words written by AI”

Not necessarily. Unless the interface explicitly defines and validates the score that way, 72% should not be reported as “72% of this essay was written by AI.” A document-level classification score cannot tell you which keystrokes came from a person, whether a brainstorming tool was used or how much editorial work followed.

2. “A high score proves misconduct or deception”

No. A detector does not know the applicable rules or the writer's process. AI assistance may be prohibited, permitted with disclosure, limited to certain tasks or required for an assignment. Even when the use would breach a rule, a classifier result alone does not establish that the use occurred. Process evidence and a fair review are needed.

3. “A low score proves the text is human-written”

No. Classifiers can miss AI-generated or AI-assisted writing. Editing, translation, short inputs, unfamiliar models and changes in writing systems can all affect an output. A low score means the submitted text did not trigger that detector strongly under the current conditions; it is not a certificate of origin.

Why two AI detectors can return different percentages

Disagreement is expected because tools can differ in their training data, supported languages, preprocessing, model architecture, thresholds and definition of the positive class. They may also treat quotations, headings, references, code and bullet lists differently. Even the same tool can produce a different result after an update or when the input boundary changes.

  • Text length: Short samples provide less evidence and may be dominated by a few conventional phrases.
  • Genre: A lab method, legal clause or standard business email may be more formulaic than a personal narrative.
  • Language and writer background: Performance observed on one population should not be assumed for another. Research has documented concern about disproportionate false classifications of non-native English writing.
  • Editing history: Translation, grammar tools, collaborative editing and extensive human revision can alter surface patterns without revealing who contributed which ideas.
  • Formatting: References, tables, copied prompts or navigation text may distort an assessment if they are included.
  • Model drift: Generative systems change, so evaluation on older outputs may not describe current performance.

Running several detectors does not automatically solve these problems. Their errors may be correlated, and a majority vote is not independent proof. If you compare tools, record what was submitted, when it was tested and what each product says its score means. Do not average unlike percentages into a new metric.

How to interpret an AI percentage responsibly

Step 1: identify the decision you are trying to make

A writer checking a draft has a different goal from a university considering a misconduct allegation. For low-stakes editing, a score might prompt a closer quality review. For a decision that could affect a grade, job or reputation, the evidence standard must be much higher and the affected person should have a fair chance to respond.

Step 2: check the input is suitable

Use enough continuous prose to meet the tool's stated requirements. Remove material that is not the writer's prose only when doing so is consistent with the review protocol, and document any removal. Do not keep trimming text until the output changes: that creates a result selected to support a preferred conclusion.

Step 3: read the result in bands, not as false precision

Treat nearby scores as broadly similar unless the provider demonstrates that the difference is meaningful. A change from 61 to 64 after a minor edit may look important on a 100-point scale, but it may not represent a reliable change in authorship evidence. Pay more attention to the method's documented limits and the context of the writing.

Step 4: examine independent evidence

Useful process evidence can include outlines, notes, source annotations, version history, tracked changes, research logs and the writer's ability to explain choices. These records are not perfect either, but they address the creation process more directly than a style classifier. Check factual accuracy and citations separately. If copied wording is a concern, use a distinct plagiarism review; similarity and AI classification answer different questions.

Step 5: ask, do not accuse

Use neutral questions: How did you develop the argument? Which sources shaped this section? What tools were used, and for what tasks? Can you show an earlier draft? Explain the relevant policy and allow time for a response. A score presented as a verdict can create harm before the evidence has been examined.

Step 6: document uncertainty and the final basis for action

Keep the submitted text, result date, tool version if available and the non-detector evidence considered. State what the score cannot establish. If action is taken, the record should show that the decision was based on the policy and a proportionate review, not on an unexplained percentage.

A worked interpretation example

Imagine a report returns a relatively high score for a 1,500-word essay. The responsible statement is not “the student used AI for this percentage of the essay.” A better note is: “The classifier returned a higher-likelihood result for the submitted text. This result is not proof of authorship, so the paper will be reviewed alongside the assignment policy, citations, draft history and the student's explanation.”

During that review, the teacher may discover that the prose follows a required template, the student's earlier drafts contain the same argument and the source notes explain the quotations. Or the student may disclose use that falls outside the stated policy. The detector did not decide either outcome; it identified a question for a human process to resolve.

How writers can use a score without writing for the detector

If you are checking your own work, do not optimise prose merely to push a number down. That can make writing less clear and can encourage superficial changes that do not improve the ideas. Instead, audit the draft for reader value:

  • Verify every factual claim against a reliable source.
  • Replace vague generalities with relevant examples, reasoning or first-hand detail you can support.
  • Make the structure serve the reader's question rather than a generic template.
  • Remove invented citations and follow the required disclosure rules for any tools used.
  • Read the final text aloud and edit for clarity, tone and accessibility.

A rewriting assistant can offer an alternative draft, but the author remains responsible for checking its meaning, facts, sources and suitability. Rewriting must not be used to misrepresent authorship or evade a school, publisher or employer policy.

Use the AI percentage as a prompt for review

You can submit a suitable passage to the AI detector and compare the output with the limitations above. Record the wording the interface uses for its score. Then move beyond the number: check provenance, sources, policy and process. For a complete review sequence, follow the responsible human-review workflow.

Sources and further reading