How We Test AI Humanizers: Wibble's Public AI Humanizer Benchmark Methodology
Every "we tested 12 AI humanizers" roundup follows the same script: a winner's table with suspiciously tidy percentages, no corpus, no detector versions, no dates, and no raw outputs you could check. You're asked to trust a number that cannot be reproduced. This page is the fix: a public draft protocol (v0.1) for the Wibble AI humanizer benchmark, published before we run it, so the results can't be quietly shaped after the fact. We call it a draft deliberately — it will keep firming up (final tool panel, sample counts, scoring rubric, and statistical plan) and every change is versioned below, but the commitments here are made in the open, before any data exists.
The order is the point. When the method is committed publicly first, the results that follow are accountable to it — same corpus, same detectors, same scoring, failures included, Wibble's own failures included. This document is the contract. To be explicit: as of July 21, 2026, no benchmark results exist yet. This article contains no scores, and any article on this site that discusses future results links back here.
Why We Publish This Protocol Before We Run It
The AI humanizer market runs on unverifiable claims. Tools advertise "100% undetectable." Review sites publish rankings that happen to favor whoever pays the highest affiliate commission. Even honest testers make silent choices — dropping a sample that "seemed unfair," rerunning a tool until it passed, quoting the best of five attempts — that turn a test into a marketing asset.
Committing to a method before you collect data is the standard guard in research: you can't tune the method to produce the answer you wanted if you've already published it. This draft protocol commits four things in advance:
- The corpus — what text gets humanized, from which models, at which lengths.
- The detector panel — which detectors score the outputs, and how access limits are handled.
- The scoring dimensions — everything measured, not just detector results.
- The publishing rules — raw inputs and outputs released with the results, failed runs included.
Once results are published, none of these can be adjusted retroactively. If the methodology changes for a future run, the change gets documented and versioned, and old results stay published against the old method.
The Test Corpus: What Goes In
Humanizers behave very differently depending on what you feed them, so a single sample type proves nothing. The corpus covers five content types:
- Academic writing — essay and literature-review style text with real in-text citations and direct quotations, because citation damage is one of the most common and least reported humanizer failures. (It's why Wibble's citation handling exists as a product, and why humanizing text without breaking citations gets its own article.)
- Technical writing — documentation-style text where precision matters and "creative" rewording breaks correctness.
- Marketing copy — where register and punch matter more than formality.
- Professional email — short-form, high-context text most tools over-rewrite.
- SEO/blog content — where keyword survival is part of the job.
Each type is generated by three source models — ChatGPT, Claude, and Gemini — because detectors respond differently to different models' statistical fingerprints, and a humanizer tuned on one model's output can fail on another's.
Each combination is produced at three lengths: 500, 1,000, and 2,000 words. Detector accuracy is known to vary with length, and 2,000 words is a common per-request ceiling for humanizers (including Wibble's own per-humanization limit).
Finally, every nondeterministic tool gets three runs per sample, and all three are scored and published. A tool that passes one run in three is not a tool that passes. Single-run testing is how marketing screenshots are made.
The Detector Panel: What We Score Against
The panel splits into two tiers, by access.
Publicly accessible detectors: GPTZero, Originality.ai, ZeroGPT, and Copyleaks. Anyone can run these — which is the reason they're the core panel. Every published score from these detectors can be independently re-checked by any reader with the raw outputs we release. Each score is recorded with the detector's stated model or version (where exposed) and the date it was run.
Turnitin is the exception. Turnitin is licensed to institutions, not individuals; there is no self-serve public access. So the rule is strict: the benchmark reports Turnitin classifications only when obtained through authorized access. If a run instead uses another detector as a stand-in to estimate Turnitin-style behavior, it is labeled a proxy, prominently, every time it appears — never presented as "Turnitin results." Roundups that print Turnitin pass rates without explaining their access are describing tests you cannot verify and they may not have run.
Background on how these specific detectors handle humanized text lives in the companion articles on whether Turnitin detects humanized AI text and whether GPTZero catches humanized text.
Scoring: Eight Dimensions, Not One Number
A detector pass is worthless if the output no longer says what you meant. So every output is scored on eight dimensions:
| Dimension | What it answers |
|---|---|
| Detector classification | How does each panel detector classify each run? |
| Meaning preservation | Do the output's claims still match the input's claims? |
| Factual drift | Did names, numbers, dates, or technical details change? |
| Citation preservation | Did authors, years, quotes, and citation formats survive intact? |
| Grammar and readability | Would a human reader accept this text, or is it word salad? |
| Consistency across runs | How much do the three runs vary — in quality and in detector results? |
| Processing time | How long does a job actually take? |
| Cost per 10,000 words | What does that volume cost at the vendor's public pricing on the test date? |
There is deliberately no weighted composite score and no single "winner" number. Composite scores are where cherry-picking hides — you can rank almost any tool first by choosing the weights. Results will be published per dimension, and readers weight what matters for their use case. A student cares about citation preservation and Turnitin-tier results; an agency cares about cost, speed, and keyword survival.
How We Keep Our Own AI Humanizer Benchmark Honest
The obvious objection: Wibble builds an AI humanizer and runs in this benchmark. That's a conflict of interest, and pretending otherwise would be the least trustworthy move available. So here is the honesty mechanism, committed now:
- Same corpus, same runs. Wibble gets the identical samples, lengths, and three-run protocol as every other tool. No special inputs, no extra attempts.
- Raw data for everything. Every input and every output — every tool, every run, Wibble included — is published with the results. You can read Wibble's worst run yourself.
- Failures are results. If Wibble gets flagged on a sample, drifts a fact, or loses a dimension to a competitor, that gets published in the same table as everything else. A benchmark that its sponsor cannot lose is an ad.
- Everything is replayable. Because the core panel is publicly accessible, anyone can re-run any published output through GPTZero, Originality.ai, ZeroGPT, or Copyleaks and check our numbers.
You also don't need to wait for the benchmark to test the sponsor. The Wibble demo takes 300 words with no account — paste AI text, humanize it, and run the output through any detector on the panel yourself:
That verify-it-yourself posture isn't a gimmick; it's the position this whole methodology is built on.
Limitations: What the Results Will and Won't Mean
Committing the method in advance also means committing to what the results cannot claim.
Results bind to a corpus, a detector version, and a date. Detectors are moving targets. Turnitin launched dedicated AI bypasser detection on August 27, 2025, then updated its detection model in February 2026 (improving recall) and again in May 2026 (a Spanish-language model targeting GPT-5- and Gemini-2.5-class output). A humanizer result from before any of those updates says little about after. Every published score is a timestamp, not a permanent property of a tool — which is also why no tool, Wibble included, should ever claim permanent or universal undetectability.
Detector scores are not authorship proof — in either direction. Detectors have documented false positives on genuine human writing and false negatives on AI text; Turnitin itself says its scores should not be the sole basis for action against a student. This benchmark measures how tools and detectors interact. It does not certify that any text is human-written, and it can't.
The corpus has edges. The first run covers long-form English across the five content types above. Other languages, code, poetry, and very short texts are out of scope until a future versioned run.
Three runs is a sample, not a guarantee. It exposes inconsistency far better than one run, but a nondeterministic tool can still surprise on run four. Consistency scores should be read as evidence, not certainty.
Status: Methodology First, Results Next
Draft protocol v0.1 — published July 2026, before any test run. This is a draft: the final vendor panel, sample counts, randomization, exclusion rules, scoring rubric, and statistical plan are still being fixed. Every change gets a version bump and a dated changelog entry here, and results will only ever be published against a stated protocol version.
This methodology is published as of July 2026. The first public run is next; its results will be released with the full raw inputs and outputs described above, and articles like our best AI humanizer comparison will link to that data rather than restating unverifiable claims.
Until that happens, one thing should be unambiguous: Wibble has published no benchmark scores. If you see "Wibble benchmark results" quoted anywhere before the run is released, you're reading something that does not exist. When the numbers arrive, you won't have to trust them — you'll be able to check them. That's the entire point of writing this document first.
Frequently Asked Questions
What is an AI humanizer benchmark?
A structured test that runs AI-generated text through humanizer tools, then scores the outputs against AI detectors and quality criteria like meaning preservation and citation accuracy. A credible benchmark publishes its corpus, detector versions, test dates, and raw outputs so anyone can reproduce the results — most published humanizer rankings do none of this.
Why publish the methodology before the results?
Because committing to the method first makes cherry-picking impossible. If the corpus, detector panel, scoring dimensions, and run counts are locked publicly in advance, the results can't be quietly tuned afterward — no dropping bad samples, no rerunning until a tool passes, no reweighting a composite score until the sponsor wins.
Which AI detectors does the benchmark use?
The core panel is GPTZero, Originality.ai, ZeroGPT, and Copyleaks — all publicly accessible, so readers can independently re-check any published score. Turnitin is licensed to institutions only, so Turnitin classifications are reported solely via authorized access; if a stand-in detector is ever used to estimate Turnitin behavior, it's explicitly labeled a proxy.
Can I reproduce the benchmark results myself?
That's the design goal. Every raw input and output — including failed runs — will be published alongside the results, and the core detector panel is publicly accessible. You can take any published output, run it through GPTZero, Originality.ai, ZeroGPT, or Copyleaks yourself, and compare what you see against the reported classification.
Isn't it a conflict of interest for Wibble to benchmark itself?
It's a real conflict, which is why it's disclosed and constrained rather than hidden. Wibble runs on the identical corpus with the identical three-run protocol as every other tool, its raw outputs are published including failures, and every score on the public detector panel can be independently re-run by readers.
When will the first benchmark results be published?
The methodology was published in July 2026; the first public run follows it, and results will be released together with the complete raw inputs and outputs. No results exist as of July 21, 2026 — no scores, pass rates, or rankings — and any 'Wibble benchmark numbers' circulating before the release are fabricated.
Do benchmark scores prove a text was written by a human?
No. Detector scores are statistical estimates with documented false positives on genuine human writing and false negatives on AI text — Turnitin itself says scores shouldn't be the sole basis for action. The benchmark measures how humanizers and detectors interact on a dated corpus; it cannot certify authorship in either direction.
Sources and verification
Paste the paragraph that got flagged
300 words free. No account. Run the output through any detector and see for yourself.
Keep reading

AI Humanizer API Comparison (2026): Pricing, Limits & Reliability
AI humanizer API comparison for 2026: documented pricing, word limits, async job models, webhooks, retries, and citation safety across providers — plus build vs buy.

Best AI Humanizers in 2026: An Honest Comparison of 12 Tools
The best AI humanizer in 2026, compared honestly: 12 tools, verified pricing, free tiers, and cost per 10,000 words — with the math shown, no fake test scores.

How to Humanize AI Text Without Breaking Citations
Humanizers rewrite author names, drift years, and fabricate quotes. How to humanize AI text without breaking citations — plus a checklist that catches damage.