New: Citations, Connectors and the Humanizer API. See what's new →

Research & Testing Standards

Last updated July 21, 2026

This page is the site-wide policy for how Wibble tests and how it reports what it finds. It governs every tested claim anywhere on this site. The detailed per-run protocol — the exact corpus, detector panel, run counts, and scoring dimensions — is published separately in the AI humanizer benchmark methodology. Read that document for the how; this page is the rules the how must obey.

Method before results

We publish the testing method before we publish any results, and results are then accountable to the pre-published method. This removes the room for quiet tuning after the fact — dropping awkward samples, rerunning a tool until it passes, or reweighting scores until a preferred tool wins. If a method changes for a future run, the change is versioned and dated, and old results stay published against the old method.

Detector access is labeled honestly

Our core detector panel uses publicly accessible detectors, so any reader can re-run a published output and check our numbers. Turnitin is different: it is licensed to institutions, with no self-serve public access. So the rule is strict — we report Turnitin classifications only when obtained through authorized access, and if another detector is ever used as a stand-in to estimate Turnitin-style behavior, it is labeled a proxy every time it appears, never presented as a Turnitin result.

Raw data ships with results

Every published result is released with its raw inputs and outputs — every tool, every run, failed runs included. A number you cannot check is a marketing asset, not a result.

Wibble's failures are published too

Wibble builds a humanizer and appears in its own tests. That conflict of interest is disclosed, and it is constrained: Wibble runs on the identical corpus and protocol as every other tool, and its failures — flagged samples, drifted facts, lost dimensions — are published in the same tables as everything else. A test its sponsor cannot lose is an ad.

Results are timestamps, not properties

Every result is bound to a specific corpus, specific detector versions, and a date. Detectors are moving targets — they update their models, and a result from before an update says little about after. No result on this site should be read as a permanent property of any tool, Wibble included.

Detector scores are not proof of authorship

AI detectors have documented false positives on genuine human writing and false negatives on AI text. Our tests measure how humanizers and detectors interact; they do not certify that any text was written by a human, and no detector score — ours or anyone else's — should be treated as proof of authorship in either direction.

Current status

As of July 21, 2026, Wibble has published no benchmark results — only the methodology. There are no Wibble pass rates, scores, or rankings anywhere; anything circulating to the contrary does not come from us. When the first results are released, they will follow every rule on this page. See also our editorial policy and corrections policy.