4,600 people sorted human text from machine text at 50 to 52%
Polish is not a detection cue, because detection barely happens. The rougher cut of my ad did win on hook rate, and I had explained why it won incorrectly.

On this page
Nobody detects AI content by spotting polish, because almost nobody detects it at all. In Jakesch, Hancock and Naaman (PNAS, March 2023), 4,600 participants across six experiments “identified the source of a self-presentation with only 50 to 52% accuracy.” In Nightingale and Farid (PNAS, February 2022), 315 participants sorting synthetic faces from real ones scored 48.2%, “close to chance performance of 50%.”
That matters to me because I built a conclusion on the opposite assumption. Meta defines hook rate as arithmetic, not a vibe: the metric “counts the number of 3-second video plays and divides it by ad impressions.” I ran two cuts of the same AI UGC ad. In the second, the creator stumbled over the opening sentence, paused, and started again. I nearly cut that clip, left it in, and that version produced a noticeably higher hook rate. I explained it to myself by saying polished AI content has become easy to recognise and the imperfection defeated recognition. The recognition half of that sentence is false. The stumble did not make my ad harder to identify as AI, and the paper I reached for to explain it turns out to argue against me on the one cue I used. What survives is a narrower mechanism with better evidence behind it, and it happens to sit inside the exact window the metric samples.
This was my test, my budget, one product, a handful of creatives, a few weeks. I published the pattern in the X post this article is built from and did not publish the hook rate figure there. I am not going to invent one here. What I can do is take the test apart, which is worth more than the number would have been.
The design fault is in the first three seconds
Hook rate samples the opening three seconds. The stumble, the pause and the restart are in the opening three seconds. The variable I believed I was testing, global polish, and the variable I actually changed, the content of the measurement window, were the same edit. Meta’s own A/B testing best practices page states the requirement I broke: “You’ll have more conclusive results for your test if your ad sets are identical except for the variable that you’re testing.”
The metric is also narrower than the argument I hung on it.
Table: How Meta defines the four metrics involved, in Meta’s own words, read 6 August 2026.
| Metric | Meta’s definition | Status |
|---|---|---|
| Hook rate | 3-second video plays divided by ad impressions | “This metric is in development” |
| Hold rate | ThruPlays divided by 3-second video plays | “This metric is in development” |
| 3-second video plays | Played at least 3 seconds, or 97% of length if shorter. “Time spent replaying the video for a single impression won’t be included.” | Stable |
| ThruPlay | Played at least 15 seconds, or 97% of length if shorter than 15 seconds | Stable |
The denominator is impressions, not reach and not video plays, so a fatiguing creative drags its own hook rate down through repeat impressions with no change to the creative at all. And replays inside an impression are excluded, so if the stumble makes someone rewind, the metric pays nothing for it.
What people actually detect, measured
The popular version of this research is “people cannot tell AI faces from real ones.” That overstates it, and the correction is worth more than the statistic. Nightingale and Farid ran a second experiment with training and trial-by-trial feedback, and accuracy rose to 59.0% (95% CI 57.7% to 60.4%), which is above chance and reliably so. At chance untrained, barely above chance trained. In their third experiment, participants rated synthetic faces 4.82 for trustworthiness against 4.48 for real faces, “only 7.7% more trustworthy” but significant at t(222) = 14.6, P < 0.001, d = 0.49.
Table: What three primary studies measured about human detection of synthetic content.
| Study | Stimuli | Accuracy | Direction of the error |
|---|---|---|---|
| Nightingale & Farid 2022, N = 315 | StyleGAN2 still faces | 48.2% | Synthetic rated more trustworthy |
| Nightingale & Farid 2022, N = 219, trained | Same, with feedback | 59.0% | Training helps a little |
| Miller et al. 2023, N = 124 | White AI still faces | No figure published | AI faces “judged as human more often than actual human faces” |
| Jakesch et al. 2023, N = 4,600 | Self-presentation text | 50 to 52% | Human roughness misread as machine |
Miller and colleagues called the third row “AI hyperrealism” and found that “people who made the most errors in this task were the most confident (a Dunning-Kruger effect).” Two ceilings travel with it. The effect was found for White AI faces specifically, because generators are trained disproportionately on White faces. And the distinguishing attributes do exist: they “permitted high accuracy using machine learning.” People have the signal and misread it, which is a fact about human heuristics rather than about how clean the output is.
The heuristic runs the other way
Jakesch maps the heuristic feature by feature, and read properly the map convicts my explanation. The paper fits two logistic regressions over the same language features: one predicting whether participants called a self-presentation AI-generated, one predicting whether it actually was. Where the two odds ratios point the same way, the cue works. Where they point opposite ways, people are not merely guessing, they are confidently backwards.
Table: Jakesch, Hancock and Naaman 2023, Table 1. Odds ratios above 1 mean more likely. Model 1 covers 38,866 judgments, model 2 covers 4,690 self-presentations. Rows marked ns were not significant in model 2.
| Feature | Rated AI | Actually AI | Reading |
|---|---|---|---|
| Nonsensical content | 1.105 | 1.233 | Machine signal, read correctly |
| Repetitive content | 1.083 | 1.470 | Machine signal, read correctly |
| Conversational words | 0.947 | 0.898 | Human signal, read correctly |
| Grammatical issues | 1.048 | 0.851 | Read as machine, actually human |
| Rare bigrams | 1.042 | 0.666 | Read as machine, actually human |
| Long words | 1.034 | 0.783 | Read as machine, actually human |
| Contractions | 0.947 | 1.134 | Read as human, actually machine |
| First-person pronouns | 0.925 | 0.992 ns | Read as human, diagnostic of nothing |
| Authentic or spontaneous words | 0.946 | 0.945 ns | Read as human, diagnostic of nothing |
Three features are functional cues and that is all: “Only nonsense, repetition, and conversational words were functional cues.” The paper is blunt about the rest. “Participants were 5% more likely to rate self-presentations with grammatical issues as AI-generated, although grammatically flawed self-presentations were, in fact, 15% less likely to be AI-generated.” “Participants often rated self-presentations with long words or rare bigrams as generated by AI, while most self-presentations with long words or rare bigrams had been written by humans.”
Read the direction twice, because it inverts the folk model. Roughness is not what people mistake for humanity. Roughness is what people mistake for a broken machine, and in this corpus roughness was a genuine human tell that they read backwards. The mirror image is on the contractions row: the single most casual-sounding feature on the list was 13% more likely to appear in the generated text, and people took it as evidence of a person. The cues they do read as human, first-person speech and spontaneous-sounding words, are not significantly associated with either source. So the heuristic is not just weak. It is wrong in a specific direction: it penalises the mess that humans actually make and rewards the ease that machines actually produce.
Then the part that should make anyone in advertising sit up. The authors trained classifiers on participants’ judgments and used them to select generated text that hit the cues people read as human. Those optimised versions “were rated as human more often than regular generated self-presentations (65.7% vs. 51.6%)” and, more sharply, were “more likely to be seen as human than self-presentations that were actually written by humans (65.7% vs. 51.7%).” Optimising for the heuristic beats being human at seeming human.
That finding is real and exploitable. It is not a licence for the edit I made. An optimiser aimed at this heuristic would smooth the grammar, shorten the words, drop the rare phrasing and push conversational register up. It would cut the stumble. I had been telling myself the stumble was the video-native form of “more human than human”, and on the only measured version of that heuristic it is closer to the opposite.
One hedge in my own favour, and it does not rescue the argument. Jakesch measured written self-presentations, and “grammatical issues” there is a crowdworker label on text, not a spoken false start in a video. I am not entitled to map one onto the other. I am also not entitled to cite the paper as support when the feature nearest to what I did carries the wrong sign, which is what I was doing.
The mechanism that does fit
It is older and I had never considered it. Corley, MacGregor and Donaldson (Cognition, 2007) recorded ERPs during spoken sentences and found the N400 effect for unpredictable words was reduced when the target was preceded by a hesitation marked by “er”. In a later recognition memory test, words preceded by disfluency were more likely to be remembered. Disfluency raises attention and improves encoding.
That is a claim about the moment of hesitation, not about the speaker’s perceived species, and it is the only mechanism I have found that joins the edit I made to the number that moved. Hook rate asks whether someone is still there at three seconds. A cue documented to raise attention, sitting at second one, is a better account of that than anything about humanness.
The ceilings are real and one of them is a citation ceiling. The experiment used lab sentences rather than paid video, it measured comprehension and memory rather than scroll behaviour, and the sample size is not reported in the PubMed abstract I can link, with the full text behind Cognition’s paywall. A reader following my citation cannot check how many people it ran on, so treat the size of the effect as unverified from here.
Why it happened, and why sample size was not the reason
Three mechanisms stack, in order of size. First, the confounded variable: I changed the measurement window and read the result as evidence about a global property.
Second, non-random impression assignment. Meta’s ad auction page lists estimated action rates as an auction input, “an estimate of whether a particular person engages with or converts from a particular ad.” If both cuts sit in one ad set, delivery routes each toward the people it predicts will engage with that specific ad, so any hook rate gap is partly the system’s own prediction of hook rate returned as impression allocation. More spend does not fix that. It feeds it. Meta’s A/B testing page says as much: “We do not recommend testing informally, such as by turning ad sets or campaigns on and off manually. This can lead to inefficient ad delivery and unreliable test results.”
Third, learning phase instability. Ad sets “exit the learning phase as soon as they can deliver stably. This usually occurs after about 50 results in the week after the ad set’s last significant edit”, and during it “performance is less stable, so your results aren’t necessarily indicative of future performance” (learning phase).
I originally blamed sample size, and at impression level that is almost certainly wrong. For a two-proportion test at 80% power and alpha 0.05 two-sided, normal approximation:
n = [1.95996*sqrt(2*p̄*q̄) + 0.84162*sqrt(p1q1 + p2q2)]^2 / (p1 - p2)^2
At a 25% baseline and a 20% relative lift that is
[1.23779 + 0.53061]^2 / 0.0025 = 1,251 impressions per arm.
Table: Impressions per arm needed to detect a hook rate difference at 80% power, alpha 0.05 two-sided, normal approximation. My arithmetic.
| Baseline hook rate | +5% relative | +10% relative | +20% relative | +30% relative |
|---|---|---|---|---|
| 20% | 25,582 | 6,509 | 1,682 | 771 |
| 25% | 19,146 | 4,861 | 1,251 | 570 |
| 30% | 14,855 | 3,762 | 963 | 437 |
Any live Meta test clears the +20% and +30% columns in hours. Treat those figures as a floor, since impressions are not independent Bernoulli trials and clustering by person and session inflates the true variance. The conclusion holds anyway: my test was almost certainly powered well enough and still could not answer the question. The defect was assignment and confounding, not n.
How to detect this in your own account
- Move the imperfection out of the first three seconds. Put the stumble at second eight and re-run. If polish itself is the driver, hook rate should not move. If it collapses back, the effect was positional.
- Compare impressions delivered to each cut. Materially unequal delivery inside one ad set is the optimiser’s fingerprint on your result.
- Check frequency, because the denominator is impressions.
- Confirm both ad sets left the learning phase, roughly 50 results in a week, before reading anything.
To design so it cannot recur: one variable, isolated outside the metric window; separate ad sets through Meta’s A/B test tool so audiences are split rather than competing; seven days minimum and thirty maximum; budget for both arms to exit learning; hypothesis and metric written down before launch. Meta’s page on how winning campaigns are determined lists three reasons its own declared winner might still have turned out more expensive than the alternative: “The length of the study was too short.” “There weren’t enough results to calculate an accurate winner.” “Best practices for study length weren’t followed or met.” Two of the three are duration.
The friction question is becoming a compliance question
Meta’s transparency post, published 3 February 2025 and updated 1 June 2026, describes two labelling regimes, and they are not the same size. For Meta’s own tools: “When an image or video is created or significantly edited with our generative AI creative features in our advertiser marketing tools, a label will appear in the three-dot menu or next to the ‘Sponsored’ label”, and when those in-house tools “result in the inclusion of an AI-generated photorealistic human, the label will appear next to the Sponsored label (not behind the three-dot menu).” For everything else, Meta “will also begin automatically detecting ads created or edited using third-party AI tools through industry-standard signals. When detected, we’ll apply an ‘AI info’ label included in About this ad”, and About this ad is the three-dot menu.
So the loud label is documented for Meta’s in-house generation, and an AI UGC ad built in a third-party tool currently earns the quiet one. I would not plan around that gap. Meta does not say what the industry-standard signals are, and my reading that they are provenance metadata of the C2PA kind is an inference, not something the post states. What that class of metadata does is documented. C2PA 2.4 defines soft bindings, computed from the content rather than the raw bits, and says they “enable digital content to be matched even if the underlying bits differ”, giving as the example “an asset rendition in a different resolution or encoding format.” A re-encode does not shake it off. Detection is moving from perception to metadata, where the viewer’s eye does not participate.
Article 50 of Regulation (EU) 2024/1689 (full text) became applicable on 2 August 2026, four days before I wrote this. Providers of systems generating synthetic audio, image, video or text “shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated”, and deployers generating a deep fake “shall disclose that the content has been artificially generated or manipulated”. I am not a lawyer, read the text yourself, but one consequence of the definitions is worth flagging. The Commission’s FAQ gives three cumulative criteria for a deep fake under Article 3(60). The second one decides this, and it is easy to read only half of. In full: “simulated persons, objects, places, entities or events need to resemble someone or something that exists, can plausibly exist or could have plausibly existed in reality.” Stop at “exists” and a fully synthetic presenter resembling no real person looks like it falls outside. The clause that matters is “can plausibly exist”, and a photorealistic synthetic person built to pass as a real one is close to the paradigm case of it. On the Commission’s own wording I read the deployer disclosure duty as reaching an AI UGC ad, with the provider’s marking duty binding the generation tool on top of that. You inherit a marked file you did not choose to mark, and you probably owe the disclosure as well. The Commission has assessed a Code of Practice on transparency of AI-generated content as adequate, so the mechanics are being standardised now rather than eventually.
Which makes the size of the label effect the question worth asking, and the two studies that measure it are routinely misreported. Altay and Gilardi (PNAS Nexus, October 2024, two preregistered experiments, N = 4,976) found that labelling headlines as AI-generated “reduced sharing and accuracy ratings by.17 points [-0.29, -0.04], P = 0.010 on the 6-point scale”, replicated at 0.11 points in Study 2, with sharing intentions alone reaching significance in neither (P = 0.27 and P = 0.085). And the line the citations leave out: “the effect of labeling headlines as AI-generated (2.66pp) was three times smaller than the effect of labeling headlines as false (9.33pp).” The direction is negative and the magnitude is small. Koning and Voorveld (Journal of Interactive Advertising, September 2025, N = 304) found an AI disclosure raised persuasion knowledge, which reduced trust in the ad and the organisation, and also reports a countervailing path through attitudinal persuasion knowledge that increased trust. Citing only the negative half misrepresents it.
What I have now
A test design that separates the variable from the metric window, and a corrected model of my own result. The win did not come from being harder to detect, because polish is not what people detect on. It did not come from feeding the humanness heuristic either, because the measured version of that heuristic reads roughness as a machine and would have deleted my stumble. What is left is attention: a disfluency effect on encoding that has been in the literature since 2007, landing by accident inside the three seconds the metric samples.
That changes the next build instead of confirming it. If the driver is disfluency, the lever is not “leave humanity in”, it is placement, and a hesitation belongs immediately before whatever the viewer most needs to remember, not at second one where I happened to put it. Step one of the recipe above tests that directly, and it can come back against me.
My original conclusion was that the advantage lies in knowing how much humanity to leave in. The sentence is wrong about the noun, and it is wrong in a more useful way than I expected. Where people do carry a shortcut for humanity, it is mapped, unreliable, and it rewards smoothness rather than friction. The friction I left in bought me something else. Calling that attention rather than humanity is the difference between a lever I can aim and a story I liked.
Sources
Every source below was opened and checked on the date shown. Links open in this tab.
- Why Perfect AI Content Feels Fake Mo RezaAli on X x.com Accessed 6 August 2026
- Hook rate Meta Business Help Center www.facebook.com Accessed 6 August 2026
- Hold rate Meta Business Help Center www.facebook.com Accessed 6 August 2026
- 3-Second Video Plays Meta Business Help Center www.facebook.com Accessed 6 August 2026
- About ThruPlay Meta Business Help Center www.facebook.com Accessed 6 August 2026
- About ad auctions Meta Business Help Center www.facebook.com Accessed 6 August 2026
- About A/B testing Meta Business Help Center www.facebook.com Accessed 6 August 2026
- A/B testing best practices Meta Business Help Center www.facebook.com Accessed 6 August 2026
- About the learning phase Meta Business Help Center www.facebook.com Accessed 6 August 2026
- How winning campaigns are determined in A/B tests without a holdout Meta Business Help Center www.facebook.com Accessed 6 August 2026
- AI-synthesized faces are indistinguishable from real faces and more trustworthy PNAS 119(8):e2120481119, Nightingale & Farid, 14 February 2022 pmc.ncbi.nlm.nih.gov Accessed 6 August 2026
- AI Hyperrealism: Why AI Faces Are Perceived as More Real Than Human Ones Psychological Science 34(12):1390-1403, Miller, Steward, Witkower, Sutherland, Krumhuber & Dawel, 13 November 2023 pubmed.ncbi.nlm.nih.gov Accessed 6 August 2026
- Human heuristics for AI-generated language are flawed PNAS 120(11):e2208839120, Jakesch, Hancock & Naaman, 7 March 2023 pmc.ncbi.nlm.nih.gov Accessed 6 August 2026
- People are skeptical of headlines labeled as AI-generated, even if true or human-made, because they assume full AI automation PNAS Nexus 3(10):pgae403, Altay & Gilardi, October 2024 pmc.ncbi.nlm.nih.gov Accessed 6 August 2026
- It's the way that you, er, say it: hesitations in speech affect language comprehension Cognition 105(3):658-668, Corley, MacGregor & Donaldson, 2007 pubmed.ncbi.nlm.nih.gov Accessed 6 August 2026
- Disclaimer! This Content Is AI-Generated: How AI-Disclosures Influence Trust in Advertisements and Organizations Journal of Interactive Advertising 25(3):240-253, Koning & Voorveld, 30 September 2025, via UvA-DARE dare.uva.nl Accessed 6 August 2026
- Gen AI Transparency in Meta's Ads Products Meta Newsroom, Pedro Pavon, 3 February 2025, updated 1 June 2026 about.fb.com Accessed 6 August 2026
- Content Credentials: C2PA Technical Specification, version 2.4, section 9.3 Soft Bindings Coalition for Content Provenance and Authenticity spec.c2pa.org Accessed 6 August 2026
- Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems Regulation (EU) 2024/1689, via artificialintelligenceact.eu artificialintelligenceact.eu Accessed 6 August 2026
- Article 50 - Transparency obligations European Commission AI Act Service Desk ai-act-service-desk.ec.europa.eu Accessed 6 August 2026
- Transparency obligations under Article 50 AI Act European Commission, Directorate-General for Communications Networks, Content and Technology digital-strategy.ec.europa.eu Accessed 6 August 2026
- Code of Practice on Transparency of AI-generated Content European Commission, Directorate-General for Communications Networks, Content and Technology digital-strategy.ec.europa.eu Accessed 6 August 2026