AI Skill Report Card

Detecting AI Generated Text

B+78·Aug 29, 2026·Source: Extension-page
13 / 15

Given a text sample, output:

  1. Verdict: AI-generated / Human-written / Mixed (with confidence %)
  2. Key evidence: 3-5 specific markers found in the text (quote them)
  3. Reasoning: why these markers indicate AI or human authorship
Verdict: Likely Human-written (72% confidence)
Evidence:
- Inconsistent punctuation and scene formatting (△ markers, dialogue tags) typical of amateur script drafts
- Irregular pacing with abrupt tonal shifts (extreme violence next to romantic tension) without smoothing transitions
- Dialect/slang usage ("老娘", "这破地方") — colloquial register AI tends to avoid or use inconsistently
- Repeated raw sentences (e.g. duplicated line "轰。半空中的幻弦被精准击中...") — a human copy-paste/editing artifact, not typical of clean AI generation
Recommendation
Add a concrete example of a 'Mixed' verdict case to fully illustrate the third category mentioned in the workflow
14 / 15

Progress:

  • Step 1: Identify language and genre/format (prose, dialogue, script, technical writing, essay)
  • Step 2: Scan for statistical/structural AI markers (see checklist below)
  • Step 3: Scan for human markers (typos, inconsistencies, personal voice, factual errors, repetition artifacts)
  • Step 4: Check for genre-appropriate authenticity (does the "messiness" match how humans actually produce this genre?)
  • Step 5: Weigh evidence, assign confidence, produce verdict with cited excerpts

AI-generated markers to check

  • Overly uniform sentence rhythm: consistent sentence length, few run-ons or fragments where genre would expect them
  • Hedging/formulaic transitions: "总的来说", "值得注意的是", "In conclusion", "It's important to note"
  • Semantic smoothness with low specificity: descriptions that sound polished but avoid concrete, idiosyncratic detail
  • Perfect grammar/punctuation in contexts where human casual writing typically has errors
  • List-like exhaustiveness: covering all "expected" angles of a topic evenly, without natural bias or omission
  • Repetition of structure across paragraphs/scenes (same sentence template reused with variables swapped)
  • Absence of true novelty: no surprising word choices, rare idioms, or genuinely unexpected narrative turns
  • Emotional whiplash without friction: tonal shifts occur too cleanly, lacking the rough transitions a rushed human writer leaves behind

Human-written markers to check

  • Genuine typos, duplicated lines, inconsistent character/scene numbering
  • Idiosyncratic slang, regional dialect, invented profanity, personal tics
  • Factual/continuity errors (name misspelled once, prop appears/disappears)
  • Uneven pacing — some scenes rushed, others over-elaborated, reflecting real drafting effort
  • Culturally/temporally specific references that are oddly precise or personal
  • Non-formulaic idea connections — ideas juxtaposed in a way a template wouldn't produce
Recommendation
Include guidance on handling adversarial cases where AI is explicitly prompted to mimic human imperfections, since this is flagged as a pitfall but not demonstrated with an example
14 / 20

Example 1: Input: A five-paragraph essay with topic sentence + three supporting points + "In conclusion" summary, uniformly ~120 words/paragraph, no personal anecdotes, generic vocabulary. Output: Verdict: AI-generated (85% confidence). Evidence: template structure, uniform paragraph length, formulaic conclusion, absence of specific/personal detail.

Example 2: Input: A screenplay draft (like 朱雀's sample) with scene-numbering inconsistencies, duplicated stage directions, raw colloquial dialogue, extreme tonal violence-to-romance shifts, dialect terms. Output: Verdict: Human-written (78% confidence). Evidence: draft artifacts (duplicate lines), inconsistent formatting, non-standard character markers (△), regional slang not smoothed into standard register.

Example 3: Input: A cover letter with perfectly parallel sentence structures, buzzword density ("passionate", "results-driven", "synergy"), no specific company/role details. Output: Verdict: AI-generated (90% confidence). Evidence: generic buzzwords, parallel templated sentences, lack of role-specific detail a real applicant would include.

Recommendation
The examples section could show fuller text excerpts rather than paraphrased descriptions to better match the 'quote them' instruction in Quick Start
  • Always quote specific text excerpts as evidence — never give a verdict without citing the text.
  • Treat length as a confidence multiplier: short samples (<100 words) cap confidence at ~65%; require hedged language.
  • Chinese text: pay special attention to particle usage (的/了/着), idiom density, and whether colloquial register is consistent or "translated-sounding."
  • Consider genre baseline: technical docs and formal reports are naturally more uniform — don't over-penalize.
  • Flag "Mixed" verdict when a human-edited AI draft or AI-polished human draft is suspected (uneven marker distribution across sections).
  • Report confidence as a range/percentage, never absolute certainty.
  • Don't assume all fluent, grammatically clean text is AI — skilled human writers exist.
  • Don't assume all typos/errors indicate human writing — AI can be prompted to inject errors.
  • Don't ignore genre conventions (e.g., legal writing is formulaic by nature, not because it's AI).
  • Don't rely solely on "burstiness"/perplexity intuition without pointing to concrete textual evidence.
  • Avoid 100% confidence verdicts — always leave room for adversarial cases (AI-assisted editing, human imitation of AI style, etc.).
0
Grade B+AI Skill Framework
Scorecard
Criteria Breakdown
Quick Start
13/15
Workflow
14/15
Examples
14/20
Completeness
15/20
Format
15/15
Conciseness
13/15