Turing Test (Imitation Game)

Context: FIT1061_MOC · the unit’s opening criterion — and the one it abandons in The AI Effect (Defining AI) What it weighs: whether a judge can reliably separate machine from human through text alone ➔ behaviour only, never machinery.

Quick Revision

  • 🎯 Test question: after a fixed text-only conversation, can the judge tell which respondent is the machine?
  • ⚠️ Key Constraint: a pass is evidence about the judge’s discrimination, not the machine’s understanding ➔ never cite “fooled ” without naming the format (duration, topic breadth, persona).

📝 Core Commitments

  • Refuses the question ➔ Turing (1950, Computing Machinery and Intelligence) calls “can machines think?” too vague to answer and substitutes a game instead of defining thinking.
  • Operationalisation is the contribution ➔ unanswerable metaphysical question ⟹ testable behavioural one.
  • The setup ➔ three players — human judge · human respondent · machine respondent; text-only channel (teletype); Turing’s suggested budget minutes.
  • The criterion ➔ judge cannot reliably identify the machine ⟹ “for any practical purpose, the machine thinks”.
  • What it refuses to count ➔ architecture, memory, understanding ➔ a -line pattern matcher and a B-parameter model are judged on identical evidence.
  • ELIZA effect ➔ the projection runs from human to machine, not the reverse; Weizenbaum (1976): “extremely short exposures to a relatively simple computer program could induce powerful delusional thinking in quite normal people.”
  • Suitcase word ➔ Mitchell (2019) — “intelligence” packs pattern recognition, abstract reasoning, language, social judgement, embodied skill, planning, creativity and self-awareness into one label, so every verdict silently drops most of them.
  • Where it lands ➔ once the test is passed on general fluency, “can machines think?” collapses into what does thinking mean?

⚖️ Framework Contrast

CriterionTest applied to a systemVerdict it tends to reachWhere it breaks down
Turing test (behavioural)can a judge distinguish it from a human by text?passed in the broad sense by 2022–25 LLMssays nothing about mechanism ➔ ELIZA-grade tricks and general fluency score alike
McCarthy task definitionwould the task require intelligence if a human did it?almost everything qualifies, brieflythe label retreats as soon as the system works
Mitchell suitcase decompositionwhich capacity — reasoning, planning, social judgement?“intelligent at X, not at Y”gives no single yes/no, so headlines ignore it
FIT1061 algorithmic movewhich of the three questions does the algorithm answer?classifies the machinery, declines the labelanswers “how does it work”, not “does it think”

🧩 Case Application Drill

⚠️ Common Mistakes

  • 💡 Headline as evidence ➔ quoting a deception percentage without the format conditions; the persona and the clock usually did the work, not the machine.
  • 💡 Test read as definition ➔ Turing explicitly declined to define thinking, so “it passed, therefore it understands” imports a claim he refused to make.
  • 💡 Unpacked suitcase ➔ asserting a system “is / is not intelligent” without naming which capacity is being claimed.