From Lovelace to LLMs: How a 19th-Century Objection Shapes Modern AI Benchmarks
Lovelace, who is widely regarded as the first computer programmer for her work on Charles Babbage’s analytical engine, wrote that the engine “has no pretensions whatever to originate anything… it can do whatever we know how to order it to perform.” Turing cited this passage in his seminal paper Computing Machinery and Intelligence to rebut Lovelace’s objection that machines cannot think. In doing so, he replaced the higher bar of “origination” with the lower bar of “surprise,” framing the test as a question of whether a machine can fool a human judge into believing it is human. Shriber notes that Turing never intended the test to be a serious benchmark for consciousness.
The imitation game quickly became a cultural touchstone. The annual Loebner Prize, which ran until 2019, awarded prizes to programs that could best fool judges in a textual Turing test. CAPTCHA challenges, ELIZA, Google’s LaMDA, and OpenAI’s ChatGPT have all been discussed in the media as passing or approaching the test. Shriber points out that the test conflates the simulation of human‑like responses with actual consciousness, and that most observers agree ChatGPT is not conscious.
Turing’s broader vision included aesthetic dimensions. He imagined a test in which a contestant would write a sonnet and explain its reasoning, a standard he considered the pinnacle of literary excellence. Lovelace, on the other hand, foresaw a collaborative creativity model in which machines could handle complex combinations while humans directed the creative process. Victorian writers debated “origination” in the context of industrial automation, and Lovelace’s notes suggest that machine‑human collaboration could produce new insights.
Shriber highlights two risks that arise from the continued focus on human‑like AI. First, the “Turing Trap,” a term coined by economist Erik Brynjolfsson, warns that an emphasis on human‑like performance can encourage the replacement of human workers. Second, Goodhart’s Law—“when a measure becomes a target it ceases to be a good measure”—suggests that optimizing for Turing‑test‑style fluency may narrow the scope of AI research. The dominant training method for large language models, reinforcement learning from human feedback (RLHF), sits on a fine line between augmenting human labor and training a replacement.
The article also discusses Generative Adversarial Networks (GANs), which consist of a generator and a discriminator. When the generator operates without the discriminator, it produces images that can be described as “machine dreaming” or latent‑space exploration. Although GANs have largely been supplanted by diffusion models used in DALL‑E and Midjourney, they illustrate how machine outputs can be shaped by human‑defined objectives while still appearing autonomous.
In sum, the Turing test remains a cultural shorthand for AI fluency, but it is no longer an active competition and is widely regarded as an inadequate measure of true intelligence or creativity. Research continues to explore alternative evaluation metrics, and regulatory attention is shifting toward safety, bias, and transparency rather than indistinguishability. The debate over how to assess machine creativity and consciousness remains unresolved, and the AI community continues to grapple with the legacy of Lovelace’s and Turing’s early arguments.