Large language models actually passed the Turing test? Recently, the University of California, San Diego announced a model named “A research report on “large language models passing the Turing test,” the research team conducted standard three-party Turing tests on OpenAI GPT-4.5, Meta LLaMa-3.1-405B, and OpenAI GPT-4o, and found that GPT-4.5 Passed the test with a 73% win rate.
Can You Tell AI and Humans Apart? GPT-4.5 Passes Turing Test, Sparking Debate
Before we begin, let’s first understand the testing process of the standard three-party Turing test. During the test, there are three roles: human tester C, human subject B, and machine A (AI). C is responsible for conversing with both B and the AI, and finally C must determine who is the machine and who is the human. A research team from the University of California, San Diego recruited 126 participants from psychology students at the university and 158 participants from the online research platform Prolific to conduct this standard three-party Turing test.
And the research team is OpenAI GPT-4.5, Meta LLaMa-3.1-405B, and OpenAI GPT-4o were all configured with NO-PERSONA and PERSONA prompts issued to these AI models. The AI models configured with the PERSONA prompt exhibited human-like behavior. The difference between NO-PERSONA and PERSONA is that NO-PERSONA only knows it is undergoing a Turing test, and its goal is to make participants believe you are human. PERSONA, on the other hand, is assigned a role, character, or personality, allowing it to behave more like a human.
The research team had each participant hold a five-minute conversation with both a human and an AI large language model at the same time, then judged which one was human based on the conversation. Everyone can take a look at the test conversation content—if it were you, could you tell who is AI and who is human just from these conversations?

Based on the responses given after testing by psychology students and participants from the Prolific recruitment platform, a conclusion was reached. PERSONA prompt’s Both GPT-4.5 and LLaMa had fairly high rates of being judged as human. GPT-4.5 PERSONA was judged as human 73% of the time (psychology students: 69%, Prolific participants: 76%), while LLaMa…-PERSONA The rate is 56% (Psychology: 45%, Prolific participants: 65%).
The traditional Turing test involves a human conversing with a machine. If the human cannot tell which side is the machine, the machine is considered to have passed the test. The research team said the experimental results represent the first empirical evidence of a large language model passing the standard three-party Turing test.
Although the research team at the University of California, San Diego believes that this test represents large language models passing the Turing test, the definition of “passing” varies depending on the testing method. After all, it isThe rate of being considered human is 73%, which does not mean everyone chose incorrectly. There is still controversy over whether test results can define a large language model as passing the Turing test. For those who are interested in learning more, you can also…Visit arXiv to read the research paper published by the research team at UC San Diego.,自行評估究竟大型語言模型通過圖靈測試的判斷是正確的還是有待商討。
Source: KOCPC Chinese