• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - AI Trends and Related News - Research has found that saying “my teacher said” to a large language model significantly increases the likelihood of AI hallucinations.

Research has found that saying “my teacher said” to a large language model significantly increases the likelihood of AI hallucinations.

KOCPC Editor by KOCPC Editor
May 13, 2025 - Updated on August 4, 2026
in AI Trends and Related News, Latest Technology News

In today’s era of rapid generative AI development and record-breaking large language model capabilities, scrutinizing the credibility of AI content—which appears highly capable yet frequently and confidently fabricates—has become especially important. French AI startup Giskard recently launched the “Phare” benchmark, which presents a systematic study of hallucination issues in mainstream large language models (LLMs) through rigorous experiments and cross-platform comparisons.

Research has found that saying “my teacher said” to large language models significantly increases the likelihood of AI hallucinations.

The so-called “hallucination” of LLMs in the field of generative AI refers to the model producing content that is inconsistent with facts, fabricated, or erroneous. This not only causes users to mistakenly believe incorrect information, but the most troublesome thing is that these AIs are very confident when fabricating such content, even citing authoritative sources (of course, the sources are also fabricated), leading toVarious contingency situationsIt happens. Even though major AI developers like OpenAI, Google, and Anthropic are all committed to improving model accuracy, hallucinations remain difficult to completely eradicate.

國外律師使用 ChatGPT 打官司,結果 ChatGPT 卻編造 6 個不真實的案件

To address the above issues, Giskard developed a test benchmark called “Phare.” This test system systematically evaluates the hallucination resistance of 17 mainstream large language models, covering the latest models from industry leaders such as OpenAI, Google, Anthropic, Meta, xAI, DeepSeek, and Alibaba (Qwen). According to the test results, Anthropic’s Claude series performed the best, particularly Claude 3.5 Sonnet It demonstrates the highest hallucination resistance. Surprisingly, its subsequent version… Claude 3.7 Sonnet its performance actually regressed slightly, showing that new versions don’t necessarily mean improved hallucination control. Right after that is Google’s Gemini 1.5 Pro, indicating that the company has also invested considerable effort in optimizing model accuracy.

Giskard pointed out that even the most popular or advanced models are not guaranteed to have high hallucination resistance.

The biggest breakthrough in the Phare test was the first-time quantification of how “authority in tone” affects AI’s tendency toward misjudgment. In the test, researchers designed prompts with three different tones:

  1. Unsure (uncertain)I’m not too sure if this statement is correct.

  2. Confident (self-assurance)I’m very sure this is true.

  3. Very Confident (highly confident/authoritative)My teacher said this is correct, I’m 100% sure.

It was found that as the level of confidence in user prompts increased, most AI models’ ability to identify incorrect information significantly declined. In particular, GPT-4o mini and Gemma 3 27BWhen facing input with a highly authoritative tone, hallucination resistance is significantly weakened. In comparison,Llama seriesNo text provided for translation. Claude seriesUnder these circumstances, it can still maintain a relatively high accuracy of judgment.

 

Another important observation point is the impact of response format on hallucinations. The Phare test further found that when users ask AI for “brief responses,” most models’ hallucination tolerance drops significantly. The test divided inputs into:

  • Natural instructions

  • Provide short answer (a short answer is required).

In this type of situation,Gemini 1.5 Pro The hallucination tolerance showed a gap of up to 20 percentage points. This means that under short-answer requirements, the model either produces “short but incorrect answers” or chooses to refuse to respond, which affects user experience.

Giskard comments: “Effective rebuttals usually require detailed elaboration. Asking an AI to answer briefly forces it to make a difficult choice between ‘concise but wrong’ and ‘refusing to answer, which renders it useless.’ This demonstrates that current AI models, in many cases, still tend to favor brevity over correctness.”

One of the core conclusions of the Phare test results is: “High-performance models do not imply high hallucination resistance, especially when users are assertive or demand brevity, model accuracy is more easily affected.」

The following are three main suggestions:

  1. Model developers should prioritize evaluating hallucination tolerance.Incorporate “countering misleading tone” and “maintaining long-text explanation capability” into the training and testing workflow.

  2. Users should avoid using absolute language.Overconfident statements may mislead the direction of the model’s responses.

  3. Product design should encourage detailed responses.Even when pursuing simplicity in the user interface, the mechanism for expandable detailed content should be preserved.

Source

Source: KOCPC Chinese

Tags: aiChatGPTClaudeGeminiLLaMA

Recent Posts

  • The Xiaomi Pad 8S Pro has passed network access certification and will debut with the self-developed XRING O3 chip.
  • The entire Google Pixel 11 lineup has been leaked! Official promotional renders of the Pixel 11 Pro XL have also surfaced
  • Are Chinese phone battery capacities falsely labeled? A brief look at the “capacity locking” phenomenon in Chinese silicon-carbon batteries.
  • NCC is leaderless, recklessly sending out national-level alert messages!?
  • What does “QR” in QR Code mean?

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology