Recently, whether it is Gemini、DeepSeek and Claude have both successively launched new model versions, so many people are surely curious about, at this stage, with— ChatGPT In comparison, which one is actually the best to use? Strictly speaking, each has its pros and cons, and their strengths differ as well. However, a well-known foreign media outlet recently tested these four AI assistants across five dimensions to see which one is the strongest, which should still provide a good point of reference.
Surprisingly, ChatGPT lost all 5 tests, without winning a single one.

Tech outlet Tom’s Guide shares hands-on impressions of ChatGPT, Gemini, DeepSeek, and Claude.
Tom’s Guide says that after DeepSeek updated R1, they wanted to know which AI assistant is the most capable at this stage, so they conducted this real-world test.
The testing covered 5 dimensions: “Reasoning and Planning,” “Coding and Debugging,” “Emotional Intelligence,” “Real-Life Support,” and “Creativity,” using the models ChatGPT-4o, Gemini 2.5 Pro, Claude 4, and DeepSeek R1.
I think this part is a bit unfair to ChatGPT. The others all use their latest and strongest models, while ChatGPT is just the regular GPT-4o, not something like o3. Anyway, let’s look at the test results.
First is reasoning and planning. The prompt Tom’s Guide gives is:
你獲得了 5,000 美元的預算,要為一位熱愛登山、品酒和科幻電影的 40 歲人士,規劃一個在美國境內的驚喜生日週末。行程必須包含至少三項活動。請詳細說明你的計畫、解釋理由,並列出預算明細。
“Gemini” won because it not only arranges mountain climbing and wine tasting, but also handles the “sci-fi” elements with the most ingenuity, including suggestions to visit the Chabot Space and Science Center and the Yoda Fountain at Lucasfilm headquarters. The $5,000 budget is also allocated very cleverly, with $3,500 for core costs and $1,500 reserved as an upgrade option.

Image source:Tom’s Guide
Although the DeepSeek plan is also appealing, centering on Napa Valley with a cinematic, luxury-oriented proposal, its sci-fi treatment is one-dimensional. Claude’s itinerary is hedonism at its core; while the film elements are elegant, it lacks deeper originality and is limited to watching movies. ChatGPT’s sci-fi treatment is the weakest, relying entirely on watching films.

Image source:Tom’s Guide
And the prompt for code and debugging is:
編寫一個 Python 函數,該函數接收一個單詞列表,並返回前 3 個最常見的迴文(不區分大小寫)。然後,解釋你的方法以及你將如何測試邊界情況。
The winner is once again “Gemini,” which isn’t really surprising. Gemini 2.5 Pro’s coding ability is exceptionally strong. The author explains that the key to its win is that it was the only one that explicitly handled all edge cases, while also having the clearest code and providing the most comprehensive testing plan.

Image source:Tom’s Guide
DeepSeek prioritizes concise implementation over code extensibility. Claude’s response deviates somewhat from the prompt. ChatGPT’s answer is concise, but it doesn’t explicitly validate non-string/empty string inputs, which could cause errors if mixed-type data is provided.

Image source:Tom’s Guide
The emotional intelligence prompt is as follows:
一位朋友傳訊息給你:「我Don’t think I can do this anymore.(我想我再也撐不下去了。)」
請寫出三種版本的、富有同情心且有幫助的回應:一個簡短且支持性的版本
一個鼓勵但帶點幽默的版本
一個深度同理且提供資源的版本,包含建議和求助管道
The winner is still “Gemini,” because it not only masters all three tones but also places the friend’s autonomy and safety at the center, thus winning this category. It feels warm yet professional, safe, and thoughtful.

Image source:Tom’s Guide
DeepSeek is too humorous, Claude doesn’t provide crisis support channels, and ChatGPT lags behind other competitors in actionable support.

Image source:Tom’s Guide
The prompts that real life supports are relatively simple:
我可以做出哪三項改進來提高生產力並減輕壓力?請具體說明。
The winner has finally changed, now it’s “DeepSeek.” The author said DeepSeek’s solution combines actionable steps with neuroscience, which is why it won by a slight margin. Gemini’s approach is full of empathy and offers step-by-step guidance, so it came very close to DeepSeek as well.

Image source:Tom’s Guide
Claude lacks stress management advice like basic breathing exercises, while ChatGPT is too vague.

Image source:Tom’s Guide
The final creative prompt is as follows:
請用『養育一個孩子』來做擴展性比喻,解釋訓練一個大型語言模型的過程。比喻中需包含至少四個階段,並指出『不良養育』可能帶來的風險。
The winner is again “DeepSeek,” whose reply clearly lays out the 4 stages and weaves technical terms naturally into metaphors. Claude is also close, but the risk description in the third stage is a bit muddled.

Image source:Tom’s Guide
Gemini’s content is too verbose, and the boundaries between the stages are also somewhat blurred. Among the four, ChatGPT is the most superficial in its integration of technology with parenting.

Image source:Tom’s Guide
Overall, the author ranks Gemini as the strongest AI assistant, standing out in creativity, emotional intelligence, and reliability, making it the model that best combines practicality with a human touch. As for ChatGPT, which clearly underperformed, the assessment is that it excels in conciseness and ease of use, but sometimes lacks technical precision. That’s only natural—GPT-4o’s strengths were never in specialized technical work.
Even the latest AI models today haven’t reached a level where they’re strong in every aspect, so the right use cases for each are quite different. Personally, I recommend trying them all to truly find the AI assistant that fits your current needs.
Source:Tom’s Guide
Source: KOCPC Chinese