As free open-weight models get stronger and stronger, many people are starting to think that instead of subscribing long-term to AI services like Claude and ChatGPT, it would be better to switch directly to local LLMs or rent their own GPUs to run models, which should save quite a bit overall. In particular, after the release of DeepSeek V4.1 Flash, many developers abroad have begun considering renting GPUs to self-host DeepSeek, replacing the Claude Code they used every day.
According to foreign media reports, a U.S. consulting firm conducted this real-world test. They rented four NVIDIA H200s and planned to have DeepSeek V4.1 Flash take over the coding work originally handled by Claude Opus 5.5, but in the end, the cost was not as low as everyone thought, and due to security issues, DeepSeek ultimately was not even used to write code.

AI-generated schematic diagram
Is DeepSeek really 80 times cheaper than Claude? A US company rented 4 H200s to self-host in a real-world test: $13,000 per month, more than double Claude’s bill.
According to foreign media Tom’s Hardware According to a report, The Call Center Doctors, a consulting firm that builds and operates call centers for other companies, recently shared its experience renting GPUs to run DeepSeek locally.
On September 27, they rented a server equipped with 4 H200s and switched their in-house Claude Code agent to DeepSeek V4.1 Flash, replacing the original Opus 5.5.
This machine’s normal on-demand price is $18.37 per hour, but they actually pay the spot price, only $9.19 per hour—half the price—though the trade-off is that the provider can reclaim the machine at any time.
And the entire setup process wasn’t as smooth as expected.
The report mentioned that every time DeepSeek V4.1 Flash is reloaded, it takes 10–15 minutes, and the first few times it crashed repeatedly due to GPU memory settings, communication issues between graphics cards, and launching too many agents at once.

The team went through several rounds of adjusting the settings, and only on the fifth attempt, by limiting the number of concurrently running agents to 128 and keeping GPU memory usage at around 90%, did the system finally stabilize.
Therefore, they believe that self-hosting large models isn’t as simple as “download the model, start it up, and it just works.” It also requires someone who understands GPUs, memory configuration, and system tuning, and can handle various issues at any time.
At first, they ran several rounds of one-minute full-load tests, each time testing only one kind of task, with zero errors throughout—it was pretty much perfect. But once they got into real programming work, problems began to appear.
AI doesn’t handle a completely new problem every time it writes code. As conversations and code grow longer, each time the model performs a task, it has to re-read the preceding content to know where it currently is.
After reviewing Claude Code usage logs from September, they found that as much as 96% of input tokens were spent rereading previous content. On average, each model call required rereading approximately 196,000 tokens.

AI-generated diagram
This also means that the bottleneck for self-hosting DeepSeek is not how many tokens it can generate per second; you also have to account for the compute resources consumed by repeatedly reading the context at scale.
After actual testing, a single machine equipped with 4 H200s can process only about 20 billion tokens per day, but on this company’s busiest day in September, Claude Code usage reached 51 billion tokens, meaning a single 4×H200 machine simply cannot meet peak demand.
Moreover, they also discovered cybersecurity issues in this experiment.
But to be clear, the problem isn’t the DeepSeek model itself, but the “sandbox” environment this company prepared for AI agents.
During the security review, the team repeatedly found ways to bypass restrictions, so the isolated environment was not secure enough, which ultimately led the company to not let DeepSeek modify production code.
Therefore, in this test, DeepSeek was mainly responsible only for “reading the code and finding issues.” The team also launched about 48–64 read-only review agents at the same time, scanned 2,377 folders, and ultimately produced 32 bug reports.
The entire test ran for about 3 hours and cost about $28 using Spot pricing; at standard On-Demand pricing, it would be about $55.
Below is the final result. If run for a full month, the cost of self-hosting DeepSeek would be far higher than the previous Claude Code subscription:
| Plan | Cost | Explanation |
|---|---|---|
| Claude Code | Approximately $5,500 USD/month | Company’s actual September bill |
| DeepSeek self-hosted (4×H200) | About 28 USD | About 3 hours of actual testing, using Spot price |
| Self-hosted DeepSeek (4×H200) | Approximately US$13,200 per month | Estimate full-month continuous rental based on On-Demand pricing. |
Of course, they didn’t actually run four H200s continuously at full load for a month; the $13,200 was just estimated based on on-demand pricing, so this test alone is not enough to conclude that “self-hosted models are necessarily more expensive than cloud AI.”
However, at least based on this hands-on test, in addition to GPU rental costs, self-hosting large models also means accounting for deployment, tuning, maintenance, and the time engineers spend. For this kind of high-usage Claude Code workflow, continuing to use cloud AI is indeed better.
Source: KOCPC Chinese