• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - AI Trends and Related News - Kimi K3, the strongest open source model, reaches the top of the front-end coding arena: netizen feedback outperforms Claude Fable 5 and GPT-5.6 Sol

Kimi K3, the strongest open source model, reaches the top of the front-end coding arena: netizen feedback outperforms Claude Fable 5 and GPT-5.6 Sol

KOCPC Editor by KOCPC Editor
July 17, 2026 - Updated on August 5, 2026
in AI Trends and Related News, Latest Technology News

Chinese AI startup Moonshot AI was officially released early this morning Kimi K3, known as “Open Frontier Intelligence”. This model has 2.8 trillion parameters, 1 million token context windows, and native multi-modal capabilities. It is the largest open source weight model in history. Within one day of publication, Kimi K3 took first place in Arena.ai’s Frontend Code Arena with 1679 points, surpassing Anthropic’s Claude Fable 5 and OpenAI’s GPT-5.6 Sol. This is the first time that an open source model has surpassed all closed source models in this type of arena. Vercel CEO Guillermo Rauch alsoConfirm, Kimi K3 ranked first in its web engineering benchmark test, achieving “comparable success rate in less time” performance.

Model architecture: sparse design with 2.8 trillion parameters

Kimi K3 adopts a Mixture of Experts architecture with a total parameter volume of 2.8 trillion, but only 16 of the 896 experts are enabled for each inference, controlling computing costs through a highly sparse approach. This design allows the model to maintain a reasonable reasoning speed while maintaining a huge knowledge capacity. Moonshot AI employs two core architectural innovations for K3. The first is Kimi Delta Attention (KDA), which can achieve up to 6.3x decoding acceleration in the context of millions of tokens. The second is Attention Residuals technology, which brings about a 25% improvement in training efficiency at an additional cost of less than 2%. The combination of these two technologies increases the overall expansion efficiency of K3 by 2.5 times compared with the previous generation K2.

In addition, K3 is equipped with Moonshot’s self-developed Mooncake decentralized reasoning architecture, which can achieve more than 90% context cache hit rate in coding scenarios, significantly reducing actual API input costs.

Benchmark Results: A Comprehensive Showdown Against the Top Closed Source Models

In the Artificial Analysis Smart Index, Kimi K3 scored 57 points, ranking fourth among 189 models, behind Claude Fable 5 and GPT-5.6 Sol, and at the same level as Claude Opus 4.8 and GPT-5.5.

The following is the key information for each benchmark test:

  • Frontend Code Arena: 1679 points, ranked 1st, surpassing Claude Fable 5th. Taking 6 first places in 7 subcategories, only losing out to Fable 5 in the gaming category.
  • GDPval v2 (agent task): Elo score 1668, a big jump from K2.6’s 1190, surpassing GLM-5.2 (1514), GPT-5.5 (1494) and Claude Opus 4.8 (1600), but still behind Claude Fable 5 (1760).
  • AutomationBench-AA: Scored 53% and ranked #1 for a Zapier Agent SaaS Workflow evaluation.
  • AA-Briefcase (knowledge work): Elo 1547, 732 points higher than K2.6, second only to Claude Fable 5.
  • Terminal-Bench 2.1: 88.3 points, slightly lower than GPT-5.6 Sol’s 88.8, but higher than Claude Fable 5’s 84.6.
  • token efficiency: Only about 132 million output tokens were used to complete all 9 evaluations of Artificial Analysis, which was 21% less than the 166 million of K2.6.

Even Elon Musk was impressed by the performance of Kimi K3:

Independent tester Paweł Huryn put the K3 through an 8-task test batteryActual measurement, the results show that K3 reaches or exceeds the performance of Opus 4.8, GPT-5.6 and Grok 4.5 in 6 of the 7 tasks completed. In the test of 21 embedded bugs, K3 found 14, Opus found 12, and no other model had more than 7.

Official display: Stunning case of independent development capabilities

In addition to benchmark test numbers, Moonshot AI demonstrated multiple K3 autonomous completion cases on its technology blog, which triggered a lot of discussion in the community. Most notably, K3 built MiniTriton, a GPU compiler from scratch. It defines its own intermediate representation on MLIR, implements a complete optimization process and PTX code generation, and its performance in the roofline benchmark is close to PyTorch’s Triton and torch.compile. More importantly, the team used this compiler to train nanoGPT and observed normal convergence, proving that K3 does not just write several independent kernels, but builds a usable compiler tool chain.

K3 even independently designed a chip to run its own architecture micro-model during a 48-hour continuous autonomous agent operation, using open source EDA tools and Nangate 45nm process library. The finished product has an area of ​​4 square millimeters, integrates 1.46 million standard cells and 0.277 MB SRAM, the clock converges at 100 MHz, and the simulation inference throughput exceeds 8,700 tokens per second. “A model designs a chip that serves the model,” Moonshot AI wrote in the announcement.

In terms of creative applications, K3 used Three.js and WebGPU programs to generate a browser-side 3D open world game, with scenes including forests, wooden cabin villages, snowy mountains and dynamic weather systems. It also showcases a 3D Long March 10 rocket launch and recovery simulation, a Game Boy Advance emulator, and a recreation of the Gargantua black hole from Star Effect.

Lets end this debate

Kimi K3 mogs opus 4.8 in every scenario , and its literally better than fable 5 in webdev and game making

we are in the era of abundance now https://t.co/OohJ6eQukq pic.twitter.com/JwQOzy9Ugl

— Chetaslua (@chetaslua) July 17, 2026

I tested Kimi K3 vs Claude Opus 4.8

Same prompt, an armory bay with lighting, props, and detail. Top is Kimi K3, bottom is Opus 4.8.

It’s not even close.

Kimi K3 built a full scene with textures, proper lighting, ammo crates, weapon racks, working detail everywhere. Opus 4.8… pic.twitter.com/jnZiEqoRfb

— Bhavy☄️ (@Bhavani_00007) July 16, 2026

Kimi K3 built a browser-based 3D martial arts

RPG featuring melee combat, quests, inventory, dynamic weather, and explorable interiors.

It modeled the game environment in Blender with improved collisions and PBR retexturing, integrated and adapted open assets, and designed… https://t.co/ioS67t6YEG pic.twitter.com/USuzuSy69V

— Chetaslua (@chetaslua) July 16, 2026

In terms of scientific research, K3 completed the reproduction of the universal I-Love-Q relationship in computational astrophysics in about two hours, a task that would normally take a senior researcher one to two weeks. It reads and cross-validates over 20 papers, evaluates over 300 state equations, generates over 3,000 lines of Python code, and comes with an interactive HTML dashboard.

Moonshot AI also revealed an interesting detail: “In the later stages of Kimi K3 development, early versions of K3 were already responsible for handling most of the team’s kernel optimization work.” K3 has been involved in development work since its birth.

Pricing strategy: sell open source models at the same price as closed source models

Kimi K3’s API pricingThe input token is USD 3 (approximately NT$97) per million, the output token is USD 15 (approximately NT$487) per million, and the cache input is discounted to USD 0.3. In comparison, GPT-5.6 Sol is priced around $5/30 USD, and Claude Fable 5 is priced at $10/50 USD. K3’s cost per mission is about $0.94, which is close to GPT-5.6 Sol’s $1.04 and about half of Opus 4.8’s $1.80.

Moonshot AI’s pricing strategy shows that they no longer sell K3 as a “cheap Chinese alternative”, but as a cutting-edge model to directly compete with Anthropic and OpenAI. This is the company’s highest-priced model in history. The output token price of the previous generation K2.6 was only $4, and the price of K3 increased by nearly 4 times.

Market Impact of Open Weights

Moonshot AI plans to release K3’s complete model weights before July 27, which will make it the world’s largest open source weight model, far exceeding DeepSeek V4 Pro’s 1.6 megabytes and GLM-5.2’s 753 billion parameters. Once the K3 weight is made public, anyone with enough hardware can provide the same model service, and the price will be reduced to the level of “service cost plus small profit”. A set of GB200 NVL72 racks (approximately $3.5 million) executing K3 at 80% utilization can produce approximately 63 billion output tokens per month at an annualized cost of approximately $2 million. If the same amount of tokens is purchased at a closed-source API price, the annual cost will be more than $11 million, and self-building is more than 6 times cheaper than renting.

Community reaction: Excitement and skepticism coexist

The publication of K3 sparked a lot of discussion on X. Positive reviews focused on the open source model’s first entry into the front-end coding arena. technology creator @learnaifaster The five highlights of K3 are compiled: No. 1 in front-end programming, tied for No. 4 with Opus in overall intelligence, leading in web browsing and spreadsheet automation, millions of token contexts, and about to be open source.

@aakashgupta From the perspective of Moonshot AI’s company, it is pointed out that this is a story of “almost being brought down by DeepSeek and then getting back on its feet.” The release of DeepSeek R1 18 months ago caused Kimi to fall from China’s third chatbot to seventh place. Now K3 has regained the first place in front-end programming. Moonshot AI completed US$2 billion in financing at a US$20 billion valuation in May. In July, the FT reported that it was raising a new round of financing at a US$31.5 billion valuation. The release of a single model increased the company’s valuation by 57%.

However, there are also more cautious voices. Bindu Reddy, head of the LiveBench benchmark, pointed out that in their hidden question test, K3 was the best open source model, but still lower than Opus 4.8, Sol and Fable. She also mentioned that K3 “spins a lot” in practice, costs about the same as Opus 4.8, and is slower.

KIMI K3 CLOSES THE GAP BUT RANKS BEHIND FRONTIER MODELS

Our benchmark, LiveBench, has a lot of hidden questions and models can’t memorize them

K3 is the best good open-source model but is below Opus 4.8, Sol and Fable

Also in practice, Kimi spins a lot and costs as much as… pic.twitter.com/QJtFeAJoL5

— Bindu Reddy (@bindureddy) July 17, 2026

Paweł Huryn’s actual test also revealed a practical problem: on the night of release, two tasks returned zero tokens and had no output at all. One of the retests received full marks after 13 hours, but the other one still failed, which means that after Kimi K3 went online, Moonshot AI’s servers could not withstand such a huge global demand (my actual test speed was really not fast).

Conclusion

The release of Kimi K3 is the first time an open source large language model has matched or exceeded the performance of closed source cutting-edge models on multiple key benchmarks. The sparse architecture of 2.8 megabytes of parameters, millions of token contexts, and relatively reasonable pricing make it the most competitive open source choice currently. However, Moonshot AI itself admits that K3 still lags behind Claude Fable 5 and GPT-5.6 Sol in terms of overall performance, and its climb to the top is mainly concentrated in the field of front-end programming code.

The open source weight release on July 27 will be the next key node. If the weights are made public as expected, K3 will directly change the way companies calculate costs between closed-source APIs and self-built services. The same GPU went from being rented to cutting-edge laboratories to earn US$80,000 per year, to becoming a service open source model and earning US$14,000 to US$34,000 per year, and the payback period was extended from less than 1 year to 2 to 5 years. This pricing war between open source and closed source may have just begun.

Source: KOCPC Chinese

Tags: aiKimiKimi K3Moonshot AIOpen source model

Recent Posts

  • The Xiaomi Pad 8S Pro has passed network access certification and will debut with the self-developed XRING O3 chip.
  • The entire Google Pixel 11 lineup has been leaked! Official promotional renders of the Pixel 11 Pro XL have also surfaced
  • Are Chinese phone battery capacities falsely labeled? A brief look at the “capacity locking” phenomenon in Chinese silicon-carbon batteries.
  • NCC is leaderless, recklessly sending out national-level alert messages!?
  • What does “QR” in QR Code mean?

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology