The Arc Pro B70 workstation graphics card launched by Intel in the first quarter of 2026 is not very well received in the market due to ecological reasons. This card is equipped with 32GB ECC GDDR6 memory and is officially priced at US$949 (approximately NT$30,800). Although the memory shortage has pushed the actual selling price to more than US$1,100 (approximately NT$35,750), it is still cheaper than NVIDIA. The 24GB configuration of RTX Pro 4000 is nearly half the price, and the price of the latter has soared over US$2,000 (approximately NT$65,000), which can be said to be a very cost-effective workstation-level graphics card.

The King of AI Price/Performance Intel Arc Pro B70
According to recent measurements by StorageReview, four B70s can be installed into the same workstation to create a 128GB shared VRAM. At a list price of approximately US$3,800 (approximately NT$123,500), a hybrid expert model with 120 billion parameters can be run. In the past, achieving the same capacity required several times the cost.

Battlemage Architecture: 32-core Xe2, 367 INT8 TOPS
Arc Pro B70 is based on Intel’s Battlemage architecture and uses the larger BMG-G31 chip, with built-in 32 Xe2 cores, 256 XMX engines, and INT8 AI computing performance of 367 TOPS. The memory configuration is 32GB GDDR6 with 256-bit bus and 608 GB/s bandwidth, and supports ECC error correction. Power consumption can be set between 160W and 290W, with the default setting being 230W.

Compared with the previous generation Arc Pro B60, Intel official data shows that the B70 has an average improvement of 44% in workstation workloads and a maximum improvement of 69% in professional applications. However, the B60 is a dual-GPU single-card design (x8/x8 split), while the B70 is changed to a single-GPU single card, which requires a complete PCIe 5.0 x16 channel, which will affect the chassis configuration when installing multiple cards.

In terms of appearance, the B70 adopts a dual-slot design, measures 267 × 99mm, and weighs about 1 kg. The single turbofan takes in air from the rear and exhausts hot air from the I/O baffle, making it suitable for densely installed workstation and server environments. The output interface is four DisplayPort 2.1, supporting up to 8K @ 120Hz. The power supply only requires an 8-pin PCIe connector, and the power requirements for the chassis are relatively friendly.

AI text generation: Llama 2 scored the highest score in the game
In the UL Procyon AI text generation benchmark test, the overall score of B70 increased by 60% compared to B50, and the first token delay was reduced by 40% to 45%. Compared to the Radeon RX 9060 XT in the AMD camp, the B70 is 224% better on the Phi model (4,152 vs. 1,281), 220% better on the Mistral, and 250% better on the Llama 3. Even against the higher-end RX 9070 XT, Intel still maintains an 83% to 100% lead. Against the NVIDIA camp, the RTX 5060 Ti lags behind the B70 by about 45%, and the RTX 5070 lags behind by 15% to 20%. The most eye-catching result appeared in the Llama 2 test: B70 scored the highest score in the game with 5,769 points, surpassing the RTX 5070 by 85% and surpassing the RX 9070 XT by 151%.
| UL Procyon: AI Text Generation | Intel Arc Pro B50 | Intel Arc Pro B70 | AMD Radeon RX 9060 XT | AMD Radeon RX 9070 | AMD Radeon RX 9070 XT | NVIDIA GeForce RTX 5060 Ti | NVIDIA GeForce RTX 5070 FE |
|---|---|---|---|---|---|---|---|
| Phi Overall Score | 2,593 | 4,152 | 1,281 | 1,933 | 2,080 | 2,870 | 3,453 |
| Phi Output Time To First Token | 0.275 s | 0.155 s | 1.473 s | 0.954 s | 0.855 s | 0.375 s | 0.323 s |
| Phi Output Tokens Per Second | 72.128 tokens/s | 104.121 tokens/s | 94.453 tokens/s | 139.187 tokens/s | 144.471 tokens/s | 120.773 tokens/s | 150.435 tokens/s |
| Phi Overall Duration | 39.179 s | 37.924 s | 39.365 s | 26.989 s | 25.587 s | 25.216 s | 20.302 s |
| Mistral Overall Score | 2,483 | 4,082 | 1,274 | 2,040 | 2,231 | 2,807 | 3,562 |
| Mistral Output Time To First Token | 0.346 s | 0.180 s | 1.827 s | 1.109 s | 0.946 s | 0.526 s | 0.433 s |
| Mistral Output Tokens Per Second | 46.799 tokens/s | 65.834 tokens/s | 65.115 tokens/s | 101.300 tokens/s | 103.348 tokens/s | 91.057 tokens/s | 120.507 tokens/s |
| Mistral Overall Duration | 59.907 s | 58.976 s | 54.516 s | 34.960 s | 33.350 s | 33.377 s | 25.496 s |
| Llama3 Overall Score | 2,427 | 4,029 | 1,150 | 1,904 | 2,070 | 2,599 | 3,125 |
| Llama3 Output Time To First Token | 0.311 s | 0.166 s | 1.632 s | 0.981 s | 0.845 s | 0.449 s | 0.379 s |
| Llama3 Output Tokens Per Second | 45.031 tokens/s | 66.340 tokens/s | 53.167 tokens/s | 87.594 tokens/s | 89.102 tokens/s | 74.709 tokens/s | 100.388 tokens/s |
| Llama3 Overall Duration | 61.926 s | 53.687 s | 62.563 s | 38.273 s | 36.742 s | 39.489 s | 29.720 s |
| Llama2 Overall Score | – | 5,769 | 1,252 | 2,047 | 2,298 | 2,576 | 3,125 |
| Llama2 Output Time To First Token | – | 0.259 s | 2.992 s | 1.926 s | 1.565 s | 0.844 s | 0.785 s |
| Llama2 Output Tokens Per Second | – | 63.666 tokens/s | 34.654 tokens/s | 59.673 tokens/s | 61.127 tokens/s | 41.386 tokens/s | 56.647 tokens/s |
| Llama2 Overall Duration | – | 44.131 s | 99.027 s | 59.100 s | 55.520 s | 71.302 s | 53.234 s |
Image generation: on par with RTX 5060 Ti
In the Stable Diffusion 1.5 FP16 test, the B70 scored 2,101 points, which is almost identical to the RTX 5060 Ti’s 2,110 points (the difference is less than 1%). Compared to the B50’s 754 points, the improvement is 179%, and the generation time is shortened from 132.6 seconds to 47.6 seconds. On the heavier Stable Diffusion XL FP16 test, the B70 was also nearly three times faster than the B50, with a slight 8% lead over the RTX 5060 Ti.
| UL Procyon: AI Image Generation (overall score: higher is better) |
Intel Arc Pro B50 | Intel Arc Pro B70 | AMD Radeon RX 9060 XT | NVIDIA GeForce RTX 5060 Ti | AMD Radeon RX 9070 | AMD Radeon RX 9070 XT | NVIDIA GeForce RTX 5070 FE |
|---|---|---|---|---|---|---|---|
| Stable Diffusion 1.5 (FP16) — Overall Score | 754 | 2,101 | 1,436 | 2,110 | 2,280 | 2,598 | 2,937 |
| Stable Diffusion 1.5 (FP16) — Overall Time | 132.585 s | 47.585 s | 69.633 s | 47.590 s | 43.858 s | 38.481 s | 34.038 s |
| Stable Diffusion 1.5 (FP16) — Image Generation Speed | 8.287 s/image | 2.974 s/image | 4.352 s/image | 2.974 s/image | 2.741 s/image | 2.405 s/image | 2.127 s/image |
| Stable Diffusion 1.5 (INT8) — Overall Score | 5,020 | 18,344 | N/A | 27,705 | N/A | N/A | 36,320 |
| Stable Diffusion 1.5 (INT8) — Overall Time | 49.795 s | 13.628 s | N/A | 9.024 s | N/A | N/A | 6.883 s |
| Stable Diffusion 1.5 (INT8) — Image Generation Speed | 6.224 s/image | 1.703 s/image | N/A | 1.128 s/image | N/A | N/A | 0.860 s/image |
| Stable Diffusion XL (FP16) — Overall Score | 748 | 2,102 | 1,124 | 1,940 | 1,805 | 2,010 | 2,473 |
| Stable Diffusion XL (FP16) — Overall Time | 790.774 s | 285.344 s | 533.736 s | 326.550 s | 332.400 s | 298.499 s | 242.606 s |
| Stable Diffusion XL (FP16) — Image Generation Speed | 49.423 s/image | 17.834 s/image | 33.359 s/image | 20.409 s/image | 20.775 s/image | 18.656 s/image | 15.163 s/image |
However, in INT8 quantization workloads, NVIDIA’s TensorRT implementation still has a clear advantage. The RTX 5070 generates nearly twice as fast as the B70 at INT8 Stable Diffusion 1.5 (0.86 seconds/picture vs. 1.7 seconds/picture). This reflects that Intel still has a gap in software optimization on the quantitative reasoning path.
vLLM multi-card actual test: four B70s beat one RTX Pro 6000
StorageReview loaded four B70s (128GB shared VRAM) into a Supermicro AS-4125GS-TNRT server (with AMD EPYC 9374F, 512GB DDR5) and ran vLLM online inference benchmarks against a single RTX Pro 6000 and DGX Spark. The reason is practical: The total price of four B70s is about the same as an RTX Pro 6000 or a DGX Spark, and it makes more sense to compare by budget rather than card count.

On the Mistral Small 24B model, the B70 stands out most. As the number of concurrency increases, B70 climbs from 450 tok/s in batch size 1 to 8,321 tok/s in batch size 32. The RTX Pro 6000 has a slight advantage at low concurrency, but the B70 starts from batch size 8 and leads by 65%. Compared to B60, B70 is about 26% faster at maximum concurrency. DGX Spark lags far behind, peaking at just 527 tok/s, and the B70 is nearly 16 times faster at maximum load.

On Llama 3.1 8B, the B70 reaches nearly 12,000 tok/s, which is 16% faster than the B60, about 85% of the RTX Pro 6000, and 4.7 times faster than the DGX Spark. On the larger Qwen3 Coder 30B and GPT-OSS-120B models, the RTX Pro 6000 maintains its lead, but the B70 is still a solid 7 to 8 percent faster than the B60 and continues to lead the DGX Spark by a wide margin.


List of other benchmark test results
In a number of professional and computing benchmark tests, the B70 is positioned in the mid-to-high-end range:
- Geekbench 6 OpenCL: 140,165 points, 100% faster than the B50, 36% faster than the RX 9060 XT, and about 24% behind the RTX 5070.
- LuxMark ray tracing rendering: Food scene 5,609 points, Hall scene 12,220 points, 128% to 137% faster than B50, 33% to 53% faster than RX 9060 XT.
- 3DMark Port Royal: 10,668 points, just 2% shy of the RTX 5060 Ti’s 10,432 points.
- Blender 4.5:Monster scene 1,524.6 samples/min, enough to handle professional visualization and content creation work.
- Topaz Video AI: AI video upscaling can reach 8.64 FPS (Gaia model), covering complete workflows such as video enhancement, noise reduction, HDR conversion and frame rate interpolation.
The above information shows that B70 also has certain capabilities in professional applications other than AI inference, but its core selling point has always been the 32GB large-capacity VRAM paired with a very competitive price.
Software Ecosystem: Intel’s Biggest Concern
StorageReview repeatedly emphasized in its reviews that the B70’s hardware is impressive, but the software ecology is the key holding it back. Intel’s LLM Scaler (a development fork of Battlemage based on vLLM) is still in beta, has limited model coverage, and some quantization paths should work but don’t yet.
Another outlet, Puget Systems, also concluded in an April reviewSimilar conclusion: “With 32GB VRAM and a price of US$950, the B70 is positioned as an AI inference card rather than a general professional workstation card.” In other words, if you want to use the B70 to run a local large-scale language model, its cost-effectiveness is indeed outstanding; but if you are looking forward to a “plug it and go” experience like NVIDIA, you can’t do it yet.
NVIDIA’s moat isn’t just CUDA, it’s documentation, examples, available container images, Q&A history on forums, and a community knowledge base that can solve most problems. AMD’s ROCm is also moving in the same direction. StorageReview believes that Intel needs to show the same level of commitment to Arc Pro, otherwise the B70 may become a card that “players want to love but teams find difficult to adopt.”
Who should buy and who should wait
For buyers who value VRAM capacity, workstation capabilities, and AI price/performance, the B70 is one of the best options on the market today to consider. There is almost no substitute for the 128GB solution with four cards at the same price. If your main need is to run large-scale language model inference locally and you are willing to spend time dealing with software compatibility issues, the hardware foundation of B70 is sufficient.
But if you need an “out of the box” experience, NVIDIA and AMD are still safer choices. Intel’s software ecosystem will take time to catch up, and there is currently no clear answer to this timetable. If Intel can put more effort into the ecosystem or actively cooperate with the open source ecosystem (AMD has been very active recently), I believe their graphics cards are still very cost-effective.