DeepSeek's move to put training on AMD was fully completed yesterday

📡 AI News 2026-08-16 4 min read

DeepSeek moved its training onto AMD, and yesterday the whole thing was fully completed. Yesterday in the AI circle, there was something that looked purely technical but was actually quite significant. The LMSYS team released an engineering report yesterday, stating that the reinforcement learning training for DeepSeek's latest-generation V4 Flash model—this step is the critical training phase where a large model goes from "being able to chat" to "being able to solve problems"—the entire pipeline has already been

💡 What You Will Learn

DeepSeek moved its training onto AMD, and yesterday the whole thing was fully completed. Yesterday in the AI circle, there was something that looked purely technical but was actually quite significant

DeepSeek's Move to Put Training on AMD — Fully Completed Yesterday

There was something in the AI world yesterday that looked purely engineering-focused, but was actually pretty significant.

The LMSYS team published an engineering report yesterday. It said that DeepSeek's latest-generation V4 Flash model's reinforcement learning training — the critical training phase where a model goes from "being able to chat" to "being able to solve problems" — has been run end-to-end on AMD's MI355X GPUs.

The report has 9 figures total, from training pipeline diagrams and architecture diagrams to precision comparison charts and reward curves, breaking down every step for you.

Why is this worth talking about? Not because AMD suddenly rose to prominence, but because domestic large-model training has been almost entirely dependent on a single thread — NVIDIA — for the past few years. GPU models, VRAM, and the surrounding software ecosystem all revolve around one company.

This time, DeepSeek has paved that path.

A Few Key Numbers

The model has 284 billion total parameters, but only activates 13 billion per response. Training used 4 nodes, with 8 AMD MI355X GPUs per node — 32 GPUs in total.

Over 100+ training steps, the "response probability" and "scoring probability" aligned, with the gap staying stable around 0.09 — no drift.

The final result on the AIME-2024 math problem set: single-attempt accuracy improved from 0.39 to 0.49, and 8-sample coverage improved from 0.53 to 0.67. The rate of truncated responses dropped from 60% to 55% — meaning the model's ability to "finish writing" also improved.

These numbers aren't mind-blowing on their own.

But they prove one thing: AMD's hardware path has gone from "it can run" to "it can reliably train models that produce results."

What's Actually Hard About This

The hard part isn't "getting the model to run" — it's "getting two different engines to produce nearly identical answers to the same problem."

You can think of reinforcement learning as "letting the model self-grade while solving problems." Get it right, and it rewards itself. Get it wrong, and it penalizes itself. Then it adjusts its reasoning based on the scores. But there's a prerequisite: the "answer generator" responsible for producing candidate responses and the "scoring engine" responsible for grading and tuning must output nearly identical probabilities for the same problem.

If there's even a slight mismatch, the model goes off track.

Why has NVIDIA's CUDA ecosystem dominated large-model training for years? Because this probability-alignment engineering between the "answer generator" and "scoring engine" has been polished to near-perfection on CUDA. Switching to AMD's ROCm stack means recalibrating every precision detail, every weight synchronization, and every round of multi-node communication from scratch.

Plenty of teams domestically and internationally have wanted to do this. But whenever it touches engineering details like reinforcement learning, mixed precision, or multi-node coordination, they almost all fall back to NVIDIA.

DeepSeek is the first to fully complete this pipeline.

What It Means for the Domestic AI Stack

It's a "hardware insurance policy."

Over the past few years, the training pace of domestic large models has been too tightly coupled to NVIDIA's supply rhythm. GPU production capacity, export controls, VRAM upgrade cycles — if any single variable hiccups, the entire training chain stalls.

Now that AMD's path has been proven to produce results, it means the next time you're buying new GPUs, negotiating prices, or building domestic inference services, you have another card in hand.

The impact on the consumer side will come more slowly, but the direction is already set. When the training side is no longer locked to a single vendor, hardware choices, pricing, and API costs on the inference side will loosen up. The next time you open a domestic AI assistant, the GPUs running behind it might no longer be a uniform model.

This isn't about money. It's about choice.

What's worth watching next is full-pipeline FP8 training, performance tuning, and validation at larger scale. Now that DeepSeek has paved this path, teams like Xiaohongshu, Kuaishou, Meituan, and Baidu — which are already using domestic GPUs for inference — will likely follow this route and migrate their training side to the AMD stack as well.

That day probably isn't too far off.

[Image suggestions] figure-1-pipeline.png (training pipeline breakdown) goes before the "What's Actually Hard About This" section; figure-2-metrics.png (data comparison) goes before "A Few Key Numbers"; figure-3-ecosystem.png (ecosystem diffusion) goes before "What It Means for the Domestic AI Stack." The three figures tell the story in sequence: how hard it is → what results were achieved → what happens next.

Related Articles
2026-09-20
Office 2021 stops receiving updates next month, but don't rush to switch to a subscription just yet
2026-09-11
CXMT scores two wins in one week: world's first LPDDR6, HBM3E also begins trial production
2026-07-13
Intel Nova Lake Leaked! 52 Cores at 700W, Z990 with TB5, Set to Battle Zen 6 in 2027
2026-08-03
I Built a Budget AI Agent Setup — Copy My Cheat Sheet
2026-09-11
NVIDIA DLSS 5 rolls out today: performance increases 5x in 6 months, but enabling it cuts frame rate in half
2026-08-12
Vera just started shipping, and NVIDIA has already leaked the next-gen CPU—three generations in three years, who can keep up with this pace

Written by our editorial team; tools listed here are tested or verified against public sources. Links point to official sites or GitHub repos for reference only — no paid placements.

💬 Comments (0)

No comments yet. Be the first!

Login to comment