Hi everyone, I just recently did an RMA and many tests to try and fix thins “hanging” issue. All the details and system information below. Any suggestions or people who had similar issues, along with possible fixes would be appreciated. Thank you!:
Date: 2026-08-05 Reporter: end user (extensive 3-week investigation)
SHORT SUMMARY
On a Ryzen 7 9800X3D + MSI MAG X870E Tomahawk WiFi (BIOS 2.AC3 / AGESA
PI-1.3.0.1b Patch A), the CPU’s reported Tctl/Tdie FREEZES at a fixed
value (observed 64-95C) after sustained combined CPU+GPU load. The frozen
value is served identically to BOTH independent read paths — the board EC
(SB-TSI: fan control + EZ Debug LED) AND software readers (SMN: HWiNFO,
MSI Center, L-Connect) — proving the freeze occurs at the common source,
the SMU telemetry function. The CPU’s control plane is unaffected (boost,
clocks, scheduling all normal; machine fully usable). Only a reboot (SMU
reset) restores the reading. Reproduces at DDR5-6000 (EXPO), 5600 (JEDEC
manual), and once at 4800 — probability scales with MEMCLK. Persists
across two BIOS/AGESA versions, three GPU driver versions, chipset driver
versions, cooler/RAM/GPU reseats, a board RMA, ASPM-off, PCIe Gen4 cap,
and Power Supply Idle Control=Typical. An identical report from another
user with the same board+CPU+RAM combination suggests a platform/AGESA
firmware issue rather than a defective unit.
SYSTEM CONFIGURATION
CPU: AMD Ryzen 7 9800X3D (8C/16T, SMT on, PBO Auto)
Board: MSI MAG X870E Tomahawk WiFi (MS-7E59) v2.0
BIOS: E7E59AMSI 2.AC3, 2026-06-24 (AGESA PI-1.3.0.1b Patch A)
- also reproduced on 2.A80 (June 2025 AGESA 1.2.0.x)
Memory: Corsair CMH32GX5M2B6000Z30 2x16GB (EXPO 6000 CL30 kit)
- reproduced at: EXPO 6000 CL30, Memory Try It 6000 CL32
(32-36-36-76 and 32-40-40-86), manual 5600 (Auto/JEDEC
timings), and once at JEDEC 4800
- current daily config: manual DDR5-5600, EXPO off
GPU: MSI RTX 5070 Ti Gaming Trio OC Plus, driver 610.88 (also .62/.74)
PCIe slot: reproduced at Gen5 (Auto) and Gen4 (forced)
Chipset: AMD chipset drivers 8.05.04.516
PSU: MSI MAG 1000W Gold ATX 3.1, native 12V-2x6 GPU cable
OS: Windows 11 Pro build 26200.8973
Cooling: be quiet! Dark Rock Pro 5 (verified mount, fresh paste);
CPU idles ~40C; all-core sustained ~87-95C (self-limiting)
Software: no sensor/RGB software resident (zero ring-0 sensor drivers);
reproduced both with and without such tools installed
SYMPTOM DETAIL
- Under sustained simultaneous CPU (all-core) + GPU (CUDA) load lasting
~5-20+ minutes, the reported CPU temperature (Tctl/Tdie) freezes at
whatever value it held at the moment of the hang (observed: 64, 69,
87, 89, 92, 94, 95C on different occasions). - After load ends, the frozen value persists indefinitely. Fan control
(EC) pins CPU-temp-sourced fans at the curve position for the frozen
value. The board’s EZ Debug CPU LED sometimes lights (inconsistent). - The die is demonstrably cool/idle while frozen (CPU usage 2-5%,
all software agrees on the same frozen number, machine performs
perfectly, benchmarks normal). - Only a reboot restores live telemetry. Sleep/resume not tested as fix.
- No OS event of any kind is logged at latch onset: zero WHEA, zero
Kernel-Power, zero thermal ACPI events, zero
Kernel-Processor-Power firmware-limit events. Machine never BSODs. - Frequency: near-deterministic after 15-20 min of synthetic combined
load (all-core AVX matmuls + CUDA training) at 6000 MT/s; less
frequent but reproduced at 5600; observed once at 4800. Real-world
workloads (gaming, GPU-only ML training) at 5600 have not triggered
it; at 6000-family configs they did (latches after gaming sessions,
after a full day of desktop use).
REPRODUCTION RECIPE (near-deterministic)
- DDR5-6000 (EXPO) for fastest repro; also works at 5600 eventually.
- Run simultaneously for 15-20 minutes:
a) all-core FP32 AVX load on all 16 threads (e.g., continuous
4096x4096 torch/oneDNN matmuls), CPU pinned at 100%
b) sustained CUDA load (e.g., YOLOv8m training, ~170-200W GPU) - Stop both loads. Observe reported Tctl remains frozen at the
load-time value while the die idles. All readers (BIOS/EC fan
control, any monitoring app) show the same frozen number. - Reboot → reading restored.
KEY EVIDENCE THAT THE FREEZE IS AT THE SMU (NOT OS/BOARD/READERS)
- TWO INDEPENDENT READ PATHS show the identical frozen value
simultaneously: the board EC (SB-TSI hardware path driving fans/LED,
no OS involvement) and software readers (SMN/driver path). Their only
common element is the SMU telemetry source. - The SMU CONTROL plane keeps working: boost to 5.18 GHz all-core during
the same sessions, normal CPPC behavior, no performance-limit events. - Survives/reproduces across: 2 BIOS/AGESA generations, 3 GPU drivers +
clean DDU reinstall, 2 chipset driver versions, presence/absence of
all monitoring software (incl. zero ring-0 driver state), ASPM
disabled, PCIe Gen4 forced, Power Supply Idle Control=Typical Current
Idle, cooler remount + repaste, RAM reseat/slot swap, GPU reseat,
and a motherboard RMA service (MSI replaced a USB3 component). - Community corroboration: independent user, same board (X870E
Tomahawk) + 9800X3D + Corsair CAS30 EXPO kit, reports the same stuck
temperature readout across all software, tied to EXPO, on an earlier
BIOS. Suggests platform-level firmware issue.
POSSIBLY-RELATED SECONDARY OBSERVATION (same I/O die)
1,221 Windows LiveKernelEvents (WER Id 1001) since 2026-07-11: GPU
video-engine timeout codes 0x141/0x117/0x193/0x1a8/0x1b0/0x1b8,
clustering during the same sustained combined-load sessions (plus a
constant low-level daily baseline), across all three NVIDIA driver
versions and both PCIe Gen settings. Intermittent CUDA failures
(“CUDA error: unknown”, cudnn EXECUTION_FAILED) with whole-GPU ~90s
dropouts (nvidia-smi loses the device, then self-recovers, usually
without an nvlddmkm System-log event). May be an independent NVIDIA
driver issue, but is noted because PCIe root and SMU share the CPU I/O
die and the symptoms co-occur under the same trigger load.