CoreFreq possibly caused issue during package update
Today I wanted to update the following packages on my system (6.17.9-2-cachyos, AMD Ryzen 9 9950X3D) while corefreq was running:
`Packages (4) imv-5.0.1-1.1 libbpf-1.6.2-1.1 libinput-1.30.0-2.1 noto-fonts-1:2025.12.01-1`
After the installation the upgrade process was stuck at
`(1/8) Reloading device manager configuration...`
From there on I had issues entering commands (huge delays typing, sending) through my ssh session or authenticating (additional sessions, sudo). ̀ uptime` showed a load of 18. Eventually I managed to reboot the machine and saw the following journal entries for the time when the issue started:
```
rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:
rcu: 10-...0: (1 GPs behind) idle=9f6c/1/0x4000000000000000 softirq=3076/3078 fqs=11519
rcu: (detected by 2, t=60122 jiffies, g=4253, q=2287 ncpus=32)
Sending NMI from CPU 2 to CPUs 10:
NMI backtrace for cpu 10
CPU: 10 UID: 0 PID: 4399 Comm: udevadm Tainted: G S OE 6.17.9-2-cachyos #1 PREEMPT(full) 81ce6cd0386df3125c20562ac914b09f65df2f37
Tainted: [S]=CPU_OUT_OF_SPEC, [O]=OOT_MODULE, [E]=UNSIGNED_MODULE
Hardware name: Micro-Star International Co., Ltd. MS-7E51/MAG X870 TOMAHAWK WIFI (MS-7E51), BIOS 1.A69 09/19/2025
RIP: 0010:native_queued_spin_lock_slowpath+0x9b/0x2b0
CoreFreq: AMD_SMN_Read(0, 59808) TryLock
Code: c0 f0 0f ba 2b 08 0f 92 c0 b9 ff 00 ff ff 23 0b 89 c6 c1 e6 08 09 f1 81 f9 00 01 00 00 73 1f 85 c9 74 0c 80 3b 00 74 07 f3 90 <80> 3b 00 75 f9 66 c7 03 01 >
RSP: 0018:ffffcf7680530d88 EFLAGS: 00000002
RAX: 0000000000000000 RBX: ffff8cc2c1629d90 RCX: 0000000000000001
RDX: ffffcf7680530e10 RSI: 0000000000000000 RDI: ffff8cc2c1629d90
RBP: 0000000000000000 R08: 00000017dee4a627 R09: 000000181a7f7027
R10: 00000081e06f4cd2 R11: ffffffffc20802d0 R12: 0000000000000000
R13: ffffffffb8949490 R14: ffffcf7680530e10 R15: 0000000000000001
FS: 00007fcd100d0880(0000) GS:ffff8cce14388000(0000) knlGS:0000000000000000
CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 000056061384d058 CR3: 000000010b123000 CR4: 0000000000f50ef0
PKRU: 55555554
Call Trace:
<IRQ>
? __pfx_match_pci_dev_by_id+0x10/0x10
_raw_spin_lock+0x2a/0x40
bus_find_device+0x53/0x170
pci_get_domain_bus_and_slot+0x7c/0x190
CCD_AMD_Family_19h_Zen4_Temp+0xca/0x380 [corefreqk 1a38744b46bc721813b1ee40277bc2f00344fdb1]
CoreFreq: AMD_SMN_Read(0, 59808) TryLock
CoreFreq: AMD_SMN_Read(0, 59808) TryLock
CoreFreq: AMD_SMN_Read(0, 59808) TryLock
CoreFreq: AMD_SMN_Read(0, 59808) TryLock
CoreFreq: AMD_SMN_Read(0, 59808) TryLock
? __pfx_Call_Raphael+0x10/0x10 [corefreqk 1a38744b46bc721813b1ee40277bc2f00344fdb1]
Entry_AMD_F17h+0x445/0x620 [corefreqk 1a38744b46bc721813b1ee40277bc2f00344fdb1]
? __pfx_Call_MSR_ACCU+0x10/0x10 [corefreqk 1a38744b46bc721813b1ee40277bc2f00344fdb1]
? __pfx_Cycle_AMD_Zen4_RPL+0x10/0x10 [corefreqk 1a38744b46bc721813b1ee40277bc2f00344fdb1]
CoreFreq: AMD_SMN_Read(0, 59808) TryLock
__hrtimer_run_queues+0xf4/0x2e0
hrtimer_interrupt+0xf9/0x3d0
__sysvec_apic_timer_interrupt+0x4c/0x170
sysvec_apic_timer_interrupt+0x6f/0x80
</IRQ>
<TASK>
asm_sysvec_apic_timer_interrupt+0x1a/0x20
RIP: 0010:_raw_spin_lock+0x1b/0x40
Code: 90 90 90 90 90 90 90 90 90 90 90 90 90 90 90 f3 0f 1e fa 0f 1f 44 00 00 65 ff 05 20 cc e3 01 b9 01 00 00 00 31 c0 f0 0f b1 0f <75> 06 c3 cc cc cc cc cc 89 >
RSP: 0018:ffffcf7685d0bce0 EFLAGS: 00000246
RAX: 0000000000000000 RBX: ffff8cc2c19f9020 RCX: 0000000000000001
RDX: 0000000000000000 RSI: ffff8cc2ccff6000 RDI: ffff8cc2c1629d90
RBP: ffff8cc2cf0a5be0 R08: 0000000000001000 R09: ffff8cc2ccff6000
R10: ffff8cce14388000 R11: ffffffffb7b07b50 R12: ffff8cc2c1629840
R13: ffffffffb898d7b0 R14: ffffffffb898feb0 R15: ffff8cc2c19f9020
? __pfx_dev_uevent+0x10/0x10
bus_to_subsys+0x32/0xa0
dev_uevent+0x1e5/0x370
uevent_show+0xa3/0x120
dev_attr_show+0x19/0x60
sysfs_kf_seq_show+0xb8/0x120
seq_read_iter+0x1a1/0x430
? security_file_permission+0x4f/0x170
CoreFreq: AMD_SMN_Read(0, 59808) TryLock
vfs_read+0x26a/0x2f0
__x64_sys_read+0x84/0xf0
do_syscall_64+0x85/0x200
? do_syscall_64+0xb6/0x200
? do_syscall_64+0xb6/0x200
entry_SYSCALL_64_after_hwframe+0x76/0x7e
RIP: 0033:0x7fcd0f8a6527
Code: f9 08 74 43 ba 04 00 00 00 48 8b 05 db 97 18 00 64 89 10 48 c7 c2 ff ff ff ff 48 83 c4 18 48 89 d0 c3 90 48 8b 44 24 20 0f 05 <48> 89 c2 48 3d 00 f0 ff ff >
RSP: 002b:00007ffc37266540 EFLAGS: 00000202 ORIG_RAX: 0000000000000000
RAX: ffffffffffffffda RBX: 000056061384c410 RCX: 00007fcd0f8a6527
RDX: 0000000000001008 RSI: 000056061384c410 RDI: 0000000000000005
RBP: 00007ffc37266690 R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000202 R12: 0000000000000005
R13: 0000000000001008 R14: ffffffffffffffff R15: 0000000000000002
</TASK>
CoreFreq: AMD_SMN_Read(0, 59808) TryLock
CoreFreq: AMD_SMN_Read(0, 59808) TryLock
CoreFreq: AMD_SMN_Read(0, 59808) TryLock
...<about 2 minutes later>
INFO: task khugepaged:226 blocked for more than 122 seconds.
Tainted: G S OE 6.17.9-2-cachyos #1
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:khugepaged state:D stack:0 pid:226 tgid:226 ppid:2 task_flags:0x200040 flags:0x00004000
Call Trace:
<TASK>
__schedule+0x5a7/0x1420
? update_load_avg+0x1f1/0x840
? update_curr+0x105/0x240
schedule+0x6e/0xe0
schedule_timeout+0x31/0x130
? ttwu_do_activate+0xda/0x250
wait_for_completion+0xc7/0x200
__flush_work.llvm.7858765524504373934+0x26f/0x300
? __pfx_wq_barrier_func+0x10/0x10
__lru_add_drain_all.llvm.4735124219749468538+0x1c7/0x210
khugepaged+0x17e/0xbd0
? __pfx_autoremove_wake_function+0x10/0x10
? __pfx_khugepaged+0x10/0x10
kthread+0x214/0x250
? __pfx_kthread+0x10/0x10
ret_from_fork+0xf4/0x1b0
? __pfx_kthread+0x10/0x10
ret_from_fork_asm+0x1a/0x30
</TASK>
INFO: task grub-probe:4414 blocked for more than 122 seconds.
Tainted: G S OE 6.17.9-2-cachyos #1
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:grub-probe state:D stack:0 pid:4414 tgid:4414 ppid:4413 task_flags:0x400100 flags:0x00004002
Call Trace:
<TASK>
__schedule+0x5a7/0x1420
schedule+0x6e/0xe0
schedule_preempt_disabled+0x15/0x30
__mutex_lock+0x395/0xac0
__lru_add_drain_all.llvm.4735124219749468538+0x34/0x210
invalidate_bdev+0x3d/0x60
blkdev_common_ioctl+0xe61/0xef0
blkdev_ioctl+0x243/0x2c0
__x64_sys_ioctl+0x73/0xc0
do_syscall_64+0x85/0x200
? iput+0x4b/0x240
? bdev_statx+0x53/0x130
? vfs_getattr_nosec+0xee/0x100
? __se_sys_newfstat+0x26b/0x2d0
? do_syscall_64+0xb6/0x200
? do_syscall_64+0xb6/0x200
entry_SYSCALL_64_after_hwframe+0x76/0x7e
RIP: 0033:0x7f490073c17f
RSP: 002b:00007ffe8dbe2e50 EFLAGS: 00000246 ORIG_RAX: 0000000000000010
RAX: ffffffffffffffda RBX: 0000000000000000 RCX: 00007f490073c17f
RDX: 0000000000000000 RSI: 0000000000001261 RDI: 0000000000000004
RBP: 00007ffe8dbe2f70 R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000004
R13: 000055e3cee63838 R14: 00007f4900992000 R15: 00007ffe8dc033d0
</TASK>
INFO: task grub-probe:4414 is blocked on a mutex likely owned by task khugepaged:226.
task:khugepaged state:D stack:0 pid:226 tgid:226 ppid:2 task_flags:0x200040 flags:0x00004000
Call Trace:
<TASK>
__schedule+0x5a7/0x1420
? update_load_avg+0x1f1/0x840
? update_curr+0x105/0x240
schedule+0x6e/0xe0
schedule_timeout+0x31/0x130
? ttwu_do_activate+0xda/0x250
wait_for_completion+0xc7/0x200
__flush_work.llvm.7858765524504373934+0x26f/0x300
? __pfx_wq_barrier_func+0x10/0x10
__lru_add_drain_all.llvm.4735124219749468538+0x1c7/0x210
khugepaged+0x17e/0xbd0
? __pfx_autoremove_wake_function+0x10/0x10
? __pfx_khugepaged+0x10/0x10
kthread+0x214/0x250
? __pfx_kthread+0x10/0x10
ret_from_fork+0xf4/0x1b0
? __pfx_kthread+0x10/0x10
ret_from_fork_asm+0x1a/0x30
</TASK>
```
There were many more lines with `CoreFreq: AMD_SMN_Read(0, 59808) TryLock` with varying IDs. Upgrading said packages without corefreq running worked fine as well as starting corefreq afterwards.
I have no idea if corefreq was the cause or not or if there is anything that can be done aside from not having corefreq running when upgrading certain packages. The only reason I am posting this here is because of the line `? __pfx_Call_Raphael+0x10/0x10 [corefreqk 1a38744b46bc721813b1ee40277bc2f00344fdb1]` and what comes after.
Please feel free to close if this is not a corefreq issue. Thank you for all your work.
关闭于 2025-12-02 3 条评论