ITADN

CoreFreq possibly caused issue during package update

#574Closedppascher 创建于 2025-12-01
P
ppaschercommented
Today I wanted to update the following packages on my system (6.17.9-2-cachyos, AMD Ryzen 9 9950X3D) while corefreq was running: `Packages (4) imv-5.0.1-1.1 libbpf-1.6.2-1.1 libinput-1.30.0-2.1 noto-fonts-1:2025.12.01-1` After the installation the upgrade process was stuck at `(1/8) Reloading device manager configuration...` From there on I had issues entering commands (huge delays typing, sending) through my ssh session or authenticating (additional sessions, sudo). ̀ uptime` showed a load of 18. Eventually I managed to reboot the machine and saw the following journal entries for the time when the issue started: ``` rcu: INFO: rcu_preempt detected stalls on CPUs/tasks: rcu: 10-...0: (1 GPs behind) idle=9f6c/1/0x4000000000000000 softirq=3076/3078 fqs=11519 rcu: (detected by 2, t=60122 jiffies, g=4253, q=2287 ncpus=32) Sending NMI from CPU 2 to CPUs 10: NMI backtrace for cpu 10 CPU: 10 UID: 0 PID: 4399 Comm: udevadm Tainted: G S OE 6.17.9-2-cachyos #1 PREEMPT(full) 81ce6cd0386df3125c20562ac914b09f65df2f37 Tainted: [S]=CPU_OUT_OF_SPEC, [O]=OOT_MODULE, [E]=UNSIGNED_MODULE Hardware name: Micro-Star International Co., Ltd. MS-7E51/MAG X870 TOMAHAWK WIFI (MS-7E51), BIOS 1.A69 09/19/2025 RIP: 0010:native_queued_spin_lock_slowpath+0x9b/0x2b0 CoreFreq: AMD_SMN_Read(0, 59808) TryLock Code: c0 f0 0f ba 2b 08 0f 92 c0 b9 ff 00 ff ff 23 0b 89 c6 c1 e6 08 09 f1 81 f9 00 01 00 00 73 1f 85 c9 74 0c 80 3b 00 74 07 f3 90 <80> 3b 00 75 f9 66 c7 03 01 > RSP: 0018:ffffcf7680530d88 EFLAGS: 00000002 RAX: 0000000000000000 RBX: ffff8cc2c1629d90 RCX: 0000000000000001 RDX: ffffcf7680530e10 RSI: 0000000000000000 RDI: ffff8cc2c1629d90 RBP: 0000000000000000 R08: 00000017dee4a627 R09: 000000181a7f7027 R10: 00000081e06f4cd2 R11: ffffffffc20802d0 R12: 0000000000000000 R13: ffffffffb8949490 R14: ffffcf7680530e10 R15: 0000000000000001 FS: 00007fcd100d0880(0000) GS:ffff8cce14388000(0000) knlGS:0000000000000000 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 CR2: 000056061384d058 CR3: 000000010b123000 CR4: 0000000000f50ef0 PKRU: 55555554 Call Trace: <IRQ> ? __pfx_match_pci_dev_by_id+0x10/0x10 _raw_spin_lock+0x2a/0x40 bus_find_device+0x53/0x170 pci_get_domain_bus_and_slot+0x7c/0x190 CCD_AMD_Family_19h_Zen4_Temp+0xca/0x380 [corefreqk 1a38744b46bc721813b1ee40277bc2f00344fdb1] CoreFreq: AMD_SMN_Read(0, 59808) TryLock CoreFreq: AMD_SMN_Read(0, 59808) TryLock CoreFreq: AMD_SMN_Read(0, 59808) TryLock CoreFreq: AMD_SMN_Read(0, 59808) TryLock CoreFreq: AMD_SMN_Read(0, 59808) TryLock ? __pfx_Call_Raphael+0x10/0x10 [corefreqk 1a38744b46bc721813b1ee40277bc2f00344fdb1] Entry_AMD_F17h+0x445/0x620 [corefreqk 1a38744b46bc721813b1ee40277bc2f00344fdb1] ? __pfx_Call_MSR_ACCU+0x10/0x10 [corefreqk 1a38744b46bc721813b1ee40277bc2f00344fdb1] ? __pfx_Cycle_AMD_Zen4_RPL+0x10/0x10 [corefreqk 1a38744b46bc721813b1ee40277bc2f00344fdb1] CoreFreq: AMD_SMN_Read(0, 59808) TryLock __hrtimer_run_queues+0xf4/0x2e0 hrtimer_interrupt+0xf9/0x3d0 __sysvec_apic_timer_interrupt+0x4c/0x170 sysvec_apic_timer_interrupt+0x6f/0x80 </IRQ> <TASK> asm_sysvec_apic_timer_interrupt+0x1a/0x20 RIP: 0010:_raw_spin_lock+0x1b/0x40 Code: 90 90 90 90 90 90 90 90 90 90 90 90 90 90 90 f3 0f 1e fa 0f 1f 44 00 00 65 ff 05 20 cc e3 01 b9 01 00 00 00 31 c0 f0 0f b1 0f <75> 06 c3 cc cc cc cc cc 89 > RSP: 0018:ffffcf7685d0bce0 EFLAGS: 00000246 RAX: 0000000000000000 RBX: ffff8cc2c19f9020 RCX: 0000000000000001 RDX: 0000000000000000 RSI: ffff8cc2ccff6000 RDI: ffff8cc2c1629d90 RBP: ffff8cc2cf0a5be0 R08: 0000000000001000 R09: ffff8cc2ccff6000 R10: ffff8cce14388000 R11: ffffffffb7b07b50 R12: ffff8cc2c1629840 R13: ffffffffb898d7b0 R14: ffffffffb898feb0 R15: ffff8cc2c19f9020 ? __pfx_dev_uevent+0x10/0x10 bus_to_subsys+0x32/0xa0 dev_uevent+0x1e5/0x370 uevent_show+0xa3/0x120 dev_attr_show+0x19/0x60 sysfs_kf_seq_show+0xb8/0x120 seq_read_iter+0x1a1/0x430 ? security_file_permission+0x4f/0x170 CoreFreq: AMD_SMN_Read(0, 59808) TryLock vfs_read+0x26a/0x2f0 __x64_sys_read+0x84/0xf0 do_syscall_64+0x85/0x200 ? do_syscall_64+0xb6/0x200 ? do_syscall_64+0xb6/0x200 entry_SYSCALL_64_after_hwframe+0x76/0x7e RIP: 0033:0x7fcd0f8a6527 Code: f9 08 74 43 ba 04 00 00 00 48 8b 05 db 97 18 00 64 89 10 48 c7 c2 ff ff ff ff 48 83 c4 18 48 89 d0 c3 90 48 8b 44 24 20 0f 05 <48> 89 c2 48 3d 00 f0 ff ff > RSP: 002b:00007ffc37266540 EFLAGS: 00000202 ORIG_RAX: 0000000000000000 RAX: ffffffffffffffda RBX: 000056061384c410 RCX: 00007fcd0f8a6527 RDX: 0000000000001008 RSI: 000056061384c410 RDI: 0000000000000005 RBP: 00007ffc37266690 R08: 0000000000000000 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000202 R12: 0000000000000005 R13: 0000000000001008 R14: ffffffffffffffff R15: 0000000000000002 </TASK> CoreFreq: AMD_SMN_Read(0, 59808) TryLock CoreFreq: AMD_SMN_Read(0, 59808) TryLock CoreFreq: AMD_SMN_Read(0, 59808) TryLock ...<about 2 minutes later> INFO: task khugepaged:226 blocked for more than 122 seconds. Tainted: G S OE 6.17.9-2-cachyos #1 "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. task:khugepaged state:D stack:0 pid:226 tgid:226 ppid:2 task_flags:0x200040 flags:0x00004000 Call Trace: <TASK> __schedule+0x5a7/0x1420 ? update_load_avg+0x1f1/0x840 ? update_curr+0x105/0x240 schedule+0x6e/0xe0 schedule_timeout+0x31/0x130 ? ttwu_do_activate+0xda/0x250 wait_for_completion+0xc7/0x200 __flush_work.llvm.7858765524504373934+0x26f/0x300 ? __pfx_wq_barrier_func+0x10/0x10 __lru_add_drain_all.llvm.4735124219749468538+0x1c7/0x210 khugepaged+0x17e/0xbd0 ? __pfx_autoremove_wake_function+0x10/0x10 ? __pfx_khugepaged+0x10/0x10 kthread+0x214/0x250 ? __pfx_kthread+0x10/0x10 ret_from_fork+0xf4/0x1b0 ? __pfx_kthread+0x10/0x10 ret_from_fork_asm+0x1a/0x30 </TASK> INFO: task grub-probe:4414 blocked for more than 122 seconds. Tainted: G S OE 6.17.9-2-cachyos #1 "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. task:grub-probe state:D stack:0 pid:4414 tgid:4414 ppid:4413 task_flags:0x400100 flags:0x00004002 Call Trace: <TASK> __schedule+0x5a7/0x1420 schedule+0x6e/0xe0 schedule_preempt_disabled+0x15/0x30 __mutex_lock+0x395/0xac0 __lru_add_drain_all.llvm.4735124219749468538+0x34/0x210 invalidate_bdev+0x3d/0x60 blkdev_common_ioctl+0xe61/0xef0 blkdev_ioctl+0x243/0x2c0 __x64_sys_ioctl+0x73/0xc0 do_syscall_64+0x85/0x200 ? iput+0x4b/0x240 ? bdev_statx+0x53/0x130 ? vfs_getattr_nosec+0xee/0x100 ? __se_sys_newfstat+0x26b/0x2d0 ? do_syscall_64+0xb6/0x200 ? do_syscall_64+0xb6/0x200 entry_SYSCALL_64_after_hwframe+0x76/0x7e RIP: 0033:0x7f490073c17f RSP: 002b:00007ffe8dbe2e50 EFLAGS: 00000246 ORIG_RAX: 0000000000000010 RAX: ffffffffffffffda RBX: 0000000000000000 RCX: 00007f490073c17f RDX: 0000000000000000 RSI: 0000000000001261 RDI: 0000000000000004 RBP: 00007ffe8dbe2f70 R08: 0000000000000000 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000004 R13: 000055e3cee63838 R14: 00007f4900992000 R15: 00007ffe8dc033d0 </TASK> INFO: task grub-probe:4414 is blocked on a mutex likely owned by task khugepaged:226. task:khugepaged state:D stack:0 pid:226 tgid:226 ppid:2 task_flags:0x200040 flags:0x00004000 Call Trace: <TASK> __schedule+0x5a7/0x1420 ? update_load_avg+0x1f1/0x840 ? update_curr+0x105/0x240 schedule+0x6e/0xe0 schedule_timeout+0x31/0x130 ? ttwu_do_activate+0xda/0x250 wait_for_completion+0xc7/0x200 __flush_work.llvm.7858765524504373934+0x26f/0x300 ? __pfx_wq_barrier_func+0x10/0x10 __lru_add_drain_all.llvm.4735124219749468538+0x1c7/0x210 khugepaged+0x17e/0xbd0 ? __pfx_autoremove_wake_function+0x10/0x10 ? __pfx_khugepaged+0x10/0x10 kthread+0x214/0x250 ? __pfx_kthread+0x10/0x10 ret_from_fork+0xf4/0x1b0 ? __pfx_kthread+0x10/0x10 ret_from_fork_asm+0x1a/0x30 </TASK> ``` There were many more lines with `CoreFreq: AMD_SMN_Read(0, 59808) TryLock` with varying IDs. Upgrading said packages without corefreq running worked fine as well as starting corefreq afterwards. I have no idea if corefreq was the cause or not or if there is anything that can be done aside from not having corefreq running when upgrading certain packages. The only reason I am posting this here is because of the line `? __pfx_Call_Raphael+0x10/0x10 [corefreqk 1a38744b46bc721813b1ee40277bc2f00344fdb1]` and what comes after. Please feel free to close if this is not a corefreq issue. Thank you for all your work.
关闭于 2025-12-02 3 条评论