In the Linux kernel, the following vulnerability has been resolved:
drm/xe: don't WARN on kernel job timeout when device already wedged
igt@xe_wedged@wedged-at-any-timeout wedges the device in mode 2
(UPON_ANY_HANG_NO_RESET) and then rebinds the driver. During unbind,
a GSC proxy kernel submission can still time out; with the device wedged
and the GuC CT stopped it can never complete, so its kernel job times out. Tile0: GT1: Kernel-submitted job timed out
WARNING: drivers/gpu/drm/xe/xe_guc_submit.c:...
at guc_exec_queue_timedout_job()
Workqueue: gt-ordered-wq drm_sched_job_timedout
Killed queues skip guc_submit_hint_wedged(), leaving 'wedged' false even
though the device is already wedged.
The timeout handler then treats the
kernel queue timeout as unexpected and taints the kernel. Honour an already-wedged device even for killed queues so the expected
teardown timeout no longer trips the WARN.
(cherry picked from commit a1c1dbd0f047bb05de6aaf6abe9103031179bf19)