В ядре Linux устранена следующая уязвимость:
ocfs2: кластер: не спать, удерживая o2hb_live_lock в o2hb_region_pin()
Серия патчей «ocfs2: кластер: исправления o2hb_region_pin()», v2. В этой серии исправлены три связанные проблемы в o2hb_region_pin(), все из
исходная реализация в коммите: 58a3158a5d17 ("ocfs2/cluster:
Закрепить/открепить регионы o2hb"):
1) Он вызывается с удержанием o2hb_live_lock (спин-блокировка), но
базовый configfs_dependent_item() спит (принимает индексный дескриптор rwsem и
закрепляет файловую систему). Это вызывает ошибку под
CONFIG_DEBUG_ATOMIC_SLEEP.
2) При вызове из обратного вызова configfs drop_item создается
инверсия порядка блокировки: родительский inode_lock -> корень configfs
inode_lock, который может блокировать отмену регистрации подсистемы.
пути, укореняющиеся -> родительский.
3) Если привязка не удалась во время выполнения o2hb_region_inc_user(),
Счетчик o2hb_dependent_users утек и частично закреплен
регионы никогда не освобождаются, оставляя области пульса
незащищенным при последующих креплениях.
Патч 1 перерабатывает o2hb_region_pin(), чтобы удалить o2hb_live_lock для каждого
вызов configfs_dependent_item() с использованием ссылки config_item на
сохранить регион живым, пока он разблокирован. Патч 2 добавляет параметр from_callback для выбора
configfs_dependent_item_unlocked() при вызове из контекста configfs,
избегая вложенности inode_lock. Патч 3 исправляет путь ошибки в o2hb_region_inc_user() для открепления и
уменьшить счетчик в случае неудачи.
Этот патч (из 3):
o2hb_region_pin() всегда вызывается с удержанной спин-блокировкой o2hb_live_lock.
(из o2hb_region_inc_user() и o2hb_heartbeat_group_drop_item()), но это
вызывает o2nm_dependent_item() -> configfs_dependent_item(), который спит: он закрепляет
файловую систему configfs и принимает корневой индекс configfs rwsem. Под
CONFIG_DEBUG_ATOMIC_SLEEP это запускает:
ОШИБКА: спящая функция вызывается из неверного контекста в файле kernel/locking/rwsem.c.
in_atomic(): 1, ... имя: mount.ocfs2
down_write
configfs_dependent_item
o2hb_region_pin
o2hb_region_inc_user
o2hb_register_callback
dlm_register_domain_handlers
...
ocfs2_dlm_init
ocfs2_mount_volume
ocfs2_fill_super
Переработана функция o2hb_region_pin() для закрепления одного региона за раз при снятой блокировке.
через спящий вызов: под o2hb_live_lock найдите следующий подходящий
региона и возьмите ссылку на config_item, чтобы сохранить его в рабочем состоянии, снимите блокировку,
вызовите o2nm_dependent_item(), затем повторно заблокируйте и запишите пин-код.
config_item_put() также выполняется со снятой блокировкой, поскольку
o2hb_region_release() также получает o2hb_live_lock и может перейти в режим сна.
список регионов может измениться, пока он разблокирован, поэтому сканирование возобновляется сверху.
после каждого контакта. Локальный пульс по-прежнему фиксирует только соответствующий регион;
глобальное сердцебиение фиксирует все подходящие регионы.
Путь открепления не изменяется: configfs_undependent_item() принимает только
спинлок и не спит.
Показать оригинальное описание (EN)
In the Linux kernel, the following vulnerability has been resolved: ocfs2: cluster: don't sleep while holding o2hb_live_lock in o2hb_region_pin() Patch series "ocfs2: cluster: o2hb_region_pin() fixes", v2. This series fixes three related issues in o2hb_region_pin(), all are from the original implementation in commit: 58a3158a5d17 ("ocfs2/cluster: Pin/unpin o2hb regions"): 1) It is called with o2hb_live_lock (a spinlock) held, but the underlying configfs_depend_item() sleeps (takes inode rwsem and pins the filesystem). This triggers BUG under CONFIG_DEBUG_ATOMIC_SLEEP. 2) When called from the configfs drop_item callback, it creates a lock order inversion: parent inode_lock -> configfs root inode_lock, which can deadlock against subsystem unregistration paths taking root -> parent. 3) If pinning fails partway through o2hb_region_inc_user(), the o2hb_dependent_users counter is leaked and partially-pinned regions are never released, leaving heartbeat regions unprotected on subsequent mounts. Patch 1 reworks o2hb_region_pin() to drop o2hb_live_lock across each sleeping configfs_depend_item() call, using a config_item reference to keep the region alive while unlocked. Patch 2 adds a from_callback parameter to select configfs_depend_item_unlocked() when called from configfs context, avoiding the inode_lock nesting. Patch 3 fixes the error path in o2hb_region_inc_user() to unpin and decrement the counter on failure. This patch (of 3): o2hb_region_pin() is always called with the o2hb_live_lock spinlock held (from o2hb_region_inc_user() and o2hb_heartbeat_group_drop_item()), but it calls o2nm_depend_item() -> configfs_depend_item(), which sleeps: it pins the configfs filesystem and takes the configfs root inode rwsem. Under CONFIG_DEBUG_ATOMIC_SLEEP this triggers: BUG: sleeping function called from invalid context at kernel/locking/rwsem.c in_atomic(): 1, ... name: mount.ocfs2 down_write configfs_depend_item o2hb_region_pin o2hb_region_inc_user o2hb_register_callback dlm_register_domain_handlers ... ocfs2_dlm_init ocfs2_mount_volume ocfs2_fill_super Rework o2hb_region_pin() to pin one region at a time with the lock dropped across the sleeping call: under o2hb_live_lock find the next eligible region and take a config_item reference to keep it alive, drop the lock, call o2nm_depend_item(), then retake the lock and record the pin. The config_item_put() is done with the lock released as well, since o2hb_region_release() also acquires o2hb_live_lock and can sleep. The region list may change while unlocked, so the scan restarts from the top after each pin. Local heartbeat still pins only the matching region; global heartbeat pins all eligible regions. The unpin path is unaffected: configfs_undepend_item() only takes a spinlock and does not sleep.