В ядре Linux устранена следующая уязвимость:
binfmt_misc: не предупреждать, когда монтирование завершено из пространства имен другого пользователя.
fsopen() записывает пространство имен пользователя вызывающего абонента в fc->user_ns и Hands
обратно обычный файловый дескриптор. Ничто не связывает задачу, которая вызывает
fsconfig(FSCONFIG_CMD_CREATE) к задаче, создавшей контекст.
fd наследуется через fork() и exec() и может передаваться через
Unix-сокет. Заполнение контекста из пространства имен другого пользователя разрешено намеренно.
vfs_cmd_create() разрешает создание с помощью mount_capable(), что для
FS_USERNS_MOUNT проверяет ns_capable(fc->user_ns, CAP_SYS_ADMIN), и это
завершается успешно для задачи, содержащей CAP_SYS_ADMIN в предке fc->user_ns.
Таким образом, непривилегированная задача может достичь WARN_ON() в bm_fill_super():
создайте пользователя и пространство имен монтирования в дочернем элементе, вызовите
fsopen("binfmt_misc"), отправьте fscontext fd родительскому элементу и позвольте
родительская проблема FSCONFIG_CMD_CREATE. Оба пространства имен взяты из простого
unshare(1) и никакие возможности нигде не нужны:
ВНИМАНИЕ: fs/binfmt_misc.c:938 по адресу bm_fill_super+0xa2/0xc0 [binfmt_misc]
ЦП: 15 UID: 1000 PID: 3243382 Связь: fswarn
Отслеживание вызова:
get_tree_keyed+0x7d/0xb0
bm_get_tree+0x34/0x90 [binfmt_misc]
vfs_get_tree+0x2a/0x100
vfs_cmd_create+0x60/0xf0
__do_sys_fsconfig+0x4b2/0x500
Дочернему элементу необходимо пространство имен монтирования, поскольку fsopen() сама включается.
may_mount(), который запрашивает CAP_SYS_ADMIN в пространстве имен пользователя, владеющем
пространство имен монтирования вызывающего абонента. fsconfig() не повторяет эту проверку. Это WARN_ON(), а не WARN_ON_ONCE(), поэтому условие может быть
вызывается в цикле, чтобы испортить ядро и затопить журнал, и это вызывает панику
ядро загружено с Panic_on_warn.
Продолжайте отказываться от крепления и перестаньте об этом предупреждать. Ничего в
bm_fill_super() зависит от совпадения двух пространств имен, он выводит
все из sb->s_user_ns.
Показать оригинальное описание (EN)
In the Linux kernel, the following vulnerability has been resolved: binfmt_misc: don't warn when the mount is completed from another user namespace fsopen() records the caller's user namespace in fc->user_ns and hands back an ordinary file descriptor. Nothing ties the task that calls fsconfig(FSCONFIG_CMD_CREATE) to the task that created the context. The fd is inherited across fork() and exec() and it can be passed over a unix socket. Completing a context from another user namespace is allowed on purpose. vfs_cmd_create() authorizes the create with mount_capable(), which for FS_USERNS_MOUNT checks ns_capable(fc->user_ns, CAP_SYS_ADMIN), and that succeeds for a task holding CAP_SYS_ADMIN in an ancestor of fc->user_ns. So an unprivileged task can reach the WARN_ON() in bm_fill_super(): create a user and a mount namespace in a child, call fsopen("binfmt_misc") there, send the fscontext fd to the parent and let the parent issue FSCONFIG_CMD_CREATE. Both namespaces come from a plain unshare(1) and no capability is needed anywhere: WARNING: fs/binfmt_misc.c:938 at bm_fill_super+0xa2/0xc0 [binfmt_misc] CPU: 15 UID: 1000 PID: 3243382 Comm: fswarn Call Trace: get_tree_keyed+0x7d/0xb0 bm_get_tree+0x34/0x90 [binfmt_misc] vfs_get_tree+0x2a/0x100 vfs_cmd_create+0x60/0xf0 __do_sys_fsconfig+0x4b2/0x500 The child needs the mount namespace because fsopen() itself gates on may_mount(), which asks for CAP_SYS_ADMIN in the user namespace owning the caller's mount namespace. fsconfig() doesn't repeat that check. It is a WARN_ON() and not a WARN_ON_ONCE(), so the condition can be raised in a loop to taint the kernel and flood the log, and it panics a kernel booted with panic_on_warn. Keep refusing the mount and stop warning about it. Nothing in bm_fill_super() depends on the two namespaces matching, it derives everything from sb->s_user_ns.