Conversation
Signed-off-by: Roxana Nicolescu <rnicolescu@ciq.com>
Signed-off-by: Roxana Nicolescu <rnicolescu@ciq.com>
jira LE-3207 feature tools_hv commit-author Shradha Gupta <shradhagupta@linux.microsoft.com> commit a9c0b33 Allow the KVP daemon to log the KVP updates triggered in the VM with a new debug flag(-d). When the daemon is started with this flag, it logs updates and debug information in syslog with loglevel LOG_DEBUG. This information comes in handy for debugging issues where the key-value pairs for certain pools show mismatch/incorrect values. The distro-vendors can further consume these changes and modify the respective service files to redirect the logs to specific files as needed. Signed-off-by: Shradha Gupta <shradhagupta@linux.microsoft.com> Reviewed-by: Naman Jain <namjain@linux.microsoft.com> Reviewed-by: Dexuan Cui <decui@microsoft.com> Link: https://lore.kernel.org/r/1744715978-8185-1-git-send-email-shradhagupta@linux.microsoft.com Signed-off-by: Wei Liu <wei.liu@kernel.org> Message-ID: <1744715978-8185-1-git-send-email-shradhagupta@linux.microsoft.com> (cherry picked from commit a9c0b33) Signed-off-by: Jonathan Maple <jmaple@ciq.com>
…nused() jira SECO-468 commit-author Luis Henriques <luis@igalia.com> commit 395b955 Add and export a new helper d_dispose_if_unused() which is simply a wrapper around to_shrink_list(), to add an entry to a dispose list if it's not used anymore. Also export shrink_dentry_list() to kill all dentries in a dispose list. Suggested-by: Miklos Szeredi <miklos@szeredi.hu> Signed-off-by: Luis Henriques <luis@igalia.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit 395b955) Signed-off-by: Roxana Nicolescu <rnicolescu@ciq.com>
jira SECO-478 RFBugFix: FUSE commit-author Miklos Szeredi <mszeredi@redhat.com> commit b4c173d Fuse allows the value of a symlink to change and this property is exploited by some filesystems (e.g. CVMFS). It has been observed, that sometimes after changing the symlink contents, the value is truncated to the old size. This is caused by fuse_getattr() racing with fuse_reverse_inval_inode(). fuse_reverse_inval_inode() updates the fuse_inode's attr_version, which results in fuse_change_attributes() exiting before updating the cached attributes This is okay, as the cached attributes remain invalid and the next call to fuse_change_attributes() will likely update the inode with the correct values. The reason this causes problems is that cached symlinks will be returned through page_get_link(), which truncates the symlink to inode->i_size. This is correct for filesystems that don't mutate symlinks, but in this case it causes bad behavior. The solution is to just remove this truncation. This can cause a regression in a filesystem that relies on supplying a symlink larger than the file size, but this is unlikely. If that happens we'd need to make this behavior conditional. Reported-by: Laura Promberger <laura.promberger@cern.ch> Tested-by: Sam Lewis <samclewis@google.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> Link: https://lore.kernel.org/r/20250220100258.793363-1-mszeredi@redhat.com Reviewed-by: Bernd Schubert <bschubert@ddn.com> Signed-off-by: Christian Brauner <brauner@kernel.org> (cherry picked from commit b4c173d) Signed-off-by: Jonathan Maple <jmaple@ciq.com>
jira SECO-478 RFBugFix: FUSE commit-author Luis Henriques <luis@igalia.com> commit 2396356 upstream-diff | conflict in fs/fuse/dir.c due to missing this piece: d701902 - fuse: return correct dentry for ->mkdir Which is a part of a larger changeset here that we're not going to take: https://lore.kernel.org/all/20250227013949.536172-1-neilb@suse.de/ | Additionally this bumps the Kernel FUSE API minor version from 41 to 44. The interface into via fuse3 currently in Rocky 10.1 is limited to API 38 anyways at 3.16.2. | There is a build conflict due to a major rewrite of the d_revalidate calls which now includes the parent directory being passed. 5be1fa8 Pass parent directory inode and expected name to ->d_revalidate() In this case we can use the dentry->i_sb because we only need the superblock for get_fuse_conn_super(). Currently userspace is able to notify the kernel to invalidate the cache for an inode. This means that, if all the inodes in a filesystem need to be invalidated, then userspace needs to iterate through all of them and do this kernel notification separately. This patch adds the concept of 'epoch': each fuse connection will have the current epoch initialized and every new dentry will have it's d_time set to the current epoch value. A new operation will then allow userspace to increment the epoch value. Every time a dentry is d_revalidate()'ed, it's epoch is compared with the current connection epoch and invalidated if it's value is different. Signed-off-by: Luis Henriques <luis@igalia.com> Tested-by: Laura Promberger <laura.promberger@cern.ch> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit 2396356) Signed-off-by: Jonathan Maple <jmaple@ciq.com> build fix: fuse: add more control over cache invalidation behaviour
jira SECO-478 BUGFIX: FUSE commit-author Miklos Szeredi <mszeredi@redhat.com> commit 0b563aa In case of FUSE_NOTIFY_RESEND and FUSE_NOTIFY_INC_EPOCH fuse_copy_finish() isn't called. Fix by always calling fuse_copy_finish() after fuse_notify(). It's a no-op if called a second time. Fixes: 760eac7 ("fuse: Introduce a new notification type for resend pending requests") Fixes: 2396356 ("fuse: add more control over cache invalidation behaviour") Cc: <stable@vger.kernel.org> # v6.9 Reviewed-by: Joanne Koong <joannelkoong@gmail.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit 0b563aa) Signed-off-by: Jonathan Maple <jmaple@ciq.com>
jira SECO-511 commit-author Chen Linxuan <chenlinxuan@uniontech.com> commit f092229 upstream-diff | There were conflicts seen while applying this patch due to the following missing commit :- 786412a ("fuse: enable fuse-over-io-uring") This commit add fuse connection device id to fdinfo of opened /dev/fuse files. Related discussions can be found at links below. Link: https://lore.kernel.org/all/CAJfpegvEYUgEbpATpQx8NqVR33Mv-VK96C+gbTag1CEUeBqvnA@mail.gmail.com/ Signed-off-by: Chen Linxuan <chenlinxuan@uniontech.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit f092229) Signed-off-by: Shreeya Patel <spatel@ciq.com>
jira SECO-518 commit-author Amir Goldstein <amir73il@gmail.com> commit 03f275a The re-factoring of fuse_dir_open() missed the need to invalidate directory inode page cache with open flag FOPEN_KEEP_CACHE. Fixes: 7de64d5 ("fuse: break up fuse_open_common()") Reported-by: Prince Kumar <princer@google.com> Closes: https://lore.kernel.org/linux-fsdevel/CAEW=TRr7CYb4LtsvQPLj-zx5Y+EYBmGfM24SuzwyDoGVNoKm7w@mail.gmail.com/ Signed-off-by: Amir Goldstein <amir73il@gmail.com> Link: https://lore.kernel.org/r/20250101130037.96680-1-amir73il@gmail.com Reviewed-by: Bernd Schubert <bernd.schubert@fastmail.fm> Signed-off-by: Christian Brauner <brauner@kernel.org> (cherry picked from commit 03f275a) Signed-off-by: Shreeya Patel <spatel@ciq.com>
cve CVE-2026-46317 commit-author Hyunwoo Kim <imv4bel@gmail.com> commit 7054335 kvm->arch.nested_mmus[] is walked under kvm->mmu_lock, including from the MMU notifier path (kvm_unmap_gfn_range() -> kvm_nested_s2_unmap()), which can run at any time. kvm_vcpu_init_nested() reallocates the array and frees the old buffer while holding only kvm->arch.config_lock, so such a walker can reference the freed array. Allocate the new array outside of mmu_lock, as the allocation can sleep. Under the lock, copy the existing entries, fix up the back pointers and reassign the array. Free the old buffer after dropping the lock, as kvfree() can sleep as well. Fixes: 4f128f8 ("KVM: arm64: nv: Support multiple nested Stage-2 mmu structures") Signed-off-by: Hyunwoo Kim <imv4bel@gmail.com> Reviewed-by: Oliver Upton <oupton@kernel.org> Link: https://patch.msgid.link/aiKIVVeIr1aAB1yp@v4bel Signed-off-by: Marc Zyngier <maz@kernel.org> Cc: stable@vger,kernel.org (cherry picked from commit 7054335) Signed-off-by: Jonathan Maple <jmaple@ciq.com>
jira KERNEL-1217 commit-author Harshitha Ramamurthy <hramamurthy@google.com> commit 93c68f1 In preparation for the upcoming page pool adoption for DQO raw addressing mode, move RX buffer management code to a new file. In the follow on patches, page pool code will be added to this file. No functional change, just movement of code. Reviewed-by: Praveen Kaligineedi <pkaligineedi@google.com> Reviewed-by: Shailend Chand <shailend@google.com> Reviewed-by: Willem de Bruijn <willemb@google.com> Signed-off-by: Harshitha Ramamurthy <hramamurthy@google.com> Reviewed-by: Jacob Keller <jacob.e.keller@intel.com> Link: https://patch.msgid.link/20241014202108.1051963-2-pkaligineedi@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org> (cherry picked from commit 93c68f1) Signed-off-by: Shreeya Patel <spatel@ciq.com>
jira KERNEL-1217 commit-author Joshua Washington <joshwash@google.com> commit 6321f5f When stopping XDP TX rings, the XDP clean function needs to be called to clean out the entire queue, similar to what happens in the normal TX queue case. Otherwise, the FIFO won't be cleared correctly, and xsk_tx_completed won't be reported. Fixes: 75eaae1 ("gve: Add XDP DROP and TX support for GQI-QPL format") Cc: stable@vger.kernel.org Signed-off-by: Joshua Washington <joshwash@google.com> Signed-off-by: Praveen Kaligineedi <pkaligineedi@google.com> Reviewed-by: Praveen Kaligineedi <pkaligineedi@google.com> Reviewed-by: Willem de Bruijn <willemb@google.com> Signed-off-by: David S. Miller <davem@davemloft.net> (cherry picked from commit 6321f5f) Signed-off-by: Shreeya Patel <spatel@ciq.com>
jira KERNEL-1217 commit-author Joshua Washington <joshwash@google.com> commit de63ac4 This patch fixes a number of consistency issues in the queue allocation path related to XDP. As it stands, the number of allocated XDP queues changes in three different scenarios. 1) Adding an XDP program while the interface is up via gve_add_xdp_queues 2) Removing an XDP program while the interface is up via gve_remove_xdp_queues 3) After queues have been allocated and the old queue memory has been removed in gve_queues_start. However, the requirement for the interface to be up for gve_(add|remove)_xdp_queues to be called, in conjunction with the fact that the number of queues stored in priv isn't updated until _after_ XDP queues have been allocated in the normal queue allocation path means that if an XDP program is added while the interface is down, XDP queues won't be added until the _second_ if_up, not the first. Given the expectation that the number of XDP queues is equal to the number of RX queues, scenario (3) has another problematic implication. When changing the number of queues while an XDP program is loaded, the number of XDP queues must be updated as well, as there is logic in the driver (gve_xdp_tx_queue_id()) which relies on every RX queue having a corresponding XDP TX queue. However, the number of XDP queues stored in priv would not be updated until _after_ a close/open leading to a mismatch in the number of XDP queues reported vs the number of XDP queues which actually exist after the queue count update completes. This patch remedies these issues by doing the following: 1) The allocation config getter function is set up to retrieve the _expected_ number of XDP queues to allocate instead of relying on the value stored in `priv` which is only updated once the queues have been allocated. 2) When adjusting queues, XDP queues are adjusted to match the number of RX queues when XDP is enabled. This only works in the case when queues are live, so part (1) of the fix must still be available in the case that queues are adjusted when there is an XDP program and the interface is down. Fixes: 5f08cd3 ("gve: Alloc before freeing when adjusting queues") Cc: stable@vger.kernel.org Signed-off-by: Joshua Washington <joshwash@google.com> Signed-off-by: Praveen Kaligineedi <pkaligineedi@google.com> Reviewed-by: Praveen Kaligineedi <pkaligineedi@google.com> Reviewed-by: Shailend Chand <shailend@google.com> Reviewed-by: Willem de Bruijn <willemb@google.com> Signed-off-by: David S. Miller <davem@davemloft.net> (cherry picked from commit de63ac4) Signed-off-by: Shreeya Patel <spatel@ciq.com>
jira KERNEL-1217 cve CVE-2025-38735 commit-author Jordan Rhee <jordanrhee@google.com> commit 75a9a46 A crash can occur if an ethtool operation is invoked after shutdown() is called. shutdown() is invoked during system shutdown to stop DMA operations without performing expensive deallocations. It is discouraged to unregister the netdev in this path, so the device may still be visible to userspace and kernel helpers. In gve, shutdown() tears down most internal data structures. If an ethtool operation is dispatched after shutdown(), it will dereference freed or NULL pointers, leading to a kernel panic. While graceful shutdown normally quiesces userspace before invoking the reboot syscall, forced shutdowns (as observed on GCP VMs) can still trigger this path. Fix by calling netif_device_detach() in shutdown(). This marks the device as detached so the ethtool ioctl handler will skip dispatching operations to the driver. Fixes: 974365e ("gve: Implement suspend/resume/shutdown") Signed-off-by: Jordan Rhee <jordanrhee@google.com> Signed-off-by: Jeroen de Borst <jeroendb@google.com> Link: https://patch.msgid.link/20250818211245.1156919-1-jeroendb@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org> (cherry picked from commit 75a9a46) Signed-off-by: Shreeya Patel <spatel@ciq.com>
jira KERNEL-1217 cve CVE-2025-71156 commit-author Ankit Garg <nktgrg@google.com> commit 3d970ed upstream-diff RLC 10 does not carry the netif_napi_set_irq_locked() call this commit's gve_add_napi() hunk lands beside (it has block->irq but not the napi-irq wiring), so enable_irq(block->irq) is added directly after netif_napi_add_locked(). Currently, interrupts are automatically enabled immediately upon request. This allows interrupt to fire before the associated NAPI context is fully initialized and cause failures like below: [ 0.946369] Call Trace: [ 0.946369] <IRQ> [ 0.946369] __napi_poll+0x2a/0x1e0 [ 0.946369] net_rx_action+0x2f9/0x3f0 [ 0.946369] handle_softirqs+0xd6/0x2c0 [ 0.946369] ? handle_edge_irq+0xc1/0x1b0 [ 0.946369] __irq_exit_rcu+0xc3/0xe0 [ 0.946369] common_interrupt+0x81/0xa0 [ 0.946369] </IRQ> [ 0.946369] <TASK> [ 0.946369] asm_common_interrupt+0x22/0x40 [ 0.946369] RIP: 0010:pv_native_safe_halt+0xb/0x10 Use the `IRQF_NO_AUTOEN` flag when requesting interrupts to prevent auto enablement and explicitly enable the interrupt in NAPI initialization path (and disable it during NAPI teardown). This ensures that interrupt lifecycle is strictly coupled with readiness of NAPI context. Cc: stable@vger.kernel.org Fixes: 1dfc2e4 ("gve: Refactor napi add and remove functions") Signed-off-by: Ankit Garg <nktgrg@google.com> Reviewed-by: Jordan Rhee <jordanrhee@google.com> Reviewed-by: Joshua Washington <joshwash@google.com> Signed-off-by: Harshitha Ramamurthy <hramamurthy@google.com> Link: https://patch.msgid.link/20251219102945.2193617-1-hramamurthy@google.com Signed-off-by: Paolo Abeni <pabeni@redhat.com> (cherry picked from commit 3d970ed) Signed-off-by: Shreeya Patel <spatel@ciq.com>
… QPL jira KERNEL-1217 cve CVE-2026-23386 commit-author Ankit Garg <nktgrg@google.com> commit fb868db upstream-diff The moved gve_unmap_packet() keeps this tree's dma_unmap_page() for the frag entries; the netmem conversion (netmem_dma_unmap_page_attrs) is not in this tree. In DQ-QPL mode, gve_tx_clean_pending_packets() incorrectly uses the RDA buffer cleanup path. It iterates num_bufs times and attempts to unmap entries in the dma array. This leads to two issues: 1. The dma array shares storage with tx_qpl_buf_ids (union). Interpreting buffer IDs as DMA addresses results in attempting to unmap incorrect memory locations. 2. num_bufs in QPL mode (counting 2K chunks) can significantly exceed the size of the dma array, causing out-of-bounds access warnings (trace below is how we noticed this issue). UBSAN: array-index-out-of-bounds in drivers/net/ethernet/drivers/net/ethernet/google/gve/gve_tx_dqo.c:178:5 index 18 is out of range for type 'dma_addr_t[18]' (aka 'unsigned long long[18]') Workqueue: gve gve_service_task [gve] Call Trace: <TASK> dump_stack_lvl+0x33/0xa0 __ubsan_handle_out_of_bounds+0xdc/0x110 gve_tx_stop_ring_dqo+0x182/0x200 [gve] gve_close+0x1be/0x450 [gve] gve_reset+0x99/0x120 [gve] gve_service_task+0x61/0x100 [gve] process_scheduled_works+0x1e9/0x380 Fix this by properly checking for QPL mode and delegating to gve_free_tx_qpl_bufs() to reclaim the buffers. Cc: stable@vger.kernel.org Fixes: a6fb8d5 ("gve: Tx path for DQO-QPL") Signed-off-by: Ankit Garg <nktgrg@google.com> Reviewed-by: Jordan Rhee <jordanrhee@google.com> Reviewed-by: Harshitha Ramamurthy <hramamurthy@google.com> Signed-off-by: Joshua Washington <joshwash@google.com> Reviewed-by: Simon Horman <horms@kernel.org> Link: https://patch.msgid.link/20260220215324.1631350-1-joshwash@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org> (cherry picked from commit fb868db) Signed-off-by: Shreeya Patel <spatel@ciq.com>
jira KERNEL-1217 commit-author Matt Olson <maolson@google.com> commit 07993df upstream-diff Conflicts from the gve_queue_config split and page_pool/ netmem/XDP refactors not in RLC 10. Kept RLC 10 struct names (qcfg/ qcfg_tx, no num_xdp_rings) and datapath (no page_pool/xsk); the gve_update_num_qpl_pages() body is applied verbatim except rx_alloc_cfg->qcfg_rx->num_queues -> ->qcfg->num_queues. buffer-mgmt and rx_dqo num_buf_states use cfg->pages_per_qpl / priv->rx_pages_per_qpl. For DQO, change QPL page registration logic to be more flexible to honor the "max_registered_pages" parameter from the gVNIC device. Previously the number of RX pages per QPL was hardcoded to twice the ring size, and the number of TX pages per QPL was dictated by the device in the DQO-QPL device option. Now [in DQO-QPL mode], the driver will ignore the "tx_pages_per_qpl" parameter indicated in the DQO-QPL device option and instead allocate up to (tx_queue_length / 2) pages per TX QPL and up to (rx_queue_length * 2) pages per RX QPL while keeping the total number of pages under the "max_registered_pages". Merge DQO and GQI QPL page calculation logic into a unified gve_update_num_qpl_pages function. Add rx_pages_per_qpl to the priv struct for consumption by both DQO and GQI. Signed-off-by: Matt Olson <maolson@google.com> Signed-off-by: Max Yuan <maxyuan@google.com> Reviewed-by: Jordan Rhee <jordanrhee@google.com> Reviewed-by: Harshitha Ramamurthy <hramamurthy@google.com> Reviewed-by: Willem de Bruijn <willemb@google.com> Reviewed-by: Praveen Kaligineedi <pkaligineedi@google.com> Signed-off-by: Joshua Washington <joshwash@google.com> Link: https://patch.msgid.link/20260225182342.1049816-2-joshwash@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org> (cherry picked from commit 07993df) Signed-off-by: Shreeya Patel <spatel@ciq.com>
jira KERNEL-1217 commit-author Matt Olson <maolson@google.com> commit a2f1918 The gVNIC device indicates a device option (MODIFY_RING) to the driver, which presents a range of ring sizes from which the user is allowed to select. But in DQO-QPL queue format, the driver ignores the "max" of this range and instead allows the user to configure the ring size in the range [min, default]. This was done because increasing the ring size could result in the number of registered pages being higher than the max allowed by the device. In order to support large ring sizes, stop ignoring the "max" of the range presented in the MODIFY_RING option. Signed-off-by: Matt Olson <maolson@google.com> Signed-off-by: Max Yuan <maxyuan@google.com> Reviewed-by: Jordan Rhee <jordanrhee@google.com> Reviewed-by: Harshitha Ramamurthy <hramamurthy@google.com> Reviewed-by: Praveen Kaligineedi <pkaligineedi@google.com> Signed-off-by: Joshua Washington <joshwash@google.com> Link: https://patch.msgid.link/20260225182342.1049816-3-joshwash@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org> (cherry picked from commit a2f1918) Signed-off-by: Shreeya Patel <spatel@ciq.com>
jira KERNEL-1217 commit-author Jordan Rhee <jordanrhee@google.com> commit 6bf1457 upstream-diff RLC 10 predates the page-pool conversion and has no gve_free_buffer(); the freed buffer goes through gve_enqueue_buf_state(rx, &rx->dqo.recycled_buf_states, buf_state), the same form this tree's rx error path uses, which is what upstream's helper reduces to on non-page-pool queues. When header split is enabled and a header-only packet is received such as a pure TCP ACK, GVE will indicate an RX SKB with a zero-length fragment. If this SKB is then hairpinned and sent back out, the GVE TX path will emit a zero-length descriptor. Hardware considers this an illegal descriptor and stops the queue, causing a TX timeout and interface reset. Fix it by not adding the zero-length skb frag. Cc: stable@vger.kernel.org Fixes: 5e37d82 ("gve: Add header split data path") Suggested-by: Praveen Kaligineedi <pkaligineedi@google.com> Co-developed-by: Ziwei Xiao <ziweixiao@google.com> Signed-off-by: Ziwei Xiao <ziweixiao@google.com> Signed-off-by: Jordan Rhee <jordanrhee@google.com> Signed-off-by: Harshitha Ramamurthy <hramamurthy@google.com> Link: https://patch.msgid.link/20260807224315.234152-2-hramamurthy@google.com Signed-off-by: Jakub Kicinski <kuba@kernel.org> (cherry picked from commit 6bf1457) Signed-off-by: Shreeya Patel <spatel@ciq.com>
jira KERNEL-1217 gve_tx_qpl_buf_init() sizes the TX buffer free list as GVE_TX_BUFS_PER_PAGE_DQO * num_entries, but the free-list links (tx_qpl_buf_next) are s16. GVE_TX_BUFS_PER_PAGE_DQO is PAGE_SIZE >> 11, so on 64K-page builds (the aarch64-64k config ships CONFIG_GVE=m) a 4096-entry TX ring yields 2048 pages and 65536 buffers: entry 32767 is seeded with a link of 32768, which reads back as -32768, and once 32768 buffers have been allocated gve_alloc_tx_qpl_buf() dereferences tx_qpl_buf_next[-32768] and returns ids that index outside the QPL. Clamp the count to S16_MAX, leaving an oversized QPL's tail unused rather than mislinked. This is not unique to this backport: upstream has the same unclamped math at its tip. Kept as a separate commit so it can be dropped in favor of the upstream fix once one lands; Fixes: 07993df ("gve: Update QPL page registration logic") Signed-off-by: Shreeya Patel <spatel@ciq.com>
cve CVE-2026-72137 commit-author Qianyu Luo <qianyuluo3@gmail.com> commit 226f4a4 nat_keepalive_send() frees the keepalive skb whenever the IPv4 or IPv6 send helper reports an error. That cleanup is only correct before the skb is handed to the output path. Once ip_build_and_send_pkt() or ip6_xmit() takes ownership, the networking stack may already have consumed the skb before returning an error, so freeing it again is unsafe. Handle the pre-handoff failure cases inside nat_keepalive_send_ipv4() and nat_keepalive_send_ipv6(), where the caller still owns the skb, and keep nat_keepalive_send() responsible only for family dispatch and the unsupported-family cleanup path. Fixes: f531d13 ("xfrm: support sending NAT keepalives in ESP in UDP states") Cc: stable@vger.kernel.org Reported-by: Yuan Tan <yuantan098@gmail.com> Reported-by: Xin Liu <bird@lzu.edu.cn> Signed-off-by: Qianyu Luo <qianyuluo3@gmail.com> Signed-off-by: Ren Wei <n05ec@lzu.edu.cn> Reviewed-by: Eyal Birger <eyal.birger@gmail.com> Signed-off-by: Steffen Klassert <steffen.klassert@secunet.com> (cherry picked from commit 226f4a4) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
cve-pre CVE-2026-53361 commit-author Kuniyuki Iwashima <kuniyu@amazon.com> commit 001a250 upstream-diff | Context difference due to backported commit by redhat that technically comes after this change. 3849005 [af_unix: Don't call wait_for_unix_gc() on every sendmsg().] We will introduce skb drop reason for AF_UNIX, then we need to set an errno and a drop reason for each path. Let's set an error only when it's needed in unix_dgram_sendmsg(). Then, we need not (re)set 0 to err. Signed-off-by: Kuniyuki Iwashima <kuniyu@amazon.com> Signed-off-by: Paolo Abeni <pabeni@redhat.com> (cherry picked from commit 001a250) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
cve CVE-2026-68376 commit-author Xin Long <lucien.xin@gmail.com> commit e0b5252 The auth_hmacs array in struct sctp_cookie is supposed to store a complete SCTP_AUTH_HMAC_ALGO parameter, which consists of a struct sctp_paramhdr followed by N HMAC identifiers. However, the array size was calculated using an extra 2 bytes instead of sizeof(struct sctp_paramhdr), which is 4 bytes. When four HMAC identifiers are configured, the HMAC-ALGO parameter stored in the endpoint is larger than the auth_hmacs buffer in the cookie. As a result, sctp_association_init() copies beyond the end of auth_hmacs when initializing the association, corrupting the adjacent auth_chunks field. This can lead to an invalid HMAC identifier being accepted and later cause an out-of-bounds read in sctp_auth_get_hmac(). Fix the array size calculation by including the full SCTP parameter header size. Fixes: 1f48564 ("[SCTP]: Implement SCTP-AUTH internals") Reported-by: Yuan Tan <yuantan098@gmail.com> Reported-by: Xin Liu <dstsmallbird@foxmail.com> Reported-by: Zihan Xi <xizh2024@lzu.edu.cn> Reported-by: Ren Wei <enjou1224z@gmail.com> Signed-off-by: Xin Long <lucien.xin@gmail.com> Link: https://patch.msgid.link/634a0de0d5de29532915e6d47c92a0cbc206e03f.1783707155.git.lucien.xin@gmail.com Signed-off-by: Paolo Abeni <pabeni@redhat.com> (cherry picked from commit e0b5252) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
cve CVE-2026-31678 commit-author Yang Yang <n05ec@lzu.edu.cn> commit 6931d21 ovs_netdev_tunnel_destroy() may run after NETDEV_UNREGISTER already detached the device. Dropping the netdev reference in destroy can race with concurrent readers that still observe vport->dev. Do not release vport->dev in ovs_netdev_tunnel_destroy(). Instead, let vport_netdev_free() drop the reference from the RCU callback, matching the non-tunnel destroy path and avoiding additional synchronization under RTNL. Fixes: a9020fd ("openvswitch: Move tunnel destroy function to oppenvswitch module.") Reported-by: Yifan Wu <yifanwucs@gmail.com> Reported-by: Juefei Pu <tomapufckgml@gmail.com> Tested-by: Ao Zhou <n05ec@lzu.edu.cn> Co-developed-by: Yuan Tan <tanyuan98@outlook.com> Signed-off-by: Yuan Tan <tanyuan98@outlook.com> Suggested-by: Xin Liu <bird@lzu.edu.cn> Signed-off-by: Yang Yang <n05ec@lzu.edu.cn> Reviewed-by: Ilya Maximets <i.maximets@ovn.org> Link: https://patch.msgid.link/20260319074241.3405262-1-n05ec@lzu.edu.cn Signed-off-by: Jakub Kicinski <kuba@kernel.org> (cherry picked from commit 6931d21) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
cve CVE-2026-72255 commit-author Haoze Xie <royenheart@gmail.com> commit c9c9b37 The br_netfilter fake rtable is embedded in struct net_bridge and is attached to bridged packets with skb_dst_set_noref(). If such a packet is queued to NFQUEUE, __nf_queue() upgrades that fake dst with skb_dst_force(). At that point the queued skb can hold a real dst reference after bridge teardown has started. The problem is not that every bridged packet needs its own dst reference. The problem is that NFQUEUE can keep the bridge private fake dst alive after unregister begins. Fix this by keeping the bridge fake dst model unchanged and pinning the bridge master device only while the packet sits in NFQUEUE. Record the bridge device in nf_queue_entry when the queued skb carries a bridge fake dst, take a device reference for the queue lifetime, and drop it when the queue entry is freed. Also make sure queued entries are reaped when that bridge device goes down, and drop the redundant nf_bridge_info_exists() test from the fake dst detection. This keeps netdev_priv(br->dev) alive until verdict completion, so the embedded fake rtable and its metrics backing storage cannot be freed out from under dst_release(). It also avoids the constant refcount bump and avoids using ipv4-specific dst helpers for IPv6 bridge traffic. Fixes: 34666d4 ("netfilter: bridge: move br_netfilter out of the core") Cc: stable@kernel.org Reported-by: Yuan Tan <yuantan098@gmail.com> Reported-by: Yifan Wu <yifanwucs@gmail.com> Reported-by: Juefei Pu <tomapufckgml@gmail.com> Reported-by: Xin Liu <bird@lzu.edu.cn> Signed-off-by: Haoze Xie <royenheart@gmail.com> Signed-off-by: Ren Wei <n05ec@lzu.edu.cn> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org> (cherry picked from commit c9c9b37) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
cve CVE-2026-46165 commit-author Ilya Maximets <i.maximets@ovn.org> commit aa69918 vports are used concurrently and protected by RCU, so netdev_put() must happen after the RCU grace period. So, either in an RCU call or after the synchronize_net(). The rtnl_delete_link() must happen under RTNL and so can't be executed in RCU context. Calling synchronize_net() while holding RTNL is not a good idea for performance and system stability under load in general, so calling netdev_put() in RCU call is the right solution here. However, when the device is deleted, rtnl_unlock() will call netdev_run_todo() and block until all the references are gone. In the current code this means that we never reach the call_rcu() and the vport is never freed and the reference is never released, causing a self-deadlock on device removal. Fix that by moving the rcu_call() before the rtnl_unlock(), so the scheduled RCU callback will be executed when synchronize_net() is called from the rtnl_unlock()->netdev_run_todo() while the RTNL itself is already released. Fixes: 6931d21 ("openvswitch: defer tunnel netdev_put to RCU release") Cc: stable@vger.kernel.org Acked-by: Eelco Chaudron <echaudro@redhat.com> Signed-off-by: Ilya Maximets <i.maximets@ovn.org> Acked-by: Aaron Conole <aconole@redhat.com> Link: https://patch.msgid.link/20260430233848.440994-2-i.maximets@ovn.org Signed-off-by: Paolo Abeni <pabeni@redhat.com> (cherry picked from commit aa69918) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
cve CVE-2026-74597 commit-author Zhiling Zou <zhilinz@nebusec.ai> commit f803c08 ip6ip6_err() clones an outer IPv6 ICMP error skb, pulls it to the quoted inner IPv6 packet, and then passes the clone to icmpv6_send(). The clone still carries the outer packet's inet6_skb_parm in skb->cb. If the outer packet had a Home Address Option, IP6CB(skb2)->dsthao remains non-zero after skb_pull(). icmpv6_send() later calls mip6_addr_swap(), which uses that stale dsthao offset against the quoted inner packet. A malformed inner destination-options header can then make the HAO lookup and address swap run past the end of the quoted packet and corrupt skb_shared_info. Clear skb2->cb[] before pulling the quoted inner IPv6 packet so the reply path does not reuse metadata left by the outer IPv6 stack. Fixes: e490d1d ("[IPV6] IP6TUNNEL: Split out generic routine in ip6ip6_err().") Cc: stable@vger.kernel.org Reported-by: Vega <vega@nebusec.ai> Signed-off-by: Zhiling Zou <zhilinz@nebusec.ai> Reviewed-by: Ido Schimmel <idosch@nvidia.com> Link: https://patch.msgid.link/fe1a5e765fbca88d69391887f0ed26a19e3e4d39.1785736562.git.zhilinz@nebusec.ai Signed-off-by: Jakub Kicinski <kuba@kernel.org> (cherry picked from commit f803c08) Signed-off-by: Shreeya Patel <spatel@ciq.com>
jira SECO-478 RFE: FUSE_IO_URING commit-author Bernd Schubert <bschubert@ddn.com> commit 92270d0 This function is needed by fuse_uring.c to clean ring queues, so make it non static. Especially in non-static mode the function name 'end_requests' should be prefixed with fuse_ Signed-off-by: Bernd Schubert <bschubert@ddn.com> Reviewed-by: Josef Bacik <josef@toxicpanda.com> Reviewed-by: Joanne Koong <joannelkoong@gmail.com> Reviewed-by: Luis Henriques <luis@igalia.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit 92270d0) Signed-off-by: Jonathan Maple <jmaple@ciq.com>
jira SECO-478 RFE: FUSE_IO_URING commit-author Joanne Koong <joannelkoong@gmail.com> commit c146284 upstream-diff | fch->lock/fch->connected changed to fc->lock/fc->connected in fuse_uring_do_register() already applied in 23e9c1c. Only the fuse_uring_create() hunk is new. Check fc->connected under fc->lock in fuse_uring_create() before attaching a new ring. Without this, a race between fuse_uring_create() and fuse_conn_destroy() can result in the ring, queue, and fpq.processing table being created after fuse_uring_abort() has already run, leading to unnecessary allocation and teardown. These are eventually cleaned up by fuse_uring_destruct() but will linger until the process exits, even with the connection aborted. Reviewed-by: Bernd Schubert <bernd@bsbernd.com> Signed-off-by: Joanne Koong <joannelkoong@gmail.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit c146284) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
jira SECO-478 cve CVE-2026-64263 RFE: FUSE_IO_URING commit-author Joanne Koong <joannelkoong@gmail.com> commit 198f45e upstream-diff | io_uring_cmd_done() uses 4 args (res2=0) vs upstream 3 args. list_move used instead of list_move_tail. fuse_uring_cancel() moves entries that are available (these have no reqs attached) to the ent_in_userspace list. ent_list_request_expired() checks the first entry on ent_in_userspace and dereferences ent->fuse_req unconditionally, which will crash on a cancelled entry that was moved to this list. Fix this by freeing the entry and dropping queue_refs directly in fuse_uring_cancel(). This is safe because cancel is the cancel handler itself - after io_uring_cmd_done(), no more cancels will be dispatched for this command, and teardown serializes with cancel via queue->lock. Since cancel now decrements queue_refs, fuse_uring_abort() must no longer gate fuse_uring_abort_end_requests() on queue_refs > 0, as cancelled entries may have already dropped queue_refs while requests are still queued. Remove the gate so abort always flushes requests and stops queues. Reported-by: Heechan Kang <gganji11@naver.com> Tested-by: Heechan Kang <gganji11@naver.com> Reviewed-by: Bernd Schubert <bernd@bsbernd.com> Fixes: 4fea593 ("fuse: optimize over-io-uring request expiration check") Cc: stable@vger.kernel.org Suggested-by: Jian Huang Li <ali@ddn.com> Suggested-by: Horst Birthelmer <horst@birthelmer.de> Signed-off-by: Joanne Koong <joannelkoong@gmail.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit 198f45e) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
jira SECO-478 cve CVE-2026-64262 RFE: FUSE_IO_URING commit-author Chris Mason <clm@meta.com> commit bea4fe9 upstream-diff | fuse_uring_send_in_task has different signature (missing io_tw_req/io_tw_token_t API in this kernel) io_uring_cmd_done() call corrected to pass 4 args (res2=0 was missing in upstream adaptation). Upstream dropped the res2 param in: ef9f603 - io_uring/cmd: drop unused res2 param from io_uring_cmd_done() When io_uring delivers task work with tw.cancel set (PF_EXITING, PF_KTHREAD fallback, or percpu_ref_is_dying on the ring context), fuse_uring_send_in_task() takes the cancel branch, assigns -ECANCELED, and falls through to fuse_uring_send(). That path only flips the entry to FRRS_USERSPACE and completes the io_uring cmd; it never discharges the ring entry's owning reference to the fuse_req that fuse_uring_add_req_to_ring_ent() handed it at dispatch time. fuse_uring_send_in_task() tw.cancel == true err = -ECANCELED fuse_uring_send(ent, cmd, err, issue_flags) ent->state = FRRS_USERSPACE list_move(&ent->list, &queue->ent_in_userspace) ent->cmd = NULL io_uring_cmd_done(-ECANCELED) /* ent->fuse_req still set, req still hashed */ The fuse_req stays linked on fpq->processing[hash] and fuse_request_end() is never invoked. The originating syscall thread blocks in D-state in request_wait_answer() until fuse_abort_conn() runs, which can be the entire connection lifetime. For FR_BACKGROUND requests fc->num_background is never decremented either, so repeated cancels inflate the counter until max_background is hit and all later background ops stall. tw.cancel does not imply a connection abort (e.g. a single io_uring worker thread exits while the fuse connection stays up), so this cannot be left for fuse_abort_conn() to clean up. Ending the req but still routing the entry through fuse_uring_send() is not enough: that leaves a req-less entry on ent_in_userspace, and ent_list_request_expired() dereferences ent->fuse_req unconditionally on the head of that list, which would then NULL-deref. Fix the cancel branch to release the entry directly. Remove it from the queue, complete the io_uring cmd, end the fuse_req, free the entry, and drop its queue_refs (waking the teardown waiter if it was the last). Fixes: c2c9af9 ("fuse: Allow to queue fg requests through io-uring") Cc: stable@vger.kernel.org Reviewed-by: Joanne Koong <joannelkoong@gmail.com> Assisted-by: kres (claude-opus-4-7) Signed-off-by: Chris Mason <clm@meta.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit bea4fe9) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
jira SECO-478 cve CVE-2026-64261 RFE: FUSE_IO_URING commit-author Bernd Schubert <bernd@bsbernd.com> commit d351da7 upstream-diff | ring->fc used instead of ring->chan->conn fuse_uring_async_stop_queues() might run when the last reference on ring->queue_refs was already dropped. In order to avoid an early destruction a reference on struct fuse_conn is now taken before starting fuse_uring_async_stop_queues() and that reference is only released when that delayed work queue terminates. Fixes: 4a9bfb9 ("fuse: {io-uring} Handle teardown of ring entries") Cc: stable@kernel.org # 6.14 Reported-by: Berkant Koc <me@berkoc.com> Signed-off-by: Bernd Schubert <bernd@bsbernd.com> Reviewed-by: Joanne Koong <joannelkoong@gmail.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit d351da7) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
…lock jira SECO-478 cve CVE-2026-64260 RFE: FUSE_IO_URING commit-author Bernd Schubert <bernd@bsbernd.com> commit b70a3ac upstream-diff | ring->fc used instead of ring->chan->conn There are several readers of queue->stopped that check the value under lock, but fuse_uring_commit_fetch() did not and actually the value was not set under the lock in fuse_uring_abort_end_requests() either. Especially in fuse_uring_commit_fetch it is important to check under a lock, because due to races 'struct fuse_req' might be freed with fuse_request_end, but another thread/cpu might already do teardown work. Cc: stable@kernel.org # 6.14 Fixes: 4a9bfb9 ("fuse: {io-uring} Handle teardown of ring entries") Reported-by: Berkant Koc <me@berkoc.com> Reported-by: xlabai <xlabai@tencent.com> Signed-off-by: Bernd Schubert <bernd@bsbernd.com> Reviewed-by: Joanne Koong <joannelkoong@gmail.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit b70a3ac) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
jira SECO-478 cve CVE-2026-64259 RFE: FUSE_IO_URING commit-author Bernd Schubert <bernd@bsbernd.com> commit 1efd3d4 upstream-diff | list_move used instead of list_move_tail and io_uring_cmd_done has extra arg (older API in this kernel) Bad userspace might try to trick us and send commit SQEs request unique / commit-id of requests that are not even send to fuse-server (io_uring_cmd_done() not called) yet. fuse_uring_commit_fetch() ends the fuse request when the ring entry has a wrong state, but that could have caused a use-after-free with the memcpy operations in fuse_uring_send_in_task(). In order to avoid such races the call of fuse_uring_add_to_pq() is moved after the copy operations and just before completing the io-uring request - malicious userspace cannot find the request anymore until all prepration work in fuse-client/kernel is completed. This also moves fuse_uring_add_to_pq() a bit up in the code to avoid a forward declaration. Also not with a preparation commit, to make it easier to back port to older kernels. Reported-by: xlabai <xlabai@tencent.com> Reported-by: Berkant Koc <me@berkoc.com> Fixes: c090c8a ("fuse: Add io-uring sqe commit and fetch support") Cc: stable@kernel.org # 6.14 Signed-off-by: Bernd Schubert <bernd@bsbernd.com> Reviewed-by: Joanne Koong <joannelkoong@gmail.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit 1efd3d4) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
…ULL deref jira SECO-478 cve CVE-2026-64258 RFE: FUSE_IO_URING commit-author Joanne Koong <joannelkoong@gmail.com> commit 1c57a69 If a copy into the userspace ring buffer fails, a request will be terminated and fuse_uring_req_end() will set ent->fuse_req to NULL but it will leave the entry on ent_w_req_queue in FRRS_FUSE_REQ state. This can lead to a NULL deref if the request expiration logic scans ent_w_req_queue in the window before the entry is moved off it. Fix this by taking the entry off ent_w_req_queue and changing its state from FRRS_FUSE_REQ to FRRS_INVALID before terminating the request. Fixes: 4fea593 ("fuse: optimize over-io-uring request expiration check") Cc: stable@kernel.org Signed-off-by: Joanne Koong <joannelkoong@gmail.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit 1c57a69) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
jira SECO-478 RFE: FUSE_IO_URING commit-author Joanne Koong <joannelkoong@gmail.com> commit 31da059 upstream-diff | fuse_copy_init uses int write vs bool write in this kernel When a background request completes via the io_uring path, the background queue gets flushed to dispatch pending background requests, but this is done before the connection-level background counters (fc->num_background, fc->active_background) are properly accounted, which may reduce effective queue depth to one. The connection-level counters are decremented in fuse_request_end(), but flush_bg_queue() flushes the /dev/fuse path queue (fc->bg_queue), not the io_uring per-queue bg one, which means pending uring background requests on the queue are never dispatched in this path. Fix this by accounting the connection-level background counters first before flushing the queue's background queue. Since fuse_request_bg_finish() clears FR_BACKGROUND, fuse_request_end() will skip the background cleanup branch entirely, which avoids any double-decrements; it will call the wake_up(&req->waitq) branch but this is effectively a no-op as background requests have no waiters on req->waitq. Reviewed-by: Bernd Schubert <bernd@bsbernd.com> Fixes: 857b026 ("fuse: Allow to queue bg requests through io-uring") Cc: stable@vger.kernel.org Signed-off-by: Joanne Koong <joannelkoong@gmail.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit 31da059) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
jira SECO-478 RFE: FUSE_IO_URING commit-author Zhenghang Xiao <kipreyyy@gmail.com> commit 7d87a5a fuse_uring_commit_fetch() error path called fuse_request_end(req) without clearing ent->fuse_req when fuse_ring_ent_set_commit() fails. The still-pending fuse_uring_send_in_task() task-work later dereferences the dangling pointer through fuse_uring_prepare_send(), causing a use-after-free. End the request with fuse_uring_req_end(), which handles all conditions already. Annotation/edition by Bernd: The UAF should be fixed by other means already and actually has to be avoided that way. Just checking for ent->fuse_req == NULL in fuse_uring_send_in_task() would be prone to race conditions, because if malicious userspace would commit requests that have passed the NULL check, but are in doing args copy, it would still trigger a use-after-free. Setting ent->fuse_req = NULL in fuse_uring_commit_fetch() still makes sense, though. Reported-by: Shuvam Pandey <shuvampandey1@gmail.com> Reported-by: Berkant Koc <me@berkoc.com> Signed-off-by: Zhenghang Xiao <kipreyyy@gmail.com> Signed-off-by: Bernd Schubert <bernd@bsbernd.com> Reviewed-by: Joanne Koong <joannelkoong@gmail.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit 7d87a5a) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
jira SECO-478 RFE: FUSE_IO_URING commit-author Joanne Koong <joannelkoong@gmail.com> commit edb310b upstream-diff | fuse_conn/fc used instead of fuse_chan/fch in this kernel fuse_block_alloc() reads fch->initialized and then fch->io_uring. fch->io_uring is set before fch->initialized, ordered by the smp_wmb() in fuse_chan_set_intialized(), but fuse_block_alloc() has no matching read barrier between the two loads. This may lead a CPU to observe fch->initialized=1 but fch->io_uring=0, and skip the check that blocks request allocation until the io-uring queues are ready. This can reintroduce the lock-order inversion deadlock that commit 3393ff9 prevents. Add an smp_rmb() barrier to pair with the smp_wmb() in fuse_chan_set_initialized() to prevent this. Fixes: 3393ff9 ("fuse: block request allocation until io-uring init is complete") Cc: stable@vger.kernel.org Reviewed-by: Bernd Schubert <bernd@bsbernd.com> Signed-off-by: Joanne Koong <joannelkoong@gmail.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit edb310b) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
jira SECO-478 RFE: FUSE_IO_URING commit-author Joanne Koong <joannelkoong@gmail.com> commit 42df916 upstream-diff | fc->lock used instead of fch->lock in this kernel fuse_uring_create_queue() initializes a fuse_ring_queue and then publishes the pointer into ring->queues[qid] with WRITE_ONCE() under the fch->lock. There are several readers that may concurrently be fetching that pointer locklessly and then deferencing it. WRITE_ONCE() doesn't ensure ordering of the queue's field initialization before the ring->queues[qid] pointer assignment. The queue must be published with smp_store_release() so the field initialization is guaranteed to happen before. Readers in paths where the read may happen concurrently with the store need to use READ_ONCE() because any race involving a plain access is undefined. Fixes: 24fe962 ("fuse: {io-uring} Handle SQEs - register commands") Cc: stable@vger.kernel.org Reviewed-by: Bernd Schubert <bernd@bsbernd.com> Signed-off-by: Joanne Koong <joannelkoong@gmail.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit 42df916) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
jira SECO-478 RFE: FUSE_IO_URING commit-author Xiang Mei <xmei5@asu.edu> commit fd10f40 upstream-diff | copy_{to,from}_user used directly instead of copy_header_{to,from}_ring helpers in this kernel; fc used instead of fch The fuse-io-uring transport copies req->in.h out to the ring in fuse_uring_copy_to_ring() and req->out.h back in fuse_uring_commit(). Both headers live inside the fuse_request slab object, whose cache (fuse_req_cachep) is created without a usercopy whitelist, so copying them directly to/from userspace trips CONFIG_HARDENED_USERCOPY and panics: usercopy: Kernel memory exposure attempt detected from SLUB object 'fuse_request' (offset 56, size 40)! kernel BUG at mm/usercopy.c:102! Oops: invalid opcode: 0000 [#1] SMP KASAN NOPTI RIP: 0010:usercopy_abort (mm/usercopy.c:90) Call Trace: __check_heap_object (mm/slub.c:8268) __check_object_size (mm/usercopy.c:197 mm/usercopy.c:258 mm/usercopy.c:223) copy_header_to_ring (fs/fuse/dev_uring.c:618) fuse_uring_prepare_send (fs/fuse/dev_uring.c:776 fs/fuse/dev_uring.c:785) fuse_uring_send_in_task (fs/fuse/dev_uring.c:1306) tctx_task_work_run (io_uring/tw.c:96) task_work_run (kernel/task_work.c:233) io_run_task_work (io_uring/tw.h:84) io_cqring_wait (io_uring/wait.c:278) __do_sys_io_uring_enter (io_uring/io_uring.c:2685) entry_SYSCALL_64_after_hwframe (arch/x86/entry/entry_64.S:121) Bounce both headers through an on-stack copy so the usercopy touches stack memory, not the slab object. Fixes: c090c8a ("fuse: Add io-uring sqe commit and fetch support") Cc: stable@vger.kernel.org Reported-by: Weiming Shi <bestswngs@gmail.com> Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Xiang Mei <xmei5@asu.edu> Reviewed-by: Bernd Schubert <bernd@bsbernd.com> Reviewed-by: Joanne Koong <joannelkoong@gmail.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit fd10f40) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
jira SECO-478 RFE: FUSE_IO_URING commit-author Bernd Schubert <bernd@bsbernd.com> commit 1f59015 upstream-diff | fc/fc->aborted used instead of fch/fch->abort_with_err; fuse_set_initialized() used instead of fuse_chan_set_initialized() in this kernel The existing condition in fuse_uring_cmd() is there only to avoid disabling io-uring for connections that already run with it, missing was a condition to refuse any IORING_OP_URING_CMD if the connection/channel didn't get enabled because of missing FUSE_INIT reply flag FUSE_OVER_IO_URING. Without the reply flag the barrier in fuse_uring_ready() doesn't work and IO could already be going on and cause deadlock states (at a minimum one between fch->bg_lock and queue->lock). The change itself is trivial, but brings behavior change, FUSE_OVER_IO_URING has to be set in the FUSE_INIT_REPLY by fuse servers to accept any IORING_OP_URING_CMD. Libfuse does that and the only non-libfuse implementation I found (fractal-fuse) also does it. Qemu patches for fuse-io-uring are not merged yet, as far as I know. Moved up is the smp_load_acquire(&fch->initialized) check, as a fuse-server implementation might try to setup io-uring before FUSE_INIT is processed and might have gotten -EOPNOTSUPP instead of -EAGAIN. Also fixed is a stale comment that explains the handling of the FUSE_OVER_IO_URING flag in early RFC versions. If there should be a report from any library or application we probably need to revert this commit. Fixes: 3393ff9 ("fuse: block request allocation until io-uring init is complete") Signed-off-by: Bernd Schubert <bernd@bsbernd.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit 1f59015) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
jira SECO-478 RFE: FUSE_IO_URING Enable FUSE io-uring support in x86_64 and aarch64 configs that have both prerequisites (CONFIG_FUSE_FS=m and CONFIG_IO_URING=y), across all variants (debug, rt, 64k). Skipped configs lacking prerequisites: - kernel-riscv64-*.config (no CONFIG_FUSE_FS or CONFIG_IO_URING) - kernel-s390x-zfcpdump-rhel.config (CONFIG_FUSE_FS is not set) The feature is gated at runtime by the enable_uring module parameter which defaults to disabled. Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
…ectl jira SECO-478 RFE: FUSE_IO_URING commit-author Chen Linxuan <chenlinxuan@uniontech.com> commit 1a7b137 This patch add a simple functional test for the 'abort' file in fusectlfs (/sys/fs/fuse/connections/ID/abort). A simple fuse daemon is added for testing. Signed-off-by: Chen Linxuan <chenlinxuan@uniontech.com> Acked-by: Shuah Khan <skhan@linuxfoundation.org> Reviewed-by: Amir Goldstein <amir73il@gmail.com> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com> (cherry picked from commit 1a7b137) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
jira SECO-478 RFE: FUSE_IO_URING On systems with only libfuse3-devel installed (no libfuse2-devel), fuse_mnt.c fails to compile because fuse3 requires API version 30+. Update fuse_mnt.c from FUSE API version 26 to 31: - getattr, truncate: add struct fuse_file_info * parameter - readdir: add enum fuse_readdir_flags parameter - filler calls: add flags argument Also update the Makefile to try pkg-config fuse3 as a fallback when pkg-config fuse is not available. Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
jira SECO-478 RFE: FUSE_IO_URING commit-author Amir Goldstein <amir73il@gmail.com> commit 9acb102 upstream-diff | kselftest_harness.h include uses relative path (../../kselftest_harness.h) because commit e6fbd17 ("selftests: complete kselftest include centralization") is not present in this tree. A FUSE mount that does not negotiate FUSE_POSIX_ACL initialises every inode with i_acl = i_default_acl = ACL_DONT_CACHE. When a fresh stat is needed (e.g. AT_STATX_FORCE_SYNC), fuse_update_get_attr() calls forget_all_cached_acls() before issuing FUSE_GETATTR. On an unfixed kernel, __forget_cached_acl() replaces ACL_DONT_CACHE with ACL_NOT_CACHED, inadvertently enabling the kernel ACL cache for that inode. This test validates the fix that preserves ACL_DONT_CACHE state in forget_cached_acl(). Signed-off-by: Amir Goldstein <amir73il@gmail.com> Signed-off-by: Christian Brauner <brauner@kernel.org> (cherry picked from commit 9acb102) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
jira SECO-478
RFE: FUSE_IO_URING
Add a kselftest for the FUSE io-uring request dispatch path. A
minimal in-memory FUSE daemon runs in a thread, negotiates
FUSE_OVER_IO_URING during FUSE_INIT, and handles requests through
FUSE_IO_URING_CMD_REGISTER / FUSE_IO_URING_CMD_COMMIT_AND_FETCH.
An atomic counter verifies requests are served through io_uring.
Test cases:
- read_file: read and verify file content
- write_and_readback: write data, read it back, compare
- readdir: list directory, verify entries
- stat_files: stat root and file, verify attributes
- concurrent_io: 4 threads doing interleaved reads and writes
- requests_via_uring: per-operation verification that read, write,
readdir, and stat each produce io_uring requests
- sustained_io: write 256KB in 4KB chunks, read back, verify pattern
- crash_recovery: fork daemon, SIGKILL with active I/O, verify
kernel health
- abort_before_ring_ready: negotiate io_uring but never register
entries, abort connection, verify no deadlock or leak
Tests SKIP when prerequisites are not met (no root, io_uring
disabled, enable_uring != Y).
Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
cve CVE-2026-52929 commit-author Wyatt Feng <bronzed_45_vested@icloud.com> commit a5f8a90 When ADD_OUT_STREAMS is denied, SCTP only shrinks the queued chunks and then lowers outcnt. That leaves removed stream metadata behind, so a later re-add can reuse a stale ext and hit a null-pointer dereference in the scheduler get path. Fix the rollback by tearing down the removed stream state the same way other stream resizes do. Unschedule the current scheduler state, drop the removed stream ext state with sctp_stream_outq_migrate(), and then reschedule the remaining streams. This keeps scheduler-private RR/FC/PRIO lists consistent while fully rolling back denied outgoing stream additions. Fixes: 637784a ("sctp: introduce priority based stream scheduler") Cc: stable@kernel.org Reported-by: Yuan Tan <yuantan098@gmail.com> Reported-by: Yifan Wu <yifanwucs@gmail.com> Reported-by: Juefei Pu <tomapufckgml@gmail.com> Reported-by: Zhengchuan Liang <zcliangcn@gmail.com> Reported-by: Xin Liu <bird@lzu.edu.cn> Signed-off-by: Wyatt Feng <bronzed_45_vested@icloud.com> Signed-off-by: Ren Wei <n05ec@lzu.edu.cn> Acked-by: Xin Long <lucien.xin@gmail.com> Link: https://patch.msgid.link/d78954ecd94954653ee299400e98d74a03a6f7d3.1780603399.git.bronzed_45_vested@icloud.com Signed-off-by: Jakub Kicinski <kuba@kernel.org> (cherry picked from commit a5f8a90) Signed-off-by: Shreeya Patel <spatel@ciq.com>
cve CVE-2026-43042 commit-author Sabrina Dubroca <sd@queasysnail.net> commit 629ec78 upstream-diff Uses a file-scope seqcount_t instead of upstream's per-netns seqcount_mutex_t to avoid a kABI-breaking struct change. Write side uses local_bh_disable() with preempt_disable_nested() for RT safety. Uses rcu_dereference_rtnl() instead of upstream's plain rcu_dereference() to stay lockdep-clean under RTNL. mpls_dump_routes() still runs under RTNL on this tree; the seqcount there is extra hardening. The RCU-protected codepaths (mpls_forward, mpls_dump_routes) can have an inconsistent view of platform_labels vs platform_label in case of a concurrent resize (resize_platform_label_table, under platform_mutex). This can lead to OOB accesses. This patch adds a seqcount, so that we get a consistent snapshot. Note that mpls_label_ok is also susceptible to this, so the check against RTA_DST in rtm_to_route_config, done outside platform_mutex, is not sufficient. This value gets passed to mpls_label_ok once more in both mpls_route_add and mpls_route_del, so there is no issue, but that additional check must not be removed. Reported-by: Yuan Tan <tanyuan98@outlook.com> Reported-by: Yifan Wu <yifanwucs@gmail.com> Reported-by: Juefei Pu <tomapufckgml@gmail.com> Reported-by: Xin Liu <bird@lzu.edu.cn> Fixes: 7720c01 ("mpls: Add a sysctl to control the size of the mpls label table") Fixes: dde1b38 ("mpls: Convert mpls_dump_routes() to RCU.") Signed-off-by: Sabrina Dubroca <sd@queasysnail.net> Link: https://patch.msgid.link/cd8fca15e3eb7e212b094064cd83652e20fd9d31.1774284088.git.sd@queasysnail.net Signed-off-by: Jakub Kicinski <kuba@kernel.org> (cherry picked from commit 629ec78) Signed-off-by: Brett Mastbergen <bmastbergen@ciq.com>
cve CVE-2026-80844 commit-author Asim Viladi Oglu Manizada <manizada@pm.me> commit 7bad4bd AH6 rearranges routing-header addresses before computing or verifying the ICV. ipv6_rearrange_rthdr() assumes that segments_left is not larger than the number of addresses described by the routing header's hdrlen field. That assumption does not hold for raw IPv6 HDRINCL packets. A packet with hdrlen equal to 2 describes one address, but can carry an arbitrary segments_left value. With segments_left equal to 255, the function moves its address pointer 4,064 bytes backwards and passes a 4,064-byte length to memmove(), resulting in an out-of-bounds access. Validate the invariant locally before modifying the routing header or performing any address-pointer arithmetic, and propagate malformed-header errors to the existing AH6 input and output error paths. Fixes: 1da177e ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Assisted-by: avom-custom-harness:gpt-5.5-qwen3.6-mod-mix Signed-off-by: Asim Viladi Oglu Manizada <manizada@pm.me> Signed-off-by: Steffen Klassert <steffen.klassert@secunet.com> (cherry picked from commit 7bad4bd) Signed-off-by: Jonathan Maple <jmaple@ciq.com>
cve CVE-2026-81000 commit-author Asim Viladi Oglu Manizada <manizada@pm.me> commit 447c930 tun_get_user() uses tun->align both as skb headroom and when choosing how much packet data to keep linear. OVS can propagate an oversized headroom request from another port to TUN or TAP. When align is larger than the usable space in a one-page skb head, SKB_MAX_HEAD(align) underflows and the result becomes negative when stored in good_linear. That value later wraps when assigned to the size_t linear variable, and tun_alloc_skb() can place skb->data outside the allocated head. Bound the headroom stored by TUN to the one-page skb-head budget and the largest non-sentinel 16-bit skb header offset. Leave one linear byte for raw TUN and a complete Ethernet header for TAP, including NET_IP_ALIGN. Also pull the raw-TUN protocol byte and the TAP Ethernet header before accessing them, so these checks remain safe for nonlinear skbs supplied by other allocation paths. Fixes: eaea34b ("net/tun: implement ndo_set_rx_headroom") Cc: stable@vger.kernel.org Signed-off-by: Asim Viladi Oglu Manizada <manizada@pm.me> Reviewed-by: Willem de Bruijn <willemb@google.com> Link: https://patch.msgid.link/20260812012139.2134643-1-manizada@pm.me Signed-off-by: Jakub Kicinski <kuba@kernel.org> (cherry picked from commit 447c930) Signed-off-by: Jonathan Maple <jmaple@ciq.com>
cve CVE-2026-68121 commit-author Asim Viladi Oglu Manizada <manizada@pm.me> commit e9c238f pppoe_sendmsg() saves a pointer to the PPPoE header before calling dev_hard_header(). Device header callbacks are allowed to reallocate the skb head, invalidating pointers into it. This can happen when a send is blocked in copy_from_user() while the first non-Ethernet port is added to an empty team device. The team's delegated GRE header callback then expands the skb head. PPPoE subsequently writes six bytes through the stale pointer into the freed head. Reload the PPPoE header through the skb's network-header offset after device header creation. pskb_expand_head() updates that offset when it relocates the head. Fixes: 1da177e ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Signed-off-by: Asim Viladi Oglu Manizada <manizada@pm.me> Reviewed-by: Vadim Fedorenko <vadim.fedorenko@linux.dev> Reviewed-by: Eric Dumazet <edumazet@google.com> Link: https://patch.msgid.link/20260722093814.3017176-1-manizada@pm.me Signed-off-by: Jakub Kicinski <kuba@kernel.org> (cherry picked from commit e9c238f) Signed-off-by: Jonathan Maple <jmaple@ciq.com>
cve CVE-2026-74469 commit-author Asim Viladi Oglu Manizada <manizada@pm.me> commit bd0e928 sctp_assoc_add_peer() increments the association's 16-bit transport_count for every new unique peer. Adding the 65,536th transport wraps the count to zero. SCTP sock_diag uses transport_count to reserve the INET_DIAG_PEERS payload, then copies one sockaddr_storage for every entry in transport_addr_list. After the wrap, a diagnostic dump reserves an empty payload and writes 8 MiB of peer addresses past the skb tail. Reject a new unique peer when transport_count has reached U16_MAX. Perform the check after the existing-peer lookup so a duplicate address continues to return its existing transport at the limit. Fixes: 8f840e4 ("sctp: add the sctp_diag.c file") Cc: stable@vger.kernel.org Signed-off-by: Asim Viladi Oglu Manizada <manizada@pm.me> Acked-by: Xin Long <lucien.xin@gmail.com> Link: https://patch.msgid.link/20260725032053.521705-1-manizada@pm.me Signed-off-by: Jakub Kicinski <kuba@kernel.org> (cherry picked from commit bd0e928) Signed-off-by: Jonathan Maple <jmaple@ciq.com>
shreeya-patel98
approved these changes
Sep 22, 2026
bmastbergen
approved these changes
Sep 22, 2026
bmastbergen
left a comment
Collaborator
There was a problem hiding this comment.
Below is the check_kernel_commits output for this PR:
% python ./check_kernel_commits.py --repo ~/ciq/kernel-src-tree --pr_branch origin/jmaple_rlc-10/6.12.0-211.56.1.el10_2 --base_branch origin/rlc-10/6.12.0-211.56.1.el10_2 --check-cves
[FIXES] PR commit 66a5f7c25455c (fuse: add more control over cache invalidation
behaviour) references upstream commit 2396356a945b, which has Fixes
tags:
6648f54f3459c fuse: move "epoch" from dentry.d_time to fuse_dentry.epoch (Miklos Szeredi)
0fa8346099b57 fuse: invalidate readdir cache on epoch bump (Jun Wu)
5a6baf2046105 fuse: fix uninit-value in fuse_dentry_revalidate() (Luis Henriques) (CVE-2026-53311)
[FIXES] PR commit b4cfa1e4b59f9 (fuse: {io-uring} Handle SQEs - register
commands) references upstream commit 24fe962c86f5, which has Fixes tags:
1dfe2a220e9cd fuse: fix uring race condition for null dereference of fc (Joanne Koong)
[CVE-MISSING] PR commit 8b6969821270d (fuse: fix missing barrier when checking
io-uring readiness) does not reference a CVE but upstream commit
edb310bc27f0 is associated with CVE-2026-80859
[CVE-MISSING] PR commit f72ebdf4c8e80 (fuse: publish io-uring queues with
release semantics) does not reference a CVE but upstream commit
42df916e5a5f is associated with CVE-2026-80858
[CVE-MISSING] PR commit c07b0a6e9ae33 (fuse: copy request headers via a stack
buffer for io-uring) does not reference a CVE but upstream commit
fd10f40af314 is associated with CVE-2026-80946
[CVE-MISSING] PR commit 043b73de49e9f (fuse: Fix the condition to enable
over-io-uring) does not reference a CVE but upstream commit
1f59015e9581 is associated with CVE-2026-90095
I haven't investigated the fixes yet, so I'm going to approve, but look into whether we want the fixes (especially the CVE-2026-53311 fix) for future releases
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
https://ciqinc.atlassian.net/browse/KERNEL-1619
Update process (This kernel CentOS base for 6.12.0-211.56.1.el10_2)
rlc-10/6.12.0-211.56.1.el10_2branch fromrocky10_2rlc-10/6.12.0-211.54.1.el10_2into new branch (skipping unneeded code)Rebase Log
BUILD
KSelfTest