two months too late (linux kernel 1day) – djnn@localhost

20 min read Original article ↗

Hey, it has been a while. :)

A coworker had been poking at USB/IP recently and I figured I’d take a look at the kernel side of it. After all, it’s just there. The protocol tunnels USB over TCP, no authentication on the wire, and the client kernel just parses whatever bytes the server sends. That is the kind of attack surface I will absolutely spend a weekend on. Moreover, USB is very stateful. It smells like bugs.

I ended up with a chain that gives a malicious server kernel RIP control against stock Debian 12 LTS. Unauthenticated, no client privileges, KASLR bypassed, lands in under a minute. Just a crafted RET_SUBMIT over TCP, and the kernel jumps wherever the server says to.

So I wrote it up, filmed the demo, drafted the advisory. As one last sanity check before emailing the kernel security team, I ran git log on drivers/usb/usbip/usbip_common.c to pin the original bug-introducing commit (you always want that line in an advisory).

What came back was a three-patch series sitting upstream, fixing the exact two bugs I had been exploiting, merged into mainline on 2026-03-25 by someone else. Two months before I finished, and only a few days before I started looking at it. (I’ve been poking at it on and off, I figured it would be fine since the bugs have been there for fifteen years. big mistake lol)

Well, fuck. Here is the patch.msgid.link if you’re curious.

The bugs were:

  • An info leak via OOB read in usbip_pad_iso(). Fixed in 591c1d972d8f and 74a2287209a8.
  • An OOB write via unvalidated number_of_packets in usbip_recv_iso(). Fixed in 1897852293fa, with a defense-in-depth follow-up a week later in 2ab833a16a82.

Chained, on Debian 12 stable (kernel 6.1 LTS), they hand the server an indirect call inside kernel space. I managed to get RIP control over the network, with usbip attach as the only client-side interaction. Mind you, it’s not a root shell on its own. It would require a bit more work, but I figured my point was proven, and did not feel like working on it more.

What is left to talk about is everything around the primitive: the heap grooming, surviving the two list_del calls that run before urb->complete is reached, the iso-length sum-check bypass, recovering the usbcore KASLR slide through the same bug class used to leak it. None of that is in Kelvin’s commit messages, so at least the writeup is not entirely redundant, I guess.

USB, a short primer

USB is a host-controller-device architecture. One host, everything else is a device. The host initiates all transfers, the device only responds when polled. Devices announce themselves through descriptors, sequences of bytes the kernel reads and acts on. If you control the bytes, you control the parser.

enumeration

Enumeration is the handshake that happens when a device connects. The bus detects a voltage change, the host resets the device, and the two of them go through a fixed sequence:

host                                   device
 |                                        |
 |---- GET_DESCRIPTOR (8 bytes) --------> |  "how big is your EP0 packet?"
 |<------- device descriptor (partial) -- |
 |                                        |
 |---- SET_ADDRESS(N) ------------------> |
 |<------- ACK -------------------------- |  device now lives at address N
 |                                        |
 |---- GET_DESCRIPTOR (18 bytes) -------> |
 |<------- device descriptor (full) ----- |  vid, pid, class, num configs
 |                                        |
 |---- GET_DESCRIPTOR (config) ---------> |
 |<------- config + interfaces + eps ---- |  full descriptor tree
 |                                        |
 |---- SET_CONFIGURATION(1) ----------->  |
 |<------- ACK -------------------------- |  device is ready

Each step depends on the one before it. If anything in the descriptor tree is malformed, the host either rejects the device, or (in some cases) crashes trying to parse it.

descriptors

device descriptor
└── configuration descriptor
    └── interface descriptor
        └── endpoint descriptor

The device descriptor identifies the vendor, product, and device class. The configuration descriptor describes one usable operating mode (a device can have several, though most have one). The interface descriptor maps to a function, for example a USB audio device might have one interface for playback and one for recording. Finally, the endpoint descriptor describes an individual communication channel: its direction, its transfer type, and the max packet size.

transfer types and frames

USB schedules transfers in fixed time slots called frames. Each frame is shared among four transfer types, and the host promises bandwidth to each one differently:

  • Control transfers are always available on endpoint 0. They carry an 8-byte setup packet followed by an optional data stage. Used for enumeration, configuration, and vendor-specific request/reply traffic.
  • Bulk transfers are best-effort but reliable. The host retries on error and there is no latency guarantee. Used by mass storage, printers and the like.
  • Interrupt transfers are periodic polls at a guaranteed interval. Used by HID-class devices (keyboards, mice). The “interrupt” name is historical; the host actually polls, not the other way around.
  • Isochronous (“iso”) transfers are real-time. Bandwidth is reserved every frame, but there is no retry on error and no acknowledgement. Think audio, video, mics.

Iso transfers carry a per-packet descriptor array, one entry per frame slot. Each descriptor has four fields:

struct iso_packet_descriptor {
    uint32_t offset;          // position of this packet's data inside transfer_buffer
    uint32_t length;          // expected size for this packet
    uint32_t actual_length;   // size actually transferred
    int32_t  status;          // 0 = OK, negative = errno
};

When the host submits an iso URB, it fills in offset and length. When the transfer comes back, the device (or, in our case, the malicious USB/IP server) writes back actual_length and status. The kernel uses those values to copy data into place inside the transfer buffer. This struct is the surface where everything in this post happens.

URBs

To actually talk to an endpoint, the kernel uses a URB (USB Request Block). It is a struct that describes one transfer: endpoint, data buffer, transfer length, flags, completion callback. The kernel hands it to the host controller driver, which sends it to the device and calls the completion callback when the response comes back. For iso transfers, the URB has a flexible array of per-packet descriptors at the end. The length of that array is fixed at URB allocation time, based on the application’s number_of_packets.

From userspace through usbdevfs, applications fill in a URB and submit it via USBDEVFS_SUBMITURB. Later, they reap the completion via USBDEVFS_REAPURB to read the data buffer and the per-packet actual_length / status. In between, the kernel sometimes reshuffles the transfer buffer to undo a wire-side bandwidth optimisation, which is what usbip_pad_iso does, and which is where one of the bugs lives.

control transfers

Control transfers are special because they always start with an 8-byte setup packet:

struct usb_ctrlrequest {
    uint8_t  bmRequestType;   // direction | type | recipient
    uint8_t  bRequest;        // request code
    uint16_t wValue;
    uint16_t wIndex;
    uint16_t wLength;         // size of optional data stage
};

bmRequestType packs three things in one byte: data direction (bit 7), request type (bits 6-5: standard, class, or vendor), and recipient (bits 4-0: device, interface, endpoint, other). Standard requests like GET_DESCRIPTOR (bRequest=0x06) or SET_ADDRESS (bRequest=0x05) are how enumeration works. Vendor-specific requests are how a device exposes proprietary control surfaces. With the upper bits of bmRequestType set to 0x40 (vendor, host-to-device, device recipient), and bRequest=0xEE, the exploit uses this as a side-channel to ship kernel addresses to its own server. More on that later.

stateful protocol

Enumeration alone is a dozen round trips, each building on the last. Once a device is in use, the host and device share a lot of context: configurations, interfaces, in-flight URBs and what responses to expect from them. A URB is submitted with a specific number_of_packets; the kernel remembers it and expects the device to match it in the response. There is a lot of surface area where “trust what the device says” and “validate it against what we sent” come apart. That is where the bugs are.

USB/IP

USB/IP is a kernel subsystem that tunnels USB over TCP. The server exports a physical USB device. The client runs usbip attach and gets a virtual USB device backed by a live network socket. The intended use case is sharing USB hardware over a network, like plugging a dongle into one machine and using it from another. So far so good.

On the client side, vhci_hcd is a virtual host controller. From the kernel’s perspective it is just a USB device going through enumeration like any other. From the server’s perspective, the server controls every byte the kernel receives, including descriptor data during enumeration and every transfer response after that. In a normal USB setup the device is a physical thing constrained to speaking USB. With USB/IP, the “device” is a process on a machine you control, and it can say anything.

The wire protocol mirrors the URB lifecycle. When the kernel submits a URB, the client sends CMD_SUBMIT to the server. The server processes it (in its head) and sends back RET_SUBMIT. For iso transfers, both messages carry per-packet descriptors:

CMD_SUBMIT (kernel -> server):
  [ header: cmd | seqnum | devid | direction | ep ]
  [ transfer_flags | buffer_length | start_frame | number_of_packets | interval ]
  [ setup[8] ]
  [ transfer_buffer (buffer_length bytes) ]
  [ iso_packet_descriptor x number_of_packets ]
         ^
         submitted with number_of_packets = 4

RET_SUBMIT (server -> kernel):
  [ header: cmd | seqnum | devid | direction | ep ]
  [ status | actual_length | start_frame | number_of_packets | error_count ]
                                                    ^
                                                    server can put anything here
  [ transfer_buffer (actual_length bytes) ]
  [ iso_packet_descriptor x number_of_packets ]

There are two bugs in how the kernel processes those iso descriptors. They are in the same file but are independent.

bug 1: info leak

usbip_pad_iso() is supposed to undo a bandwidth optimisation done over the wire. The iso descriptors carry an offset and an actual_length; the function walks them backwards and memmoves each packet’s data into its declared offset slot. Per iteration:

actualoffset -= urb->iso_frame_desc[i].actual_length;
memmove(urb->transfer_buffer + urb->iso_frame_desc[i].offset,
        urb->transfer_buffer + actualoffset,
        urb->iso_frame_desc[i].actual_length);

actualoffset is a signed int. It starts at urb->actual_length (server-controlled) and gets decremented by each packet’s actual_length (also server-controlled). There is no lower-bound check. A response whose per-packet lengths sum to more than the buffer total drives actualoffset negative, and the memmove source pointer ends up below the start of transfer_buffer.

The bytes that come back are whatever was just before the URB’s transfer buffer in kernel heap (80 bytes on Debian 12 6.1, 72 on 6.12, depending on where urb->complete sits in the kmalloc-256 slot). When the URB is reaped, userspace gets those bytes through the normal REAPURB ioctl.

Repeatable per malicious response. On a busy system the leaked bytes can contain kernel pointers that the kernel later tries to interpret as list entries, panicking in __list_del_entry_valid_or_report. On a freshly-booted minimal initramfs the leaked bytes are usually nothing interesting and the kernel keeps running.

bug 2: OOB write

usbip_recv_iso() reads the per-packet descriptor array from the socket and unpacks it into the URB’s iso_frame_desc array:

np = urb->number_of_packets;   /* already clobbered from the wire by usbip_pack_pdu() */
for (i = 0; i < np; i++)
    usbip_pack_iso(buff + i * ..., &urb->iso_frame_desc[i], 0);

By the time this runs, urb->number_of_packets has already been overwritten with whatever the server sent. There is no check that the new value matches what the URB was submitted with.

Funny thing: this code has been there since the beginning, but the original 2008 USB/IP client was not vulnerable. It used its own stored number_of_packets value and just ignored the server’s. A patch in April 2011 added the field to the RET_SUBMIT wire format, for interop with another USB/IP client (the Windows one), and in doing so started writing the server’s value into the URB without checking it. The commit is 1325f85fa49f by Arjan Mels. It was there for fifteen years, and got fixed mere days before I looked into it by chance. Talk about unlucky lol

Anyway, struct urb with 4 iso descriptors is exactly 256 bytes: 192 for the base struct, plus 4 entries of 16 bytes each. That is one kmalloc-256 allocation. If the server returns number_of_packets = 16, the kernel happily writes 16 entries starting at urb->iso_frame_desc[0]. Entries [4] through [15] overflow 192 bytes past the end of the allocation, into whatever sits next to it in the slab.

If that adjacent slot happens to be another URB, entry [15] lands right on top of urb2->complete:

urb1  (256 bytes, kmalloc-256):
  +184  complete              <- function pointer, called on URB completion
  +192  iso_frame_desc[0]
  ...
  +432  iso_frame_desc[15]    <- this is urb2+176 if urb1 and urb2 are adjacent

urb2  (next slab slot):
  +176  iso[15].offset        (benign)
  +180  iso[15].length        (benign)
  +184  iso[15].actual_length -> urb2->complete[0:4]
  +188  iso[15].status        -> urb2->complete[4:8]

The server packs a 64-bit kernel address across the actual_length and status fields of entry [15]. When the kernel completes urb2, it calls urb2->complete(urb2) and jumps there.

That’s about it, pretty much. Turning it into an actual RIP hijack is where it gets interesting.

trying it on Debian 13

So I built the chain in a lab kernel and it worked, right ? Then I tried to demonstrate it against stock Debian trixie, because that is what people actually run.

148 attempts, zero adjacent URB pairs. Something was systematically defeating adjacency. Actually, it was two things, in order of importance:

CONFIG_SLAB_BUCKETS=y, on by default in Debian since kernel 6.10. It routes same-size allocations from different call sites to different sub-caches. The grooming primitive I was trying to use (spraying add_key payloads that land in kmalloc-256) was draining one bucket while URBs were landing in another, so the spray had no effect.

usbdevfs_submit_iso() allocates a struct async between consecutive URB submissions. So “two back-to-back SUBMITURB ioctls produce two adjacent URB allocations” was already a fragile assumption under freelist randomisation. With an intervening allocation in the same cache, the two URBs are not consecutive freelist picks at all.

what about IBT? FineIBT? kCFI?

I went in expecting that even with perfect heap grooming, the blocker would be control-flow integrity. A corrupted urb->complete pointing at dump_stack should fail an indirect-call signature check on a kernel with kCFI on. So even if I got the write right, the call would never go through.

Looking at the actual Debian config, that turned out to be wrong on both targets:

# Debian 12 bookworm (linux-image-6.1.0-47-amd64), the demo target:
# CONFIG_X86_KERNEL_IBT is not set
# (no CONFIG_CFI_CLANG line)
# (no CONFIG_FINEIBT line)
CONFIG_CC_HAS_IBT=y
CONFIG_ARCH_SUPPORTS_CFI_CLANG=y

So there is no CFI check at all on the demo kernel. The corrupted urb->complete runs unchecked. CC_HAS_IBT just says the toolchain can emit IBT instructions, and ARCH_SUPPORTS_CFI_CLANG says x86 supports kCFI, neither of which says the kernel enabled them. Plus the Debian kernel is built with gcc, and CONFIG_CFI_CLANG requires Clang, so kCFI is not available on this build path regardless.

The blocker on Debian 13 is purely the allocator. If you could land the adjacent URB, the indirect call would just go through. And on Debian 12, where the chain actually lands, the call check is a non-event because it is not enabled.

pivot to Debian 12 (6.1 LTS)

CONFIG_SLAB_BUCKETS landed in 6.10. Debian 13 ships 6.12. Therefore, I just decided to target Debian 12 LTS. After all, it’s still used a lot out there, so it made sense to prove my point.

Without SLAB_BUCKETS the add_key spray drains the same kmalloc-256 cache the URBs come from, so urb1 and urb2 actually end up on the same slab page. SLUB’s per-page freelist is shuffled, so adjacency is still probabilistic, about 1 in 16 per attempt, but the runner just retries until it hits. Usually lands under a minute.

Surviving the actual write turned out to be the fiddly part. The kernel does not just call urb->complete(urb) and crash usefully. Between the OOB write and the complete callback, it goes through usb_hcd_unlink_urb_from_ep (which does a list_del on urb->urb_list), and then through usb_hcd_giveback_urb (which dereferences urb->dev->bus->busnum for kcov, plus a few other fields). The OOB write hits all of urb2[0..192], so every one of those fields is corrupted. Each is a place the kernel can panic before our hijacked pointer ever fires.

Three issues, three fixes, all encoded into the same crafted RET_SUBMIT:

  1. urb2->urb_list.prev has to keep pointing back at urb1’s urb_list, or the list_del for urb1 fails its next->prev != entry integrity check and panics with list_del corruption, but we can restore it. (so we do)
  2. urb2->urb_list.next has to point at something whose prev field will satisfy the same check when the kernel later deletes urb2. By then, the kernel has rewritten urb2->prev itself (during list_del(urb1)), so we plant a fake list_head in the adjacent anchor_list slot at urb2+40, with its prev set back to urb2+24. Neat, right ?
  3. urb2->dev has to remain a valid struct usb_device *, because usb_hcd_giveback_urb dereferences it before calling complete. We leak its address from the same probe URB we already use for urb2, and write it back.

There is a fourth one that took me longer than I want to admit: usbip_recv_iso has a sum check after parsing the iso descriptors. If the sum of all iso[i].actual_length values does not match urb->actual_length, it logs an error and tears down the connection via usbip_event_add(VDEV_EVENT_ERROR_TCP). Our pointer-bearing entries contain bytes that, treated as int32s, sum to nonsense. So we set urb->actual_length=0 in the header and use one of the unused iso entries to balance the sum back to zero. The byte it lands on is urb->pipe, but that field is not read until after urb->complete fires in usb_hcd_giveback_urb, so corrupting it is harmless to the chain.

With those four constraints satisfied, urb1’s benign list_del survives, control reaches urb1->complete (the legitimate async_completed), the kernel goes back to the rx loop, the server sends a normal RET_SUBMIT for urb2, and when urb2 is given back the kernel calls urb2->complete(urb2), which is now the value we wrote.

To be precise about what this proves: it’s RIP control, not machine control. The OOB write puts an arbitrary 64-bit kernel address into urb->complete, and the kernel calls through it in kernel mode. On a vulnerable server, an attacker can divert kernel execution one time per exploit attempt. Turning that into a root shell, persistence, or ring-0 read/write that survives the call requires a follow-on gadget chain, typically something along the lines of commit_creds(prepare_kernel_cred(NULL)) to bump the calling task’s uid to 0 and return to userspace.

holy shit

At this point I had basically already shipped it in my head. The chain was landing reliably on 6.1, the demo was filmed, the article was almost done, the draft advisory was sitting in a notes file ready to go to linux-distros. I was doing the last fact-check pass, going through every detail one more time, when I realised I had not actually pinned the commit that introduced number_of_packets into the RET_SUBMIT wire format. You always want that line in an advisory (“present in mainline since commit X by Y, …”), it looks nice in the blog post. So I ran:

git log --follow drivers/usb/usbip/usbip_common.c | grep -i number_of_packets | head

I expected one match from 2011. But. I got five:

2ab833a16a82 2026-04-02 usbip: validate number_of_packets in usbip_pack_ret_submit()
1897852293fa 2026-03-25 usb: usbip: fix integer overflow in usbip_recv_iso()
74a2287209a8 2026-03-25 usb: usbip: fix OOB read/write in usbip_pad_iso()
591c1d972d8f 2026-03-25 usb: usbip: validate iso frame actual_length in usbip_recv_iso()
1325f85fa49f 2011-04-05 staging: usbip: bugfix add number of packets for isochronous frames

Ah shit.

The bottom one is the one I was initially looking for. April 5 2011, Arjan Mels, adding number_of_packets to RET_SUBMIT for Windows interop. Then I read the subject lines on the four commits above it.

  • fix OOB read/write in usbip_pad_iso. That is bug 1.
  • fix integer overflow in usbip_recv_iso. That is bug 2.
  • validate iso frame actual_length in usbip_recv_iso. That is the root cause of bug 1.
  • validate number_of_packets in usbip_pack_ret_submit. That is the root cause of bug 2.

I stared at the screen for a few seconds. Then I opened the patches and read them properly. Same bugs, exact same primitives. All four commits merged into mainline on March 25, two months before my first RIP-control demo against Debian 6.1.

The commit message on 74a2287209a8 even goes one further and points out a third primitive I had not noticed: iso_frame_desc[i].offset is also unchecked, giving a fully controlled OOB write into the slab independent of the np overflow (“confirmed with offset=400 on a 392-byte buffer, 64-byte write”). So the dev who fixed it found one more attack surface in the same area than my catalog did. Fair play.

I don’t know how they found it. There is no Reported-by: tag on any of the four commits, and the commit messages just describe the bugs. Could be a manual KASAN run, could be a code audit powered by Anthropic, could be something else. Doesn’t matter much in the end.

Respect, though. The game is the game, always.

I should probably add a The Wire gif here, right?;D

demo

Demo recording is here.

The demo runs the full chain against stock linux-image-6.1.0-47-amd64-unsigned from deb.debian.org, and the VM boots with KASLR on. We:

  • download the deb,
  • extract the modules,
  • load the USB stack,
  • attach vhci_hcd to the malicious server,
  • run the leak phase to recover the usbcore slide,
  • then run the OOB write phase until adjacency lands.

The proof is the kernel call trace showing the hijacked function executing inside __usb_hcd_giveback_urb. Once RIP control is achieved, I just call async_completed (via the slide we leaked) to prove the code is working. Then we dump_stack.

Once the fixes actually land in mainline, I might release the PoC code on github or something and update this article.

so what

I lost the race, basically. Fair play, it’s just how it works. A few takeaways though:

If a fifteen-year-old bug looks easy to fuzz with KASAN and syzkaller, it is on a dart-board, and has been for a while. You are not the only person looking (duh). Maintainers, fuzzing infrastructure operators and curious researchers all triage these daily. The gap between “a fuzzer or auditor reports a bug” and “someone writes the patch” is in the order of days. Run git log --follow first. I missed it by a few weeks; somebody else will miss it by less.

Also, on mitigations: SLAB_BUCKETS was the only thing keeping mainline Debian 13 from getting popped here. Not KASLR, not the freelist hardening, not control-flow integrity.

Finally, USB was designed around physical proximity. The security model is: if it is plugged in, you trust it. That made some kind of sense in 1996. The problem is that everything built on top of USB inherited that assumption, and USB/IP just puts a TCP socket where the cable used to be. Same trust model, except now the “device” can be anywhere and say anything. There is no authentication in the protocol because the original protocol never needed it. These bugs were a symptom of that.

Whatever, there will be other bugs. The hacking never stops.

See you next time :)~

djnn.sh

offensive security & software engineering


2026-05-19