← all posts
Intel CET
The CPU thought we were an attacker
hardware control-flow integrity
A shadow stack is a second copy of every return address, kept
by the hardware, that software is not allowed to write. It exists to make
return-oriented programming impossible. It also, it turns out, makes a naive
trampoline injector impossible — for exactly the same reason.
When a probe relocates a CALL out of the function it is
patching, the obvious implementation synthesises the call so the callee
returns straight back to native code and never sees the trampoline. We did
that on purpose, to stay invisible to a stack walker or Go's garbage
collector. It looked like this:
push rax
movabs rax, native_return
xchg rax, [rsp]
jmp qword [rip+0] ← the indirect forms end in `ret`
The push touches the normal stack only. The ret
then consumes a shadow-stack entry that belongs to someone else.
From the processor's point of view that is indistinguishable from an exploit,
so it does what it is designed to do: SEGV_CPERR, and not at the
probe — at the return of the function the probed one called.
It only bites a function that makes calls. That one fact
disguised the whole bug.
Here is the part that cost real time. Basic-block analysis treats
CALL as a terminator, so a leaf function's terminator is a
RET and nothing gets synthesised. Our multithreaded reproduction
happened to probe a non-leaf function; the single-threaded one happened to
probe a leaf. That produced a completely convincing story — "it only crashes
with threads" — and we chased concurrency for a while. Threads were never the
variable. Leaf-versus-non-leaf was.
The fix
When the shadow stack is armed, emit the real instruction and
jump back afterwards, so the hardware pushes both stacks and the callee's
ret matches. It costs the transparency we were protecting — the
callee now sees a trampoline return address — so it is applied only when CET
is actually armed, checked at runtime. Every other target keeps the fully
transparent path.
Verified on real silicon that enforces it — a bare-metal Sapphire Rapids
box, because a virtualized instance withholds CET entirely:
mt_shstk, 4 threads entry 4 records DEAD → 97,541 records ALIVE
real FRR bgpd, shadow stacks armed on 4 threads, BGP session up:
1 record then DEAD → 3 records, correct, session still up
This is the failure mode no bytecode instrumenter ever meets, and every
binary patcher eventually does. It is also the clearest example of why we run
on hardware that enforces the thing, rather than an emulator that
approximates it — an emulator arms from the binary's notes, not the running
CPU, and would have told us a comforting lie.
← all posts