18 min
Why anti-cheats walk your call stack
Hiding a module leaves the problem of explaining how its code got called. I look at what anti-cheats can learn from a stack and why spoofing one return address only goes so far.
A thread can reach NtReadVirtualMemory without its injected module appearing in the loader’s list. It can call GetAsyncKeyState from code that never existed on disk. Those are problems for module enumeration and memory scanning, but the calls still leave return addresses on the stack.
That is why I find the stack so useful. At the moment of a Windows API call, its x64 frames lead back to RtlUserThreadStart, and each return address gives me something to check. Which module owns it? Is that memory file-backed or privately allocated? Does a call instruction precede the return site? A plausible address has to survive all of those questions, then fit with the frames around it.
Manual mapping, direct syscalls, and kernel execution remove many familiar artifacts. Stack walking gives anti-cheats and endpoint detection and response platforms something else to inspect. BattlEye, Easy Anti-Cheat, and modern EDR products use it as a primary heuristic for injected, mapped, or otherwise illegitimate code. They can investigate an origin that should not exist without matching a payload against a signature database.
The implementation details here come mainly from secret.club’s 2020 reverse engineering of BattlEye’s shellcode and independent analyses of EAC’s kernel driver. I like comparing the two because they show where the difficulty moves. Checking one return address is straightforward. Capturing a stack that hostile code cannot interfere with, and deciding whether the whole chain makes sense, takes more work.
How Windows reconstructs a caller
On x86, unwinding follows frame pointers. Read the saved EBP, follow it to the preceding frame, and repeat. Windows x64 uses metadata. Every non-leaf function in a Portable Executable image, or PE image, has a RUNTIME_FUNCTION entry in .pdata describing its compiled stack layout. A leaf function can go without one if it never modifies non-volatile registers or touches RSP.
typedef struct _RUNTIME_FUNCTION {
DWORD BeginAddress; // RVA of function start
DWORD EndAddress; // RVA past function end
DWORD UnwindData; // RVA to UNWIND_INFO
} RUNTIME_FUNCTION;
UnwindData points to UNWIND_INFO, which records the operations in the function’s prologue. Each has a corresponding UNWIND_CODE entry.
UWOP_PUSH_NONVOLpushes a non-volatile register.UWOP_ALLOC_SMALLallocates up to 128 bytes of stack space.UWOP_ALLOC_LARGEhas two forms, 136 bytes to 512K-8 with op info 0, or 512K to 4 GB minus 8 with op info 1.UWOP_SET_FPREGestablishes a frame pointer register.
The unwinder reverses those operations to calculate the frame size and locate its return address. It does not need an RBP chain or have to execute the function to learn what its prologue did. The linker already recorded that work in .pdata.
Two functions do most of the traversal. RtlLookupFunctionEntry finds the RUNTIME_FUNCTION for an instruction pointer. RtlVirtualUnwind reads the unwind codes and computes the caller’s register context. When there is no entry, the unwinder treats the address as a leaf function and takes the return address directly from [RSP].
RtlCaptureContext(&context);
while (context.Rip != 0) {
PRUNTIME_FUNCTION fn = RtlLookupFunctionEntry(
context.Rip, &imageBase, NULL
);
if (!fn) {
// leaf function - return address is at [RSP]
context.Rip = *(ULONG64*)context.Rsp;
context.Rsp += 8;
} else {
RtlVirtualUnwind(
UNW_FLAG_NHANDLER, imageBase,
context.Rip, fn, &context,
&handlerData, &establisherFrame, NULL
);
}
capturedStack[depth++] = context.Rip;
}
The convenience wrapper, RtlCaptureStackBackTrace, is enough when all I want is a trace. Anti-cheats also use the full unwind path because frame sizes and module ownership matter to their checks. A list of addresses tells only part of the story.
Does this frame belong here?
The first check is whether a return address points into memory that Windows loaded as an image. That memory has type MEM_IMAGE and corresponds to a file on disk. VirtualAlloc produces MEM_PRIVATE memory, while MapViewOfFile produces MEM_MAPPED memory. A PE manually mapped into a VirtualAlloc allocation therefore leaves return addresses on MEM_PRIVATE pages.
For injected payloads, shellcode, and manually mapped modules using that memory, this check is an immediate problem. It also explains the appeal of module stomping. Overwriting the .text section of a legitimately loaded Dynamic Link Library lets the attacker’s addresses resolve to a real, file-backed module. The address passes this test because the code occupies the module’s memory.
A frame also has to lead somewhere. A well-formed user mode stack unwinds to RtlUserThreadStart or BaseThreadInitThunk. If the unwinder cannot parse a frame and stops early, that points to corrupted data, a pivot into attacker-controlled memory, or synthetic frames that do not join up correctly. Repairing the immediate return address does little good when the rest of the chain falls apart.
A real address can still be a fake return site
Return address spoofing often borrows a Jump-Oriented Programming gadget, or JOP gadget, from a real module. A common version uses jmp qword ptr [rbx] as the return address. After the monitored function returns, the gadget redirects execution through a register to the caller’s fixup code.
BattlEye recognizes that particular choice. Its exception handler inspects the opcode at the return address for FF 23.
const auto spoof = *(_WORD *)caller_function == 0x23FF; // jmp qword ptr [rbx]
Another register or instruction can avoid that signature. The more interesting check looks immediately before the return address for a call. A legitimate return site follows the instruction that put it on the stack. A gadget selected because it redirects control usually lacks that relationship, so checking the preceding instruction catches most basic spoofing without needing a signature for every possible gadget.
Even convincing individual frames can describe an impossible sequence. A trace showing KERNELBASE!PathReplaceGreedy calling KERNELBASE!SystemTimeToTzSpecificLocalTimeEx uses real functions but claims a path that legitimate code never takes. Correct frame sizes and real module addresses do not make those functions call one another.
This is the expensive part. A defender needs a model of which functions can call which others, and someone has to maintain it. Most anti-cheats do not perform that analysis at the time of writing. Elastic’s EDR does. Anti-cheat vendors tend to adopt EDR techniques a few years later, so I would not assume that checking frame geometry will remain enough.
The stack has boundaries too
The Thread Environment Block, or TEB, records each thread’s stack boundaries in StackBase and StackLimit. An RSP outside those bounds indicates a stack pivot, commonly into a heap allocation or global data holding a Return-Oriented Programming chain, or ROP chain.
PTEB teb = NtCurrentTeb();
if (rsp < teb->NtTib.StackLimit || rsp > teb->NtTib.StackBase) {
// stack pivot detected
}
Windows Defender Exploit Guard hooks VirtualAlloc and VirtualProtect so it can run this check on each call. There is very little work involved, which makes it an appealing check to repeat. A ROP chain has to stay on the thread’s real stack to pass it, and that constraint cuts down the attacker’s options.
Capturing a stack from the kernel
The checks so far can run in user mode. A kernel driver can make acquiring the evidence much harder to interfere with, which is where I think stack-based detection becomes particularly difficult to evade.
One approach queues an Asynchronous Procedure Call, or APC, to the target thread. The APC runs in that thread’s context and can access its entire stack. Calling RtlCaptureStackBackTrace there captures the chain at that moment. Both BattlEye and EAC use APC-based capture.
There is a catch in the scheduling. A thread running with kernel APCs disabled makes a queued APC wait until it enables them again. The capture is useful when it runs, but it still depends on the target reaching a state where the APC can execute. Cheat developers tend to find that limitation quickly.
Non-Maskable Interrupts, or NMIs, remove that dependency. They exist for signals software must not be able to block. Anti-cheats register a callback with KeRegisterNmiCallback and periodically send NMIs to every CPU core. When one arrives, the callback captures the interrupted processor state, including RIP, RSP, and the full register context. An interrupted RIP outside every loaded kernel module points to unsigned code executing there.
This is sampling, so one interrupt can miss the code of interest. Repeating inexpensive samples across an entire match changes the odds for a thread that spends meaningful time in unsigned memory. Disabling APCs does not stop NMIs, and the usual flags or callback manipulation do not prevent them from firing. Moving execution into the kernel does not make it exempt from observation.
Checking the return from a syscall
EAC also watches the point where a syscall returns to user mode. Its driver sets the InstrumentationCallback field in EPROCESS through NtSetInformationProcess. The kernel checks that field on return and redirects execution to the callback when it is set, placing the original return address in R10.
// in the instrumentation callback
PVOID returnAddr = (PVOID)teb->InstrumentationCallbackPreviousPc; // R10
if (!isInModule(returnAddr, "ntdll.dll") &&
!isInModule(returnAddr, "win32u.dll")) {
// direct syscall detected
}
The callback checks whether that address belongs to ntdll.dll or win32u.dll, the modules expected to contain the syscall instruction on a legitimate path. A direct syscall from another module fails that check.
I find this useful because the monitor does not need an inline hook on an ntdll function or a modification to the System Service Descriptor Table, or SSDT. The kernel performs the redirection after every syscall return. User mode code has very little room to interfere with a decision made at that privilege level.
Sensitive operations with a stack attached
The Event Tracing for Windows Threat Intelligence provider, or ETW TI provider, exposes kernel level events for executable memory allocation, cross-process writes, thread context changes, and APC insertion. Its GUID is {09997CFF-4065-4844-BE7E-2A2DEEAD0D02}. Access is restricted to the Protected Process Light level, or PPL, with Antimalware PPL or Early Launch Anti-Malware signed drivers able to consume its events.
Enabling EVENT_ENABLE_PROPERTY_STACK_TRACE attaches a full call stack captured through RtlWalkFrameChain. An event for NtWriteVirtualMemory, NtProtectVirtualMemory, or NtCreateThreadEx then includes the path that invoked it. The defender gets both the sensitive operation and the context needed to judge its origin.
Elastic’s EDR uses this to attach more than 20 behavioral indicators to ETW events. Those include direct syscalls with no ntdll frame, ROP gadgets in the chain, executable memory without module backing, and proxy calls through unexpected modules. Anti-cheats can use the same infrastructure. Their implementations are less publicly documented, so the mechanism is easier to describe than any one vendor’s coverage.
Where BattlEye puts its breakpoints
secret.club’s analysis makes BattlEye’s user mode implementation unusually visible. During competitive matches, BattlEye streams a shellcode payload of roughly 8 KB into the game process. secret.club calls it “shellcode8kb.” It installs a Vectored Exception Handler, or VEH, then places INT3 breakpoints on these functions.
GetAsyncKeyState, GetCursorPos, IsBadReadPtr, NtUserGetAsyncKeyState, GetForegroundWindow, CallWindowProcW, NtUserPeekMessage, NtSetEvent, sqrtf, __stdio_common_vsprintf_s, CDXGIFactory::TakeLock, TppTimerpExecuteCallback
When a breakpoint fires, the VEH reads the caller address from the top of the stack and checks four conditions.
NtQueryVirtualMemoryfails for the caller address, suggesting a hidden region and a hook.- The memory is uncommitted, pointing to Virtual Address Descriptor manipulation.
- The memory is executable but has a type other than
MEM_IMAGE, pointing to injected or manually mapped code. - The return address contains
FF 23, the encoding of namazso’sjmp qword ptr [rbx]gadget.
Any match produces report ID 0x31, sent to BattlEye’s servers with 32 bytes of the caller’s code and the associated memory metadata.
The function list is my favorite part of this implementation. Input calls are an obvious place to watch, but cheats also use ordinary math and string formatting. That is why sqrtf and __stdio_common_vsprintf_s appear alongside the mouse and keyboard functions. CDXGIFactory::TakeLock covers interaction with DirectX. A cheat still has to do mundane work, and BattlEye has put tripwires where that work is likely to take it.
The implementation has mundane failures too. secret.club found that the signature locating CDXGIFactory::TakeLock included CC padding bytes from one particular compilation. Different builds break it. A good choice of place to monitor still depends on reliably finding the function.
EAC does not need the thread to cooperate
EAC’s kernel driver goes beyond queuing APCs. It copies raw kernel thread stacks asynchronously, then scans those copies for code pointers into non-paged memory that belongs to no loaded kernel module. It reads the stack through the thread’s KTHREAD structure and examines the copied data offline.
That avoids the scheduling limitation of an APC. The thread can disable APCs or stay non-alertable, but its stack is still memory the driver can read. The scan can also find saved return addresses left behind by functions that have already returned. Finishing the call does not necessarily remove its evidence.
EAC combines this copying with instrumentation callbacks for user mode syscall origins and proper unwinding through RtlLookupFunctionEntry and RtlVirtualUnwind. Together they cover user and kernel transitions. That breadth is why I consider it the most thorough stack-based detection implementation in the anti-cheat space.
How far spoofing gets
A basic return address spoof replaces the monitored call’s return address with a JOP gadget in a legitimate module and saves the real address in a non-volatile register. On return, the gadget redirects through that register to the real caller. The fake address belongs to ntdll.dll or kernel32.dll, so module backing validation accepts it.
The other checks still apply. An FF 23 gadget matches BattlEye’s signature, and a gadget without a preceding call fails the structural check. namazso’s original work shows that choosing another gadget can avoid BattlEye’s particular signature. It does not repair the missing relationship between the claimed return address and a real call.
klezVirus’s SilentMoonwalk takes on the larger problem by constructing whole frames from .pdata metadata. Synthetic mode sizes artificial frames using each function’s UNWIND_INFO, letting RtlVirtualUnwind produce a coherent trace through real functions in real modules. Desync mode replaces problematic frames in the existing stack with legitimate ones. It looks for UWOP_PUSH_NONVOL entries that restore RBP during unwinding so the frame pointer chain stays intact.
Both modes use ROP to restore the original stack after the monitored call. During the call, the trace contains legitimate modules and plausible function sequences. SilentMoonwalk is the most complete public user mode stack spoofing implementation available, and it shows how much more work is involved once a monitor looks past the first frame.
I still put weight on the word “plausible.” A synthetic trace can have matching frame sizes and valid module addresses while describing a sequence no real execution would produce. A defender maintaining legitimate call graphs can test that sequence, even when the unwind succeeds.
Changing the code’s location
Module stomping approaches the backing check by changing where the code lives. It loads a legitimate DLL and overwrites its .text section with the payload. The attacker’s return addresses then resolve to that module because the code physically occupies its pages.
A memory-to-disk comparison exposes the replacement. PE-sieve compares the bytes and flags a module when more than roughly 10,000 bytes differ in its code section, since legitimate hooks change far fewer. Some implementations restore the original bytes after execution. The write still changes the working set entry’s SharedOriginal flag, which tools such as Moneta inspect. Restoring code bytes does not restore every property of the page.
Indirect syscalls make a narrower change. They jump to the syscall instruction inside ntdll.dll so the kernel sees a return address in ntdll instead of in the attacker’s own code. That passes the instrumentation callback’s module check.
It also skips the inline hooks that EDR products place at ntdll function entries. A monitor can check whether execution entered through the real prologue, beyond checking that the eventual syscall address falls inside ntdll. Passing the return-address check leaves that question unanswered.
Why I expect the stack to keep mattering
The difficulty for evasion is having to satisfy all of these checks throughout a session. A defender can find a gadget pattern in one call, an implausible sequence in another, or a changed page after module stomping. Direct syscalls can leave non-ntdll return addresses in KTRAP_FRAME. A single anomaly anywhere across the frames and threads is enough to give the defender something to investigate.
The attacker has to keep every inspected path consistent, including at moments chosen by NMI sampling, instrumentation callbacks, and ETW TI capture. User mode code cannot reach the privilege level where those checks run.
Intel Control-flow Enforcement Technology, or CET, adds hardware enforcement to the same problem. Its shadow stacks, available since Windows 11, hold a copy of return addresses that user mode software cannot modify. As CET-capable hardware becomes more widely deployed, software-only spoofing has less room to rewrite the history of a call. The stack was already a useful record. Giving it a hardware-backed copy makes that record harder to argue with.