A string encryption helper is a poor trade if calling it bluescreens the machine. Most obfuscation libraries assume user mode, and that assumption gets into everything. Decryption buffers use malloc. Templates pull in Standard Template Library containers. Global objects register destructors through atexit. Somewhere underneath it all is the expectation that the code will never run above interrupt request level 0.
Kernel drivers do not get those conveniences. They have no C runtime or C++ exception infrastructure, and a helper that touches paged memory at DISPATCH_LEVEL can bring down the system. I wanted string encryption, mixed Boolean arithmetic, and control flow flattening in ring 0, so I rebuilt the parts I needed around stack storage and compile-time keys. That became Kernelcloak.
The obfuscation was only half of the job. The other half was making sure that hiding a string or a branch did not quietly introduce an allocation, a missing runtime symbol, or an IRQL violation.
What I wanted to hide
A loaded driver gives an analyst plenty to work with. Its import table lists kernel API calls such as MmCopyVirtualMemory, KeStackAttachProcess, and IoCreateDevice. Often that is enough to work out its purpose in a few minutes. Plaintext literals in .rdata supply the device names, registry paths, and process targets. Give Hex-Rays ordinary structured control flow and, within an hour of opening the binary, the analyst may be reading something close to the original source.
Anti-cheats and endpoint detection and response products use the same visibility for automated checks. Easy Anti-Cheat and BattlEye scan loaded modules’ code sections against known signatures. If successive builds produce the same code, a signature taken from one build can keep identifying the others.
I wanted each build to look different enough that matching one sample would not settle the next. Compile-time diversity matters more here than making any individual expression look impressive. Manual analysis should also take longer, but I do not expect obfuscation to make it impossible. Someone patient enough still gets to read the instructions the CPU runs.
The kernel sets the rules
Interrupt request level, or IRQL, is the first constraint. PASSIVE_LEVEL is IRQL 0 and permits nearly everything. DISPATCH_LEVEL is where deferred procedure calls run and where code holding a spinlock executes. At that level, code cannot access paged memory, wait on dispatcher objects, or call most Zw/Nt system services. Touching a paged address can produce IRQL_NOT_LESS_OR_EQUAL. There is no useful cleanup path after a bugcheck.
User-mode string encryption usually decrypts into a heap buffer because malloc is available. The kernel has ExAllocatePool2, but a pool allocation is a heavyweight dependency for a helper that might run on a hot path. It also carries a tag visible in tools such as PoolMon. Removing a plaintext string and adding a visible allocation is an awkward bargain. I treat runtime allocation in an obfuscation helper as something to redesign.
The missing C runtime creates a separate problem. There is no malloc, free, printf, or std::string to fall back on, and no STL containers. A dependency buried several templates deep is still a dependency. The linker does not care that I only wanted the string encryption part of the library.
Kernelcloak’s core/types.h implements the small subset I needed, including is_same, enable_if, conditional, move, forward, and exchange. Where the library does need memory management, I use explicit ExAllocatePool2 and ExFreePoolWithTag calls with RAII scope guards. The obfuscation helpers work with fixed-size stack arrays.
Cleanup has to belong to a scope
A global encrypted string object is convenient in user mode. Its constructor decrypts the string and its destructor cleans up during DLL unload, with registration handled by atexit or the CRT’s onexit table. The kernel has no equivalent registration mechanism. Either the cleanup never happens, or the compiler emits references to runtime functions that do not exist and the build fails to link.
That ruled out helpers that combine static storage duration with automatic cleanup. A local object has a scope the compiler can see. A global object would need a runtime that the driver does not have.
C++ exceptions are out for the same reason. Kernel drivers use structured exception handling through __try and __except, rather than try and catch. MSVC’s /kernel flag disables the C++ exception infrastructure. A library that throws or catches will not compile under it. Many user-mode libraries fail on these constraints before there is any obfuscated code to inspect.
Decrypting strings onto the stack
String encryption was the easiest part to port. I kept the compile-time encryption and changed where the plaintext lives.
The encrypted bytes sit in a static constexpr array in the binary’s const data section. At the call site, a __forceinline lambda decrypts them into the current stack frame. The caller uses the result within that scope. There is no pool allocation, shared global state, or destructor registration to arrange.
Kernelcloak exposes this as KC_STR(). The XOR key comes from __COUNTER__ and __LINE__, giving each string instance its own encryption within a translation unit. Inlining removes the call overhead and keeps the decrypted value on the stack:
auto name = KC_STR("\\Device\\MyDriver");
IoCreateDevice(driver, 0, &name, FILE_DEVICE_UNKNOWN, 0, FALSE, &device);
// decrypted string lives on the stack - gone when this scope exits
The .rdata section contains ciphertext. The plaintext is confined to the stack and the enclosing scope’s lifetime.
I also added KC_STR_LAYERED() for cases where one XOR pass was not enough. It applies a rolling XOR pass, then XTEA over 8-byte segments with a 128-bit key for 32 rounds, half of the standard 64-round XTEA cycle count. A compile-time Fisher-Yates shuffle reorders the result. Six independent keys come from one __COUNTER__ seed.
KC_STR_LAYERED_HOLDER adds runtime re-keying. An interlocked counter tracks accesses, and after every N calls the holder re-encrypts the string with fresh keys seeded from rdtsc entropy. N defaults to 1000. Two memory dumps taken at different points can therefore contain different ciphertext for the same logical string.
At the other end is KC_STACK_STR(), which stores no ciphertext array at all. Each character gets its own template instantiation and holds its XOR’d value in volatile storage. The compiler emits mov instructions with obfuscated immediates instead of string data. That costs several instructions per character, so I consider it practical for short device names and registry paths. It would be an expensive way to hide a paragraph.
All three variants stay within the stack and registers, which makes them safe at any IRQL.
The optimizer would like its addition back
Mixed Boolean arithmetic, or MBA, replaces an ordinary operation with an equivalent expression. For example, a + b becomes (a ^ b) + 2 * (a & b), a - b becomes (a & ~b) - (~a & b), and a & b becomes ~(~a | ~b). The result is the same, but the disassembly no longer presents the original operation directly. Self-canceling noise adds and subtracts the same value through different algebraic paths to make the pattern harder to recognize.
This ports cleanly to kernel mode because it only needs arithmetic. Kernelcloak has three variants of each operation and picks one at compile time with (__COUNTER__ * 0x45D9F3Bu ^ __LINE__) % 3. Two calls to KC_ADD(a, b) at different source locations can produce different expansions. Recognizing one form does not automatically recover the others.
MSVC complicates this. Its optimizer can recognize that the elaborate expression is just a + b and helpfully turn the whole thing back into one add. From the compiler’s perspective, it has done a good job.
I use volatile intermediates and _ReadWriteBarrier() inside immediately invoked lambdas to keep those operations in the output. _ReadWriteBarrier is a compiler-only fence, so it blocks instruction reordering and constant folding without emitting a CPU memory barrier. Microsoft has deprecated it in favor of volatile under /volatile:iso or explicit barriers such as KeMemoryBarrier(), though it still works in current MSVC. The volatile values force stores and loads through memory that the optimizer cannot simply erase:
// without protection - compiler reduces this to add
int result = (a ^ b) + 2 * (a & b);
// KC_ADD - volatile barriers prevent constant folding
int result = KC_ADD(a, b);
KC_INT() does related work for stored values. It XORs integers and pointers with a compile-time key on write and decodes them on read. Operators including +=, -=, ++, --, and comparisons are overloaded so the source still reads normally. Each instantiation derives its key from (__COUNTER__ + 1) * 0x45D9F3Bu ^ __LINE__ * 0x1B873593u, giving instances separate keys.
The generated MBA operations use register arithmetic and stack-local volatile stores. There are no allocations or API calls, and no IRQL-sensitive dependencies hidden underneath the macros.
Making control flow harder to recover
Control flow graph flattening has the largest effect of the source-level techniques I use. A decompiler is good at recovering familiar branches and loops. Flattening removes the relationships that make that recovery straightforward.
Each basic block becomes a case in a switch inside a dispatch loop. A state variable records which block runs next, including the conditions that choose it. All blocks return to the same dispatcher, so the function’s original branching structure is encoded in state transitions.
This is the unflattened example:
void process_request(int type) {
if (type == 1) {
allocate_buffer();
validate_header();
} else if (type == 2) {
flush_queue();
}
send_response();
}
IDA can recover the two branches, the fallthrough, and the final call without much trouble. The flattened form does the same work:
void process_request(int type) {
unsigned int state = 0xA3F1;
while (state != 0) {
switch (state ^ 0x5E2D) {
case 0xFDDC:
state = (type == 1) ? 0x7B2C : (type == 2) ? 0x1D8E : 0x44F0;
break;
case 0x2501:
allocate_buffer();
validate_header();
state = 0x44F0;
break;
case 0x43A3:
flush_queue();
state = 0x44F0;
break;
case 0x1ADD:
send_response();
state = 0;
break;
}
}
}
The XOR between the state variable and the switch means the assigned state values do not directly match the case labels. Recovering execution order requires resolving that mapping for each transition.
Writing all of this by hand would make the source unpleasant to maintain, so I wrapped the state machine in Kernelcloak macros:
KC_FLAT_FUNC(process_request, int type) {
KC_FLAT_BLOCK(entry) {
KC_FLAT_IF(type == 1, handle_type1, check_type2);
}
KC_FLAT_BLOCK(check_type2) {
KC_FLAT_IF(type == 2, handle_type2, finish);
}
KC_FLAT_BLOCK(handle_type1) {
allocate_buffer();
validate_header();
KC_FLAT_GOTO(finish);
}
KC_FLAT_BLOCK(handle_type2) {
flush_queue();
KC_FLAT_GOTO(finish);
}
KC_FLAT_BLOCK(finish) {
send_response();
}
} KC_FLAT_END;
Block labels are hashed at compile time with keyed Fowler-Noll-Vo 1a, or FNV-1a, and a per-function seed. Transitions use an encryption key derived from __COUNTER__, unique to each function. The binary gets the hashed and encrypted values, never the label strings. Three dead code blocks before the default case add noise to the switch structure.
The split form, KC_FLAT_FUNC_HEAD with KC_FLAT_ENTER, exists for a less interesting but necessary reason. Local variables need to be declared before the dispatch loop. Separating the function signature from the entry point leaves somewhere to put them.
Keeping the dispatcher safe
The dispatcher itself is a while loop and a switch over a stack-local integer. Its conditional jump and indirect branch are safe at any IRQL. The surrounding storage is what needs attention.
State variables stay on the stack. Dispatcher internals need no pool allocations, globals, or data section entries. XOR keys are compile-time constants embedded as instruction immediates, appearing as operands in xor instructions rather than separate data.
Kernelcloak’s compile-time pseudorandom number generator uses __TIME__, __COUNTER__, and __LINE__. That gives translation units and builds different keys without runtime initialization. constexpr FNV-1a also finishes label hashing during compilation. There are no labels to process when the driver loads.
Depending on case density, MSVC usually turns the switch into a jump table or a comparison chain. Jump tables live in .rdata, which is non-paged by default in a kernel image. The I/O manager places sections without a PAGE prefix in non-paged memory.
Each state transition adds one XOR and one indirect jump. Executing a function with 10 blocks adds 10 of each. That can matter on a microsecond-sensitive hot path, even if it effectively disappears in ordinary driver logic. Being legal at DISPATCH_LEVEL does not make extra work free.
Dead code that stays in the binary
The optimizer also wants to remove code that does nothing. KC_JUNK() keeps its noise in place by writing the sentinel values 0xDEADC0DE and 0xBAADF00D to dummy volatile stack variables. The stores survive optimization and leave extra blocks for an analyst to separate from the real work. They still touch only the stack.
Opaque predicates add conditions derived from __rdtsc() and stack address entropy. (rdtsc | 1) & 1 is always true because the OR sets the bit that the AND tests. x * (x + 1) & 1 == 0 is always true because two consecutive integers have an even product. Their results are known at runtime but are harder to prove statically, particularly when the inputs come from volatile storage.
Reading the time stamp counter gives symbolic analysis another unknown to account for. These predicates and the dead stores work at any IRQL, which was the requirement for including them in the first place.
What the analyst gets
In IDA’s graph view, flattening is hard to miss. A tree of branches and convergence points turns into a radial web. One dispatcher connects to every block, and every block connects back. It advertises the use of flattening while hiding the structure of the function underneath it.
Hex-Rays struggles more with the result. It tries to recover if and else chains, while loops, and for iterations from the control flow graph. Given the dispatch loop, it produces a large while(true) with nested switch cases and a state variable threaded through XOR expressions it cannot simplify. More complex functions can produce no pseudocode at all. When output does appear, it has little useful resemblance to the original logic.
Binary Ninja’s Medium Level Intermediate Language, or MLIL, does somewhat better. It can sometimes propagate constants through the XOR operations and recover state values. Runtime-derived opaque predicates still leave symbolic unknowns, so the result is partial. The analyst gets some transitions back and has to resolve the rest.
Symbolic execution tools such as angr and Triton can trace the dispatcher and try to rebuild the state graph. They have to solve the XOR chain at each transition, model predicates involving rdtsc and stack entropy, and distinguish dead blocks from real ones. A small function with perhaps 5 to 8 blocks remains tractable. Larger functions make that recovery much more expensive. That was the tradeoff I wanted.
A motivated analyst can still recover the logic. The loop-and-switch pattern is recognizable, and IDA plugins such as D-810 target it directly. Per-function keys, state values, and label hashes mean a deobfuscation script has to handle those parameters instead of matching one fixed signature. The work scales with the number of obfuscated functions instead of ending after one solve.
Why I stopped at source-level transforms
Macros and constexpr templates have a ceiling. They can express only what the preprocessor and compile-time evaluation allow, and the optimizer gets the final decision about the emitted binary.
LLVM-pass obfuscators such as Obfuscator-LLVM, or O-LLVM, Hikari, and their descendants work on intermediate representation. They can transform every basic block in a module, mix bogus control flow with real flow, and substitute instructions beyond the reach of macros. My flattener only changes blocks explicitly wrapped in KC_FLAT_BLOCK. An LLVM pass can cover the whole module automatically.
That comes with a toolchain cost. O-LLVM needs a custom Clang build, while Windows kernel drivers normally use MSVC. Teams have made clang-cl work with Windows Driver Kit headers, but it is a non-standard setup with its own compatibility burden.
I chose a header-only library because it fits into an existing WDK build. Nobody has to adopt a compiler fork to use it. A more capable transform is not much use if the team never gets it through the build pipeline.
Where I apply it
Flattening belongs in selected functions. The extra XOR and indirect jump at every transition add up in DPC routines, interrupt service routine handlers, and filter callbacks that run for every I/O request. In IOCTL handlers, initialization routines, and other code that does not run thousands of times a second, that overhead is negligible. I flatten sensitive logic and leave performance-critical plumbing alone.
Cloakwork and Kernelcloak also include anti-debug and anti-VM checks, including KdDebuggerEnabled, the CPUID hypervisor bit, and hardware breakpoint detection through DR0 to DR3. Those are speed bumps. An analyst can NOP out a KdDebuggerEnabled check in thirty seconds. String encryption, MBA, and flattening do the more useful work because they change the structure of the binary the analyst has to understand.
Different builds need different seeds
Kernelcloak is the kernel-mode library and Cloakwork is its user-mode counterpart. Both are header-only and MIT-licensed, and fit the standard MSVC toolchain and WDK without modification.
Their compile-time PRNG depends on __TIME__, __COUNTER__, and __LINE__. That creates an awkward interaction with reproducible builds. Identical timestamps produce identical obfuscation output and identical signatures. The timestamp has to change between builds for diversity to work, which happens by default unless the pipeline has been configured for reproducibility.
Against automated signature matching, per-build keys and compile-time diversity are generally enough to keep one sample from becoming a lasting answer. Against a dedicated reverse engineer with a week and a debugger, the result is a delay. I am comfortable with that limit. The useful question is how much work the analysis takes, provided the driver still does its job without crashing.