RESEARCH / WASM SECURITY
When the Wasm boundary becomes a native stack boundary.
WASMageddon is the chosen name of a Wasmer Singlepass compiler flaw that turns a valid, attacker-influenced WebAssembly execution into a host-process control-flow primitive. The finding was disclosed to affected vendors in Q1 2026.
The technical record below stays close to the vendor report. It follows the complete chain: a result-bearing if/else cleanup bug, cumulative native-stack drift, Linux mmap geometry, Wasmer coroutine-stack reuse, allocator grooming, a scheduling race, return-slot corruption, ROP pivoting, signal-wrapped process stability, and two independent delivery surfaces.
This research article explains the root cause and exploitation techniques, but in certain parts deliberately omits all the information needed for weaponizing the exploit.
COORDINATED DISCLOSURE
Disclosure chronology
This chronology records the coordination sequence and distinguishes the upstream runtime disclosure from downstream remediation of related or differently demonstrated vectors.
Had been notified before this research was complete. After that report, the issue was fixed through a different vector; at that point, we had not yet found the RCE vector described in this research.
A coordinated response channel was opened for the Wasmer upstream runtime and affected downstream integrations.
Wasmer, ecosystems and certain direct downstreams were notified and entered coordinated disclosure.
Wasmer publicly released its remediation for the upstream runtime.
Certain directly affected downstream projects completed their fixes following the coordinated disclosure.
Remaining downstream communications were managed by the relevant upstream maintainers and holders.
00 / TECHNICAL RECORD
Summary
Affected upstream range:
Wasmer >= 2.1.0 and < 7.0.0
That is Wasmer v2.1.0 through the v6.x line; Wasmer v7.0.0 and later are not included in this affected range.
Wasmer Singlepass miscompiles a result-bearing Wasm if/else cleanup path. The compiler releases memory-backed value-stack locations and then adjusts the native stack in the wrong direction. Inside a loop this becomes a deterministic downward drift of the host rsp inside the Wasmer coroutine stack.
In embedders that execute attacker-controlled Wasm with Singlepass and reusable coroutine/fiber stacks, this compiler bug can become a sandbox escape and host process RCE. The exploit primitive is a native stack write stream: generated Singlepass code repeatedly writes attacker-chosen qwords while rsp drifts downward, crosses a guard page, and corrupts a live return path on a neighboring coroutine stack.
This upstream report intentionally does not include target-specific PoC bundles or reproduction commands. High-profile downstream targets have been contacted under coordinated disclosure, and full end-to-end RCE PoCs have been presented to the affected parties for their own environments.
INDEX / TECHNICAL RECORD
Technical record
- 0. Affected component
- 1. The compiler bug
- 2. Turning the bug into a qword write stream
- 3. Coroutine stack layout
- 4. Holder/trigger choreography
- 5. Allocator grooming
- 6. The race condition
- 7. Warmup and staggered execution
- 8. Calibrating the overwrite
- 10. Control-flow impact
- 11. ASLR notes and the two-stage design
- 12. Signal wrapping and post-write stability
- 13. Downstream impact and coordinated disclosure
- 14. Results
- 15. Mitigation
00 / TECHNICAL RECORD
1. Affected component
Affected upstream component:
file: codegen.rs
range: Wasmer >= 2.1.0 and < 7.0.0
The issue is in Singlepass native code generation. It is independent of any particular downstream product, host application, or guest framework. Downstream exploitability depends on how an embedder exposes Wasm execution, how coroutine stacks are allocated and reused, and whether attacker-controlled Wasm can create enough concurrent/suspended executions to shape stack adjacency.
01 / TECHNICAL RECORD
2. The compiler bug
The vulnerable code is in the Wasmer Singlepass compiler:
wasmer/lib/compiler-singlepass/src/codegen.rs
The normal non-value release path restores native stack space:
let delta_stack_offset = self.machine.round_stack_adjust(delta_stack_offset);
if delta_stack_offset != 0 {
self.machine.restore_stack(delta_stack_offset as u32)?;
}
The value release path performs the opposite operation:
let delta_stack_offset = self.machine.round_stack_adjust(delta_stack_offset);
if delta_stack_offset != 0 {
self.machine.adjust_stack(delta_stack_offset as u32)?;
}
That value path is reached from the Operator::Else code path:
let stack_depth = frame.value_stack_depth;
self.release_locations_value(stack_depth)?;
At this point, Singlepass has removed memory-backed value-stack locations from its model. The native stack should be restored. Instead, it is adjusted downward. On x86_64, this is equivalent to moving rsp deeper into the stack.
Minimal exploit-relevant shape:
loop:
push controlled i64 values
enter if (result i64)
take then branch
reach else-transition cleanup with live result
release_locations_value()
-> BUG: native stack adjusted downward
drop values
call a host import
repeat
ASCII control-flow sketch:
Wasm if (result i64)
|
v
Operator::Else
|
v
release_locations_value(stack_depth)
|
v
memory-backed value locations removed
|
v
round_stack_adjust(delta)
|
v
adjust_stack(delta) restore_stack(delta)
(buggy value path) (correct sibling path)
| |
v v
RSP moves downward RSP restored
|
v
loop repeats
|
v
deterministic native stack drift
The observed trigger shape drifts by 16 bytes per iteration. The granularity is a property of the exact generated code shape. Changing the import, adding preamble stores, or changing local pressure can alter the phase.
This is native stack corruption caused by code generation. It is not a Wasm linear-memory bounds error and not an application-level Wasm bug.
INTERACTIVE FIGURE / COMPILER
One wrong instruction becomes a write stream
Follow the exact transition described in the report. The compiler drops the value locations, then the vulnerable path subtracts native stack space again instead of restoring it.
The logical model and native stack begin at the same level.
02 / TECHNICAL RECORD
3. Turning the bug into a qword write stream
The exploit primitive does not only move rsp. It writes controlled qwords while rsp moves. In a vulnerable embedder, those qwords can be selected so that the first-stage overwrite changes a live return path on another coroutine stack.
The trigger loop must be emitted so that:
ATTACK_VALUE0 = attacker-chosen qword
ATTACK_VALUE1 = attacker-chosen qword
...
ATTACK_VALUE7 = attacker-chosen qword
Conceptually:
iteration i:
generated native code stores/pushes:
attacker-chosen qword
attacker-chosen qword
attacker-chosen qword
...
at the current native stack location.
release_locations_value() then moves RSP downward again.
After enough iterations, the qword stream exits the trigger stack, crosses the guard page, and lands in the holder stack.
Approximate landing model:
landing(i) = trigger_stack_hi - base_frame - phase - (16 * i)
The exact constants are target- and build-dependent. They are measured for a given embedder and runtime, not guessed.
03 / TECHNICAL RECORD
4. Coroutine stack layout
Wasmer/corosensei executions use separate fiber stacks. The useful observed shape is a 1 MB writable stack with a guard page. Native stacks grow from high addresses to low addresses.
Required adjacency:
higher addresses
+------------------------------------------+
| trigger fiber stack |
| active Singlepass drift loop |
| qword stream = pivot gadget |
| |
| RSP drifts downward |
| | |
| v |
+------------------------------------------+
| guard page |
| ---p |
+------------------------------------------+
| holder fiber stack |
| suspended host import frame |
| live return slot |
| holder locals containing ROP chain |
+------------------------------------------+
lower addresses
The trigger must be above the holder. If the trigger is below the holder, or if there is no adjacent lower holder stack, the drift cannot hit the intended return slot.
Representative target slot:
holder_stack_hi - target_specific_offset
Holder stack, simplified:
holder_stack_hi
|
| top runtime setup / host frames
|
| exact-path post-write tail surface
|
| holder local block
| ROP qwords
| path/data pointers
| stack-overflow tail constants
|
| live return slot <-- target-specific offset
|
| older/stale frame data
|
holder_stack_lo
Representative model:
offset from holder stack_hi contents role
--------------------------- -------- ----
hi - 0x000 .. hi - 0x200 setup / host frames should survive
hi - 0x2d0 approx exact-path tail surface wrap target
hi - 0x300 .. hi - 0x650 holder locals ROP stack
hi - target_offset live return slot overwrite
below target_offset older/stale frames avoid early poison
The holder locals are not created by the bug. They are constants from the holder guest module. The bug only connects the holder's return path to those locals.
04 / TECHNICAL RECORD
5. Holder/trigger choreography
The exploitability demonstration uses two guest-execution roles:
holder:
creates the future ROP stack in its locals
enters a host import / host-call path
suspends on a fiber stack
eventually resumes through a return slot
trigger:
runs the Singlepass drift loop
writes pivot qwords while RSP drifts
must run on the adjacent upper fiber stack
Timeline:
time ->
holder:
init locals
|
v
call host import/query
|
v
suspended on fiber stack ............................
|
v
resumes from import
|
v
returns via ret slot
|
v
pivot into locals
trigger:
allocated on upper adjacent stack
|
v
run else-drift loop
|
v
qword stream crosses guard
|
v
overwrite holder ret slot
The holder must still be suspended when the trigger crosses into it. If the holder returns too early, the target frame is gone.
The more normal concurrent Wasm activity an embedder permits, the more opportunities an attacker has to win useful stack adjacency. Organic load can therefore improve exploit reliability rather than only acting as noise.
05 / TECHNICAL RECORD
6. Allocator grooming
The most important exploit engineering step is not the ROP chain. It is making the runtime give us the right two fiber stacks at the right time.
The desired stack pair:
[ lower holder stack ][ guard ][ upper trigger stack ]
The desired state:
lower holder:
suspended
has live host-return frame
has ROP constants in locals
upper trigger:
active
running drift loop
qword stream contains pivot
The runtime will not hand this pair to us deterministically. The exploit abuses normal allocator and scheduler behavior:
- Wasmer creates/reuses coroutine stacks.
- Concurrent guest execution puts pressure on the stack pool.
- Host imports suspend some executions while others continue.
- Warmup calls create a population of live and recently-used stacks.
- Staggered phases bias which stacks are likely to be adjacent at final trigger time.
This is allocator grooming in the old-school sense: we are not inventing a fake memory layout; we are forcing the real allocator to walk into a useful one.
Stack-pool intuition:
Before grooming:
stack pool / mmap population:
S0 S1 S2 S3 S4 S5 S6 S7
? ? ? ? ? ? ? ?
After holder warmup:
S0 S1 S2 S3 S4 S5 S6 S7
H H H H H H H H
| | | | | | | |
suspended holders occupy many candidate lower stacks
During trigger wave:
S0 S1 S2 S3 S4 S5 S6 S7
H H H H H H H H
T T T T T T T
Some trigger stacks land directly above suspended holder stacks.
The downstream end-to-end demonstrations use staggered warmup to create enough stack pressure while keeping the final holder/trigger shape aligned. The exact counts and timings are downstream-specific and are intentionally omitted from this upstream report.
INTERACTIVE FIGURE / MMAP + ALLOCATOR
Race to shape, then make it useful for adjacency
Warmup keeps holder-shaped executions alive, the trigger wave reuses the same stack population, and the final race matters only when the allocator places the active trigger above a suspended holder.
H = holder-shaped execution, T = trigger-shaped execution. The highlighted pair is a useful placement, not a deterministic slot index.
06 / TECHNICAL RECORD
7. The race condition
The exploit is a race over three events:
Event A: holder is suspended with live return frame
Event B: trigger is scheduled on adjacent upper stack
Event C: trigger drift reaches holder ret before holder frame disappears
The right interleaving:
CPU / worker scheduling stack state
----------------------- -----------
1. holder starts
2. holder initializes locals H has ROP locals
3. holder calls host import H suspends
4. runtime schedules trigger T obtains upper stack
5. trigger runs drift loop T writes down
6. qwords cross guard H ret overwritten
7. trigger returns/traps host resumes H
8. holder returns through ret pivot fires
Exploitability depends on the embedder scheduler, host-call behavior, and degree of concurrent Wasm execution. In the downstream targets tested under coordinated disclosure, this race was reachable with ordinary external inputs.
07 / TECHNICAL RECORD
8. Warmup and staggered execution
The important exploitability concepts:
armed holders:
warmup holders use the same relevant holder shape. This prevents the
warmup from grooming a stack population that disappears when the real
holder shape is introduced.
phase counts:
a large but tapering population of stack users can preserve stack pressure
near the final trigger without simply overdriving the host.
waves:
work is released in groups that match the empirically useful scheduling
width for the target. Bigger waves are not automatically better; they can
destroy stability.
depth:
forces nested execution enough to exercise the coroutine stack pool
without making every warmup a final-trigger-scale event.
barrier:
waits for warmup phases to shape the allocator before the final
holder/trigger sequence.
nested execution:
keeps the final stack population close to the holder/trigger topology
rather than introducing unrelated flat traffic.
phase sketch:
Phase P1: broad task population
workers/stacks:
H H H H H H H H | H H H H H H H H | ... many holders
effect:
create broad stack population and initial pool pressure
Phase P2: reuse and reshuffle
H H H H H H H H | H H H H H H H H | ...
effect:
reuse and reshuffle stacks after P1, keep holders live enough to matter
Phase P3: adjacency density
H H H H H H H H | H H H H H H H H | ...
effect:
keep adjacency density while reducing total churn
Phase P4: final pre-trigger population
H H H H H H H H | H H H H H H H H
effect:
final pre-trigger population, less likely to overdrive than a huge wave
Barrier:
wait until P4 has completed the intended allocator disturbance
Final holder/trigger graph:
fire final trigger against the shaped stack pool
Downstream end-to-end validation used bounded attempts because the race remains probabilistic even after grooming. The goal of grooming is not 100 percent deterministic placement; the goal is a reproducible probability of control-flow transfer while keeping the target process healthy enough to observe the result.
08 / TECHNICAL RECORD
9. Calibrating the overwrite
The primitive is deterministic at the code-generation level, but the useful landing point is target-specific. The embedder, compiler version, build flags, host-call path, and coroutine stack layout all affect:
drift iteration count
qword stream phase
live return-slot offset
safe pre-ret corruption window
post-pivot recovery path
The calibration question is never only:
where is the ret slot?
It is:
where does the qword stream land,
which qword in the stream reaches ret,
and what did it corrupt before ret?
So if many qwords above the ret slot are corrupted first, the holder may die before executing the pivot. A stable exploit therefore puts the pivot on the ret slot as early as possible in the useful overwrite window.
10 / TECHNICAL RECORD
10. Control-flow impact
The cross-stack overwrite is narrow. It does not try to lay down a full payload with the drifting write stream. It changes a live return path so that execution continues through attacker-controlled data already staged in the suspended holder frame.
High-level control transfer:
holder ret slot
|
v
attacker-selected first-stage control value
|
v
RSP moves into holder local block
|
v
attacker-controlled native control-flow chain staged in holder locals
The cross-stack overwrite is narrow. It does not try to lay down the full ROP chain. It only changes one return slot. The holder, before suspension, already created the long chain in its own local frame.
Downstream demonstrations used benign file-write markers to prove host-process code execution. The same primitive is not inherently limited to a file write; payload size and shape are determined by the staged local-frame data and the available native gadgets or dynamically computed addresses in the target process.
Control transfer sketch:
before overwrite:
holder stack:
[ live ret slot ] -> legitimate Wasmer/libwasm return
[ locals ] -> ROP constants, inert until pivoted to
after overwrite:
[ live ret slot ] -> attacker-selected first-stage control value
[ locals ] -> ROP constants
resume:
ret
-> attacker-selected first-stage control value
-> adjust RSP into locals
-> native control-flow chain
11 / TECHNICAL RECORD
11. ASLR notes and the two-stage design
Some downstream demonstrations used fixed addresses because the target release binary exposed useful non-PIE surfaces. That should not be confused with a fundamental limitation of the bug.
This section deliberately does not include all the techniques for ASLR/PIE bypass, just the general idea. End-to-end PoC was achieved internally and proven.
JIT code pointers
embedder / runtime text pointers
heap pointers
coroutine guard addresses
stack-top derived addresses
Representative derivations from that class of work:
jit_base = leaked_jit_pointer - target_delta
runtime_text_base = leaked_text_pointer - target_delta
stack_guard = leaked guard page
stack_top = stack_guard + stack_mapping_size
The Stage 2 design relays computed qwords back to the trigger through the embedder's ordinary guest-to-guest or host-mediated call surface:
holder leaks addresses
|
v
holder computes dynamic gadget/stack values
|
v
holder sends qwords to trigger
|
v
trigger loads those qwords from linear memory
|
v
trigger uses them as the drift write stream
INTERACTIVE FIGURE / ASLR + PIE
ASLR moves targets, not the write primitive
The compiler bug still produces the same downward qword stream. The difference is how the stream gets useful addresses: a non-PIE proof can use known targets, while an ASLR-aware chain first leaks pointers, derives bases, and relays computed qwords to the trigger.
base = leak − delta; stack_top = guard + sizeMeasured stock proof: fixed non-PIE addresses make the target values known. ASLR does not remove the compiler primitive.
12 / TECHNICAL RECORD
12. Signal wrapping and post-write stability
The hard part after control transfer is not only demonstrating native execution. It is leaving the runtime in a state where the embedder receives a normal Wasmer trap/error instead of dying on corrupted native control flow.
The relevant Wasmer behavior is the signal wrapper for two broad cases: out-of-bounds Wasm linear-memory access and guard-page stack exhaustion.
Wasmer's signal handler only successfuly wraps and recovers if the current stack pointer is inside the active coroutine stack:
if !self.coro_trap_handler.stack_ptr_in_bounds(sp) {
return false;
}
For signal faults, classification depends on the fault address:
fault address inside same coroutine stack -> TrapCode::StackOverflow
fault address outside same coroutine stack -> TrapCode::HeapAccessOutOfBounds
The stability tail tries to create this situation:
active fiber stack
high
|
| RSP after payload marker / post-control work
| |
| v
| tail moves RSP down
|
| ...
|
| lower guard page <-- deliberate fault target
|
low
Expected exact path:
SP still belongs to active fiber stack
fault classifies as stack overflow
Wasmer returns wrapped runtime error
process survives
Miss-path danger:
active fiber stack
high
|
| exact-path RSP
|
| miss-path RSP is already much lower
| |
| v
| same tail subtracts too much
|
| lower guard page
|
+--------------------------
| below guard / unrelated memory
Result:
signal may escape Wasmer wrapping
process may die before or after the observable marker
The stable downstream PoCs used target-specific post-control tails so the embedder remained healthy after the observable marker was written. Those target-specific constants are omitted here.
13 / TECHNICAL RECORD
13. Downstream impact and coordinated disclosure
We validated the compiler bug against high-profile downstream targets that embed Wasmer Singlepass and expose attacker-controlled Wasm execution. Those downstream reports contain target-specific exploit delivery, calibration, and end-to-end RCE PoCs.
1. Singlepass emits native code that moves RSP in the wrong direction.
2. The drift can be repeated inside a loop.
3. The generated code can write attacker-chosen qwords while RSP drifts.
4. In embedders with reusable adjacent coroutine stacks, those qwords can cross
a guard page and corrupt a live return path on another coroutine stack.
5. Downstream end-to-end RCE has been demonstrated and shared with affected
downstream vendors under coordinated disclosure.
The downstream demonstrations covered distinct operational surfaces. This is important because the compiler bug is upstream, while delivery and reliability depend on each embedder's scheduler, host-call surface, binary hardening, and deployment model.
INTERACTIVE FIGURE / DELIVERY
Two ways the payload reaches a node
Same bug, same malicious Wasm, two delivery routes. The only thing that changes is how many machines end up running it.
Step 1 of 4: in both routes the attacker only has to send something the chain normally accepts.
Step 2 of 4: this is where the routes split. A simulation stays on the one RPC node that received it. A block is passed to every node, including validators behind NAT, because they fetch blocks over connections they opened themselves.
Step 3 of 4: wherever the input lands, the node executes the malicious Wasm and the compiler bug fires.
Step 4 of 4: RPC simulation compromises one node. Block propagation compromises every node that processes the block.
14 / TECHNICAL RECORD
14. Results
The upstream vulnerability has been validated at three levels:
compiler bug:
Singlepass uses adjust_stack() where the value cleanup path must restore
stack space.
native primitive:
repeated result-bearing if/else cleanup creates deterministic downward
native stack drift with attacker-chosen qword writes.
downstream impact:
full end-to-end host-process RCE PoCs were produced for high-profile
downstream targets and presented to those affected parties.
PUBLIC VALIDATION BOUNDARY
How impact was validated
The research was validated on fresh Linux/amd64 environments using official, unmodified downstream containers. The validation harness exercised the two separate delivery surfaces described above and recorded three independent signals: whether the harmless proof marker was written, whether the node remained healthy, and whether the execution log matched the expected runtime path.
The public article intentionally does not publish contract emitters, Docker launch scripts, calibrated stack offsets, exact gadget addresses or runner commands. Those are not needed to understand the vulnerability or assess a deployment, and the relevant reproduction material was supplied during coordinated disclosure.
public success condition:
proof marker written
node remains healthy
runtime identity matches the intended binary
public failure interpretation:
clean miss = no proof, node healthy
unstable miss = process failure before a valid proof
stable proof = proof plus healthy node
A file write after a node crash is evidence of memory corruption, but it is not the same as a stable host-process RCE result in an operational node.
15 / TECHNICAL RECORD
15. Mitigation
Compiler fix:
release_locations_value() must restore native stack space.
It must not adjust the stack downward after releasing value locations.
Recommended fix:
- Apply upstream patches
- Upgrade to a fixed Wasmer release or to Wasmer 7.x
ECOSYSTEM SCOPE
The process behind the node is part of the threat model.
In a blockchain deployment, the vulnerable runtime may sit inside a validator, RPC provider or other node-side service. A committed transaction can make the execution path relevant to validators and full nodes. A simulation request can instead target a selected RPC process or validator endpoint one node at a time.
The same reasoning applies outside blockchains. Any Web2 service that embeds an affected Wasmer Singlepass configuration and accepts untrusted or attacker-influenced Wasm inherits the runtime-level question.
ACKNOWLEDGEMENTS
Credits
We thank SEAL911 and pcaversaccio for their assistance throughout the coordinated disclosure process. We also thank MajorExcitement for the careful proofreading and editorial review of this article.