Processes and Fork-Join
Every initial/always block in Verilog already runs concurrently with every other one — that's the "concurrent hardware description" idea from Verilog's Introduction page. What plain Verilog lacks is a way to spin up multiple concurrent processes from within a single procedural block and control exactly how the surrounding code waits for them. SystemVerilog's fork/join (and its two variants) is that control.
fork/join: wait for every branch
initial begin
fork
begin
@(posedge clk);
$display("branch 1 done");
end
begin
#100;
$display("branch 2 done");
end
join
$display("both branches finished");
end
Every begin...end block inside fork/join starts running at the same simulation time, genuinely in parallel with each other — not one after another. join (plain) blocks the surrounding code until every forked branch has completed, whichever finishes last. Here, "both branches finished" only prints once both the clock edge and the 100-time-unit delay have happened, however long the slower of the two takes.
join_any: wait for the first branch
fork
begin
@(posedge clk);
$display("clock edge arrived first");
end
begin
#1000;
$display("timeout reached first");
end
join_any
join_any continues past the fork/join_any block as soon as any one branch completes, without waiting for the others (which keep running in the background). This is the standard pattern for a timeout watchdog: race the "wait for the expected event" branch against a "wait a fixed timeout" branch, and whichever happens first determines whether the test saw the expected behavior in time or hung.
join_none: don't wait at all
fork
begin
@(posedge clk);
$display("this runs whenever its event eventually happens");
end
join_none
$display("this prints immediately, without waiting for the forked branch");
join_none starts the forked branch(es) and immediately continues past the fork block without waiting for anything — the forked code keeps running independently in the background. This is the shape a driver typically uses to kick off a background monitor process at the start of a test without blocking the main test sequence on it.
disable: terminating a running process early
fork : watchdog_block
begin
@(posedge clk);
disable watchdog_block; // cancel the timeout branch below, since the event arrived
end
begin
#1000;
$display("TIMEOUT — event never arrived");
end
join
A named block (fork : watchdog_block) can be terminated early with disable block_name from anywhere inside it — used here so the successful branch can cancel the now-unnecessary timeout branch instead of leaving it running to completion pointlessly. This combines naturally with join_any too: after join_any continues past whichever branch finished first, any branches still running can be explicitly disabled if they're no longer needed.
wait fork and disable fork: two more ways to manage every child process at once
initial begin
fork
begin #10; $display("A done"); end
begin #20; $display("B done"); end
join_none
$display("kicked off A and B, not waiting yet");
wait fork; // now block until every descendant process finishes
$display("A and B are both done");
end
wait fork blocks the calling process until every process it has spawned (directly or indirectly), at any nesting depth, has completed — the tool for a join_none (or join_any) block that eventually does need to know everything it launched has finished, without having to restructure it into a plain join from the start. disable fork is the destructive counterpart: instead of waiting for descendant processes, it immediately terminates every one of them — the process-control equivalent of disable block_name, but scoped to "every child process spawned from here," not one specific named block.
The process class: fine-grained control over one specific process
process mon_proc;
initial begin
fork
begin
mon_proc = process::self();
forever @(posedge clk) $display("monitoring...");
end
join_none
#500;
mon_proc.kill(); // terminate that one specific process, from outside it
end
process::self() returns a handle to the currently-executing process — stashing that handle (as mon_proc above) lets a different part of the testbench control that specific process later: .kill() terminates it, .suspend()/.resume() pause and continue it, and .status() reports whether it's currently running, waiting, finished, or killed. This is more targeted than disable fork (which kills every descendant indiscriminately) or a named disable block_name (which requires the process to be inside a specific named block) — a stored process handle can reach one specific process directly, from anywhere that handle is visible.
The fork-in-a-loop variable-capture gotcha
// BUGGY: every spawned process ends up printing the same final value
for (int i = 0; i < 3; i++) begin
fork
$display("i = %0d", i);
join_none
end
// prints "i = 3" three times, not "i = 0", "i = 1", "i = 2"
// FIXED: give each iteration its own copy of the loop variable
for (int i = 0; i < 3; i++) begin
automatic int j = i;
fork
$display("j = %0d", j);
join_none
end
// prints "j = 0", "j = 1", "j = 2", in some order
Spawned processes don't actually start executing until the parent thread hits a blocking statement — so a tight for loop with no delay inside it runs to completion before any of the join_none-spawned branches actually execute, and by then the loop variable i has already reached its final value (3), which every branch reads. Declaring a fresh automatic variable inside the loop body (j above) forces a brand-new copy to be created on every iteration, so each spawned process captures its own iteration's value instead of all of them sharing the one loop variable. This is one of the most common real-world fork/join_none bugs, and shows up constantly when spawning one process per transaction, per DUT instance, or per array element in a loop.
fork/join cannot appear inside a function
Every example on this page lives inside an initial block or a task — never a function. That's not a stylistic choice: a function must return in zero simulation time, and fork/join (even join_none, which doesn't block the calling process, still consumes time for the forked branches themselves) is fundamentally about time-consuming concurrency. A function body containing fork/join in any form is a compile-time error — reach for a task instead whenever a piece of reusable code needs to spawn concurrent processes.
Why this matters for testbenches specifically
A layered testbench needs a driver, a monitor, and a scoreboard all running at once, for the entire duration of a test — fork/join_none is exactly how a test's top-level sequence launches all of them without becoming a single, serialized piece of code that runs the driver, then the monitor, then the scoreboard, one after another (which would defeat the entire point of them monitoring and reacting to the same signals in real time). Testbench and UVM both build their own component/phase machinery on top of exactly this concurrency mechanism, even though a UVM testbench rarely writes raw fork/join directly — the phasing mechanism does it on the user's behalf.
What's next
Concurrent processes often need to coordinate with each other — one process signaling another that data is ready, or several processes sharing a limited resource. The next page covers the three primitives SystemVerilog provides for exactly that: semaphore, mailbox, and event.