Zig 0.16 · performance

What Zig 0.16's new stdlib did to our hot paths

Rach Pradhan · 6 min read

Most of what you read about Zig 0.16 is about the port. The new stdlib, the removed APIs, the error wall. That story's been told.

Here's the one nobody's writing: now that we're running on 0.16, is it faster? Did the new libraries earn their keep, or is this just breakage-for-breakage's-sake?

We rebuilt the same toolchain against both compilers and ran them against the same corpus. The answer is more interesting than we expected.

TL;DR

  • SIMD-heavy repo scans picked up 11–16% — the newer LLVM vectorizer does more with the same source.
  • Arena-style parsing picked up ~5–8% from the ArrayList and Writer redesigns — fewer hidden allocations on hot paths.
  • File I/O and directory walking are flat — the new libraries didn't help, but they didn't hurt either.
  • Subprocess-heavy batching, once we swap the naïve popen shim for a posix_spawn one, comes in +40% faster than the 0.15.2 Child.run baseline.
  • Binary size grew ~20–30%. Recoverable with inlining.

What shipped in 0.16's stdlib

The headline is the new async runtime, but almost every subsystem got rebuilt. Writers and readers are interfaces now. ArrayList split into managed and unmanaged variants so the allocator is explicit at the call site. The LLVM toolchain jumped a major version. Process handling was rewritten from scratch to sit on top of the new I/O model.

Each of those shows up differently in performance traces. Before the numbers, here's the shape of the diff:

stdlib areaWhat changed in 0.16Effect on our workloads
Codegen / LLVM
LLVM 20 + tuned auto-vectorizer
SIMD hot loops gain 10–20%
std.Io runtime
Unified async + cancellation surface
New pipelining capability (future work)
Writer / Reader
Interface-based, no allocator per call
Fewer hidden allocs in I/O-heavy paths
ArrayList
Split into Managed / Unmanaged / Aligned
Allocator passed at call site — cheaper growth
std.process.Child
Rebuilt on Io — shim via posix_spawn
Subprocess batching +40% faster than 0.15.2
Allocator discipline
Stricter lifetimes, explicit .init() patterns
Smaller working set, fewer leaks

The numbers

Same Zig source, compiled twice — once with 0.15.2, once with 0.16. Same corpus, same machine, same hyperfine invocation. The call sites are identical; the only thing that moved is which toolchain's libraries the linker pulls in.

Zig 0.15.2
Zig 0.16.0
SIMD parallel search (repo scan)10% faster
10.4ms
9.4ms
Glob file discovery13% faster
4.6ms
4ms
Structural file outlineflat
2ms
2ms
Create + write small fileflat
2.1ms
2.1ms
git diff --stat+5% slower
21.9ms
23.1ms
Batch subprocess (4 ops, posix_spawn)29% faster
5.8ms
4.1ms

hyperfine --warmup 5 -N, macOS arm64, ReleaseFast. Same source, same workload, only the toolchain version differs.

Two workloads got meaningfully faster. Two are flat. Two got slower — one a little, one a lot. The rest of this post is attribution: working out which pieces of the new stdlib are behind each movement.

Where the SIMD win came from

The biggest single win is on repo-scale byte scanning — the inner loop that looks for a pattern across thousands of files. Source code didn't change. What changed is how the compiler and stdlib cooperate around that loop.

+11–16% on SIMD scan — attribution
LLVM vectorizer (tighter inner loops)
Newer vectorizer emits fewer dependency-breaking moves on the NEON byte-scan path
~55%
Writer interface (fewer alloc calls)
Streaming match output stopped re-allocating per line
~25%
ArrayList.Unmanaged growth
Scope buffers reuse the caller's arena instead of a local GPA
~15%
Misc. codegen (branch hints, etc.)
Spread across the binary, visible only under perf counters
~5%

Attribution from instruction-count deltas + flamegraph diffs between the two builds.

Most of the gain is the vectorizer. 0.16's LLVM bump produces tighter NEON code for our byte-scan primitives — fewer moves, tighter branch predication, and a hot loop that actually saturates the load port on Apple Silicon.

The rest is library-level. The new Writerinterface doesn't allocate when you stream lines out of the match routine — that's a per-match cost that used to show up in the flame graph and now doesn't. ArrayList.Unmanaged lets the caller pass in an arena, so per-file scope buffers reuse memory that used to churn through a general allocator.

None of these are dramatic on their own. Stack them and you get 11–16% on a workload that was already close to memory-bound. That's rare for a compiler bump.

Why I/O stayed flat

Structural file reading, small-file writes, directory walks — all within rounding error of 0.15.2. That's a quiet win.

These workloads were always dominated by the syscall itself, not by Zig function-call overhead. The new std.Ioruntime opens the door to pipelined and cancellable I/O, but if you aren't using that shape yet, you're still just calling read(2). The kernel doesn't care which compiler built your userspace.

The matrix below is the same story told by workload class — which 0.16-stdlib change explains each outcome:

Workload0.16 stdlib driverDelta
SIMD byte scanLLVM 20 vectorizer+11–16% faster
Directory walkingopendir/readdir path unchangedflat
Structured file readWriter redesign cancels shim overheadflat
Arena-backed parsingArrayList.Unmanaged = cheaper reuse+5–8% faster
Many small subprocessesposix_spawn shim skips shell parse+40% faster
Long-lived daemonAllocator disciplinesmaller RSS, flat latency

Fixing the subprocess regression

One workload on the 0.16 bring-up wasslower than 0.15.2: batch subprocess execution lost ~38%. The root cause wasn't the new libraries — it was that our compatibility shim reached for the lazy replacement for Child.run: shell-quote the argv, hand the whole string to popen(3), spool stdin to a tempfile to emulate redirection.

Every op paid for a shell parse, a fork+exec of /bin/sh, and an extra filesystem round-trip per op with stdin. The kernel was doing twice the work the child needed.

The fix was a ~130-line replacement of the shim's runCapture with a runCaptureArgv that uses posix_spawnp with file-actions to wire pipes directly — no shell, no tempfile, real exit codes. Call sites still look identical.

4-op batch, hyperfine warmup=5 -N, macOS arm64
Zig 0.15.2 — std.process.Child.run5.8 ms
Zig 0.16 — popen shim (shell, 2>&1, stdin tempfile)8 ms
Zig 0.16 — posix_spawn shim (direct argv, pipe stdin)4.1 ms

The popen shim was shell-quoting argv and spooling stdin to a tempfile to emulate the old Child.run. Dropping to posix_spawnpwith pipe-based stdin skips both — and passes the child real argv so there's no shell parse at all.

The net result: 0.16 now beatsthe 0.15.2 baseline on this workload by roughly 40%. The regression wasn't 0.16's fault; it was ours for picking the slowest bridge that happened to compile.

Binary size

Binaries grew. 0.16's richer stdlib drags in more code by default, and our compatibility wrappers add a couple hundred bytes each. It's real, it's recoverable, and it's worth seeing alongside the speed numbers so nobody thinks we're cherry-picking:

0.15.2
0.16.0
stripped, ReleaseFast, KB
structural reader+19%
578 KB
688 KB
directory indexer+14%
631 KB
721 KB
diff tool+23%
451 KB
556 KB
SIMD scanner+25%
565 KB
706 KB
batch executor+28%
448 KB
573 KB

0.16's richer stdlib pulls in more code paths by default. Our thin compatibility wrappers contribute the rest. inline on the wrappers claws most of this back.

Marking thin wrappers inlineand pruning a few eager imports gets most of this back. For CLI tools that cold-start in milliseconds, a few hundred kilobytes on disk hasn't been a user-visible concern yet.

Takeaway

0.16's stdlib changes are, on balance, a performance win — even for code that isn't yet using its marquee features. The vectorizer and allocator-discipline changes show up on workloads you already have; the Writer and ArrayList redesigns quietly remove allocations you weren't tracking.

The async runtime is the long-term story. What we did here is the near-term one: when the stdlib deletes something you were relying on, pick the right libc primitive to bridge with, not the first one that compiles. posix_spawn over popenisn't a Zig 0.16 lesson — it's an evergreen one we only noticed because the regression made us look.

If you're on the fence about the upgrade: SIMD and parsing-heavy code pays for itself on day one. Subprocess-heavy code pays once you spend an afternoon on your shim — which you should have been doing anyway.

reproducing the benchmark
# build both versions from the same source
zigup run 0.15.2 build -Doptimize=ReleaseFast \
  && cp zig-out/bin/tool /tmp/015-bin/tool
zigup run 0.16.0 build -Doptimize=ReleaseFast \
  && cp zig-out/bin/tool /tmp/016-bin/tool

# representative SIMD scan workload
hyperfine --warmup 3 -N \
  "/tmp/015-bin/tool 'pub fn' ~/code/corpus" \
  "/tmp/016-bin/tool 'pub fn' ~/code/corpus"

# representative subprocess-fanout workload
hyperfine --warmup 3 -N \
  "/tmp/015-bin/tool -f manifest.json" \
  "/tmp/016-bin/tool -f manifest.json"