Most of what you read about Zig 0.16 is about the port. The new stdlib, the removed APIs, the error wall. That story's been told.
Here's the one nobody's writing: now that we're running on 0.16, is it faster? Did the new libraries earn their keep, or is this just breakage-for-breakage's-sake?
We rebuilt the same toolchain against both compilers and ran them against the same corpus. The answer is more interesting than we expected.
TL;DR
- SIMD-heavy repo scans picked up 11–16% — the newer LLVM vectorizer does more with the same source.
- Arena-style parsing picked up ~5–8% from the ArrayList and Writer redesigns — fewer hidden allocations on hot paths.
- File I/O and directory walking are flat — the new libraries didn't help, but they didn't hurt either.
- Subprocess-heavy batching, once we swap the naïve
popenshim for aposix_spawnone, comes in +40% faster than the 0.15.2Child.runbaseline. - Binary size grew ~20–30%. Recoverable with inlining.
What shipped in 0.16's stdlib
The headline is the new async runtime, but almost every subsystem got rebuilt. Writers and readers are interfaces now. ArrayList split into managed and unmanaged variants so the allocator is explicit at the call site. The LLVM toolchain jumped a major version. Process handling was rewritten from scratch to sit on top of the new I/O model.
Each of those shows up differently in performance traces. Before the numbers, here's the shape of the diff:
The numbers
Same Zig source, compiled twice — once with 0.15.2, once with 0.16. Same corpus, same machine, same hyperfine invocation. The call sites are identical; the only thing that moved is which toolchain's libraries the linker pulls in.
hyperfine --warmup 5 -N, macOS arm64, ReleaseFast. Same source, same workload, only the toolchain version differs.
Two workloads got meaningfully faster. Two are flat. Two got slower — one a little, one a lot. The rest of this post is attribution: working out which pieces of the new stdlib are behind each movement.
Where the SIMD win came from
The biggest single win is on repo-scale byte scanning — the inner loop that looks for a pattern across thousands of files. Source code didn't change. What changed is how the compiler and stdlib cooperate around that loop.
Attribution from instruction-count deltas + flamegraph diffs between the two builds.
Most of the gain is the vectorizer. 0.16's LLVM bump produces tighter NEON code for our byte-scan primitives — fewer moves, tighter branch predication, and a hot loop that actually saturates the load port on Apple Silicon.
The rest is library-level. The new Writerinterface doesn't allocate when you stream lines out of the match routine — that's a per-match cost that used to show up in the flame graph and now doesn't. ArrayList.Unmanaged lets the caller pass in an arena, so per-file scope buffers reuse memory that used to churn through a general allocator.
None of these are dramatic on their own. Stack them and you get 11–16% on a workload that was already close to memory-bound. That's rare for a compiler bump.
Why I/O stayed flat
Structural file reading, small-file writes, directory walks — all within rounding error of 0.15.2. That's a quiet win.
These workloads were always dominated by the syscall itself, not by Zig function-call overhead. The new std.Ioruntime opens the door to pipelined and cancellable I/O, but if you aren't using that shape yet, you're still just calling read(2). The kernel doesn't care which compiler built your userspace.
The matrix below is the same story told by workload class — which 0.16-stdlib change explains each outcome:
| Workload | 0.16 stdlib driver | Delta |
|---|---|---|
| SIMD byte scan | LLVM 20 vectorizer | +11–16% faster |
| Directory walking | opendir/readdir path unchanged | flat |
| Structured file read | Writer redesign cancels shim overhead | flat |
| Arena-backed parsing | ArrayList.Unmanaged = cheaper reuse | +5–8% faster |
| Many small subprocesses | posix_spawn shim skips shell parse | +40% faster |
| Long-lived daemon | Allocator discipline | smaller RSS, flat latency |
Fixing the subprocess regression
One workload on the 0.16 bring-up wasslower than 0.15.2: batch subprocess execution lost ~38%. The root cause wasn't the new libraries — it was that our compatibility shim reached for the lazy replacement for Child.run: shell-quote the argv, hand the whole string to popen(3), spool stdin to a tempfile to emulate redirection.
Every op paid for a shell parse, a fork+exec of /bin/sh, and an extra filesystem round-trip per op with stdin. The kernel was doing twice the work the child needed.
The fix was a ~130-line replacement of the shim's runCapture with a runCaptureArgv that uses posix_spawnp with file-actions to wire pipes directly — no shell, no tempfile, real exit codes. Call sites still look identical.
The popen shim was shell-quoting argv and spooling stdin to a tempfile to emulate the old Child.run. Dropping to posix_spawnpwith pipe-based stdin skips both — and passes the child real argv so there's no shell parse at all.
The net result: 0.16 now beatsthe 0.15.2 baseline on this workload by roughly 40%. The regression wasn't 0.16's fault; it was ours for picking the slowest bridge that happened to compile.
Binary size
Binaries grew. 0.16's richer stdlib drags in more code by default, and our compatibility wrappers add a couple hundred bytes each. It's real, it's recoverable, and it's worth seeing alongside the speed numbers so nobody thinks we're cherry-picking:
0.16's richer stdlib pulls in more code paths by default. Our thin compatibility wrappers contribute the rest. inline on the wrappers claws most of this back.
Marking thin wrappers inlineand pruning a few eager imports gets most of this back. For CLI tools that cold-start in milliseconds, a few hundred kilobytes on disk hasn't been a user-visible concern yet.
Takeaway
0.16's stdlib changes are, on balance, a performance win — even for code that isn't yet using its marquee features. The vectorizer and allocator-discipline changes show up on workloads you already have; the Writer and ArrayList redesigns quietly remove allocations you weren't tracking.
The async runtime is the long-term story. What we did here is the near-term one: when the stdlib deletes something you were relying on, pick the right libc primitive to bridge with, not the first one that compiles. posix_spawn over popenisn't a Zig 0.16 lesson — it's an evergreen one we only noticed because the regression made us look.
If you're on the fence about the upgrade: SIMD and parsing-heavy code pays for itself on day one. Subprocess-heavy code pays once you spend an afternoon on your shim — which you should have been doing anyway.
# build both versions from the same source zigup run 0.15.2 build -Doptimize=ReleaseFast \ && cp zig-out/bin/tool /tmp/015-bin/tool zigup run 0.16.0 build -Doptimize=ReleaseFast \ && cp zig-out/bin/tool /tmp/016-bin/tool # representative SIMD scan workload hyperfine --warmup 3 -N \ "/tmp/015-bin/tool 'pub fn' ~/code/corpus" \ "/tmp/016-bin/tool 'pub fn' ~/code/corpus" # representative subprocess-fanout workload hyperfine --warmup 3 -N \ "/tmp/015-bin/tool -f manifest.json" \ "/tmp/016-bin/tool -f manifest.json"