Crafter.Graphics/examples/SponzaBench/README.md
catbot 619e39369d test(bench): SponzaBench harness + #40→HEAD perf measurement (#155)
Headless benchmark around the native Sponza RT scene: times setup and a
measured Render() loop over the full multi-mesh atrium, prints BENCH
metrics, and exits. Includes run-bench.sh and a README documenting the
methodology and the measured net gain from #40 to current master.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 19:35:54 +00:00

131 lines
6.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# SponzaBench — measuring the post-#40 performance work
This example exists to answer issue **#155**: *starting with #40 a lot of
performance-related issues were merged — what is the net performance gain,
measured in a representative scene (Sponza)?*
It is a headless benchmark harness built around the **native Vulkan**
Sponza ray-tracing scene. Same asset bundle and camera as
[`examples/Sponza`](../Sponza), but instead of opening an interactive
window it times the work and prints machine-readable `BENCH …` lines, then
exits.
## What it measures
* **setup** — process start through the first command submission: asset
decompression, BLAS build per mesh group, the multi-instance TLAS, the
RT pipeline, GPU memory placement and the descriptor writes. This is the
window most of the post-#40 *native* perf work acts on.
* **frames** — a warmup followed by a measured loop calling
`Window::Render()`, reporting per-frame wall-clock stats and throughput.
Unlike the interactive Sponza example — which is single-material on native
because of the hit-shader dynamic-`descriptor_heap`-index driver fault
(see `examples/Sponza/README.md`) — SponzaBench's closest-hit shades from
barycentric coordinates and samples **no** texture, so it can build the
**full multi-mesh atrium** (25 mesh groups, ~262 k triangles) as one
multi-instance TLAS. The albedo is still decompressed and uploaded during
setup (to keep that path in the measurement) but is not bound.
## Running
```bash
cd examples/SponzaBench
crafter-build # native Vulkan
# from the produced bin dir:
VK_LOADER_LAYERS_DISABLE='~all~' \
CRAFTER_PRESENT_IMMEDIATE=1 \
BENCH_WARMUP=200 BENCH_FRAMES=2000 BENCH_MESHES=25 ./SponzaBench
```
`run-bench.sh <bindir> <label> <meshes> <reps> [--cold]` runs it N times
and reports the median / min / max of each metric.
Environment knobs:
| var | effect |
|---|---|
| `CRAFTER_PRESENT_IMMEDIATE=1` | uncapped present mode — without it FIFO pins frame time to the compositor's vblank (~60 Hz) and steady-state throughput can't be seen. |
| `VK_LOADER_LAYERS_DISABLE='~all~'` | disable the validation layer. The engine enables Khronos validation **and GPU-assisted validation** unconditionally; that adds large, version-dependent overhead and must be off for a representative measurement. |
| `BENCH_WARMUP` / `BENCH_FRAMES` | warmup and measured frame counts (default 200 / 2000). |
| `BENCH_MESHES` | cap on mesh groups loaded (default: all 25). A small cap shrinks GPU work so per-frame **CPU** cost dominates; the full scene is GPU-traversal bound. |
| `BENCH_CLEAN_EXIT=1` | exit via `std::exit` (runs `atexit`, which writes the #69 pipeline cache) instead of the default hard `_Exit`. Used once to seed `pipeline_cache.bin` for a warm-cache measurement. |
## Results — #40 vs current master
Measured on this repo's CI box: **NVIDIA RTX 4090**, driver `610.43.02`,
1280×720, validation disabled, `IMMEDIATE` present, 200 warmup + 2000
measured frames, **median of 9 runs**. "#40" is commit `1451e3a` (the #40
merge); "HEAD" is current master. The *same* SponzaBench sources were
built against each library revision (a git worktree at #40).
### Full atrium — 25 meshes / ~262 k triangles (GPU-traversal bound)
| metric | #40 | HEAD | Δ |
|---|---:|---:|---:|
| setup (cold) ms | 322.7 | 316.6 | **1.9 %** |
| setup (warm pipeline cache, #69) ms | 322.7¹ | 313.9 | **2.7 %** |
| throughput fps | 7 560 | 7 637 | **+1.0 %** |
| frame time p50 ms | 0.1311 | 0.1298 | **1.0 %** |
| peak host RSS (cold) MB | 336.0 | 344.2 | +2.4 % |
| peak host RSS (warm cache) MB | 336.0¹ | 332.0 | **1.2 %** |
¹ #40 predates the disk pipeline cache (#69), so its setup is always the
cold-compile path.
### Light scene — 1 mesh (deliberately CPU-bound)
| metric | #40 | HEAD | Δ |
|---|---:|---:|---:|
| throughput fps | 15 022 | 15 237 | **+1.4 %** |
| frame time p50 ms | 0.0662 | 0.0650 | **1.8 %** |
## Interpretation
**The net measured gain in a static Sponza RT scene is small: ~12 %
faster frames, ~23 % faster setup, and roughly flat host memory.** That is
an honest result, and the reason is *what Sponza exercises*, not that the
perf work was ineffective:
* **Most post-#40 PRs don't touch this workload.** The largest block is UI
/ text rendering (shaped-run cache, font-atlas dirty uploads, the UI
compute-shader rewrites — #46#57, #61, #122#129, #132). Sponza has no
UI, so those contribute **zero**. A second block optimises **per-frame
dynamic uploads** (TLAS dirty-tracking #118, deforming-mesh refit #119,
the staging ring #120). A *static* scene builds its TLAS once and never
re-uploads, so these don't fire in steady state either. WebGPU-only
(#130/#131/#133) and Win32-only (#134) PRs don't apply to a native
Vulkan build at all.
* **The PRs that *do* apply act on setup and memory, not frame time.**
Device-local placement (#65/#72/#73/#75), staging release
(#66/#67/#114), the pipeline cache (#69) and the deferred-deletion queue
(#101/#116) move setup cost and peak memory — which is exactly where the
measured ~23 % setup change and the warm-cache RSS drop show up. The
pipeline cache itself saves ~3 ms here (one RT pipeline) and, more
visibly, ~12 MB of peak RSS by skipping the cold shader-compile
allocations.
* **Steady-state frame time is GPU-traversal bound.** With 262 k triangles
the per-frame CPU work (barrier scoping #48/#115, cached heap-bind
structs #42/#43) is hidden behind GPU traversal, so it can't move the
frame time. Shrinking the scene until it is CPU-bound (the light scene
above) surfaces the per-frame CPU savings — and even then they are only
~1.8 %, because that CPU path was already cheap.
**Takeaway:** the post-#40 work is real but concentrated in UI/text and
per-frame dynamic-upload paths; a static, UI-less RT scene is the wrong
workload to see most of it. To quantify the UI/text gains a separate
benchmark over a text-heavy `UIRenderer` scene (or an animated/deforming
scene for the dynamic-upload PRs) would be needed.
## Caveats / notes
* The stock `examples/Sponza` **native** path does not currently compile
against the (unpinned) Vulkan-Headers `main`: the
`VkResourceDescriptorDataEXT` union no longer has the
`pCombinedImageSampler` member the example uses. SponzaBench sidesteps
this by not binding a combined image+sampler. (Pre-existing, unrelated to
#155 — worth a follow-up to port Sponza to the split sampled-image +
sampler-heap model the engine's own UI renderer already uses.)
* Numbers are CPU/GPU specific. Re-run `run-bench.sh` locally for your
hardware; the *deltas* between revisions are the point, not the absolute
figures.