8 changed files with 633 additions and 1 deletions
131
examples/SponzaBench/README.md
Normal file
131
examples/SponzaBench/README.md
Normal file
|
|
@ -0,0 +1,131 @@
|
||||||
|
# SponzaBench — measuring the post-#40 performance work
|
||||||
|
|
||||||
|
This example exists to answer issue **#155**: *starting with #40 a lot of
|
||||||
|
performance-related issues were merged — what is the net performance gain,
|
||||||
|
measured in a representative scene (Sponza)?*
|
||||||
|
|
||||||
|
It is a headless benchmark harness built around the **native Vulkan**
|
||||||
|
Sponza ray-tracing scene. Same asset bundle and camera as
|
||||||
|
[`examples/Sponza`](../Sponza), but instead of opening an interactive
|
||||||
|
window it times the work and prints machine-readable `BENCH …` lines, then
|
||||||
|
exits.
|
||||||
|
|
||||||
|
## What it measures
|
||||||
|
|
||||||
|
* **setup** — process start through the first command submission: asset
|
||||||
|
decompression, BLAS build per mesh group, the multi-instance TLAS, the
|
||||||
|
RT pipeline, GPU memory placement and the descriptor writes. This is the
|
||||||
|
window most of the post-#40 *native* perf work acts on.
|
||||||
|
* **frames** — a warmup followed by a measured loop calling
|
||||||
|
`Window::Render()`, reporting per-frame wall-clock stats and throughput.
|
||||||
|
|
||||||
|
Unlike the interactive Sponza example — which is single-material on native
|
||||||
|
because of the hit-shader dynamic-`descriptor_heap`-index driver fault
|
||||||
|
(see `examples/Sponza/README.md`) — SponzaBench's closest-hit shades from
|
||||||
|
barycentric coordinates and samples **no** texture, so it can build the
|
||||||
|
**full multi-mesh atrium** (25 mesh groups, ~262 k triangles) as one
|
||||||
|
multi-instance TLAS. The albedo is still decompressed and uploaded during
|
||||||
|
setup (to keep that path in the measurement) but is not bound.
|
||||||
|
|
||||||
|
## Running
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd examples/SponzaBench
|
||||||
|
crafter-build # native Vulkan
|
||||||
|
# from the produced bin dir:
|
||||||
|
VK_LOADER_LAYERS_DISABLE='~all~' \
|
||||||
|
CRAFTER_PRESENT_IMMEDIATE=1 \
|
||||||
|
BENCH_WARMUP=200 BENCH_FRAMES=2000 BENCH_MESHES=25 ./SponzaBench
|
||||||
|
```
|
||||||
|
|
||||||
|
`run-bench.sh <bindir> <label> <meshes> <reps> [--cold]` runs it N times
|
||||||
|
and reports the median / min / max of each metric.
|
||||||
|
|
||||||
|
Environment knobs:
|
||||||
|
|
||||||
|
| var | effect |
|
||||||
|
|---|---|
|
||||||
|
| `CRAFTER_PRESENT_IMMEDIATE=1` | uncapped present mode — without it FIFO pins frame time to the compositor's vblank (~60 Hz) and steady-state throughput can't be seen. |
|
||||||
|
| `VK_LOADER_LAYERS_DISABLE='~all~'` | disable the validation layer. The engine enables Khronos validation **and GPU-assisted validation** unconditionally; that adds large, version-dependent overhead and must be off for a representative measurement. |
|
||||||
|
| `BENCH_WARMUP` / `BENCH_FRAMES` | warmup and measured frame counts (default 200 / 2000). |
|
||||||
|
| `BENCH_MESHES` | cap on mesh groups loaded (default: all 25). A small cap shrinks GPU work so per-frame **CPU** cost dominates; the full scene is GPU-traversal bound. |
|
||||||
|
| `BENCH_CLEAN_EXIT=1` | exit via `std::exit` (runs `atexit`, which writes the #69 pipeline cache) instead of the default hard `_Exit`. Used once to seed `pipeline_cache.bin` for a warm-cache measurement. |
|
||||||
|
|
||||||
|
## Results — #40 vs current master
|
||||||
|
|
||||||
|
Measured on this repo's CI box: **NVIDIA RTX 4090**, driver `610.43.02`,
|
||||||
|
1280×720, validation disabled, `IMMEDIATE` present, 200 warmup + 2000
|
||||||
|
measured frames, **median of 9 runs**. "#40" is commit `1451e3a` (the #40
|
||||||
|
merge); "HEAD" is current master. The *same* SponzaBench sources were
|
||||||
|
built against each library revision (a git worktree at #40).
|
||||||
|
|
||||||
|
### Full atrium — 25 meshes / ~262 k triangles (GPU-traversal bound)
|
||||||
|
|
||||||
|
| metric | #40 | HEAD | Δ |
|
||||||
|
|---|---:|---:|---:|
|
||||||
|
| setup (cold) ms | 322.7 | 316.6 | **−1.9 %** |
|
||||||
|
| setup (warm pipeline cache, #69) ms | 322.7¹ | 313.9 | **−2.7 %** |
|
||||||
|
| throughput fps | 7 560 | 7 637 | **+1.0 %** |
|
||||||
|
| frame time p50 ms | 0.1311 | 0.1298 | **−1.0 %** |
|
||||||
|
| peak host RSS (cold) MB | 336.0 | 344.2 | +2.4 % |
|
||||||
|
| peak host RSS (warm cache) MB | 336.0¹ | 332.0 | **−1.2 %** |
|
||||||
|
|
||||||
|
¹ #40 predates the disk pipeline cache (#69), so its setup is always the
|
||||||
|
cold-compile path.
|
||||||
|
|
||||||
|
### Light scene — 1 mesh (deliberately CPU-bound)
|
||||||
|
|
||||||
|
| metric | #40 | HEAD | Δ |
|
||||||
|
|---|---:|---:|---:|
|
||||||
|
| throughput fps | 15 022 | 15 237 | **+1.4 %** |
|
||||||
|
| frame time p50 ms | 0.0662 | 0.0650 | **−1.8 %** |
|
||||||
|
|
||||||
|
## Interpretation
|
||||||
|
|
||||||
|
**The net measured gain in a static Sponza RT scene is small: ~1–2 %
|
||||||
|
faster frames, ~2–3 % faster setup, and roughly flat host memory.** That is
|
||||||
|
an honest result, and the reason is *what Sponza exercises*, not that the
|
||||||
|
perf work was ineffective:
|
||||||
|
|
||||||
|
* **Most post-#40 PRs don't touch this workload.** The largest block is UI
|
||||||
|
/ text rendering (shaped-run cache, font-atlas dirty uploads, the UI
|
||||||
|
compute-shader rewrites — #46–#57, #61, #122–#129, #132). Sponza has no
|
||||||
|
UI, so those contribute **zero**. A second block optimises **per-frame
|
||||||
|
dynamic uploads** (TLAS dirty-tracking #118, deforming-mesh refit #119,
|
||||||
|
the staging ring #120). A *static* scene builds its TLAS once and never
|
||||||
|
re-uploads, so these don't fire in steady state either. WebGPU-only
|
||||||
|
(#130/#131/#133) and Win32-only (#134) PRs don't apply to a native
|
||||||
|
Vulkan build at all.
|
||||||
|
* **The PRs that *do* apply act on setup and memory, not frame time.**
|
||||||
|
Device-local placement (#65/#72/#73/#75), staging release
|
||||||
|
(#66/#67/#114), the pipeline cache (#69) and the deferred-deletion queue
|
||||||
|
(#101/#116) move setup cost and peak memory — which is exactly where the
|
||||||
|
measured ~2–3 % setup change and the warm-cache RSS drop show up. The
|
||||||
|
pipeline cache itself saves ~3 ms here (one RT pipeline) and, more
|
||||||
|
visibly, ~12 MB of peak RSS by skipping the cold shader-compile
|
||||||
|
allocations.
|
||||||
|
* **Steady-state frame time is GPU-traversal bound.** With 262 k triangles
|
||||||
|
the per-frame CPU work (barrier scoping #48/#115, cached heap-bind
|
||||||
|
structs #42/#43) is hidden behind GPU traversal, so it can't move the
|
||||||
|
frame time. Shrinking the scene until it is CPU-bound (the light scene
|
||||||
|
above) surfaces the per-frame CPU savings — and even then they are only
|
||||||
|
~1.8 %, because that CPU path was already cheap.
|
||||||
|
|
||||||
|
**Takeaway:** the post-#40 work is real but concentrated in UI/text and
|
||||||
|
per-frame dynamic-upload paths; a static, UI-less RT scene is the wrong
|
||||||
|
workload to see most of it. To quantify the UI/text gains a separate
|
||||||
|
benchmark over a text-heavy `UIRenderer` scene (or an animated/deforming
|
||||||
|
scene for the dynamic-upload PRs) would be needed.
|
||||||
|
|
||||||
|
## Caveats / notes
|
||||||
|
|
||||||
|
* The stock `examples/Sponza` **native** path does not currently compile
|
||||||
|
against the (unpinned) Vulkan-Headers `main`: the
|
||||||
|
`VkResourceDescriptorDataEXT` union no longer has the
|
||||||
|
`pCombinedImageSampler` member the example uses. SponzaBench sidesteps
|
||||||
|
this by not binding a combined image+sampler. (Pre-existing, unrelated to
|
||||||
|
#155 — worth a follow-up to port Sponza to the split sampled-image +
|
||||||
|
sampler-heap model the engine's own UI renderer already uses.)
|
||||||
|
* Numbers are CPU/GPU specific. Re-run `run-bench.sh` locally for your
|
||||||
|
hardware; the *deltas* between revisions are the point, not the absolute
|
||||||
|
figures.
|
||||||
16
examples/SponzaBench/closesthit.glsl
Normal file
16
examples/SponzaBench/closesthit.glsl
Normal file
|
|
@ -0,0 +1,16 @@
|
||||||
|
#version 460
|
||||||
|
#extension GL_EXT_ray_tracing : enable
|
||||||
|
|
||||||
|
// Benchmark closest-hit: shade from the hit's barycentric coordinates.
|
||||||
|
// SponzaBench deliberately avoids the descriptor-heap texture sample the
|
||||||
|
// interactive Sponza example uses — the point here is to measure the
|
||||||
|
// ray-tracing traversal + frame-loop cost, not texturing, and a
|
||||||
|
// texture-free hit keeps the host setup portable across library
|
||||||
|
// revisions (issue #155).
|
||||||
|
hitAttributeEXT vec2 hitAttrs;
|
||||||
|
layout(location = 0) rayPayloadInEXT vec3 hitValue;
|
||||||
|
|
||||||
|
void main() {
|
||||||
|
vec3 bary = vec3(1.0 - hitAttrs.x - hitAttrs.y, hitAttrs.x, hitAttrs.y);
|
||||||
|
hitValue = bary;
|
||||||
|
}
|
||||||
306
examples/SponzaBench/main.cpp
Normal file
306
examples/SponzaBench/main.cpp
Normal file
|
|
@ -0,0 +1,306 @@
|
||||||
|
// SponzaBench — headless benchmark around the native Sponza RT scene.
|
||||||
|
//
|
||||||
|
// Identical setup to examples/Sponza (native branch): load the Sponza
|
||||||
|
// .cmesh + albedo .ctex, build BLAS/TLAS, wire an RTPass. Instead of
|
||||||
|
// opening an interactive window and idling in StartSync(), it times two
|
||||||
|
// things and exits:
|
||||||
|
//
|
||||||
|
// * setup — process start through the first command submission
|
||||||
|
// (asset decompress + BLAS/TLAS build + pipeline + memory
|
||||||
|
// placement + descriptor writes), which the bulk of the
|
||||||
|
// post-#40 perf work targets (device-local placement,
|
||||||
|
// staging release, pipeline cache, decompress, ...).
|
||||||
|
// * frames — a warmup then a measured loop calling Window::Render(),
|
||||||
|
// reporting per-frame wall-clock stats. Run with
|
||||||
|
// CRAFTER_PRESENT_IMMEDIATE=1 to uncap from vblank.
|
||||||
|
//
|
||||||
|
// Tunables (env): BENCH_WARMUP (default 200), BENCH_FRAMES (default 2000).
|
||||||
|
// Output lines are prefixed "BENCH " for easy scraping.
|
||||||
|
|
||||||
|
#include "vulkan/vulkan.h"
|
||||||
|
|
||||||
|
import Crafter.Graphics;
|
||||||
|
import Crafter.Asset;
|
||||||
|
import Crafter.Math;
|
||||||
|
import Crafter.Event;
|
||||||
|
import std;
|
||||||
|
|
||||||
|
using namespace Crafter;
|
||||||
|
namespace fs = std::filesystem;
|
||||||
|
using Clock = std::chrono::steady_clock;
|
||||||
|
|
||||||
|
namespace {
|
||||||
|
struct RGBA8 { std::uint8_t r, g, b, a; };
|
||||||
|
|
||||||
|
void RequireAssets(const fs::path& mesh, const fs::path& tex) {
|
||||||
|
if (fs::exists(mesh) && fs::exists(tex)) return;
|
||||||
|
std::println(std::cerr, "[SponzaBench] missing asset(s): {} / {}",
|
||||||
|
mesh.string(), tex.string());
|
||||||
|
std::abort();
|
||||||
|
}
|
||||||
|
|
||||||
|
// Peak resident set size (high-water mark) in kB, from /proc/self/status.
|
||||||
|
long PeakRSSkB() {
|
||||||
|
std::ifstream st("/proc/self/status");
|
||||||
|
std::string line;
|
||||||
|
while (std::getline(st, line)) {
|
||||||
|
if (line.rfind("VmHWM:", 0) == 0) {
|
||||||
|
long kb = 0;
|
||||||
|
std::sscanf(line.c_str(), "VmHWM: %ld", &kb);
|
||||||
|
return kb;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return -1;
|
||||||
|
}
|
||||||
|
|
||||||
|
std::size_t EnvSize(const char* name, std::size_t fallback) {
|
||||||
|
if (const char* v = std::getenv(name)) {
|
||||||
|
char* end = nullptr;
|
||||||
|
unsigned long parsed = std::strtoul(v, &end, 10);
|
||||||
|
if (end != v) return static_cast<std::size_t>(parsed);
|
||||||
|
}
|
||||||
|
return fallback;
|
||||||
|
}
|
||||||
|
|
||||||
|
double Ms(std::chrono::nanoseconds ns) {
|
||||||
|
return std::chrono::duration<double, std::milli>(ns).count();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
int main() {
|
||||||
|
const std::size_t warmup = EnvSize("BENCH_WARMUP", 200);
|
||||||
|
const std::size_t frames = EnvSize("BENCH_FRAMES", 2000);
|
||||||
|
|
||||||
|
const auto tProcStart = Clock::now();
|
||||||
|
|
||||||
|
const fs::path texPath = "tex_0.ctex";
|
||||||
|
RequireAssets("mesh_0.cmesh", texPath);
|
||||||
|
|
||||||
|
// scene.txt (from the asset bundle): line 1 albedoCount, line 2
|
||||||
|
// meshCount, then per-mesh albedo index. We render every mesh group —
|
||||||
|
// unlike the interactive Sponza example, which is single-material on
|
||||||
|
// native because of the hit-shader dynamic-heap-index driver fault.
|
||||||
|
// SponzaBench's hit shader samples no texture, so it can build the
|
||||||
|
// full multi-mesh atrium as one multi-instance TLAS.
|
||||||
|
std::uint32_t albedoCount = 0, meshCount = 0;
|
||||||
|
{
|
||||||
|
std::ifstream manifest("scene.txt");
|
||||||
|
if (!manifest) { std::println(std::cerr, "[SponzaBench] missing scene.txt"); std::abort(); }
|
||||||
|
manifest >> albedoCount >> meshCount;
|
||||||
|
}
|
||||||
|
// BENCH_MESHES caps the number of mesh groups loaded (default: all).
|
||||||
|
// A small cap (e.g. 1) shrinks GPU work so per-frame CPU cost dominates;
|
||||||
|
// the full scene is GPU (traversal) bound.
|
||||||
|
meshCount = std::min<std::uint32_t>(meshCount, EnvSize("BENCH_MESHES", meshCount));
|
||||||
|
|
||||||
|
CompressedTextureAsset loadedTex = LoadCompressedTexture(texPath);
|
||||||
|
std::println("[SponzaBench] scene: {} meshes, {} albedos, {}x{} probe albedo",
|
||||||
|
meshCount, albedoCount, loadedTex.sizeX, loadedTex.sizeY);
|
||||||
|
|
||||||
|
Device::Initialize();
|
||||||
|
Window window(1280, 720, "SponzaBench");
|
||||||
|
VkCommandBuffer cmd = window.StartInit();
|
||||||
|
|
||||||
|
DescriptorHeapVulkan descriptorHeap;
|
||||||
|
descriptorHeap.Initialize(/*images*/ 2, /*buffers*/ 1, /*samplers*/ 0);
|
||||||
|
|
||||||
|
VkSpecializationMapEntry raygenEntry = { .constantID = 0, .offset = 0, .size = sizeof(std::uint16_t) };
|
||||||
|
VkSpecializationInfo raygenSpec = {
|
||||||
|
.mapEntryCount = 1, .pMapEntries = &raygenEntry,
|
||||||
|
.dataSize = sizeof(std::uint16_t), .pData = &descriptorHeap.bufferStartElement,
|
||||||
|
};
|
||||||
|
|
||||||
|
auto imgSlots = descriptorHeap.AllocateImageSlots(1);
|
||||||
|
auto bufSlots = descriptorHeap.AllocateBufferSlots(1);
|
||||||
|
|
||||||
|
std::array<VulkanShader, 3> shaders {{
|
||||||
|
{ "raygen.spv", "main", VK_SHADER_STAGE_RAYGEN_BIT_KHR, &raygenSpec },
|
||||||
|
{ "miss.spv", "main", VK_SHADER_STAGE_MISS_BIT_KHR, nullptr },
|
||||||
|
{ "closesthit.spv", "main", VK_SHADER_STAGE_CLOSEST_HIT_BIT_KHR, nullptr },
|
||||||
|
}};
|
||||||
|
ShaderBindingTableVulkan shaderTable;
|
||||||
|
shaderTable.Init(shaders);
|
||||||
|
|
||||||
|
std::array<VkRayTracingShaderGroupCreateInfoKHR, 1> raygenGroups {{ {
|
||||||
|
.sType = VK_STRUCTURE_TYPE_RAY_TRACING_SHADER_GROUP_CREATE_INFO_KHR,
|
||||||
|
.type = VK_RAY_TRACING_SHADER_GROUP_TYPE_GENERAL_KHR,
|
||||||
|
.generalShader = 0, .closestHitShader = VK_SHADER_UNUSED_KHR,
|
||||||
|
.anyHitShader = VK_SHADER_UNUSED_KHR, .intersectionShader = VK_SHADER_UNUSED_KHR,
|
||||||
|
} }};
|
||||||
|
std::array<VkRayTracingShaderGroupCreateInfoKHR, 1> missGroups {{ {
|
||||||
|
.sType = VK_STRUCTURE_TYPE_RAY_TRACING_SHADER_GROUP_CREATE_INFO_KHR,
|
||||||
|
.type = VK_RAY_TRACING_SHADER_GROUP_TYPE_GENERAL_KHR,
|
||||||
|
.generalShader = 1, .closestHitShader = VK_SHADER_UNUSED_KHR,
|
||||||
|
.anyHitShader = VK_SHADER_UNUSED_KHR, .intersectionShader = VK_SHADER_UNUSED_KHR,
|
||||||
|
} }};
|
||||||
|
std::array<VkRayTracingShaderGroupCreateInfoKHR, 1> hitGroups {{ {
|
||||||
|
.sType = VK_STRUCTURE_TYPE_RAY_TRACING_SHADER_GROUP_CREATE_INFO_KHR,
|
||||||
|
.type = VK_RAY_TRACING_SHADER_GROUP_TYPE_TRIANGLES_HIT_GROUP_KHR,
|
||||||
|
.generalShader = VK_SHADER_UNUSED_KHR, .closestHitShader = 2,
|
||||||
|
.anyHitShader = VK_SHADER_UNUSED_KHR, .intersectionShader = VK_SHADER_UNUSED_KHR,
|
||||||
|
} }};
|
||||||
|
|
||||||
|
PipelineRTVulkan pipeline;
|
||||||
|
pipeline.Init(cmd, raygenGroups, missGroups, hitGroups, shaderTable);
|
||||||
|
|
||||||
|
// One Mesh + RenderingElement3D per mesh group → BLAS each, then a
|
||||||
|
// single multi-instance TLAS. Pointers handed to RenderingElement3D
|
||||||
|
// live in the static vectors, so reserve up-front: a reallocation
|
||||||
|
// would dangle every registered element.
|
||||||
|
static std::vector<Mesh> meshes;
|
||||||
|
static std::vector<RenderingElement3D> renderers;
|
||||||
|
meshes.reserve(meshCount);
|
||||||
|
renderers.reserve(meshCount);
|
||||||
|
std::uint64_t totalVerts = 0, totalIdx = 0;
|
||||||
|
for (std::uint32_t i = 0; i < meshCount; ++i) {
|
||||||
|
const fs::path mp = std::format("mesh_{}.cmesh", i);
|
||||||
|
if (!fs::exists(mp)) continue;
|
||||||
|
CompressedMeshAsset loaded = LoadCompressedMesh(mp);
|
||||||
|
totalVerts += loaded.vertexCount;
|
||||||
|
totalIdx += loaded.indexCount;
|
||||||
|
meshes.emplace_back();
|
||||||
|
meshes.back().Build(loaded, cmd);
|
||||||
|
|
||||||
|
renderers.emplace_back();
|
||||||
|
RenderingElement3D& r = renderers.back();
|
||||||
|
r.instance = {
|
||||||
|
.transform = {},
|
||||||
|
.instanceCustomIndex = 0,
|
||||||
|
.mask = 0xFF,
|
||||||
|
.instanceShaderBindingTableRecordOffset = 0,
|
||||||
|
.flags = VK_GEOMETRY_INSTANCE_FORCE_OPAQUE_BIT_KHR,
|
||||||
|
.accelerationStructureReference = meshes.back().blasAddr,
|
||||||
|
};
|
||||||
|
MatrixRowMajor<float, 4, 3, 1>::Identity()
|
||||||
|
.Store(reinterpret_cast<float*>(r.instance.transform.matrix));
|
||||||
|
RenderingElement3D::Add(&r);
|
||||||
|
}
|
||||||
|
std::println("[SponzaBench] built {} BLAS, {} verts, {} idx total",
|
||||||
|
renderers.size(), totalVerts, totalIdx);
|
||||||
|
|
||||||
|
Image2D<RGBA8> albedo;
|
||||||
|
albedo.Create(loadedTex.sizeX, loadedTex.sizeY, /*mipLevels*/ 1, cmd,
|
||||||
|
VK_FORMAT_R8G8B8A8_UNORM,
|
||||||
|
VK_IMAGE_USAGE_TRANSFER_DST_BIT | VK_IMAGE_USAGE_SAMPLED_BIT,
|
||||||
|
VK_IMAGE_LAYOUT_SHADER_READ_ONLY_OPTIMAL);
|
||||||
|
albedo.Update(loadedTex, cmd, VK_IMAGE_LAYOUT_SHADER_READ_ONLY_OPTIMAL);
|
||||||
|
// The albedo is uploaded to keep the texture-decompress + staging
|
||||||
|
// setup path in the measured "setup" window, but is not bound — the
|
||||||
|
// closest-hit shades from barycentrics, so only the TLAS + per-frame
|
||||||
|
// output image need descriptors.
|
||||||
|
|
||||||
|
for (std::uint32_t f = 0; f < Window::numFrames; ++f)
|
||||||
|
RenderingElement3D::BuildTLAS(cmd, f);
|
||||||
|
|
||||||
|
window.FinishInit();
|
||||||
|
|
||||||
|
VkDeviceAddressRangeKHR tlasRanges[Window::numFrames];
|
||||||
|
VkImageDescriptorInfoEXT outImgInfos[Window::numFrames];
|
||||||
|
for (std::uint32_t f = 0; f < Window::numFrames; ++f) {
|
||||||
|
tlasRanges[f] = { .address = RenderingElement3D::tlases[f].address };
|
||||||
|
outImgInfos[f] = {
|
||||||
|
.sType = VK_STRUCTURE_TYPE_IMAGE_DESCRIPTOR_INFO_EXT,
|
||||||
|
.pView = &window.imageViews[f],
|
||||||
|
.layout = VK_IMAGE_LAYOUT_GENERAL,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
std::vector<VkResourceDescriptorInfoEXT> resources;
|
||||||
|
std::vector<VkHostAddressRangeEXT> destinations;
|
||||||
|
resources.reserve(Window::numFrames * 2);
|
||||||
|
destinations.reserve(Window::numFrames * 2);
|
||||||
|
for (std::uint32_t f = 0; f < Window::numFrames; ++f) {
|
||||||
|
resources.push_back({
|
||||||
|
.sType = VK_STRUCTURE_TYPE_RESOURCE_DESCRIPTOR_INFO_EXT,
|
||||||
|
.type = VK_DESCRIPTOR_TYPE_ACCELERATION_STRUCTURE_KHR,
|
||||||
|
.data = { .pAddressRange = &tlasRanges[f] },
|
||||||
|
});
|
||||||
|
destinations.push_back({
|
||||||
|
.address = descriptorHeap.resourceHeap[f].value
|
||||||
|
+ descriptorHeap.BufferByteOffset(bufSlots.firstElement),
|
||||||
|
.size = Device::descriptorHeapProperties.bufferDescriptorSize,
|
||||||
|
});
|
||||||
|
resources.push_back({
|
||||||
|
.sType = VK_STRUCTURE_TYPE_RESOURCE_DESCRIPTOR_INFO_EXT,
|
||||||
|
.type = VK_DESCRIPTOR_TYPE_STORAGE_IMAGE,
|
||||||
|
.data = { .pImage = &outImgInfos[f] },
|
||||||
|
});
|
||||||
|
destinations.push_back({
|
||||||
|
.address = descriptorHeap.resourceHeap[f].value
|
||||||
|
+ descriptorHeap.ImageByteOffset(imgSlots.firstElement),
|
||||||
|
.size = Device::descriptorHeapProperties.imageDescriptorSize,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
Device::vkWriteResourceDescriptorsEXT(Device::device,
|
||||||
|
static_cast<std::uint32_t>(resources.size()),
|
||||||
|
resources.data(), destinations.data());
|
||||||
|
for (std::uint32_t f = 0; f < Window::numFrames; ++f) {
|
||||||
|
descriptorHeap.resourceHeap[f].FlushDevice();
|
||||||
|
}
|
||||||
|
|
||||||
|
window.descriptorHeap = &descriptorHeap;
|
||||||
|
RTPass rtPass(&pipeline);
|
||||||
|
window.passes.push_back(&rtPass);
|
||||||
|
|
||||||
|
// ── Setup complete. Everything above is the per-run, one-time cost. ──
|
||||||
|
const auto tSetupDone = Clock::now();
|
||||||
|
const double setupMs = Ms(tSetupDone - tProcStart);
|
||||||
|
|
||||||
|
// ── Warmup: let the pipeline fill (frames-in-flight), caches warm. ──
|
||||||
|
for (std::size_t i = 0; i < warmup; ++i) window.Render();
|
||||||
|
|
||||||
|
// ── Measured loop: per-frame wall-clock around Render(). ────────────
|
||||||
|
std::vector<double> frameMs;
|
||||||
|
frameMs.reserve(frames);
|
||||||
|
const auto tLoopStart = Clock::now();
|
||||||
|
for (std::size_t i = 0; i < frames; ++i) {
|
||||||
|
const auto a = Clock::now();
|
||||||
|
window.Render();
|
||||||
|
const auto b = Clock::now();
|
||||||
|
frameMs.push_back(Ms(b - a));
|
||||||
|
}
|
||||||
|
const auto tLoopEnd = Clock::now();
|
||||||
|
// Drain so the loop's wall-clock isn't charged the GPU tail of the
|
||||||
|
// last few in-flight frames inconsistently across runs.
|
||||||
|
vkQueueWaitIdle(Device::queue);
|
||||||
|
|
||||||
|
const long peakRSS = PeakRSSkB();
|
||||||
|
|
||||||
|
// ── Stats. throughput uses the loop wall-clock (steady-state FPS);
|
||||||
|
// per-frame percentiles use the sorted per-frame samples. ─────────
|
||||||
|
const double loopWallMs = Ms(tLoopEnd - tLoopStart);
|
||||||
|
std::sort(frameMs.begin(), frameMs.end());
|
||||||
|
auto pct = [&](double p) {
|
||||||
|
if (frameMs.empty()) return 0.0;
|
||||||
|
std::size_t idx = static_cast<std::size_t>(p * (frameMs.size() - 1));
|
||||||
|
return frameMs[idx];
|
||||||
|
};
|
||||||
|
double sum = 0.0; for (double v : frameMs) sum += v;
|
||||||
|
const double meanMs = frameMs.empty() ? 0.0 : sum / frameMs.size();
|
||||||
|
|
||||||
|
std::println("");
|
||||||
|
std::println("BENCH setup_ms {:.3f}", setupMs);
|
||||||
|
std::println("BENCH frames {}", frames);
|
||||||
|
std::println("BENCH loop_wall_ms {:.3f}", loopWallMs);
|
||||||
|
std::println("BENCH fps {:.1f}", frames / (loopWallMs / 1000.0));
|
||||||
|
std::println("BENCH frame_mean_ms {:.4f}", meanMs);
|
||||||
|
std::println("BENCH frame_min_ms {:.4f}", pct(0.0));
|
||||||
|
std::println("BENCH frame_p50_ms {:.4f}", pct(0.50));
|
||||||
|
std::println("BENCH frame_p99_ms {:.4f}", pct(0.99));
|
||||||
|
std::println("BENCH frame_max_ms {:.4f}", pct(1.0));
|
||||||
|
std::println("BENCH peak_rss_kb {}", peakRSS);
|
||||||
|
|
||||||
|
// The interactive Sponza example never tears Device/Window down (it
|
||||||
|
// loops in StartSync forever); doing so here races GPU-still-in-flight
|
||||||
|
// resources at static-destruction time. The measurement is already
|
||||||
|
// printed, so skip destructors entirely with a hard exit.
|
||||||
|
//
|
||||||
|
// BENCH_CLEAN_EXIT=1 instead exits via std::exit, which runs atexit
|
||||||
|
// handlers (on builds that have the disk pipeline cache, #69, this is
|
||||||
|
// where it is written) before the destructor teardown. Used once to
|
||||||
|
// seed pipeline_cache.bin for a warm-cache setup measurement; the
|
||||||
|
// subsequent destructor race is irrelevant since the data is flushed.
|
||||||
|
std::cout.flush();
|
||||||
|
if (std::getenv("BENCH_CLEAN_EXIT")) std::exit(0);
|
||||||
|
std::_Exit(0);
|
||||||
|
}
|
||||||
11
examples/SponzaBench/miss.glsl
Normal file
11
examples/SponzaBench/miss.glsl
Normal file
|
|
@ -0,0 +1,11 @@
|
||||||
|
#version 460
|
||||||
|
#extension GL_EXT_ray_tracing : enable
|
||||||
|
|
||||||
|
layout(location = 0) rayPayloadInEXT vec3 hitValue;
|
||||||
|
|
||||||
|
void main() {
|
||||||
|
// Soft sky gradient based on ray direction Y. The actual ray dir
|
||||||
|
// isn't accessible without an extra payload field; use a flat warm
|
||||||
|
// tone that matches Sponza's interior lighting.
|
||||||
|
hitValue = vec3(0.10, 0.08, 0.06);
|
||||||
|
}
|
||||||
60
examples/SponzaBench/project.cpp
Normal file
60
examples/SponzaBench/project.cpp
Normal file
|
|
@ -0,0 +1,60 @@
|
||||||
|
import std;
|
||||||
|
import Crafter.Build;
|
||||||
|
namespace fs = std::filesystem;
|
||||||
|
using namespace Crafter;
|
||||||
|
|
||||||
|
// SponzaBench — a headless benchmark harness around the native Sponza
|
||||||
|
// ray-tracing scene. Same asset bundle and shaders as examples/Sponza,
|
||||||
|
// but main.cpp times setup + a fixed measured frame loop instead of
|
||||||
|
// opening an interactive window. Native (Vulkan) only — the point is to
|
||||||
|
// measure the native renderer's per-frame and setup cost across library
|
||||||
|
// revisions (issue #155: net perf gain from #40 to now).
|
||||||
|
constexpr std::string_view kSponzaGitUrl = "https://github.com/jimmiebergmann/Sponza.git";
|
||||||
|
constexpr std::string_view kSponzaCommitSHA = "222338979d32f4f4818466291bdbc29f192b86ba";
|
||||||
|
constexpr std::uint16_t kAlbedoSize = 1024u;
|
||||||
|
|
||||||
|
extern "C" Configuration CrafterBuildProject(std::span<const std::string_view> args) {
|
||||||
|
std::vector<std::string> graphicsArgs(args.begin(), args.end());
|
||||||
|
Configuration* graphics = LocalProject({
|
||||||
|
.projectFile = "../../project.cpp",
|
||||||
|
.args = graphicsArgs,
|
||||||
|
});
|
||||||
|
|
||||||
|
Configuration cfg;
|
||||||
|
cfg.path = "./";
|
||||||
|
cfg.name = "SponzaBench";
|
||||||
|
cfg.outputName = "SponzaBench";
|
||||||
|
cfg.type = ConfigurationType::Executable;
|
||||||
|
ApplyStandardArgs(cfg, args);
|
||||||
|
cfg.dependencies = { graphics };
|
||||||
|
|
||||||
|
std::array<fs::path, 0> ifaces = {};
|
||||||
|
std::array<fs::path, 1> impls = { "main" };
|
||||||
|
cfg.GetInterfacesAndImplementations(ifaces, impls);
|
||||||
|
|
||||||
|
fs::path sponzaRoot = GitFetch({
|
||||||
|
.url = std::string(kSponzaGitUrl),
|
||||||
|
.commit = std::string(kSponzaCommitSHA),
|
||||||
|
});
|
||||||
|
std::string bundleKey = std::format("{}|{}", kSponzaCommitSHA, kAlbedoSize);
|
||||||
|
auto bundleHash = std::hash<std::string>{}(bundleKey);
|
||||||
|
fs::path bundleDir = fs::path("build") / std::format("sponza-bundle-{:016x}", bundleHash);
|
||||||
|
|
||||||
|
if (auto err = BuildOBJBundle(
|
||||||
|
sponzaRoot / "sponza.obj",
|
||||||
|
sponzaRoot / "sponza.mtl",
|
||||||
|
bundleDir,
|
||||||
|
kAlbedoSize); !err.empty()) {
|
||||||
|
std::println(std::cerr, "Sponza bundle error: {}", err);
|
||||||
|
std::exit(1);
|
||||||
|
}
|
||||||
|
|
||||||
|
for (const auto& entry : fs::directory_iterator(bundleDir)) {
|
||||||
|
if (entry.is_regular_file()) cfg.files.push_back(entry.path());
|
||||||
|
}
|
||||||
|
|
||||||
|
cfg.shaders.emplace_back(fs::path("raygen.glsl"), std::string("main"), ShaderType::RayGen);
|
||||||
|
cfg.shaders.emplace_back(fs::path("closesthit.glsl"), std::string("main"), ShaderType::ClosestHit);
|
||||||
|
cfg.shaders.emplace_back(fs::path("miss.glsl"), std::string("main"), ShaderType::Miss);
|
||||||
|
return cfg;
|
||||||
|
}
|
||||||
52
examples/SponzaBench/raygen.glsl
Normal file
52
examples/SponzaBench/raygen.glsl
Normal file
|
|
@ -0,0 +1,52 @@
|
||||||
|
#version 460
|
||||||
|
#extension GL_EXT_ray_tracing : enable
|
||||||
|
#extension GL_EXT_shader_image_load_formatted : enable
|
||||||
|
#extension GL_EXT_shader_explicit_arithmetic_types_int16 : enable
|
||||||
|
#extension GL_EXT_descriptor_heap : enable
|
||||||
|
#extension GL_EXT_nonuniform_qualifier : enable
|
||||||
|
|
||||||
|
// Specialization constant set from descriptorHeap.bufferStartElement —
|
||||||
|
// shared with closesthit.glsl. The TLAS lives at descriptor_heap slot
|
||||||
|
// `bufferStart` (it's an SSBO-typed entry), the per-frame output image
|
||||||
|
// at heap slot 0.
|
||||||
|
layout(constant_id = 0) const uint16_t bufferStart = 0us;
|
||||||
|
layout(descriptor_heap) uniform accelerationStructureEXT topLevelAS[];
|
||||||
|
layout(descriptor_heap) uniform writeonly image2D image[];
|
||||||
|
|
||||||
|
layout(location = 0) rayPayloadEXT vec3 hitValue;
|
||||||
|
|
||||||
|
void main() {
|
||||||
|
uvec2 pixel = gl_LaunchIDEXT.xy;
|
||||||
|
uvec2 resolution = gl_LaunchSizeEXT.xy;
|
||||||
|
vec2 uv = (vec2(pixel) + 0.5) / vec2(resolution);
|
||||||
|
vec2 ndc = uv * 2.0 - 1.0;
|
||||||
|
|
||||||
|
// Camera positioned to look down the Sponza atrium axis. Sponza-OBJ
|
||||||
|
// from McGuire's archive is roughly 30 units wide × 13 tall × 18 deep,
|
||||||
|
// axis-aligned, with the floor near y=0 and the atrium centered on
|
||||||
|
// origin. -X faces the long end, so we sit inside looking +X.
|
||||||
|
vec3 origin = vec3(-10.0, 5.0, 0.0);
|
||||||
|
float aspect = float(resolution.x) / float(resolution.y);
|
||||||
|
float fov = radians(70.0);
|
||||||
|
float tanHalf = tan(fov * 0.5);
|
||||||
|
vec3 direction = normalize(vec3(
|
||||||
|
ndc.x * aspect * tanHalf,
|
||||||
|
-ndc.y * tanHalf,
|
||||||
|
1.0));
|
||||||
|
|
||||||
|
// Rotate +Z forward → +X forward (90° about Y).
|
||||||
|
direction = vec3(direction.z, direction.y, -direction.x);
|
||||||
|
|
||||||
|
traceRayEXT(
|
||||||
|
topLevelAS[bufferStart],
|
||||||
|
gl_RayFlagsNoneEXT,
|
||||||
|
0xff,
|
||||||
|
0, 0, 0,
|
||||||
|
origin,
|
||||||
|
0.001,
|
||||||
|
direction,
|
||||||
|
10000.0,
|
||||||
|
0);
|
||||||
|
|
||||||
|
imageStore(image[0], ivec2(pixel), vec4(hitValue, 1.0));
|
||||||
|
}
|
||||||
40
examples/SponzaBench/run-bench.sh
Executable file
40
examples/SponzaBench/run-bench.sh
Executable file
|
|
@ -0,0 +1,40 @@
|
||||||
|
#!/usr/bin/env bash
|
||||||
|
# Repeated-measurement harness for SponzaBench. Runs the binary in $BINDIR
|
||||||
|
# REPS times for a given mesh cap and reports the median (and min/max) of
|
||||||
|
# each BENCH metric. Validation is force-disabled via the loader (it adds
|
||||||
|
# huge, version-dependent overhead) and present mode is uncapped so frame
|
||||||
|
# time reflects real work rather than the compositor's vblank.
|
||||||
|
#
|
||||||
|
# Usage: run-bench.sh <bindir> <label> <meshes> <reps> [--cold]
|
||||||
|
# --cold deletes pipeline_cache.bin before EACH run (cold-compile path).
|
||||||
|
set -u
|
||||||
|
BINDIR="$1"; LABEL="$2"; MESHES="$3"; REPS="$4"; COLD="${5:-}"
|
||||||
|
cd "$BINDIR" || exit 1
|
||||||
|
|
||||||
|
declare -A vals
|
||||||
|
metrics="setup_ms fps frame_mean_ms frame_p50_ms frame_p99_ms peak_rss_kb"
|
||||||
|
for m in $metrics; do vals[$m]=""; done
|
||||||
|
|
||||||
|
for ((i=0;i<REPS;i++)); do
|
||||||
|
[ "$COLD" = "--cold" ] && rm -f pipeline_cache.bin
|
||||||
|
out=$(VK_LOADER_LAYERS_DISABLE='~all~' CRAFTER_PRESENT_IMMEDIATE=1 \
|
||||||
|
BENCH_WARMUP=200 BENCH_FRAMES=2000 BENCH_MESHES="$MESHES" \
|
||||||
|
./SponzaBench 2>/dev/null)
|
||||||
|
for m in $metrics; do
|
||||||
|
v=$(echo "$out" | awk -v k="$m" '$1=="BENCH" && $2==k {print $3}')
|
||||||
|
vals[$m]="${vals[$m]} $v"
|
||||||
|
done
|
||||||
|
done
|
||||||
|
|
||||||
|
med() { tr ' ' '\n' | grep -v '^$' | sort -n | awk '{a[NR]=$1} END{print a[int((NR+1)/2)]}'; }
|
||||||
|
lo() { tr ' ' '\n' | grep -v '^$' | sort -n | head -1; }
|
||||||
|
hi() { tr ' ' '\n' | grep -v '^$' | sort -n | tail -1; }
|
||||||
|
|
||||||
|
echo "## $LABEL (meshes=$MESHES reps=$REPS ${COLD})"
|
||||||
|
printf "%-15s %12s %12s %12s\n" metric median min max
|
||||||
|
for m in $metrics; do
|
||||||
|
echo -n "$(printf '%-15s' "$m") "
|
||||||
|
printf "%12s %12s %12s\n" \
|
||||||
|
"$(echo "${vals[$m]}" | med)" "$(echo "${vals[$m]}" | lo)" "$(echo "${vals[$m]}" | hi)"
|
||||||
|
done
|
||||||
|
echo
|
||||||
|
|
@ -1072,8 +1072,24 @@ void Window::CreateSwapchain()
|
||||||
swapchainCI.imageSharingMode = VK_SHARING_MODE_EXCLUSIVE;
|
swapchainCI.imageSharingMode = VK_SHARING_MODE_EXCLUSIVE;
|
||||||
swapchainCI.queueFamilyIndexCount = 0;
|
swapchainCI.queueFamilyIndexCount = 0;
|
||||||
// VK_PRESENT_MODE_FIFO_KHR ("v-sync") is guaranteed available per spec, so
|
// VK_PRESENT_MODE_FIFO_KHR ("v-sync") is guaranteed available per spec, so
|
||||||
// it is used unconditionally — no present-mode enumeration is needed.
|
// it is the default. For benchmarking we want frame pacing uncapped from
|
||||||
|
// the display refresh, so when CRAFTER_PRESENT_IMMEDIATE is set in the
|
||||||
|
// environment we enumerate the surface's present modes and prefer
|
||||||
|
// IMMEDIATE (then MAILBOX) when offered, falling back to FIFO otherwise.
|
||||||
swapchainCI.presentMode = VK_PRESENT_MODE_FIFO_KHR;
|
swapchainCI.presentMode = VK_PRESENT_MODE_FIFO_KHR;
|
||||||
|
if (std::getenv("CRAFTER_PRESENT_IMMEDIATE")) {
|
||||||
|
std::uint32_t presentModeCount = 0;
|
||||||
|
vkGetPhysicalDeviceSurfacePresentModesKHR(Device::physDevice, vulkanSurface, &presentModeCount, nullptr);
|
||||||
|
std::vector<VkPresentModeKHR> presentModes(presentModeCount);
|
||||||
|
vkGetPhysicalDeviceSurfacePresentModesKHR(Device::physDevice, vulkanSurface, &presentModeCount, presentModes.data());
|
||||||
|
bool haveImmediate = false, haveMailbox = false;
|
||||||
|
for (VkPresentModeKHR m : presentModes) {
|
||||||
|
if (m == VK_PRESENT_MODE_IMMEDIATE_KHR) haveImmediate = true;
|
||||||
|
if (m == VK_PRESENT_MODE_MAILBOX_KHR) haveMailbox = true;
|
||||||
|
}
|
||||||
|
if (haveImmediate) swapchainCI.presentMode = VK_PRESENT_MODE_IMMEDIATE_KHR;
|
||||||
|
else if (haveMailbox) swapchainCI.presentMode = VK_PRESENT_MODE_MAILBOX_KHR;
|
||||||
|
}
|
||||||
// Setting oldSwapChain to the saved handle of the previous swapchain aids in resource reuse and makes sure that we can still present already acquired images
|
// Setting oldSwapChain to the saved handle of the previous swapchain aids in resource reuse and makes sure that we can still present already acquired images
|
||||||
swapchainCI.oldSwapchain = oldSwapchain;
|
swapchainCI.oldSwapchain = oldSwapchain;
|
||||||
// Setting clipped to VK_TRUE allows the implementation to discard rendering outside of the surface area
|
// Setting clipped to VK_TRUE allows the implementation to discard rendering outside of the surface area
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue