perf(ui): hoist loop-invariant per-item math out of the per-pixel loop (#125)

The cooperative-load section already streams each item into shared memory
once per workgroup, but the per-pixel inner loop then recomputed values
that are constant for the whole item/glyph. The compiler can't hoist them
itself — they read shared memory at a varying index.

Precompute them once per item at load time into new shared slots:

  - ui-images / ui-text / ui-fused (images+text phases): invRectSize =
    1.0/rect.zw, turning the per-pixel `(sp-rect.xy)/rect.zw` vec2 divide
    into a multiply.
  - ui-text / ui-fused (text phase): the SDF AA `band` (a vec2 divide +
    two maxes depending only on the glyph's uv span and rect size).

In ui-fused the new slots (s_inv, s_band) are reused across phases like
the existing s_v* scratch, so the LDS bump is per-category, not summed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
catbot 2026-06-18 13:22:07 +00:00
commit ac9180076c
3 changed files with 45 additions and 20 deletions

View file

@ -16,6 +16,11 @@ shared vec4 s_rect[UI_CHUNK];
shared vec4 s_uv[UI_CHUNK];
shared vec4 s_tint[UI_CHUNK];
shared uvec4 s_slots[UI_CHUNK];
// 1.0/rect.zw, precomputed once per item at cooperative-load time so the
// per-pixel inner loop multiplies instead of recomputing a vec2 reciprocal of
// the (loop-invariant, but shared-mem + varying-index, so non-hoistable by the
// compiler) rect size for every pixel × item.
shared vec2 s_invRectSize[UI_CHUNK];
shared uint s_keep[UI_CHUNK];
shared uint s_order[UI_CHUNK];
shared uint s_count;
@ -43,6 +48,7 @@ void main() {
s_uv[lid] = uiImageHeap[pc.hdr.itemBuffer].items[idx].uv;
s_tint[lid] = uiImageHeap[pc.hdr.itemBuffer].items[idx].tint;
s_slots[lid] = uiImageHeap[pc.hdr.itemBuffer].items[idx].slots;
s_invRectSize[lid] = 1.0 / s_rect[lid].zw;
keep = uiAabbOverlapsTile(s_rect[lid].xy, s_rect[lid].xy + s_rect[lid].zw,
tileMin, tileMax);
}
@ -67,7 +73,7 @@ void main() {
if (sp.x < lo.x || sp.y < lo.y) continue;
if (sp.x >= hi.x || sp.y >= hi.y) continue;
vec2 t = (sp - s_rect[c].xy) / s_rect[c].zw;
vec2 t = (sp - s_rect[c].xy) * s_invRectSize[c];
vec2 uv = mix(s_uv[c].xy, s_uv[c].zw, t);
uint texSlot = s_slots[c].x;