Commit Graph
9967 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 aa38fafd77 DrawEngine: Don't drop large batches of points, lines and rectangles
Software transform expands each point, line and rectangle to four
vertices, and gave up on the whole draw when that didn't fit
VERTEX_BUFFER_MAX. Batching only counted input vertices, so a batch over
16384 points (or 32768 line or rectangle vertices) vanished silently,
whether it came from one PRIM or several merged ones. Count the expanded
vertices when batching, and submit a PRIM too big on its own in parts.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 12:04:53 -06:00
Henrik Rydgård ae460c9e1a Merge pull request #22390 from hrydgard/gpu-review-leftovers
Claude code review of GPU: Fix leftover findings
2026-09-29 15:43:31 -06:00
Henrik RydgårdandClaude Opus 5.5 b9501df529 TextureReplacer: Fix stale and shared lookup results
- Reloading the ini clears the per-key lookup caches, which could keep
  saying "no replacement" for textures the new ini replaces.
- With ignoreAddress, hash ranges were skipped when sizing the
  replacement, though ComputeHash applies them. cache_ is now keyed by
  the full key, so its lookups hit too.
- Textures sharing files but differing in size, hash range or filtering
  no longer share one ReplacedTexture (the first one's settings won).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:28:21 -06:00
Henrik RydgårdandClaude Opus 5.5 b8789af997 GPU: Draw frames displayed from RAM in non-buffered mode
They were uploaded but never drawn, and the block transfer hack drew
whatever source was left over. Draw them straight into the backbuffer
pass, without post shaders, which would need to bind their own targets.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:17:01 -06:00
Henrik RydgårdandClaude Opus 5.5 7e553b0732 GE debugger: Lock when adding command breakpoints
Debuggers add them from their own threads, racing ClearTempBreakpoints
on the emu thread.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:17:00 -06:00
Henrik RydgårdandClaude Opus 5.5 035c98d14f D3D11: Re-read the device and context on DeviceRestore
The restored draw context can be a new device, so the cached pointers
could go stale.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:16:59 -06:00
Henrik RydgårdandClaude Opus 5.5 ba37f327b0 D3D11: Check Map() results
Map fails after device removal, leaving pData garbage. Skip the upload or
draw instead of writing through it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:16:59 -06:00
Henrik RydgårdandClaude Opus 5.5 bc581349fd GPU: Install the draw engines' invalidation callback from BeginFrame
The draw engine is created on the loader thread while the UI thread may
already be rendering and invoking the callback.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:16:42 -06:00
Henrik RydgårdandClaude Opus 5.5 759b494b6a GPU: Check post shaders once per host frame, after resizes
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:16:41 -06:00
Henrik RydgårdandClaude Opus 5.5 4749872f2e Vulkan: Clear a texture level when hardware scaling fails
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:16:41 -06:00
Henrik RydgårdandClaude Opus 5.5 4c2086aac5 GPU: Keep spline/bezier tessellation from reaching zero
Reduce only the larger factor when over the vertex limit, and stop at 1.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:16:40 -06:00
Henrik RydgårdandClaude Opus 5.5 9d3022642b SoftGPU: Fix bin queue bookkeeping and dependency tracking
- Flush() on an empty queue now trims the state and CLUT rings, since its
  callers flush because one is full and push right after.
- BinQueue::Full() uses >=, so an overshoot can't go unnoticed.
- IsExactSelfRender compares against the target the queued draws were
  binned for, not gstate, which already has the next one during a flush.
- A depth test without depth writes marks the depth buffer as read.
- The DarkStalkers untextured sprite recomputes the binner state around it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:16:40 -06:00
Henrik RydgårdandClaude Opus 5.5 7de7b2d76b TextureCache: Use XXH3 instead of the quick hash for texture data
With xxhash 0.8.4, XXH3 is faster than StableQuickTexHash on ARM64
(about 37 vs 24 GB/s on a Snapdragon X), and its scalar path, which
RISC-V builds get, is about as fast as the quick hash's. It doesn't
collide the way the quick hash does (#8249).

Texture replacement still uses its own hash setting, so texture packs
are unaffected.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:10:28 -06:00
Henrik RydgårdandClaude Opus 5.5 74553bfa1a D3D11: Note that the state object caches are deliberately never trimmed
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 12:43:12 -06:00
Henrik RydgårdandClaude Opus 5.5 93f57a4f38 GPU: Delete copy operations on classes that own resources
These own GPU objects, memory or refcounts in their destructors (or assert
there that they were torn down), so a copy would double-free. Nothing copies
them today; this keeps it that way. The manager base classes cover every
backend's subclass.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 39049a67fd GLES: Free textures on device lost, and unsubmitted step data at exit
- The texture and fragment test caches dropped their GLRTexture objects on
  DeviceLost without queueing them for deletion, leaking them on every
  Android background/resume. The deleter already skips the GL calls when
  the context is gone.
- GLRenderManager::ThreadEnd cleared unsubmitted init and render steps
  without freeing the data they own. Run them through the dry run instead,
  which now also frees stereo matrices and shader code.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 b72927bbeb Texture replacement: Plug leaks on error paths
- Release the file reference when loading a level fails or finds nothing,
  since only a loaded level takes ownership of it.
- Free the PNG image when the size changed since the header was read.
- Delete the VFS when a pack without an ini has no hash-named textures.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 e108e41675 GPU: Fix assorted leaks, null derefs and small rendering bugs
- Put the anisotropy level in the sampler key, so changing it applies on
  Vulkan and D3D11.
- Release CLUT textures at shutdown on GLES and D3D11.
- Fix the depth readback viewport, which squeezed the image whenever the
  read rectangle was smaller than the fbo.
- Test the computed depth, not the unset result, in the equal-depth clear
  check.
- Don't read back a CLUT from a framebuffer without an fbo.
- Tolerate null entries when releasing post-shader objects and CLUT
  textures after a failed creation.
- ImGe: Don't crash on a framebuffer without an fbo, or on GetVFB under the
  software renderer.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 71bc3187db Vulkan: Reload the shader cache after a device restore
DeviceLost saves the cache and clears everything, and the next save wrote
back only what was drawn since, so each Android background/resume cycle
shrank the cache.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 71210f3fa2 GE recorder: Don't let a second request hijack a recording
RecordNextFrame now refuses while a recording is active and during frame
dump playback (which asserted on the next replay). The callback handoff to
the CPU thread is locked, and a failed file open no longer crashes.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 d6f8615d2a SoftGPU: Fix block transfer and self-render overlap tracking
- The block transfer overlap check passed the stride in pixels where bytes
  are expected, so it only covered part of the rectangle.
- A selfrender/selfdepth flush in UpdateState dropped the current draw's
  pending writes and reads, so later transfers didn't wait for it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 507bb0801b GE: Complete lists dropped on error, flush before immediate draws
- A list dropped for a bad pc or a GE error stayed RUNNING: its ID was never
  freed and sceGeListSync on it never returned. Complete it like a finished
  one.
- FlushImm switches to through mode and another vertex decoder, so flush the
  queued draws first even when the immediate flags match.
- Clear leftover temporary GE breakpoints when setting or clearing the next
  break, so a step that never got there doesn't trip later.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 16ea2cbf88 GPU: Fix draw engine buffer overruns and stale vertex data
- Flush before the queued draws would decode more than VERTEX_BUFFER_MAX
  vertices. The batch was limited by index count, which doesn't bound a
  sparse index range, and DecodeVerts silently stopped while DecodeInds
  still emitted indices for the undecoded draws.
- Give TestBoundingBox its own scratch buffer. It used offsets in decoded_,
  which can hold decoded vertices that aren't flushed yet.
- Read 32-bit indices the way the PSP does, ignoring the upper 16 bits.
  IndexConverter and the fast bounding box test used all 32, so a game
  setting them indexed far past the decoded vertices.
- D3D11: Flush in FinishDeferred like the other backends, since indices
  are still read from PSP memory at flush time (#10095).
- Don't JIT new vertex decoders once the code space is full.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 24017a1ba4 GE debugger: Fix the stepping request handshake
- Wait on actionComplete instead of a bare condition variable wait, which
  could miss the wakeup and hang until resume.
- Serialize requesters, so two debuggers can't overwrite each other's
  action, and make SetCmdValue/FlushDrawing wait too.
- Give up and withdraw the request when stepping ends, instead of waiting
  forever (this deadlocked game shutdown against the Win32 GE debugger).
- Run requests during CPU stepping, which already accepted them.
- Clear the stepping state on Core_Resume from GE stepping and on
  GPU_Shutdown.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:49:10 -06:00
Henrik RydgårdandClaude Opus 5.5 24b71386ff Document that shader cache key changes need a CACHE_VERSION bump
The OpenGL and Vulkan shader caches store raw shader IDs (and, for Vulkan,
pipeline keys) on disk. Add the rule to AGENTS.md and point to it from the
persisted types.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 a9eacff1e1 SoftGPU: Keep the compile flushes on builds without the JIT
Skipping them is only a speed-up, but it changes how states get batched and
optimized, which changes the rendered output (NBA 2K13 and Virtua Tennis
frame dumps). Keep the old behaviour until that's understood.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 a9c782bfcf SoftGPU: Recompute the raster state after a JIT cache clear
A code space clear frees the functions the current state points to, but
the state was kept as long as the GE registers didn't change. Track the
clear generations and recompute, also when a compile during the state
computation clears the caches.

Also skip the binner flush for compiles on builds without the software
JIT, where Compile() does nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 45d997c3ea Texture replacement: Fix DDS mip double free and VFS lifetime
- DDS files with mips stored level 0's file reference in level 1 too, so it
  was freed twice.
- The VFS was deleted on config changes (and ini reloads) while load tasks
  still used it. Now each cached texture waits for its task and releases its
  file references through the old VFS first, and reloads afterwards.
- A failed ini reload turns replacement off instead of leaving it on without
  a VFS.
- Release file references on purge.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 4df1b9c2c3 Vulkan: Clear pipelines before shaders when use flags change
The pipeline deletion callbacks block on in-flight compiles, which use the
shader module promises that the shaders' deletion callbacks free. Queueing
the shaders first freed the promises under a pending compile.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 b912b5823b GPU: Fix texture cache and framebuffer cache lifetime bugs
- Delete TexCacheEntry objects dropped on rehash instead of leaking them.
- Don't leave a released null entry in cache_ when the framebuffer match
  returns before the slot is refilled.
- Reset clutRenderAddress_ in Clear(), which releases the dynamic CLUT FBOs.
- Don't cache a null texture in drawPixelsCache_ when creation fails.
- Fix the reversed subtraction in the failed-FBO retry check.
- Remove the never-taken buffered-rendering early-out in UpdateRenderSize.
  Taking it would leave existing VFBs without an fbo.
- Include smoothedDepal in the depal shader cache key, and print/parse the
  debug IDs as 64-bit.
- Release depal pipelines through Draw2DPipeline::Release so the shader
  source isn't leaked, and make that null-safe.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 8b4e93021e Vulkan: Cache compiled SPIR-V so shaders skip glslang on later runs
GLSLtoSPV takes an optional SPIRVCache, keyed on a 32-bit hash of the
source, stage and variant, plus the source length. A changed shader
simply misses. thin3d's shaders and the other fixed ones use a global
cache in PSP/SYSTEM/CACHE/vulkan_spirv.cache, loaded on first use and
saved after graphics init, when a game's cache is saved, and at
shutdown; it's flushed once it reaches 32 entries, about twice what a
session compiles, so outdated ones don't pile up. Game shaders keep
theirs in the .vkshadercache, ahead of the shader IDs so that the
compiles on load find it (version 60), and only what the session used
is saved.

A cold glslang costs about 40ms before its first shader here, and
0.3-0.9ms per shader after that.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 16:03:52 -06:00
Henrik Rydgård 91a34056d0 Merge pull request #22373 from hrydgard/video-texture-clamp
Clamp sampling of video textures and direct-displayed video to the 480x272 frame
2026-09-28 13:46:00 -06:00
Henrik RydgårdandClaude Opus 5.5 3f44da709b Clamp sampling of video textures and direct-displayed video to the 480x272 frame
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:22:01 -06:00
Henrik RydgårdandClaude Opus 5.5 40a70b004c Savestate: Reset video frame tracking and ME busy time on load
Neither is serialized, and both went stale on load. The ME busy time was
measured against the pre-load clock, so loading an earlier state made the
next SAS/codec job wait until the old time came around, freezing the game.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 09:36:48 -06:00
Henrik RydgårdandClaude Opus 5.5 f96a58fb3e Savestate: Don't change the running game's state when saving
Some DoState code meant for after a load ran on every save:
- scePower reset the bus frequency a game set (and with a locked CPU
  speed, applied the current setting to the clock).
- sceDisplay reset the lag sync baseline, and could schedule lag sync in
  the measuring pass only, which failed the save.
- GPUState dirtied the texture, sceUmd notified the UI, and sceMpeg
  dropped a pending ringbuffer fix-up for an old state.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 09:34:06 -06:00
Henrik RydgårdandClaude Opus 5.5 4bfaab058b GPU: Reject savestates with display list ids out of range
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 09:34:05 -06:00
Henrik RydgårdandClaude Opus 5.5 ad25588651 GPU: Allow for main RAM contention in the movie blit cost
The blit rates were measured with nothing else running. In a game, threads
waking up and SAS mixing on the Media Engine compete with the GE for main RAM:
Star Wars: Lethal Alliance's movie blit takes 8.65ms alone and 10.3ms in the
game. We don't model that load, so RAM texture fetches get a fixed 1.17x for a
typical one. With it, that game's long movie plays at 30 fps as on hardware.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 17:37:53 -06:00
Henrik RydgårdandClaude Opus 5.5 8f66065f53 GPU: Add a clear cost, disabled for now
Full-screen clears measured on a PSP (pspautotests gpu/timing/blittiming):
0.49ms on a 16-bit framebuffer whatever is cleared, 0.69ms on 8888, 1.02ms on
8888 with depth. Charging them may help games that spin hard on an empty
screen, but it's off (chargeClearTime) until tried on some. The video blit
cost moves into the same function, now EstimateFillCycles.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 17:18:42 -06:00
Henrik RydgårdandClaude Opus 5.5 242ef0a984 GPU: Track video frames in GPUCommon, so they expire for the blit cost too
The blit cost remembered only the last buffer a decoder wrote into, forever.
Move the texture cache's video list (with its ageing out a few flips after
the last write) into GPUCommon, so the texture cache, the blit cost and
SoftGPU all share one. That also counts both of a double-buffered player's
frames, which exposed that a clear drawn with texturing still enabled was
being charged as a blit - skip clears and draws without texture coordinates.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 17:00:11 -06:00
Henrik RydgårdandClaude Opus 5.5 7a675b42d5 Model the movie blit by texture format, location and width; scale costs with the clock
Measured on a PSP (pspautotests gpu/timing/blittiming), a full-screen blit
from an unswizzled texture costs what the texture fetch costs: 16-bit formats
half of 32-bit, VRAM a fifth of RAM, and rectangles wider than ~128 texels
~7.5x as much as narrow strips, from texture cache thrashing. Framebuffer
format, filtering and blending don't matter. Ys I & II draws its movie as one
full-width sprite from a 565 texture in RAM, which takes 33ms - that, not the
decode, is what holds it to 30 fps.

All of the ME and GE costs speed up with the clock (2/3 as long at 333/166),
since the whole system runs from the one PLL.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 16:40:11 -06:00
Henrik RydgårdandClaude Opus 5.5 a0b812bc31 Charge hardware-measured time for movie decode, colour conversion and blit
Movie players like the one in Star Wars: Lethal Alliance present every decoded
frame after a single vblank wait, with no clock or timestamp check, so the
frame rate depends on decode, CSC, ATRAC decode and the GE blit adding up to
more than a vblank. We charged nearly nothing for any of them, so such movies
ran at 60 fps until the ringbuffer's slack ran out.

Costs measured on a PSP with a copy of that player (pspautotests
video/mpeg/playertiming), for a 480x272 frame:

- sceVideocodecDecode: 3.4ms (sceMpegAvcDecode 5.8ms less sceMpegAvcCsc 2.4ms)
- sceMpegBaseCscAvc: 2.4ms, was a flat 4ms
- sceAudiocodecDecode, ATRAC3+ only: 2.5ms per frame
- GE: 9.7ms for a through-mode rectangle blit from a decoded video frame,
  charged by area, only for textures in the buffer a decoder last wrote.

GE time also now carries across stall address updates. Before, a list sent
in stalled chunks only had its last chunk's time counted, so sceGeDrawSync
returned 39us after a blit that takes 9.7ms. This affects every game that
builds its lists incrementally, so GE-timing-sensitive games need checking.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 15:23:08 -06:00
Ygor Dreyer a422b818b7 Hash full swizzled CLUT4 glyph atlases 2026-09-24 17:10:10 -03:00
sum2012 a8b85eec8b Add ForceEnableGPUReadback compat
To solve gpu readback issue
2026-09-23 22:06:28 +08:00
Henrik RydgårdandClaude Opus 5 4aef060293 Carry "this is video" across block copies
videos_ only learns about the CSC output, so the texture we actually draw is
just an ordinary 512x512 8888 texture whose contents happen to be different
every frame: hash, miss, throw it in the secondary cache, rebuild, forever.

So track the copy. NotifyVideoCopy marks the destination as video when the
source is, and the copy funnels call it: sceDmacMemcpy, sceKernelMemcpy, and
the four replaced memcpy/memmove variants. It sits outside their "is either
side VRAM" gate, since a RAM-to-RAM copy of a frame is still a frame.

Being a video texture also gets it forced linear filtering and keeps it out of
texture upscaling, which is what you want for a video either way.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-22 16:10:28 -06:00
Henrik RydgårdandClaude Opus 5 aa7b1de28c Drop the texture replacer's "video" option
A `video = true` in textures.ini opted a pack into replacing and dumping video
textures. Both halves are a bad deal. Dumping writes a file per decoded frame,
which fills a disk rather than producing anything a pack can use, and replacing
means a hash lookup on content that is different every frame and will never be
found twice.

It was also the only reason the texture cache still hashed video textures at
all, so it cost every game that has ever played a cutscene, not just the packs
that set it.

Video textures are now never replaced and never dumped, and skipHash is simply
isVideo. An existing ini keeping the key is harmless - unknown options are
ignored.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-22 16:10:16 -06:00
Henrik RydgårdandClaude Opus 5 14bfd332dc Stop stacking duplicate video entries, and stop hashing video textures
NotifyWriteFormattedFromMemory appends to videos_ unconditionally. A game
blitting its decoded frame to the display buffer does that every displayed
frame while it waits for the next one to decode, so the same two or three
addresses come back over and over: Death Jr pushes 733 entries where there are
two distinct buffers, Tekken 6 around 53,000 where there are three. IsVideo()
walks that vector linearly on every texture. Refresh the matching entry instead
of appending a new one - Death Jr now holds 2 entries and Tekken 6 holds 18,
peak size 2 and 3.

The other half is the two TODOs that were already sitting there. A video
texture is new every frame by definition, so re-hashing it only confirms what
the VIDEO flag already said, and the secondary cache has nothing to offer a
frame that will never recur. Skip both, and with them the secondary lookup that
would otherwise key off a hash we no longer compute.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-22 16:10:00 -06:00
sum2012 f5b6b45b5b Remove IgnoreEnqueue hack 2026-09-22 20:19:57 +08:00
Henrik RydgårdandClaude Opus 5 f28896e1ed Vertex decoder: JIT morph on arm64
The arm64 vertex JIT had no morph support (the table entries were commented
out), so every morphed format fell back to the step functions. Add the same
morph steps x86 has: texcoords (plain and prescaled), normals, positions and
the four color formats.

Each step sums its frames the way its step function does as compiled: fused
where that's plain C++, which the compiler turns into fused multiply-adds,
and separate multiply and add where it's CrossSIMD. The unit test confirms
they match bit for bit. Morph formats decode about 2.5-3x faster than with
the steps.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 16:51:00 -06:00
Henrik Rydgård 7b95ff2808 Merge pull request #22325 from hrydgard/vertex-decoder-jit-match
Vertex decoder: New test, make the JITs match the C++ decoder closely
2026-09-21 16:28:27 -06:00
Henrik RydgårdandClaude Opus 5 66b59c73a9 Vertex decoder: bring the riscv64 and loongarch64 JITs in line
Enables TestVertexJitMatchesSteps on both. riscv64 now passes all 216000
jitted formats and loongarch64 all 204000, against 28906 on arm64.

riscv64:
- Jit_PosFloat didn't clean NaN or infinity at all, it just copied the
  three words. Clamp to +-FLT_MAX like the x86 JIT does.
- Jit_PosFloatThrough was missing the truncation of Z to an integer.
- The morph helpers started the sum from the first product rather than
  from +0.0, which rounds differently and lets a -0 term through, and
  rounded that first product towards zero where the steps round to
  nearest. The rest of the sum stays fused, since the compiler contracts
  the steps into fused multiply-adds.
- The texcoord prescale and 5551 color morph paths read morph weights
  from tables that GetMorphValueUsage never asked to be filled in, so
  they used whatever an earlier vertex type had left there.
- The non-Zbb bounds update compared the wrong way around, so through
  mode texcoord bounds came out inverted.

loongarch64:
- Jit_PosFloat had a TODO to clean NaN and infinity, and didn't.
- Jit_PosFloatThrough was missing the same Z truncation.
- Jit_WriteMorphColor narrowed with the logical saturating shifts, so a
  negative channel became a huge unsigned value and saturated to 255
  instead of clamping to 0, and it rounded where the steps truncate. It
  also read the packed color back sign-extended, so any alpha above 0x7F
  compared as larger than 0xFF000000 and claimed full alpha.
- The three packed color morph formats are rewritten. The LSX versions
  built each channel with a chain of inserts, shifts and shuffles that
  didn't survive being run; 4444 also broadcast its scale from the mask
  register. They now follow the steps channel by channel. Note the
  accumulator has to be an LSX scratch register - F4-F7 alias V4-V7,
  which hold the skin matrix for the whole vertex.
- The vertex bounds were loaded with a signed halfword load, so the
  0xFFFF they start at became -1 and no texcoord was ever below it.

PrescaleUV now fuses on these two as well - both JITs fuse it, and the
compiler would have contracted the plain expression there anyway.

The test tolerates a small relative difference on decoded floats, scaled
by the morph count since each term rounds once. Every bug above was
orders of magnitude larger than that.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 13:48:56 -06:00