Commit Graph
147 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 16ea2cbf88 GPU: Fix draw engine buffer overruns and stale vertex data
- Flush before the queued draws would decode more than VERTEX_BUFFER_MAX
  vertices. The batch was limited by index count, which doesn't bound a
  sparse index range, and DecodeVerts silently stopped while DecodeInds
  still emitted indices for the undecoded draws.
- Give TestBoundingBox its own scratch buffer. It used offsets in decoded_,
  which can hold decoded vertices that aren't flushed yet.
- Read 32-bit indices the way the PSP does, ignoring the upper 16 bits.
  IndexConverter and the fast bounding box test used all 32, so a game
  setting them indexed far past the decoded vertices.
- D3D11: Flush in FinishDeferred like the other backends, since indices
  are still read from PSP memory at flush time (#10095).
- Don't JIT new vertex decoders once the code space is full.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5 66b59c73a9 Vertex decoder: bring the riscv64 and loongarch64 JITs in line
Enables TestVertexJitMatchesSteps on both. riscv64 now passes all 216000
jitted formats and loongarch64 all 204000, against 28906 on arm64.

riscv64:
- Jit_PosFloat didn't clean NaN or infinity at all, it just copied the
  three words. Clamp to +-FLT_MAX like the x86 JIT does.
- Jit_PosFloatThrough was missing the truncation of Z to an integer.
- The morph helpers started the sum from the first product rather than
  from +0.0, which rounds differently and lets a -0 term through, and
  rounded that first product towards zero where the steps round to
  nearest. The rest of the sum stays fused, since the compiler contracts
  the steps into fused multiply-adds.
- The texcoord prescale and 5551 color morph paths read morph weights
  from tables that GetMorphValueUsage never asked to be filled in, so
  they used whatever an earlier vertex type had left there.
- The non-Zbb bounds update compared the wrong way around, so through
  mode texcoord bounds came out inverted.

loongarch64:
- Jit_PosFloat had a TODO to clean NaN and infinity, and didn't.
- Jit_PosFloatThrough was missing the same Z truncation.
- Jit_WriteMorphColor narrowed with the logical saturating shifts, so a
  negative channel became a huge unsigned value and saturated to 255
  instead of clamping to 0, and it rounded where the steps truncate. It
  also read the packed color back sign-extended, so any alpha above 0x7F
  compared as larger than 0xFF000000 and claimed full alpha.
- The three packed color morph formats are rewritten. The LSX versions
  built each channel with a chain of inserts, shifts and shuffles that
  didn't survive being run; 4444 also broadcast its scale from the mask
  register. They now follow the steps channel by channel. Note the
  accumulator has to be an LSX scratch register - F4-F7 alias V4-V7,
  which hold the skin matrix for the whole vertex.
- The vertex bounds were loaded with a signed halfword load, so the
  0xFFFF they start at became -1 and no texcoord was ever below it.

PrescaleUV now fuses on these two as well - both JITs fuse it, and the
compiler would have contracted the plain expression there anyway.

The test tolerates a small relative difference on decoded floats, scaled
by the morph count since each term rounds once. Every bug above was
orders of magnitude larger than that.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 13:48:56 -06:00
Henrik RydgårdandClaude Opus 5 c58baedd75 Vertex decoder: fix the jit-match test on MSVC
The test didn't compile: VERTS is captured by reference into testFormat, and
MSVC won't use a captured constexpr as an array bound. Make it static.

On arm64 it then failed on the UV prescale steps, by one ULP. The arm64 JIT and
the NEON handwritten decoders fuse the multiply-add, and the steps only match
that when the compiler contracts a * b + c - which clang does and MSVC doesn't,
in Debug or Release. Spell out which one happens instead of relying on it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 12:13:09 -06:00
Henrik RydgårdandClaude Opus 5 a1ed24ae2d Vertex decoder: fix the handwritten decoders, and test them too
The two handwritten SIMD decoders are used with or without the JIT, and had
drifted from the step functions:
- The God of War one ignored g_DoubleTextureCoordinates, giving HD Remaster
  games the wrong UVs, and passed NaN and infinite positions through. Those
  now come out finite like everywhere else.
- The GTA one expanded 5551 colors wrongly on NEON (a left shift where the
  SSE version shifts right).
- On ARM64, both now fuse the UV scale and offset, like the JIT and the steps
  as the compiler builds them.

Both are back in the unit test, which now also feeds NaN and infinity to plain
float positions and checks they come out finite, and runs only on x86-64 and
arm64 for now.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 11:29:40 -06:00
Henrik RydgårdandClaude Opus 5 63a6c11214 Vertex decoder: make the JIT match the step functions
Add a unit test that decodes every vertex format through both the step
functions and the JIT and requires identical output, side effects included.
Only skinning may differ by rounding, since arm64 accumulates the bone
matrices with fused multiply-adds.

What it found, and fixed:
- x86 morph colors rounded to nearest where the steps truncate, and applied
  the scale before the weight, which rounds differently.
- x86 through-mode u16 UV bounds compared signed, so texcoords above 32767
  scrambled the bounds for everyone running the x86 JIT.
- x86 morph sums could produce -0 where the steps produce +0.
- Step_NormalS16Morph scaled by 1/32768 twice, giving near-zero normals.
- Step_PosFloatThrough lost the truncation of Z to an integer that the JITs
  and the vertex reader did before it moved into the decoder.
- SetVertexType never reset skinInDecode.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 11:29:40 -06:00
Henrik RydgårdandClaude Opus 5 0b5a8f537e Decide on the vertex decoder JIT with CoreParameter, not the CPU core
The vertex decoder JIT was enabled only when g_Config.iCpuCore was one of the
JIT cores, which tied an unrelated GPU path to the CPU setting. Headless leaned
on that by forcing iCpuCore to the interpreter after ApplyToConfig(), to keep
the decoder JIT off in the tests.

Add CoreParameter::bUseVertexDecoderJit instead. Standalone and libretro set it
wherever the host can JIT, so IR interpreter users on such hosts now get the
decoder JIT too. Headless sets it false, which keeps the test results as they
were and lets --cpu go through ApplyToConfig() like every other option.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 10:31:58 -06:00
Henrik Rydgård e062c90bf7 Try to fix the test difference (basically by replicating an old misfeature of headless...) 2026-08-11 22:36:47 +02:00
Henrik Rydgård c3eafeb005 Remove obsolete vertex decoder code and options 2026-07-14 17:26:00 +02:00
Henrik Rydgård cd40e2d4f0 Delete more code related to hardware skinning 2026-07-14 17:17:40 +02:00
Henrik Rydgård 61e1ef8f7a Remove the "Software skinning" option. Now always on. 2026-07-14 17:02:57 +02:00
Henrik Rydgård 28b2b1793b Pipeline checker: Add support for checking for identical pipelines with different vertex formats
This turned out not to be very fruitful.
2026-07-10 19:07:27 +02:00
Henrik Rydgård 195455a7f4 Break out VertexReader, prepare VertexReader for CrossSIMD use in software transform 2026-06-04 12:45:17 +02:00
Henrik Rydgård e6185ba6bb Clean out NaNs and INFs from vertex coordinates
This enables a behavior seen on the real PSP where 0 * NaN == 0 in the
GPU (NOTE: This means the vertex transform pipeline).

However, this
doesn't touch INFs unfortunately, and we need to modify those too...

This, together with #21715, finally fixes #20204 .
2026-06-02 10:52:39 +02:00
Henrik Rydgård 12d280839c Microoptimizations in VertexDecoder interpreter 2026-02-17 11:59:30 +01:00
Henrik Rydgård 991d7bdfab VertexDecoder: Refactor away lowerbound/upperbound parameters 2026-02-05 13:26:17 +01:00
Henrik Rydgård 47f6b9975e Optimize the color alpha computation for color morph 2026-01-20 13:51:24 +01:00
Henrik Rydgård 064ad64a55 More minor optimizations 2026-01-20 13:51:24 +01:00
Henrik Rydgård 73e41750f7 More non-JIT morph optimizations 2026-01-20 13:51:24 +01:00
Henrik Rydgård 3378faca5e Optimize morph code (non-JIT vertex decoders) 2026-01-20 13:51:24 +01:00
Henrik Rydgård 5e16bf907b VertexDecoder: Improve logging for missing formats. Add missing convert function.
The missing function is mainly used in D3D11, which can be used on
Windows for ARM64. It's not necesssary for the other backends, which is
why it used to be missing in the ARM64 vertex decoder.

Also fix a minor memory leak in AtracCtx2.
2025-09-24 10:52:09 -06:00
Lin Runze 5b406f00fd loongarch: Initial VertexJIT support and bug fix 2025-07-17 00:33:56 +08:00
Henrik Rydgård 9ecc135601 Vertex decoder C++ fallbacks: Make sure to never read from the destination
This can have devastating effects on performance on some architectures.

Will help, but not fix, #20171 on some hardware (there's more
optimization work that needs doing).
2025-05-26 11:18:45 +02:00
Henrik Rydgård 4e0b6ac3ec More misc cleanup 2025-05-14 15:14:03 +02:00
Henrik Rydgård dd38e8d012 GPU constify 2025-05-14 09:39:14 +02:00
Henrik Rydgård c6691de64c Remove excessive error reporting 2025-03-02 02:28:42 +01:00
Henrik Rydgård 449e3e1360 Very minor soft transform optimization 2024-12-28 18:45:03 +01:00
Henrik Rydgård c91169e702 Restore removed <algorithm> includes.
Turns out these were needed after all. For some reason, on Windows and
Mac, <algorithm> gets auto-included by something else so I don't notice
when it's missing, and MSVC's include dependency tracker doesn't see it
either.
2024-12-19 09:53:07 +01:00
Henrik Rydgård df6ed8cfc9 Do some cleanup of #includes in GPU 2024-12-18 13:57:26 +01:00
Henrik Rydgård e74101a2fb applySkinInDecode belongs in the VertexTypeID, not in the options. 2024-12-17 18:24:18 +01:00
Henrik Rydgård 2c283fbb07 Minor cleanups, crashfixes 2024-10-14 23:57:19 +02:00
Henrik Rydgård 42c32c5afc VertexDecoder: Don't read loop counts from memory. Improves codegen 2024-09-10 17:53:19 +02:00
Henrik Rydgård 43c68c4277 VertexDecoder: Remove member function pointers from decoding 2024-07-22 14:06:15 +02:00
Henrik Rydgård fd9daf7594 Fix some minor issues found by --sanitize. Add --sanitizeub.
Unfortunately the ub (undefined behavior) sanitizer has some bugs, it doesn't
understand pointers to member functions, so can't use it in-game (due to the
vertex decoder).

Thanks Nemoumbra for the reminder.
2024-07-22 11:37:18 +02:00
Henrik Rydgård e01ca5b057 Logging API change (refactor) (#19324)
* Rename LogType to Log

* Explicitly use the Log:: enum when logging. Allows for autocomplete when editing.

* Mac/ARM64 buildfix

* Do the same with the hle result log macros

* Rename the log names to mixed case while at it.

* iOS buildfix

* Qt buildfix attempt, ARM32 buildfix
2024-07-14 14:42:59 +02:00
Henrik Rydgård 8d6e96d04e Use binary search to find IR block offsets 2024-06-07 09:28:27 +02:00
Henrik Rydgård c794f4bd41 Add an unrelated comment and some casts 2024-06-05 08:35:09 +02:00
Henrik Rydgård 6ce087430b JIT-less vertex decoder: SSE/NEON-optimize ComputeSkinMatrix 2024-06-04 12:29:16 +02:00
Henrik Rydgård 9ac7054b01 Vertex decoder (non-JIT): Optimize 16-bit color decoders. 2024-06-04 10:35:31 +02:00
Henrik Rydgård 7a32507ab7 Add a decode counter to vertex decoders in _DEBUG mode 2024-06-02 10:25:05 +02:00
Henrik Rydgård fb599cd0a6 Only use the optimized decoders if SSE or NEON is available. 2024-05-11 14:18:42 +02:00
Henrik Rydgård 4a66f8978b Fix the GoW optimized vertex decoder, add NEON optimizations 2024-05-11 13:27:11 +02:00
Henrik Rydgård bafff7f5db Temporarily disable the custom GoW vertex decoder, it needs some work. 2024-05-11 11:11:48 +02:00
Henrik Rydgård 3526416173 Add another handwritten vertex decoder 2024-05-11 10:00:39 +02:00
Henrik Rydgård 81f1b3fd95 Make handwritten vertex decoders work with non-compiled vertex decoding 2024-05-11 10:00:35 +02:00
Henrik Rydgård 3e11e54405 Remove obsolete flag 2024-05-11 10:00:35 +02:00
Herman Semenov b57dab2812 [GPU] Make static and const methods if possible 2024-04-05 17:04:31 +03:00
Henrik Rydgård e3177ac870 Make some global string pointers const, not just the strings.
Minor cleanup.
2023-12-29 14:09:45 +01:00
Henrik Rydgård f86189c951 Show vertex decoders separately in profiles 2023-12-19 12:25:54 +01:00
Herman Semenov 315340fc62 Using const reference for C++17 range-based loop and freq used objects 2023-12-13 17:33:01 +01:00
Henrik Rydgård 71aaad23fb Fix issue with zero-vertex draw calls. Though, should maybe just filter them out earlier. 2023-12-10 12:21:07 +01:00