70 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5 320cfd7741 CrossSIMD: fill in the scalar fallback, fix two LSX bugs
LoadConvertU8, StoreConvertToU8 and LoadTranspose existed in the SSE2,
NEON and LSX implementations but not in the scalar one, so anything using
them wouldn't build on a target without SIMD.

The two LSX bugs were found by the new CrossSIMD unit test, run under
qemu-loongarch64:

- Vec4F32::operator[] had a switch with no breaks, so every index fell
  through to the default and returned lane 3.
- StoreConvertToU8 narrowed with the logical (unsigned) saturating shifts,
  which turn a negative value into a huge unsigned one and saturate it to
  255. It should clamp to 0, as the packs/packus pair in the SSE version
  does. Narrow signed->signed and then signed->unsigned instead.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-19 12:35:09 -06:00
Henrik RydgårdandClaude Opus 5 8e3b4245c2 CrossSIMD: fix the scalar fallback, which nothing compiled until now
The riscv64 cross build is the first target to take the non-SIMD path,
and it didn't compile:

- LoadF24x4 called LoadR24x3_One, which doesn't exist. It should shift
  all four lanes, like the SSE/NEON/LSX versions do.
- isnan/isinf were unqualified, and <cmath> wasn't included.
- WithLane3From and AnyCompareBitsSet were missing entirely.

LoadF24x3_One also left lane 3 as zero, where all three SIMD versions set
it to 1.0f - a behavioural bug that would only have shown up once someone
ran this path.

Verified by flipping TEST_FALLBACK in the header: the whole tree builds,
and with the scalar path in use the unit tests and pspautotests both pass
in full (342/342, software renderer, which leans on this code heavily).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-19 12:19:24 -06:00
Henrik Rydgård 54d5c1d5e3 Improve some initialization in CrossSIMD.h. Add comments about CLUT variants 2026-06-09 10:15:08 +02:00
Henrik Rydgård f65b3f7a6b Fix LoadU8Norm in CrossSIMD
Caused breakage in the previous PR.
2026-06-08 23:06:41 +02:00
Henrik Rydgård 61a2d78930 Buildfix 2026-06-04 14:02:12 +02:00
Henrik Rydgård 195455a7f4 Break out VertexReader, prepare VertexReader for CrossSIMD use in software transform 2026-06-04 12:45:17 +02:00
Henrik Rydgård c697496364 Other fixes 2026-06-04 11:24:15 +02:00
Henrik Rydgård 1cc0062120 Add Dot3 and Dot4 instructions to Vec4F32 2026-06-04 11:24:15 +02:00
Henrik Rydgård bf196aded6 Fix critical bug in LoadS8Norm for SSE 2026-06-03 15:21:53 +02:00
Henrik Rydgård 7f4c51c141 Merge pull request #21752 from hrydgard/clean-nans-sw
Clean out NaNs and INFs from PosFloat vertex coordinates
2026-06-03 10:38:13 +02:00
Henrik Rydgård bb573e6e0c Well, it builds, but doesn't work yet. 2026-06-02 12:33:53 +02:00
Henrik Rydgård e6185ba6bb Clean out NaNs and INFs from vertex coordinates
This enables a behavior seen on the real PSP where 0 * NaN == 0 in the
GPU (NOTE: This means the vertex transform pipeline).

However, this
doesn't touch INFs unfortunately, and we need to modify those too...

This, together with #21715, finally fixes #20204 .
2026-06-02 10:52:39 +02:00
Henrik Rydgård 2a5c2fa477 Check if the viewport transform matches the clip space. If so we can skip the near clip plane. 2026-05-30 19:07:59 +02:00
Henrik Rydgård 328e1b681c More CrossSIMD ops 2026-05-28 10:56:33 +02:00
Henrik Rydgård 82a888fd8e Reimplement the fast culling using CrossSIMD 2026-05-25 18:16:34 +02:00
Henrik Rydgård ef78791f5d Add some more required functions to CrossSIMD 2026-05-25 18:16:31 +02:00
Henrik Rydgård 250abe0d56 Loongarch64 build fixes 2026-05-25 15:41:24 +02:00
Henrik Rydgård 8b6951c803 Apply vldi fix 2026-05-25 15:40:22 +02:00
copilot-swe-agent[bot]andhrydgard c05e1cfd3c Address LSX CrossSIMD review feedback
Agent-Logs-Url: https://github.com/hrydgard/ppsspp/sessions/6187ce4f-6c64-4b50-b82c-49799ea8cfd2

Co-authored-by: hrydgard <[email protected]>
2026-05-25 15:40:22 +02:00
Henrik Rydgård 37f6fda9c9 Add a CrossSIMD loongarch64 LSX implementation. Note: This is untested! 2026-05-25 15:40:22 +02:00
Henrik Rydgård 7d987cd78b Make fast_matrix_mul_4x4 inlineable 2026-01-30 14:10:32 +01:00
Eric Warmenhoven 58eaf56c5e libretro: build fix
gcc9 and the version of msvc19 that the libretro builders use don't
have _mm_storeu_si32. Using _m_cvtsi128_si32 followed by memcpy is
picked up by compilers that support _mm_storeu_si32 as being the same
and they generate the same assembly, so this is no harm to them.
2026-01-28 00:15:35 -05:00
Henrik Rydgård 47f6b9975e Optimize the color alpha computation for color morph 2026-01-20 13:51:24 +01:00
Henrik Rydgård 3378faca5e Optimize morph code (non-JIT vertex decoders) 2026-01-20 13:51:24 +01:00
Henrik Rydgård a8f89e1544 Correct rounding in StoreConvertToU8, correct Vec4F32::Store3. 2026-01-20 13:51:24 +01:00
Henrik Rydgård fc09669ba7 Add DualSense support. Remove trying to support bluetooth devices (not reliable without more work) 2025-07-03 11:11:34 +02:00
Henrik Rydgård fd88f79d07 CrossSIMD: Fix more no-simd fallbacks. The depth rasterizer now works in TEST_FALLBACK mode. 2025-02-10 11:51:22 -06:00
Henrik Rydgård 8907a4c001 More fixes, improve unit test 2025-02-06 10:26:46 -06:00
Henrik Rydgård fc0b1eef4c Fix 4x4 2025-02-06 10:26:46 -06:00
Henrik Rydgård 07e892e1a2 Fix some type mismatches in CrossSIMD.h that not all compilers complain about.
Fixes #19905
2025-02-01 10:30:02 -06:00
Henrik Rydgård c78fa60431 Add better way to check if CrossSIMD has been natively implemented 2025-01-28 10:56:52 +01:00
Henrik Rydgård acd5b24924 Complete CrossSIMD non-simd fallback (although buggy, it seems). Minor ARM64 opt. 2025-01-28 10:54:43 +01:00
Henrik Rydgård 2aaa1e5379 CrossSIMD: Expand the no-simd path 2025-01-28 10:54:43 +01:00
Henrik Rydgård 74501b06b6 CrossSIMD: Add more no-simd fallback types 2025-01-28 10:54:43 +01:00
Henrik Rydgård 9d164b71fb Some new CrossSIMD operations 2025-01-28 10:54:43 +01:00
Henrik Rydgård 9ccf47a6b3 Fix type error in minor ARM64 optimization, which some compilers don't like.
Fixes #19905
2025-01-22 09:54:20 +01:00
Henrik Rydgård a5116e1590 Use vmaxvq_s32 to implement AnyZeroSignBit more efficiently on ARM64 2025-01-19 17:49:09 +01:00
Henrik Rydgård 206d4d1fea Implement the low-quality depth raster mode, default to it on Android/iOS.
I really can't tell much of a difference in practice...
2024-12-31 11:19:38 +01:00
Henrik Rydgård f85d7db5b1 Comment fixes, buildfix 2024-12-31 02:39:58 +01:00
Henrik Rydgård dee5fe6990 Fix issue in Midnight Club where Z now wrapped around at a distance, after removing the clamp. Might as well cull. 2024-12-31 02:30:05 +01:00
Henrik Rydgård f5cc41caab More CrossSIMD (breaking change) 2024-12-31 01:59:11 +01:00
Henrik Rydgård c3ac798545 More crosssimd 2024-12-31 01:59:08 +01:00
Henrik Rydgård ef934df0f2 Compute and cull by triangle area early before writing 4-groups of triangles 2024-12-29 01:21:16 +01:00
Henrik Rydgård e8786fc401 Cull 4-groups of triangles early 2024-12-29 01:12:53 +01:00
Henrik Rydgård 66926e28c1 Simplify / fix through mode triangles 2024-12-28 18:51:37 +01:00
Henrik Rydgård eec7853efe More SIMD: Add some matrix operations to CrossSIMD (#19773)
* More CrossSIMD functionality

* Use the new SIMD API for the matrix multiplies
2024-12-28 18:45:14 +01:00
Henrik Rydgård 8c069917b5 Depth raster optimizations: Merge viewport into projection matrix, prepare for further SIMD-ification (#19769)
* Remove some old redundant reports

* Fix scissor off by one

* More CrossSIMD

* Move the viewport scale out to the proj matrix

* DepthRaster: Also merge the viewport translation into the projection matrix.

* Depth raster: Do the triangle clipping in batches of 4 triangles

* Cleanup
2024-12-28 10:36:39 +01:00
Henrik Rydgård cc040ab251 Depth raster: Bugfix, minor opt (#19768)
* Correct two errors in CrossSIMD.h, thanks hiroyuki177

Fixes #19767

* DepthRaster: Merge offset into viewport X/Y
2024-12-26 20:23:52 +01:00
Henrik Rydgård 0e59cd7641 DepthRaster: Fix typo breaking LESS depth comparison mode on x86(64) 2024-12-24 23:26:32 +01:00
Henrik Rydgård 8e747dc948 CrossSIMD: Add a multiply-as-16bit function to Vec4S32. This can be implemented quickly on SSE2. 2024-12-22 18:53:10 +01:00