LoadConvertU8, StoreConvertToU8 and LoadTranspose existed in the SSE2,
NEON and LSX implementations but not in the scalar one, so anything using
them wouldn't build on a target without SIMD.
The two LSX bugs were found by the new CrossSIMD unit test, run under
qemu-loongarch64:
- Vec4F32::operator[] had a switch with no breaks, so every index fell
through to the default and returned lane 3.
- StoreConvertToU8 narrowed with the logical (unsigned) saturating shifts,
which turn a negative value into a huge unsigned one and saturate it to
255. It should clamp to 0, as the packs/packus pair in the SSE version
does. Narrow signed->signed and then signed->unsigned instead.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The riscv64 cross build is the first target to take the non-SIMD path,
and it didn't compile:
- LoadF24x4 called LoadR24x3_One, which doesn't exist. It should shift
all four lanes, like the SSE/NEON/LSX versions do.
- isnan/isinf were unqualified, and <cmath> wasn't included.
- WithLane3From and AnyCompareBitsSet were missing entirely.
LoadF24x3_One also left lane 3 as zero, where all three SIMD versions set
it to 1.0f - a behavioural bug that would only have shown up once someone
ran this path.
Verified by flipping TEST_FALLBACK in the header: the whole tree builds,
and with the scalar path in use the unit tests and pspautotests both pass
in full (342/342, software renderer, which leans on this code heavily).
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
This enables a behavior seen on the real PSP where 0 * NaN == 0 in the
GPU (NOTE: This means the vertex transform pipeline).
However, this
doesn't touch INFs unfortunately, and we need to modify those too...
This, together with #21715, finally fixes#20204 .
gcc9 and the version of msvc19 that the libretro builders use don't
have _mm_storeu_si32. Using _m_cvtsi128_si32 followed by memcpy is
picked up by compilers that support _mm_storeu_si32 as being the same
and they generate the same assembly, so this is no harm to them.
* Remove some old redundant reports
* Fix scissor off by one
* More CrossSIMD
* Move the viewport scale out to the proj matrix
* DepthRaster: Also merge the viewport translation into the projection matrix.
* Depth raster: Do the triangle clipping in batches of 4 triangles
* Cleanup