LoadConvertU8, StoreConvertToU8 and LoadTranspose existed in the SSE2,
NEON and LSX implementations but not in the scalar one, so anything using
them wouldn't build on a target without SIMD.
The two LSX bugs were found by the new CrossSIMD unit test, run under
qemu-loongarch64:
- Vec4F32::operator[] had a switch with no breaks, so every index fell
through to the default and returned lane 3.
- StoreConvertToU8 narrowed with the logical (unsigned) saturating shifts,
which turn a negative value into a huge unsigned one and saturate it to
255. It should clamp to 0, as the packs/packus pair in the SSE version
does. Narrow signed->signed and then signed->unsigned instead.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>