The contract is only that a bad lane becomes something that yields zero
when multiplied by zero, by whatever route is cheapest. SSE2 clamps to
+-FLT_MAX, NEON, LSX and the scalar fallback zero the lane - both fine.
Check that property, and that good lanes are untouched.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
CrossSIMD has four independent implementations and only one is compiled
per machine, so each architecture has to check its own copy against
hand-worked results. Running this under qemu gives us coverage of the
targets we have no hardware for - it's already caught two LSX bugs.
Covers Vec4S32 and Vec4F32 arithmetic, comparisons and the mask helpers,
the lane accessors and shuffles, transpose, the various loads (including
the 24-bit and normalizing ones), NaN/Inf handling, and the matrix
routines the old test already had.
Verified on three of the four implementations: NEON natively, the scalar
fallback via TEST_FALLBACK, and LSX under qemu-loongarch64. SSE2 is left
to CI.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>