VFPU: SSE2 version of the exact vdot

Same shape as the NEON one, which now shares its rounding tail. SSE2 has
no per-lane shift, so the alignment shift is a multiply by a power of two
built from float bits and converted by truncation; the unsigned maxima
use the 16-bit instructions, since every value involved fits in 15 bits.
Nothing depends on the host rounding mode or flush-to-zero.

Checked against the reference by VFPUDot and on 300M more inputs offline,
also with MXCSR set to round toward zero with FTZ and DAZ.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
Henrik RydgårdandClaude Opus 5.5 committed 2026-09-23 08:18:25 -06:00
1 parent 649561764d
commit cd005b9041
2 files changed
+79 -21

No files matched your search

+1 -1
View File
@@ -1986,7 +1986,7 @@ bool TestFastVec() {
return true;
}
// vfpu_dot's SIMD version against the reference, on inputs chosen to make trouble: close
// vfpu_dot's SIMD versions against the reference, on inputs chosen to make trouble: close
// exponents, cancelling products, ties, zeroes and subnormals, the overflow and underflow edges,
// and inf and NaN.
bool TestVFPUDot() {