Same shape as the NEON one, which now shares its rounding tail. SSE2 has
no per-lane shift, so the alignment shift is a multiply by a power of two
built from float bits and converted by truncation; the unsigned maxima
use the 16-bit instructions, since every value involved fits in 15 bits.
Nothing depends on the host rounding mode or flush-to-zero.
Checked against the reference by VFPUDot and on 300M more inputs offline,
also with MXCSR set to round toward zero with FTZ and DAZ.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>