Files
ppsspp/unittest
Henrik RydgårdandClaude Opus 5.5 649561764d VFPU: NEON version of the exact vdot
The four lanes are computed together: exponents, 24x24-bit products with
round-to-odd, alignment by truncation and a signed horizontal sum. One
pairwise maximum finds both the alignment exponent and any inf or NaN,
which go to the reference. The final rounding is branch-free and in
integers, since the host rounding mode may be the game's.

About three times the throughput of the reference on Apple M-series
(5.2 vs 15.2 ns per call). VFPUDot checks it against the reference on
four million inputs picked to cover cancellation, ties, subnormals and
the overflow edges; a billion more matched offline.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 08:18:25 -06:00
..
2026-09-18 11:59:48 -06:00
2026-09-18 11:59:48 -06:00