softgpu: Truncate, don't round, in the linear sampler JIT

Jit_GetTexelCoordsQuad converted s*w*256 with CVTPS2DQ, which rounds. The C++
reference casts with (int), and the nearest JIT paths use CVTTPS2DQ - the quad
path just never got converted when the others did. The JIT's sample point sat up
to 1/512 texel further along than the interpreter's, so roughly one pixel in
sixteen picked a different frac_u/frac_v, and at exact texel boundaries a
different texel.

That mismatch is visible wherever the two paths coexist: x86-64 desktop runs the
JIT, 32-bit x86 and UWP have no sampler JIT at all, and even within one x86-64
session the first draws with a new SamplerID run the C++ path while later ones
run the JIT. CVTPS2DQ also honors MXCSR's rounding mode, so anything that left a
non-default mode in the render thread would have changed rasterized output.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01Vd8ntC2brCUtCrDJMqLbs8
This commit is contained in:
Henrik RydgårdandClaude Opus 5 committed 2026-08-29 11:55:23 +02:00
1 parent 74dd222ea3
commit a007976258
1 file changed
+5 -2
+5 -2
View File
@@ -2814,8 +2814,11 @@ bool SamplerJitCache::Jit_GetTexelCoordsQuad(const SamplerID &id) {
MULPS(sReg, M(constWidthHeight256f_));
}
// And now, convert to integers for all later processing.
CVTPS2DQ(sReg, R(sReg));
// And now, convert to integers for all later processing. Must truncate, not round: the C++
// reference casts with (int), and the nearest paths above use CVTTPS2DQ. CVTPS2DQ also follows
// MXCSR's rounding mode, which would let anything that leaves a non-default mode in this thread
// change rasterized pixels.
CVTTPS2DQ(sReg, R(sReg));
// Now adjust X and Y...
X64Reg tempXYReg = regCache_.Alloc(RegCache::VEC_TEMP0);