From a0079762586f2b877dd9794501aba73d2c54f4ed Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Henrik=20Rydg=C3=A5rd?= Date: Sat, 29 Aug 2026 01:00:22 +0200 Subject: [PATCH] softgpu: Truncate, don't round, in the linear sampler JIT Jit_GetTexelCoordsQuad converted s*w*256 with CVTPS2DQ, which rounds. The C++ reference casts with (int), and the nearest JIT paths use CVTTPS2DQ - the quad path just never got converted when the others did. The JIT's sample point sat up to 1/512 texel further along than the interpreter's, so roughly one pixel in sixteen picked a different frac_u/frac_v, and at exact texel boundaries a different texel. That mismatch is visible wherever the two paths coexist: x86-64 desktop runs the JIT, 32-bit x86 and UWP have no sampler JIT at all, and even within one x86-64 session the first draws with a new SamplerID run the C++ path while later ones run the JIT. CVTPS2DQ also honors MXCSR's rounding mode, so anything that left a non-default mode in the render thread would have changed rasterized output. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01Vd8ntC2brCUtCrDJMqLbs8 --- GPU/Software/SamplerX86.cpp | 7 +++++-- 1 file changed, 5 insertions(+), 2 deletions(-) diff --git a/GPU/Software/SamplerX86.cpp b/GPU/Software/SamplerX86.cpp index 46172d19e7..16a12bd23e 100644 --- a/GPU/Software/SamplerX86.cpp +++ b/GPU/Software/SamplerX86.cpp @@ -2814,8 +2814,11 @@ bool SamplerJitCache::Jit_GetTexelCoordsQuad(const SamplerID &id) { MULPS(sReg, M(constWidthHeight256f_)); } - // And now, convert to integers for all later processing. - CVTPS2DQ(sReg, R(sReg)); + // And now, convert to integers for all later processing. Must truncate, not round: the C++ + // reference casts with (int), and the nearest paths above use CVTTPS2DQ. CVTPS2DQ also follows + // MXCSR's rounding mode, which would let anything that leaves a non-default mode in this thread + // change rasterized pixels. + CVTTPS2DQ(sReg, R(sReg)); // Now adjust X and Y... X64Reg tempXYReg = regCache_.Alloc(RegCache::VEC_TEMP0);