mirror of
https://github.com/hrydgard/ppsspp.git
synced 2026-10-01 14:58:14 +00:00
Enables TestVertexJitMatchesSteps on both. riscv64 now passes all 216000 jitted formats and loongarch64 all 204000, against 28906 on arm64. riscv64: - Jit_PosFloat didn't clean NaN or infinity at all, it just copied the three words. Clamp to +-FLT_MAX like the x86 JIT does. - Jit_PosFloatThrough was missing the truncation of Z to an integer. - The morph helpers started the sum from the first product rather than from +0.0, which rounds differently and lets a -0 term through, and rounded that first product towards zero where the steps round to nearest. The rest of the sum stays fused, since the compiler contracts the steps into fused multiply-adds. - The texcoord prescale and 5551 color morph paths read morph weights from tables that GetMorphValueUsage never asked to be filled in, so they used whatever an earlier vertex type had left there. - The non-Zbb bounds update compared the wrong way around, so through mode texcoord bounds came out inverted. loongarch64: - Jit_PosFloat had a TODO to clean NaN and infinity, and didn't. - Jit_PosFloatThrough was missing the same Z truncation. - Jit_WriteMorphColor narrowed with the logical saturating shifts, so a negative channel became a huge unsigned value and saturated to 255 instead of clamping to 0, and it rounded where the steps truncate. It also read the packed color back sign-extended, so any alpha above 0x7F compared as larger than 0xFF000000 and claimed full alpha. - The three packed color morph formats are rewritten. The LSX versions built each channel with a chain of inserts, shifts and shuffles that didn't survive being run; 4444 also broadcast its scale from the mask register. They now follow the steps channel by channel. Note the accumulator has to be an LSX scratch register - F4-F7 alias V4-V7, which hold the skin matrix for the whole vertex. - The vertex bounds were loaded with a signed halfword load, so the 0xFFFF they start at became -1 and no texcoord was ever below it. PrescaleUV now fuses on these two as well - both JITs fuse it, and the compiler would have contracted the plain expression there anyway. The test tolerates a small relative difference on decoded floats, scaled by the morph count since each term rounds once. Every bug above was orders of magnitude larger than that. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>