28 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5 28b5d0ee5c RISC-V, LoongArch64: fix caller-saved FPRs, div by zero, and more
RiscVRegCache::FlushBeforeCall claimed the caller-saved X and F registers
"match between X0 and F0". They don't: X0-X4 are zero/ra/sp/gp/tp and are
never allocated, but F0-F4 (ft0-ft4) are caller-saved and are in the FPR
allocation order. A block that spilled into them and then called out (for
vsin or vcos, say) had the callee destroy registers the cache still
thought were live and dirty. LoongArch and arm64 get this right.

The LoongArch signed div-by-zero sequence applied its ADDI_D(-1) after
the jump target, so it ran on both paths: a negative numerator ended up
with a quotient of 0 instead of 1, and a non-negative one borrowed into
the remainder in the high half, giving hi = num - 1. Insert the -1 into
the low word rather than subtracting, and branch around the other case.

FCvtScaledWS loaded 0x7FFFFFFF into SCRATCHF1 and then immediately
overwrote it with the multiplier, while the FSELs picked SCRATCHF2 - a
register the function never writes. A NAN input produced whatever was
left there.

Also:
- RISC-V had no SyscallUnresolved case, falling through to INVALIDOP,
  which is a live assert in release builds. Added it, matching arm64.
- applyRoundingMode_ fed fcr31 straight to FSRM, which only takes three
  bits - the flags at 2-6 could turn a mode of 0 into 4. Mask with 3
  first, as MIPS.cpp and LoongArch64Asm.cpp do.
- OverwriteExit lacked arm64's assert that the exit fits the 8-byte hole,
  which matters more here since QuickJ is 4, 8 or 12 bytes.
- Both backends allocated the FMin/FMax temp GPR inside the conditional
  NAN path, where a spill store would only run on one side while the
  regcache assumed it always ran. Hoisted above the branch.

Both backends build now, and pspautotests and the unit tests pass under
qemu on each - but nothing in the suite reaches the paths above, so the
two worth checking were checked by running the emitted sequences
directly. The div-by-zero one is wrong for all eight numerators tried and
right for all eight after. The rounding one turns a guest mode of 0 into
RMM whenever fcr31 has its inexact flag set, so 0.5 converts to 1 rather
than 0. The rest are still by inspection: they need register pressure, an
unresolved import, or a debug build to reach.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-22 08:39:25 -06:00
Henrik Rydgård 67ddf899ba Plumb through the PC value for syscalls, so we can get better diagnostics for unresolved ones. 2026-08-15 19:14:13 +02:00
Henrik Rydgård d3e9398cb3 Split Core_EnableStepping into Core_Break and Core_Resume 2024-11-03 17:53:42 +01:00
Unknown W. Brackets 15cb782f85 riscv: Implement Zfa encoding.
Not yet enabled/detected.
2023-12-29 09:42:23 -08:00
Unknown W. Brackets 053831bf4d HLE: Add mechanics for sliced replacements. 2023-12-16 09:08:58 -08:00
Unknown W. Brackets 9b2fa46861 IR: Add mini native jit MIPS block profiler. 2023-09-24 23:04:29 -07:00
Unknown W. Brackets e02426cbbf arm64jit: Implement some system ops. 2023-09-03 21:16:08 -07:00
Unknown W. Brackets 6a75e6712e riscv: Use automapping for special cases too. 2023-08-20 12:42:11 -07:00
Unknown W. Brackets cc4bc406d5 riscv: Cleanup VfpuCtrlToReg meta, use auto-map. 2023-08-20 12:42:11 -07:00
Unknown W. Brackets e40ae60029 riscv: Mark normalized32 after mapping.
It's less confusing to separate it.
2023-08-20 12:42:11 -07:00
Unknown W. Brackets f9bf7de701 riscv: Use a single reg cache. 2023-08-20 12:42:11 -07:00
Unknown W. Brackets a23ade8f75 riscv: Map IR regs based on metadata.
Only doing this in places without GPR/FPR mix or FPR/Vec overlap for now.
2023-08-20 12:42:11 -07:00
Unknown W. Brackets 718a1b3944 riscv: Centralize MarkDirty flagging. 2023-08-19 16:15:49 -07:00
Unknown W. Brackets 4e41f83ecc riscv: Centralize IR reg cache metadata checks.
These are all largely the same between backends.
2023-08-17 23:03:31 -07:00
Unknown W. Brackets 2b36e0a625 irjit: ZeroFpCond -> FpCondFromReg.
We already have a zero reg, so this is more useful and symmetrical.
2023-08-13 10:40:47 -07:00
Unknown W. Brackets 4b9011e475 riscv: Reduce call bloat using temps. 2023-08-08 23:17:32 -07:00
Unknown W. Brackets 93e3d35f5d irjit: Move more to IRNativeBackend, split. 2023-08-06 00:16:43 -07:00
Unknown W. Brackets a5671bc716 riscv: Add simple debug log of missed ops. 2023-07-30 00:02:10 -07:00
Unknown W. Brackets 6d4fb949c2 riscv: Implement float compare ops. 2023-07-29 19:02:15 -07:00
Unknown W. Brackets 23e9dffc68 riscv: Implement vec4 shuffle and init. 2023-07-25 20:33:56 -07:00
Unknown W. Brackets df313bd296 riscv: Fix rounding mode setting. 2023-07-25 20:33:56 -07:00
Unknown W. Brackets 05360d5c7a riscv: Implement simplest float ops. 2023-07-25 20:33:56 -07:00
Unknown W. Brackets 7071884a47 riscv: Handle rounding mode and ctrl transfers. 2023-07-25 20:33:56 -07:00
Unknown W. Brackets a8edf5fa24 riscv: Reduce bloat in jit fallbacks. 2023-07-25 19:42:04 -07:00
Unknown W. Brackets 4100767b5e riscv: Optimize SetConst a bit. 2023-07-23 18:01:00 -07:00
Unknown W. Brackets 34bfe93ea5 riscv: Fix block lookup issues. 2023-07-23 18:01:00 -07:00
Unknown W. Brackets c2da7d18bb riscv: Stub out more IR compilation categories. 2023-07-23 18:01:00 -07:00
Unknown W. Brackets bf7a6eb2cd riscv: Add jit for some initial instructions. 2023-07-23 18:01:00 -07:00