135 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 d1c04d70a6 IR interpreter: Use threaded dispatch with GCC and Clang
Every op ended with a break back to one shared indirect jump, which the
CPU has to predict for every op in the program. With labels as values,
each op jumps through a table from its own site instead, which predicts
much better: an integer-heavy benchmark runs about 13% faster on an M1.
Other compilers keep the switch, and ops missing from the table fall back
to it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 11:47:17 -06:00
Henrik RydgårdandClaude Opus 5.5 7eb371b231 IR interpreter: Merge a conditional exit with the ExitToConst after it
Blocks ending in a branch dispatch a conditional exit and then the
fallthrough ExitToConst. One op now returns either target, reading the
second from the ExitToConst, which stays behind unexecuted.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 11:47:17 -06:00
Henrik RydgårdandClaude Opus 5.5 f92340c092 IR: Compute vrot's sine and cosine in one call
A new FSinCos op writes both from one argument reduction. arm64 and x64 get
both back from a single call, packed in one double; RISC-V and LoongArch make
the two calls.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:22:41 -06:00
Henrik RydgårdandClaude Opus 5.5 6b78138c08 VFPU: Compile vh2f in the IR, and call vfpu_h2f for it everywhere
vh2f always went to the interpreter in the IR. A new FHalfToFloat op
converts the lower or upper half of a word, and the native backends
call vfpu_h2f for it like FSin. The legacy arm64 JIT now makes the same
call instead of computing the conversion inline; vh2f is rare, and the
call is much less code.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 eaf55c467b IR: Fold ClampToZero into the 31-bit packs
Vec4ClampToZero and Vec2ClampToZero only ever fed Vec4Pack31To8 and
Vec2Pack31To16, for vi2uc and vi2us. The packs now clamp negative lanes
to zero themselves, which saves an op and a vector temp, and lets x64
clamp with PACKUSWB's saturation after an arithmetic shift.

While at it, RISC-V compiles Vec2Unpack16To31, Vec2Pack31To16 and
Vec4Pack32To8, and LoongArch Vec2Unpack16To31, Vec2Pack31To16 and the
non-LSX Vec4Pack32To8, all of which went to the IR interpreter.
LoongArch's Vec2Pack32To16 and Vec2Unpack16To32 now take their scalar
path with LSX too, instead of falling back.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 b832ceab83 IR: Add FExp2 and FLog2 for vexp2, vlog2 and vrexp2
These were always interpreted. They're now IR ops that the native
backends compile to calls to vfpu_exp2 and vfpu_log2, like FSin and
FAsin. vrexp2 is FNeg followed by FExp2, which is how vfpu_rexp2
computes it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 13:34:19 -06:00
Henrik RydgårdandClaude Opus 5.5 a2b10778be VFPU: vsqrt, vrsq, vrcp and vnrcp are exact in every backend
These went through the host's sqrt and division everywhere except the
interpreter's vrcp and vnrcp (vsqrt and vrsq there only behind
USE_VFPU_SQRT, now gone). They now always give the PSP's bits: the IR
gets FVSqrt (FSqrt stays the FPU's IEEE sqrt.s), and FRSqrt and FRecip,
which only the VFPU emits, become vfpu_rsqrt and vfpu_rcp; the IR
interpreter and the x64, arm64, RISC-V and LoongArch backends call them.
The old JITs call them directly, the ARM ones keeping the lanes in
callee-saved registers across the calls. cpu/vfpu/exact now passes on every core.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 12:24:42 -06:00
Henrik RydgårdandClaude Fable 5.1 2d14cad065 FPU: sqrt.s of a negative gives the PSP's positive NaN
x86 (SQRTSS and libm alike) returns 0xffc00000 for it, the PSP 0x7fc00000.
-0 stays -0 and a NaN input comes through as it is, so only a negative
input needs the sign cleared: a compare and an xor in the x64 JITs, a
branch in the interpreters. cpu/fpu/roundmode.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:25:41 -06:00
Henrik RydgårdandClaude Fable 5.1 fc9dbf5ff5 FPU: saturate float-to-int in the interpreters
The C cast is undefined past the int32 range, and x86 makes it INT_MIN, so
round/trunc/ceil/floor/cvt.w.s of anything from 2^31 up gave 0x80000000 on
x86 hosts while the PSP saturates to 0x7fffffff (cpu/fpu/roundmode). Route
all of them through SaturatedFloatToInt, which also covers NaN and inf, and
drop the special cases that did.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:25:41 -06:00
Henrik RydgårdandClaude Fable 5.1 b16cd9a734 VFPU: vsgn of a denormal is zero
In the interpreter, the IR interpreter, every backend's FSign and the x86
JIT's own vsgn. Recorded in cpu/vfpu/specials.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 e0f09c55d5 MIPS: INT_MIN / -1 leaves HI at zero, like the hardware
Every path special-cased the one signed division overflow and set the
remainder to -1. cpu/cpu_alu/cpu_div, recorded on a PSP, says it's 0.
The classic arm64 JIT was the only one that got it right, by not
special-casing it at all.

arm64, RISC-V and LoongArch all produce INT_MIN with remainder 0 natively,
so their fixup blocks go away. x86 has to keep the check (IDIV traps) but
sets HI to 0 now, and so do both interpreters.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:02:20 -06:00
Henrik RydgårdandClaude Opus 5 029d17edfc IR: fix FpCondFromReg operand slot, and some smaller IR bugs
FpCondFromReg's meta is "_G", so the register is in src1 - which is where
every native backend, PropagateConstants and ReorderLoadStore read it.
But IRInterpret read it from dest, and Comp_VecDo3 wrote it to dest to
match. The other emitter passes (0, MIPS_REG_ZERO), so both fields are
zero there and the disagreement stayed hidden.

The result was that vsge/vslt, which save fpcond to IRTEMP_0 and restore
it afterwards, restored r0 (always zero) on the native IR JITs instead of
the saved value, losing any c.cond.s result live across them. Fixed both
the emitter and the interpreter to use src1.

Vec4Pack31To8's SSE2 path computed (v >> 24) << 1, which is (v >> 23)
with bit 23 forced to zero - the scalar and NEON paths both do
(v >> 23) & 0xFF. Shift left first instead. The pspautotests inputs all
happen to have bit 23 clear after the clamp, so this wasn't caught.

Also:
- Evaluate() didn't mask constant-folded shift amounts to 5 bits, though
  the neighbouring one-immediate path does.
- ApplyMemoryValidation only invalidated its address-check cache for ops
  with a 'G' destination, while Interpret and CallReplacement can write
  any GPR - both are barriers, so drop the whole cache when we see one.
- ReduceVec4Flush indexed isVec4 with (src2 & 3) instead of (src2 & ~3)
  for Vec4Scale; a missed optimization rather than a miscompile.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-19 10:31:33 -06:00
Henrik RydgårdandClaude Opus 5 3fa67f22ed Interpreter: honor the guest's FPU rounding mode and flush-to-zero
Every JIT backend puts the host FPU into the mode fcr31 asks for (bits 0-1 and
24) before running emulated code, and takes it back out before calling any host
code. The plain interpreter did none of that, so all its float math rounded to
nearest with denormals intact no matter what the game had set - cpu/fpu/fpu
fails under -i and passes under the JIT on exactly this.

Move the helpers the IR interpreter already had for this out of IRInterpreter
and into MIPS.cpp as ApplyHostRoundingMode/RestoreHostRoundingMode, and use them
around the interpreter's run loop and single step, restoring around syscalls and
replacement functions, which are host code. ctc1 re-applies immediately, since
the interpreter has no block boundary to defer it to.

round.w.s changes with it: it was floorf(x + 0.5f), which is half-away-from-zero
rather than the half-to-even every JIT produces, and the add would now pick up
the guest's rounding mode on top of that. round_ieee_754 is both correct and
mode-independent, and is what cvt.w.s already used for the same rounding.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01SfY7iFJEjmRXf1XGrTs4MF
2026-08-27 16:43:06 +02:00
Henrik Rydgård 67ddf899ba Plumb through the PC value for syscalls, so we can get better diagnostics for unresolved ones. 2026-08-15 19:14:13 +02:00
Henrik Rydgård e9a3449ede More MIPSState * plumbing (manual) 2026-08-12 14:02:19 +02:00
Henrik Rydgård cc90546301 Prep for plumbing the MIPS context pointer into the interpreter. 2026-08-12 11:25:06 +02:00
Henrik Rydgård 55a255b042 Fix the ARM version of Vec4Pack32To8. 2026-03-26 10:49:32 -06:00
Henrik Rydgård 5a5630d130 More NEON/SSE in IRInterpreter 2026-03-26 10:49:32 -06:00
Henrik Rydgård 78739104b1 Additional micro-optimizations (verified) in the IRInterpreter
Turns out that u8 promotes to int, causing signed indexing arithmetic
which is completely unnecessary.
2026-03-26 10:49:32 -06:00
Henrik Rydgård b233745640 Add NEON versions of a few more IRInterpreter instructions
Buildfix
2026-03-26 10:49:32 -06:00
Henrik Rydgård 936c344a52 misc 2026-02-09 16:41:51 +01:00
Henrik Rydgård 20e3a0cc70 Use temporaries to improve codegen in more IR vector ops 2025-10-17 09:26:15 +02:00
Henrik Rydgård d23a224e7a Use temporaries to try to improve codegen for IROp::Vec4Blend 2025-10-17 09:26:15 +02:00
Henrik Rydgård c3dfddebd7 IR interpreter: Improve code gen for the main interpreter loop
Thanks to fp64 for the idea of using unreachable markers to avoid the
range check on the switch!

Additionally, use it in a few more places.
2025-10-15 21:15:30 +02:00
Henrik Rydgård a72fc6f79c Support showing (and sorting by) IR profiler info in the JIT viewer
This will help future work on #19143 .
2025-10-09 17:02:52 -06:00
Henrik Rydgård db0cc01a81 IR interpreter on ARM64: Cleanup FPCR set/get, support floating point control on Windows 2025-10-09 11:14:47 -06:00
Henrik Rydgård e93c80db4e Cleaning up our SIMD header includes, using the new header 2024-12-19 16:08:48 +01:00
Henrik Rydgård 96c4a10e8c Add two new core states, rename RUNNING to RUNNING_CPU and similar for stepping. 2024-12-01 21:04:21 +01:00
Henrik Rydgård 7992ff4627 Make CBreakpoints an object 2024-11-25 00:22:53 +01:00
Henrik Rydgård d3e9398cb3 Split Core_EnableStepping into Core_Break and Core_Resume 2024-11-03 17:53:42 +01:00
Nemoumbra ff5877e993 Renamed the IR instruction, new UI button added 2024-09-14 19:46:05 +03:00
Nemoumbra 25f6b01d86 Added the initialization code + UI bindings + logs 2024-09-14 19:46:05 +03:00
Nemoumbra e3b09bea59 Ported the MIPSLogger's UI + basic integration of MIPSTracer 2024-09-14 19:46:05 +03:00
Nemoumbra a6be0517dc New IR instruction added 2024-09-14 19:46:04 +03:00
Henrik Rydgård 9fb97add3f Bugfixes 2024-07-26 14:22:31 +02:00
Henrik Rydgård d3e6f19b6d Comments, log, cleanup 2024-07-22 01:15:35 +02:00
Henrik Rydgård 982a83d867 IRInterpreter: Optimize variable shifts (no need to mask by 0x1f) 2024-06-24 09:30:21 +02:00
Henrik Rydgård 06315ae6ee IRInterpreter: Slight optimization for fmul
Just put stuff in temporaries, allows for better codegen
2024-06-24 09:12:57 +02:00
Henrik Rydgård 06e636bfdc Build and comment fixes 2024-06-19 20:24:59 +02:00
Henrik Rydgård e64d768113 Implement the same for ARM64 2024-06-19 20:00:36 +02:00
Henrik Rydgård 0080f71ca4 Implement FPU rounding mode support in the IR interpreter for x86/x64 2024-06-19 18:09:38 +02:00
Henrik Rydgård c9ca3904d3 Combine move-from-gpr and float cast. 2024-06-08 22:59:48 +02:00
Henrik Rydgård 0c246297d2 Create an IR op for a FPRtoGPR + shift-right-8, very common 2024-06-07 21:26:20 +02:00
Henrik Rydgård da88011805 Specialize a few arithmetic instructions for the interpreter. 2024-06-07 19:32:37 +02:00
Henrik Rydgård becc145099 Improve code generation for some IRInterpreter ops 2024-06-02 10:25:04 +02:00
Henrik Rydgård d4e3597ddb Minor codegen improvement 2024-06-02 10:23:44 +02:00
Henrik Rydgård 3b5c71170c IRInterpreter: Various SIMD optimization. Move out the reverse-bits implementation 2024-06-01 20:29:03 +02:00
Henrik Rydgård 49b0af20ca IRInterpreter: Reorder some ops towards the end, trying to keep "hot" ops together 2024-06-01 18:06:31 +02:00
Henrik Rydgård fae846e52a Remove the count parameter from IRInterpret. This is a good speed boost! 2024-05-10 23:31:24 +02:00
Henrik Rydgård 092179c42d More IR interpreter tweaks 2024-05-10 18:41:55 +02:00