126 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 7eb371b231 IR interpreter: Merge a conditional exit with the ExitToConst after it
Blocks ending in a branch dispatch a conditional exit and then the
fallthrough ExitToConst. One op now returns either target, reading the
second from the ExitToConst, which stays behind unexecuted.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 11:47:17 -06:00
Henrik RydgårdandClaude Opus 5.5 c6aea794f6 IR interpreter: Skip ReduceVec4Flush
It avoids flushes a native backend would need, at the cost of extra
instructions (a copy of each scalar before a Vec4Scale, for instance) that
the interpreter only has to dispatch.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 11:47:17 -06:00
Henrik RydgårdandClaude Opus 5.5 dab5d82227 IR: Forward stores and loads to later loads in the block
OptimizeLoadsAfterStores only dropped a load right after a store of the same
reg. Now a load of anything the block stored or loaded before becomes a reg
move (with the extension for 8/16-bit loads), as long as nothing in between
may have changed the memory, the address reg, or the reg holding the value.
Only a store through the same base at a disjoint range is known not to alias,
and constant addresses outside RAM are left alone.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:56:03 -06:00
Henrik RydgårdandClaude Opus 5.5 498511fd51 IR: Drop everything after a conditional exit folded to always taken
Only the ExitToConst right after it was skipped before.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:39:24 -06:00
Henrik RydgårdandClaude Opus 5.5 a99a82cc6b IR: Keep more constants known in PropagateConstants
- A constant stays known after it's written out for a read (by a store,
  MovZ, a multiply...), so later uses still fold. Whatever an op writes is
  forgotten after its inputs are written, and setting a reg to the value it
  already holds isn't written twice.
- A conditional exit that isn't taken keeps the constants known.
- The saturating and min/max FP ops, FSign and the 31-bit Vec2 pack/unpack
  no longer flush every GPR constant.
- A load through its own base is folded (lui v0, hi; lw v0, lo(v0)).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:35:11 -06:00
Henrik RydgårdandClaude Opus 5.5 25e8f1d57c IR: Fix lwl/lwr pairing into the base reg, and sp validation past barriers
- An lwl/lwr pair was combined into one load even when the first half loads
  into the base register, which changes the address of the second half.
- ApplyMemoryValidation shared one sp check across the block even past an
  Interpret or CallReplacement, which may change sp.
- Drop a duplicate FSqrt meta entry, and name Load8Ext correctly.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:35:11 -06:00
Henrik RydgårdandClaude Opus 5.5 1dbb62d570 IR: Fix PurgeTemps miscompiles
- Copy propagation through an FPR temp didn't stop when an instruction
  rewrote the temp in place (it compared an FPR number against the +32
  offset reg), so later reads lost that write.
- A read of the temp in both operands only had src1 replaced, yet the copy
  into the temp was still removed.
- The replacement matched operands by number without checking their type,
  so a StoreFloat whose GPR address had the temp's number got its address
  replaced (IRVTEMP_PFX_S and IRTEMP_0 are both 192).
- A write to lanes 1-3 of a Vec4 temp wasn't noticed.
- IRReadsFromFPRs stopped after the F operands, missing Vec4Scale's vector.
- Exits and barriers didn't count as reading everything, so a write to a
  real reg could be moved above an exit.
- Load32Linked and Store32Conditional were removed when their reg was
  overwritten unread, losing LLBIT and the store.

Also fixes an off-by-one in the vec src3 read check. The unit test now
reports every failing case instead of stopping at the first.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:35:11 -06:00
Henrik RydgårdandClaude Opus 5.5 f92340c092 IR: Compute vrot's sine and cosine in one call
A new FSinCos op writes both from one argument reduction. arm64 and x64 get
both back from a single call, packed in one double; RISC-V and LoongArch make
the two calls.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:22:41 -06:00
Henrik RydgårdandClaude Opus 5.5 6b78138c08 VFPU: Compile vh2f in the IR, and call vfpu_h2f for it everywhere
vh2f always went to the interpreter in the IR. A new FHalfToFloat op
converts the lower or upper half of a word, and the native backends
call vfpu_h2f for it like FSin. The legacy arm64 JIT now makes the same
call instead of computing the conversion inline; vh2f is rare, and the
call is much less code.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 eaf55c467b IR: Fold ClampToZero into the 31-bit packs
Vec4ClampToZero and Vec2ClampToZero only ever fed Vec4Pack31To8 and
Vec2Pack31To16, for vi2uc and vi2us. The packs now clamp negative lanes
to zero themselves, which saves an op and a vector temp, and lets x64
clamp with PACKUSWB's saturation after an arithmetic shift.

While at it, RISC-V compiles Vec2Unpack16To31, Vec2Pack31To16 and
Vec4Pack32To8, and LoongArch Vec2Unpack16To31, Vec2Pack31To16 and the
non-LSX Vec4Pack32To8, all of which went to the IR interpreter.
LoongArch's Vec2Pack32To16 and Vec2Unpack16To32 now take their scalar
path with LSX too, instead of falling back.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 b832ceab83 IR: Add FExp2 and FLog2 for vexp2, vlog2 and vrexp2
These were always interpreted. They're now IR ops that the native
backends compile to calls to vfpu_exp2 and vfpu_log2, like FSin and
FAsin. vrexp2 is FNeg followed by FExp2, which is how vfpu_rexp2
computes it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 13:34:19 -06:00
Henrik RydgårdandClaude Opus 5.5 a2b10778be VFPU: vsqrt, vrsq, vrcp and vnrcp are exact in every backend
These went through the host's sqrt and division everywhere except the
interpreter's vrcp and vnrcp (vsqrt and vrsq there only behind
USE_VFPU_SQRT, now gone). They now always give the PSP's bits: the IR
gets FVSqrt (FSqrt stays the FPU's IEEE sqrt.s), and FRSqrt and FRecip,
which only the VFPU emits, become vfpu_rsqrt and vfpu_rcp; the IR
interpreter and the x64, arm64, RISC-V and LoongArch backends call them.
The old JITs call them directly, the ARM ones keeping the lanes in
callee-saved registers across the calls. cpu/vfpu/exact now passes on every core.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 12:24:42 -06:00
Henrik RydgårdandClaude Opus 5 029d17edfc IR: fix FpCondFromReg operand slot, and some smaller IR bugs
FpCondFromReg's meta is "_G", so the register is in src1 - which is where
every native backend, PropagateConstants and ReorderLoadStore read it.
But IRInterpret read it from dest, and Comp_VecDo3 wrote it to dest to
match. The other emitter passes (0, MIPS_REG_ZERO), so both fields are
zero there and the disagreement stayed hidden.

The result was that vsge/vslt, which save fpcond to IRTEMP_0 and restore
it afterwards, restored r0 (always zero) on the native IR JITs instead of
the saved value, losing any c.cond.s result live across them. Fixed both
the emitter and the interpreter to use src1.

Vec4Pack31To8's SSE2 path computed (v >> 24) << 1, which is (v >> 23)
with bit 23 forced to zero - the scalar and NEON paths both do
(v >> 23) & 0xFF. Shift left first instead. The pspautotests inputs all
happen to have bit 23 clear after the clamp, so this wasn't caught.

Also:
- Evaluate() didn't mask constant-folded shift amounts to 5 bits, though
  the neighbouring one-immediate path does.
- ApplyMemoryValidation only invalidated its address-check cache for ops
  with a 'G' destination, while Interpret and CallReplacement can write
  any GPR - both are barriers, so drop the whole cache when we see one.
- ReduceVec4Flush indexed isVec4 with (src2 & 3) instead of (src2 & ~3)
  for Vec4Scale; a missed optimization rather than a miscompile.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-19 10:31:33 -06:00
Henrik Rydgård ae14ebb6ac IRWriter: Remove the confusing and inefficient AddConstant 2026-08-15 18:31:20 +02:00
Henrik Rydgård e51c58716b Fix minor code issues flagged by PVS-Studio and reported by alphrixus. 2024-10-10 14:10:30 +02:00
Henrik Rydgård e01ca5b057 Logging API change (refactor) (#19324)
* Rename LogType to Log

* Explicitly use the Log:: enum when logging. Allows for autocomplete when editing.

* Mac/ARM64 buildfix

* Do the same with the hle result log macros

* Rename the log names to mixed case while at it.

* iOS buildfix

* Qt buildfix attempt, ARM32 buildfix
2024-07-14 14:42:59 +02:00
Henrik Rydgård 5bfc0259be Optimize away the redundant vector loads in the Wipeout games. 2024-06-08 23:22:58 +02:00
Henrik Rydgård c9ca3904d3 Combine move-from-gpr and float cast. 2024-06-08 22:59:48 +02:00
Henrik Rydgård 0abcd00372 IR Interpreter: Don't split up lone lwr/swr/lwl/swl instructions, the interpreter handles them faster. 2024-06-08 22:28:11 +02:00
Henrik Rydgård 0c246297d2 Create an IR op for a FPRtoGPR + shift-right-8, very common 2024-06-07 21:26:20 +02:00
Henrik Rydgård bd0beb68a4 Add new IR optimization pass, OptimizeLoadsAfterStores 2024-06-07 19:32:37 +02:00
Henrik Rydgård da88011805 Specialize a few arithmetic instructions for the interpreter. 2024-06-07 19:32:37 +02:00
Henrik Rydgård bdf7f5f40f IRInterpreter: Fix issue where we could accidentally optimize out CallReplacement ops. 2024-05-26 00:12:48 +02:00
Unknown W. Brackets 5e813e6bd6 irjit: Correct bad Vec4 overlap handling. 2023-10-14 20:54:40 -07:00
Unknown W. Brackets cd46f0b4cb irjit: Cache IR metadata lookups.
This improves compilation performance, because all those lookups were
adding up.
2023-09-30 15:56:53 -07:00
Unknown W. Brackets 69b8fb9bc2 irjit: Correct alignment checks against SP. 2023-09-23 13:15:06 -07:00
Unknown W. Brackets 739e474957 irjit: Improve dot and store vec4 interaction. 2023-09-01 22:35:59 -07:00
Unknown W. Brackets 2ca638868c irjit: Remove Vec4Scale/Vec4Dot aliasing in pass. 2023-09-01 22:35:59 -07:00
Unknown W. Brackets 96c90c19c6 irjit: Reduce unnecessary shuffles.
If it's never used as a Vec4 again, it's often not worth it.
2023-09-01 22:35:59 -07:00
Unknown W. Brackets 5f84887dea irjit: Add a pass to keep Vec4s in Vec4s. 2023-09-01 22:35:59 -07:00
Henrik Rydgård ae0af175de Merge pull request #18004 from unknownbrackets/x86-jit-minor
x86jit: Small cleanup and tweaks, downcountInRegster, GetFPRLaneCount bugfix
2023-08-29 08:01:47 +02:00
Unknown W. Brackets 14d871730f x86jit: Optimize slow memory. 2023-08-28 21:16:38 -07:00
Unknown W. Brackets 8e8605935e irjit: Detect subtract-to-zero as constant. 2023-08-28 21:09:56 -07:00
Unknown W. Brackets 4e90bd5070 x86jit: Use NEG more automatically for Sub x,0,y. 2023-08-28 21:09:56 -07:00
Unknown W. Brackets f263698897 irjit: Cleanup temp purging on exit.
We were sometimes considering it read by exit and not purging.
2023-08-27 12:26:05 -07:00
Unknown W. Brackets 552cd88938 irjit: Skip some work in PurgeTemps. 2023-08-20 08:59:47 -07:00
Unknown W. Brackets 57123e8f9e irjit: Reserve some arrays that churn.
Improves IR compile time by around 20-30%.
2023-08-20 08:59:47 -07:00
Henrik Rydgård 2cdcc413b7 Merge pull request #17898 from unknownbrackets/irjit-vfputemps
irjit: Cleanup/purge FPU/VFPU temps
2023-08-13 21:08:00 +02:00
Unknown W. Brackets 2b36e0a625 irjit: ZeroFpCond -> FpCondFromReg.
We already have a zero reg, so this is more useful and symmetrical.
2023-08-13 10:40:47 -07:00
Unknown W. Brackets fa53b80574 irjit: Cleanup/purge FPU/VFPU temps.
A lot of cases are followed by an FMov that just moved the temp to a
regular register, from VFPU instructions playing safe about overlaps.
2023-08-13 10:14:10 -07:00
Unknown W. Brackets 31ff23746c irjit: Prefer ExitToConst over ExitToReg. 2023-08-08 23:14:01 -07:00
Unknown W. Brackets 3f97545f99 irjit: Reduce exits from constants.
This reduces bloat a bit, and may help with common funcs that enter short
loops.
2023-08-08 23:05:14 -07:00
Unknown W. Brackets 5f9a8fd1a1 irjit: Rename IRRegCache to IRImmRegCache.
For clarity, since it's not a native regcache.
2023-08-08 23:05:14 -07:00
Henrik Rydgård e9431d0d1e Merge pull request #17859 from unknownbrackets/irjit-vec4
irjit: Use Vec4 a bit more
2023-08-06 23:05:33 +02:00
Unknown W. Brackets 79ca880ac7 irjit: Implement vqmul, add Vec4Blend.
Should be useful more places.
2023-08-06 13:38:00 -07:00
Unknown W. Brackets a32889d3ca irjit: Consistently dirty vfpuctrl in IR. 2023-08-06 08:36:19 -07:00
Unknown W. Brackets b03398a46c Merge pull request #17815 from unknownbrackets/riscv-jit
riscv: Spill registers more intelligently
2023-07-30 14:49:37 -07:00
Unknown W. Brackets f870271011 riscv: Spill registers more intelligently. 2023-07-30 14:24:12 -07:00
Unknown W. Brackets 5db6b11ef2 irjit: Cleanup self-fmovs.
These were sometimes getting emitted.
2023-07-30 14:16:17 -07:00
Henrik Rydgård 180bda6f6b Merge pull request #17799 from unknownbrackets/irjit-lsu
Add ll/sc to IR and x86jit
2023-07-30 09:15:55 +02:00