142 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 1cdc432d6a IR: Let vrot's FSinCos write straight into an [s, c] pair
When the cosine lane follows the sine lane, FSinCos can write both in place
instead of going through a temp and two FMovs.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:42:03 -06:00
Henrik RydgårdandClaude Opus 5.5 f92340c092 IR: Compute vrot's sine and cosine in one call
A new FSinCos op writes both from one argument reduction. arm64 and x64 get
both back from a single call, packed in one double; RISC-V and LoongArch make
the two calls.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:22:41 -06:00
Henrik RydgårdandClaude Opus 5.5 6fc4eb19df VFPU: Fix vrot with the angle in a destination lane
The cosine is then taken of what vrot wrote to that lane: the sine, or zero.
The IR looked at the sine lane instead of the lane holding the angle, and the
legacy JITs ignored the overlap. The assembler refuses such a vrot, so those
now leave it to the interpreter, and don't pair one with the vrot before it.

Covered by the new cpu/vfpu/vrot test.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:22:41 -06:00
Henrik RydgårdandClaude Opus 5.5 d1a8ce0dcc IR: Compile vmscl on any size and transposition
Only 4x4 with source and destination transposed alike, and a scale
outside the destination, compiled; most vmscl in games are transposed
or 3x3. The rest now multiply element by element, with the scale
copied first, and only a partly overlapping source still falls back.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 9def690781 IR: Compile vt4444, vt5551 and vt5650
They always went to the interpreter. Each channel is now a shift, mask
and shift in the existing integer ops, so every backend handles them.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 6b78138c08 VFPU: Compile vh2f in the IR, and call vfpu_h2f for it everywhere
vh2f always went to the interpreter in the IR. A new FHalfToFloat op
converts the lower or upper half of a word, and the native backends
call vfpu_h2f for it like FSin. The legacy arm64 JIT now makes the same
call instead of computing the conversion inline; vh2f is rare, and the
call is much less code.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 eaf55c467b IR: Fold ClampToZero into the 31-bit packs
Vec4ClampToZero and Vec2ClampToZero only ever fed Vec4Pack31To8 and
Vec2Pack31To16, for vi2uc and vi2us. The packs now clamp negative lanes
to zero themselves, which saves an op and a vector temp, and lets x64
clamp with PACKUSWB's saturation after an arithmetic shift.

While at it, RISC-V compiles Vec2Unpack16To31, Vec2Pack31To16 and
Vec4Pack32To8, and LoongArch Vec2Unpack16To31, Vec2Pack31To16 and the
non-LSX Vec4Pack32To8, all of which went to the IR interpreter.
LoongArch's Vec2Pack32To16 and Vec2Unpack16To32 now take their scalar
path with LSX too, instead of falling back.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 b832ceab83 IR: Add FExp2 and FLog2 for vexp2, vlog2 and vrexp2
These were always interpreted. They're now IR ops that the native
backends compile to calls to vfpu_exp2 and vfpu_log2, like FSin and
FAsin. vrexp2 is FNeg followed by FExp2, which is how vfpu_rexp2
computes it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 13:34:19 -06:00
Henrik RydgårdandClaude Opus 5.5 a2b10778be VFPU: vsqrt, vrsq, vrcp and vnrcp are exact in every backend
These went through the host's sqrt and division everywhere except the
interpreter's vrcp and vnrcp (vsqrt and vrsq there only behind
USE_VFPU_SQRT, now gone). They now always give the PSP's bits: the IR
gets FVSqrt (FSqrt stays the FPU's IEEE sqrt.s), and FRSqrt and FRecip,
which only the VFPU emits, become vfpu_rsqrt and vfpu_rcp; the IR
interpreter and the x64, arm64, RISC-V and LoongArch backends call them.
The old JITs call them directly, the ARM ones keeping the lanes in
callee-saved registers across the calls. cpu/vfpu/exact now passes on every core.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 12:24:42 -06:00
Henrik RydgårdandClaude Fable 5.1 613422c15f IR: vi2s/vi2us with the destination inside the source
The second pack read a lane the first one had just written
(vi2s.q C002, C000). Pack into temps when the outputs overlap the inputs.
Found by the corrected cpu/vfpu/overlap.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 783892ed25 VFPU: Writing an RNG state register keeps 0x3F8 in the top bits
mtvc to RCX0-7 keeps the low 20 bits of the value and forces the top
twelve to 0x3F8 (cpu/vfpu/vrnd), which is also the form vrnds and vrndi
leave them in. We masked with 0x3FFFFFFF. The generator itself matches
hardware for every seed and state in that test.

Interpreter, IR and the arm64 JIT. The x86 and ARM32 JITs don't mask
control register writes at all and are left as they were.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 12:49:06 -06:00
Henrik RydgårdandClaude Fable 5.1 591731cb8e IR: A swizzle naming a lane past the op's size isn't "within size"
IsPrefixWithinSize only looked at prefix positions past the size, not at
in-size positions naming a lane past it, so vadd.p with [z, w, ...] was
compiled with whatever the prefix code substitutes. Hardware writes those
result lanes as zero (cpu/vfpu/prefix_ctrl), which the interpreter does,
so send them there.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:37:48 -06:00
Henrik RydgårdandClaude Opus 5 315d128141 VFPU: make vmfvc see a pending prefix, and actually write its result
Comp_Vmfvc read vfpuCtrl[] straight from the context in all four
backends, while the mfvc path in Comp_Mftv flushes first, with the
comment "In case we have a saved prefix" - so "vpfxs X" followed by
"vmfvc sN, $128" returned the stale value in memory. The IR frontend
flushes only for the three prefix registers, which is the tighter form,
so vmfvc does the same there.

On ARM and ARM64 the fix alone wouldn't have been observable: those two
map the destination with no flags, leaving isDirty false, so the flush
dropped the loaded value without storing it. x86 already passes
MAP_DIRTY | MAP_NOINIT. That part is a fix of its own, but the two are
inseparable in this function - a vmfvc that reads the right value and
then throws it away is no better.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 11:32:24 -06:00
Henrik RydgårdandClaude Opus 5 029d17edfc IR: fix FpCondFromReg operand slot, and some smaller IR bugs
FpCondFromReg's meta is "_G", so the register is in src1 - which is where
every native backend, PropagateConstants and ReorderLoadStore read it.
But IRInterpret read it from dest, and Comp_VecDo3 wrote it to dest to
match. The other emitter passes (0, MIPS_REG_ZERO), so both fields are
zero there and the disagreement stayed hidden.

The result was that vsge/vslt, which save fpcond to IRTEMP_0 and restore
it afterwards, restored r0 (always zero) on the native IR JITs instead of
the saved value, losing any c.cond.s result live across them. Fixed both
the emitter and the interpreter to use src1.

Vec4Pack31To8's SSE2 path computed (v >> 24) << 1, which is (v >> 23)
with bit 23 forced to zero - the scalar and NEON paths both do
(v >> 23) & 0xFF. Shift left first instead. The pspautotests inputs all
happen to have bit 23 clear after the clamp, so this wasn't caught.

Also:
- Evaluate() didn't mask constant-folded shift amounts to 5 bits, though
  the neighbouring one-immediate path does.
- ApplyMemoryValidation only invalidated its address-check cache for ops
  with a 'G' destination, while Interpret and CallReplacement can write
  any GPR - both are barriers, so drop the whole cache when we see one.
- ReduceVec4Flush indexed isVec4 with (src2 & 3) instead of (src2 & ~3)
  for Vec4Scale; a missed optimization rather than a miscompile.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-19 10:31:33 -06:00
Henrik Rydgård a810597ac0 Implement MMIO for the JIT (by falling back to the interpreter for load/stores from kernel addresses)
Fixes the VSH in JIT mode (but NOT ir)
2026-08-25 00:22:40 +02:00
Henrik Rydgård 0b90e42214 IR Interpreter with fastmemory off: Validate alignment of memory accesses 2026-08-18 10:59:08 +02:00
Henrik Rydgård ae14ebb6ac IRWriter: Remove the confusing and inefficient AddConstant 2026-08-15 18:31:20 +02:00
Henrik Rydgård e01ca5b057 Logging API change (refactor) (#19324)
* Rename LogType to Log

* Explicitly use the Log:: enum when logging. Allows for autocomplete when editing.

* Mac/ARM64 buildfix

* Do the same with the hle result log macros

* Rename the log names to mixed case while at it.

* iOS buildfix

* Qt buildfix attempt, ARM32 buildfix
2024-07-14 14:42:59 +02:00
Henrik Rydgård 5e3abe3def IRInterpreter compiler: Reject all vec2ops where the prefix is unknown
Fixes #19164

May help #19172
2024-05-22 23:19:14 +02:00
Henrik Rydgård 5d8a0b3ac7 Merge pull request #18266 from unknownbrackets/ir-vtfm
irjit: Fix vhtfm instruction
2023-09-29 09:43:06 +02:00
Unknown W. Brackets c92148ee2c irjit: Fix vhtfm instruction. 2023-09-28 21:16:54 -07:00
Unknown W. Brackets e104a28b71 irjit: Handle VDet. 2023-09-24 23:03:25 -07:00
Unknown W. Brackets c85886c11e irjit: Use enum for rounding modes. 2023-09-01 22:29:24 -07:00
Unknown W. Brackets 35fe15d718 x86jit: Do not use Vec4Dot for vdot.t.
It was much slower to do so in LittleBigPlanet.
2023-08-27 12:39:21 -07:00
Unknown W. Brackets 74e5e43fdc jit: Skip known prefix writes.
If we already know what's in memory and it's default, we can skip
overwriting with default values.  This is common, actually.
2023-08-22 23:26:31 -07:00
Unknown W. Brackets 82fb41cba0 irjit: Implement vtfm 4x4 using dots. 2023-08-20 13:50:02 -07:00
Unknown W. Brackets bd1d93ae6f irjit: Cleanup Write() calls with extra const.
Some instructions, such as Vec4Blend, are encoded requiring the const
field, and this interface was designed when we used a pool.
2023-08-19 16:23:42 -07:00
Henrik Rydgård 1b2cffe632 Address feedback 2023-08-14 11:06:20 +02:00
Henrik Rydgård ff6e118fff Get rid of a lot of ifdefs around presentation mode. Instead, set things dynamically. 2023-08-14 11:02:29 +02:00
Unknown W. Brackets 159b41a0fa irjit: Fuse unaligned svl.q/svr.q together.
They're almost never used outside paired, which we can do on most
platforms easily.
2023-08-13 18:10:40 -07:00
Unknown W. Brackets 5729de90d2 irjit: Use more partial Vec4s / Vec4Blend. 2023-08-13 18:10:40 -07:00
Unknown W. Brackets 2e6dbab5fa irjit: Add flag to prefer Vec4, use for add/sub.
This will improve things when using SIMD.
2023-08-13 18:10:40 -07:00
Unknown W. Brackets e0be6858b8 irjit: Implement vcrs.t.
As used in Jeanne d'Arc.
2023-08-13 18:10:12 -07:00
Unknown W. Brackets 217a1837ed irjit: Allow typical prefixes in vdiv/vasin/etc.
Some of these behave strangely, but there are some common usages that work
fine.
2023-08-13 18:10:07 -07:00
Unknown W. Brackets 23c79f8e7f irjit: Implement vsge/vslt.
These are not ideal especially for SIMD, but they do work.
Improves performance in Silent Hill on RISC-V by like 20%.
2023-08-13 10:40:47 -07:00
Unknown W. Brackets 5d20f2aabd irjit: Simplify VecDo3. 2023-08-13 10:40:47 -07:00
Henrik Rydgård 2342c4522c Merge pull request #17875 from unknownbrackets/riscv-jit
RISC-V: Implement a few more ops
2023-08-09 09:30:15 +02:00
Unknown W. Brackets 28c58c1d24 irjit: Allow more forms of vmidt.
Mildly worth it.
2023-08-08 23:17:32 -07:00
Unknown W. Brackets e73c203984 irjit: Fix Vec4Shuffle overlap issue. 2023-08-08 23:00:39 -07:00
Henrik Rydgård e9431d0d1e Merge pull request #17859 from unknownbrackets/irjit-vec4
irjit: Use Vec4 a bit more
2023-08-06 23:05:33 +02:00
Unknown W. Brackets 3dc71cff75 irjit: Keep a couple more ops in Vec4. 2023-08-06 13:46:24 -07:00
Unknown W. Brackets 6a1dbd4cde irjit: Allow Vec4 to be used with masks. 2023-08-06 13:46:24 -07:00
Unknown W. Brackets 2b964fd3b0 irjit: Handle more common Vec4 prefix cases. 2023-08-06 13:38:00 -07:00
Unknown W. Brackets 79ca880ac7 irjit: Implement vqmul, add Vec4Blend.
Should be useful more places.
2023-08-06 13:38:00 -07:00
Unknown W. Brackets 85ee7c85c1 irjit: Allow masked vneg.q. 2023-08-06 13:38:00 -07:00
Unknown W. Brackets a29a35b91a irjit: Fix mfvc eating prefixes.
It doesn't and shouldn't, which is why it's marked as not.
2023-08-06 08:28:25 -07:00
Unknown W. Brackets 5db6b11ef2 irjit: Cleanup self-fmovs.
These were sometimes getting emitted.
2023-07-30 14:16:17 -07:00
Unknown W. Brackets e228748449 irjit: Add FCvtScaledSW to safely scale vi2f. 2023-07-29 18:30:15 -07:00
Unknown W. Brackets a5a2671af3 irjit: Implement vf2ix.
Used in LittleBigPlanet when playing intro movies.
2023-07-29 18:01:08 -07:00
Unknown W. Brackets d97790e28e irjit: Fix vi2us/vi2s with non-consecutive.
Vec2ClampToZero and similar assume consecutive.
2023-03-15 21:30:35 -07:00