Blocks ending in a branch dispatch a conditional exit and then the
fallthrough ExitToConst. One op now returns either target, reading the
second from the ExitToConst, which stays behind unexecuted.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
It avoids flushes a native backend would need, at the cost of extra
instructions (a copy of each scalar before a Vec4Scale, for instance) that
the interpreter only has to dispatch.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
OptimizeLoadsAfterStores only dropped a load right after a store of the same
reg. Now a load of anything the block stored or loaded before becomes a reg
move (with the extension for 8/16-bit loads), as long as nothing in between
may have changed the memory, the address reg, or the reg holding the value.
Only a store through the same base at a disjoint range is known not to alias,
and constant addresses outside RAM are left alone.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- A constant stays known after it's written out for a read (by a store,
MovZ, a multiply...), so later uses still fold. Whatever an op writes is
forgotten after its inputs are written, and setting a reg to the value it
already holds isn't written twice.
- A conditional exit that isn't taken keeps the constants known.
- The saturating and min/max FP ops, FSign and the 31-bit Vec2 pack/unpack
no longer flush every GPR constant.
- A load through its own base is folded (lui v0, hi; lw v0, lo(v0)).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- An lwl/lwr pair was combined into one load even when the first half loads
into the base register, which changes the address of the second half.
- ApplyMemoryValidation shared one sp check across the block even past an
Interpret or CallReplacement, which may change sp.
- Drop a duplicate FSqrt meta entry, and name Load8Ext correctly.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Copy propagation through an FPR temp didn't stop when an instruction
rewrote the temp in place (it compared an FPR number against the +32
offset reg), so later reads lost that write.
- A read of the temp in both operands only had src1 replaced, yet the copy
into the temp was still removed.
- The replacement matched operands by number without checking their type,
so a StoreFloat whose GPR address had the temp's number got its address
replaced (IRVTEMP_PFX_S and IRTEMP_0 are both 192).
- A write to lanes 1-3 of a Vec4 temp wasn't noticed.
- IRReadsFromFPRs stopped after the F operands, missing Vec4Scale's vector.
- Exits and barriers didn't count as reading everything, so a write to a
real reg could be moved above an exit.
- Load32Linked and Store32Conditional were removed when their reg was
overwritten unread, losing LLBIT and the store.
Also fixes an off-by-one in the vec src3 read check. The unit test now
reports every failing case instead of stopping at the first.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A new FSinCos op writes both from one argument reduction. arm64 and x64 get
both back from a single call, packed in one double; RISC-V and LoongArch make
the two calls.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
vh2f always went to the interpreter in the IR. A new FHalfToFloat op
converts the lower or upper half of a word, and the native backends
call vfpu_h2f for it like FSin. The legacy arm64 JIT now makes the same
call instead of computing the conversion inline; vh2f is rare, and the
call is much less code.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Vec4ClampToZero and Vec2ClampToZero only ever fed Vec4Pack31To8 and
Vec2Pack31To16, for vi2uc and vi2us. The packs now clamp negative lanes
to zero themselves, which saves an op and a vector temp, and lets x64
clamp with PACKUSWB's saturation after an arithmetic shift.
While at it, RISC-V compiles Vec2Unpack16To31, Vec2Pack31To16 and
Vec4Pack32To8, and LoongArch Vec2Unpack16To31, Vec2Pack31To16 and the
non-LSX Vec4Pack32To8, all of which went to the IR interpreter.
LoongArch's Vec2Pack32To16 and Vec2Unpack16To32 now take their scalar
path with LSX too, instead of falling back.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
These were always interpreted. They're now IR ops that the native
backends compile to calls to vfpu_exp2 and vfpu_log2, like FSin and
FAsin. vrexp2 is FNeg followed by FExp2, which is how vfpu_rexp2
computes it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
These went through the host's sqrt and division everywhere except the
interpreter's vrcp and vnrcp (vsqrt and vrsq there only behind
USE_VFPU_SQRT, now gone). They now always give the PSP's bits: the IR
gets FVSqrt (FSqrt stays the FPU's IEEE sqrt.s), and FRSqrt and FRecip,
which only the VFPU emits, become vfpu_rsqrt and vfpu_rcp; the IR
interpreter and the x64, arm64, RISC-V and LoongArch backends call them.
The old JITs call them directly, the ARM ones keeping the lanes in
callee-saved registers across the calls. cpu/vfpu/exact now passes on every core.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
FpCondFromReg's meta is "_G", so the register is in src1 - which is where
every native backend, PropagateConstants and ReorderLoadStore read it.
But IRInterpret read it from dest, and Comp_VecDo3 wrote it to dest to
match. The other emitter passes (0, MIPS_REG_ZERO), so both fields are
zero there and the disagreement stayed hidden.
The result was that vsge/vslt, which save fpcond to IRTEMP_0 and restore
it afterwards, restored r0 (always zero) on the native IR JITs instead of
the saved value, losing any c.cond.s result live across them. Fixed both
the emitter and the interpreter to use src1.
Vec4Pack31To8's SSE2 path computed (v >> 24) << 1, which is (v >> 23)
with bit 23 forced to zero - the scalar and NEON paths both do
(v >> 23) & 0xFF. Shift left first instead. The pspautotests inputs all
happen to have bit 23 clear after the clamp, so this wasn't caught.
Also:
- Evaluate() didn't mask constant-folded shift amounts to 5 bits, though
the neighbouring one-immediate path does.
- ApplyMemoryValidation only invalidated its address-check cache for ops
with a 'G' destination, while Interpret and CallReplacement can write
any GPR - both are barriers, so drop the whole cache when we see one.
- ReduceVec4Flush indexed isVec4 with (src2 & 3) instead of (src2 & ~3)
for Vec4Scale; a missed optimization rather than a miscompile.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
* Rename LogType to Log
* Explicitly use the Log:: enum when logging. Allows for autocomplete when editing.
* Mac/ARM64 buildfix
* Do the same with the hle result log macros
* Rename the log names to mixed case while at it.
* iOS buildfix
* Qt buildfix attempt, ARM32 buildfix