104 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 f92340c092 IR: Compute vrot's sine and cosine in one call
A new FSinCos op writes both from one argument reduction. arm64 and x64 get
both back from a single call, packed in one double; RISC-V and LoongArch make
the two calls.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:22:41 -06:00
Henrik RydgårdandClaude Opus 5.5 6b78138c08 VFPU: Compile vh2f in the IR, and call vfpu_h2f for it everywhere
vh2f always went to the interpreter in the IR. A new FHalfToFloat op
converts the lower or upper half of a word, and the native backends
call vfpu_h2f for it like FSin. The legacy arm64 JIT now makes the same
call instead of computing the conversion inline; vh2f is rare, and the
call is much less code.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 9b4d4ee35b VFPU: Build sin, cos, asin, exp2 and log2 results as integer bits
A float made from a constant NaN bit pattern can come out quieted: MSVC
turned the 0x7F800001 that vcos, vexp2 and vlog2 return into 0x7FC00001,
which failed cpu/vfpu/exact on Windows. Each function now computes its
result's bits in integers and converts once at the end, like vrcp and
friends already did. Fixed-point results become floats by shifting,
which is exact since the VFPU keeps 22 significant bits. Identical to
the previous code over every 32-bit input.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 12:59:04 -06:00
Henrik RydgårdandClaude Opus 5.5 b9144d4fca VFPU: Fold the result exponent into the rcp, rsqrt and sqrt coefficients
Each of the three is now a range check, a fast path and a few special
cases. The fast path reads its segment from a table whose constant term
already includes the result's exponent bits, so what remains is two
multiplies, some shifts and one exponent adjustment, all in integers.
Identical to the previous code over every 32-bit input.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 12:24:43 -06:00
Henrik RydgårdandClaude Opus 5.5 d6b3d590eb VFPU: Compute log2, sin/cos and asin without correction tables, too
They use the same quadratic interpolator as rcp and friends, with three
twists. sin indexes the quarter wave from the top, and asin and sin work
in a per-segment exponent whose 4-ulp truncation also applies to results
in a lower binade. log2 truncates exponent + log2(1.m) toward zero to 22
significant bits, and where that step is coarser than 2^-24 the datapath
drops coefficient bits to match; that also covers the region just below
1.0 that needed a special case.

vfpu_sincos now reduces the angle once. With every table gone, so are
the asset folder, the loader, InitVFPU and the fallbacks for tables that
failed to load. All seven functions are bit-exact with the table-based
code over every 32-bit input.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 10:07:01 -06:00
Henrik RydgårdandClaude Opus 5.5 dade876474 VFPU: Compute vrcp, vrsq, vsqrt and vexp2 without correction tables
The four share one quadratic interpolator: 128 segments picked by the top
7 bits of the input, each with a constant, a linear and a squared-term
coefficient, and a squarer on the top 10 bits of the rest that rounds t^2
up to a multiple of 256. 128 small coefficient sets per function
replace the 1 MB of delta tables. Derived from the output of the
table-based code, and bit-exact with it over every 32-bit input.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 10:06:53 -06:00
Henrik RydgårdandClaude Opus 5.5 cd005b9041 VFPU: SSE2 version of the exact vdot
Same shape as the NEON one, which now shares its rounding tail. SSE2 has
no per-lane shift, so the alignment shift is a multiply by a power of two
built from float bits and converted by truncation; the unsigned maxima
use the 16-bit instructions, since every value involved fits in 15 bits.
Nothing depends on the host rounding mode or flush-to-zero.

Checked against the reference by VFPUDot and on 300M more inputs offline,
also with MXCSR set to round toward zero with FTZ and DAZ.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 08:18:25 -06:00
Henrik RydgårdandClaude Opus 5.5 649561764d VFPU: NEON version of the exact vdot
The four lanes are computed together: exponents, 24x24-bit products with
round-to-odd, alignment by truncation and a signed horizontal sum. One
pairwise maximum finds both the alignment exponent and any inf or NaN,
which go to the reference. The final rounding is branch-free and in
integers, since the host rounding mode may be the game's.

About three times the throughput of the reference on Apple M-series
(5.2 vs 15.2 ns per call). VFPUDot checks it against the reference on
four million inputs picked to cover cancellation, ties, subnormals and
the overflow edges; a billion more matched offline.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 08:18:25 -06:00
Henrik RydgårdandClaude Fable 5.1 ed6d13e15f VFPU: bit-exact vh2f and vf2h
vf2h truncates the mantissa (no rounding), gives a signed zero below 2^-14
and a signed inf from 65536 up, and keeps the low ten mantissa bits of a
NaN, so one with those clear becomes inf. vh2f flushes a subnormal half to
a signed zero and keeps inf/NaN mantissa bits unshifted. The x86 JIT's own
vh2f gets the subnormal flush; the other backends go through the
interpreter. Recorded in cpu/vfpu/specials.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 783892ed25 VFPU: Writing an RNG state register keeps 0x3F8 in the top bits
mtvc to RCX0-7 keeps the low 20 bits of the value and forces the top
twelve to 0x3F8 (cpu/vfpu/vrnd), which is also the form vrnds and vrndi
leave them in. We masked with 0x3FFFFFFF. The generator itself matches
hardware for every seed and state in that test.

Interpreter, IR and the arm64 JIT. The x86 and ARM32 JITs don't mask
control register writes at all and are left as they were.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 12:49:06 -06:00
Henrik RydgårdandClaude Opus 5 675ad0acd6 Interpreter: assorted fixes found while reviewing Core/MIPS
WriteMMIO_U32 was missing the return after the kernel-mode check, so a
user-mode write raised the exception and then went through anyway. The
other five MMIO accessors already return here; this one lost it when the
GPIO/syscon branches were added.

Int_Vrot scanned all four entries of dregs, but GetVectorRegs only fills
the first n - so a vrot with vs == 0 (S000) matched a lane that isn't
there and took the cosine from it. The IR backend gets this right via
IsOverlapSafe, so the two disagreed.

The breakpoint checks in MIPSInterpret and RunUntilDowncountZeroWithChecks
read instr->flags without checking for null, which MIPSGetInstruction
returns for the eight primary opcodes (and many subops) that don't decode.
With a memcheck or register breakpoint active, landing on one of those
crashed instead of raising ILLEGAL.

Also: GetVectorOverlap decoded the second vector with size1 (no callers
today), and RegisterFunction left most of its AnalyzedFunction
uninitialized, including the size that ends up in knownfuncs.ini.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-19 10:02:46 -06:00
Henrik Rydgård c9c886d78f Plumb a MIPS context into ReadVector etc. 2026-08-12 14:02:19 +02:00
Thomas Lamb d1656d425a Improve debugger assemble/disassemble compatiblity
Update armips for VFPU instruction parsing changes:
- Kingcom/armips#255
- Kingcom/armips#256

Remove brackets from VFPU instruction `vpfx*` params

Rename VFPU instruction `vuc2i.s` to `vuc2ifs.s`

Add brackets to VFPU instruction `vpfxd` saturation operations (i.e. [0:1] & [-1:1])

Rename FPU instructions `c.$OP` to `c.$OP.s`

Fix VFPU instruction `vuc2ifs.s` `vd` size in disassembly

Replace `CC[imm3]` with `imm3` in VFPU instructions `vcmov*` & `bv*` disassembly
2026-07-07 20:30:51 -04:00
fp64 2e290e0133 Fix VFPU dot bug
Fix rounding-overflows-into-next-exponent bug in vdot pointed out in
https://github.com/hrydgard/ppsspp/issues/21070#issuecomment-4692749931
(unless I messed up again).
2026-06-12 23:28:57 +03:00
fp64 3f154fb24e Implement (hopefully) bitwise-exact vmul/vdiv
See https://github.com/hrydgard/ppsspp/issues/21070#issuecomment-4621393701
for details.

The implementation is not wired to anything currently, this is
just for reference purposes.

The alternative FTZ logic from
https://github.com/hrydgard/ppsspp/issues/21070#issuecomment-4642348708
is not implemented (i.e. this uses original double-based logic).

Also fixes space->tab for `vfpu_dot`.
2026-06-11 16:40:06 +03:00
fp64 ef74666ac0 Fix style 2026-06-07 13:18:01 +03:00
fp64 cf7ff0044d Implement (hopefully) accurate vdot instruction
Hopefully bitwise-exact to PSP.
See https://github.com/hrydgard/ppsspp/issues/21070#issuecomment-4640382516
for details.

Again, massive thanks to danzel for the data.

SIMD version not implemented.

Didn't touch USE_VFPU_DOT, etc., so needs to be enabled if you want
to test it.
2026-06-07 12:24:44 +03:00
fp64 ac22c45526 Fix vfpu_sin/vfpu_cos bug
Was looking over VFPU stuff, and noticed this.
Presumably, on x86 the code was already doing exactly this, but still, ouch.

Did retest the code on all relevant (|x|>2^32) available inputs, all matches (except NaN payloads, as usual, but that is unrelated).
2026-01-03 07:15:42 +02:00
oltolm 02e767866a fix compiler warnings 2025-02-22 14:15:15 +01:00
Henrik Rydgård e93c80db4e Cleaning up our SIMD header includes, using the new header 2024-12-19 16:08:48 +01:00
Henrik Rydgård 3e198c53b2 More include cleanup 2024-12-18 13:57:26 +01:00
Herman Semenov 3c66f149d3 [Common/Core/Windows] Removed excess check pointer before delete or free() 2024-09-17 11:34:42 +02:00
Henrik Rydgård e01ca5b057 Logging API change (refactor) (#19324)
* Rename LogType to Log

* Explicitly use the Log:: enum when logging. Allows for autocomplete when editing.

* Mac/ARM64 buildfix

* Do the same with the hle result log macros

* Rename the log names to mixed case while at it.

* iOS buildfix

* Qt buildfix attempt, ARM32 buildfix
2024-07-14 14:42:59 +02:00
Henrik Rydgård 64ee5675b8 Minor unrelated cleanup 2023-10-06 15:39:59 +02:00
Henrik Rydgård 0d06af87b6 Interpreter: Optimize ReadVector/WriteVector by removing voffset lookups
Drops these functions down the ranking of top functions by quite a bit in GTA,
speedup at most 0.5% though. But enough of these small ones and they
start adding up.

Not sure why GTA falls back to the interpreter for these so much though.
I guess some "uneaten" prefix..
2023-10-05 19:11:34 +02:00
Henrik Rydgård 60a304f29b Turn the ifs inside out 2023-10-05 18:59:56 +02:00
Henrik Rydgård f21523ff74 WriteVector: Pluck transpose out of the loop 2023-10-05 18:56:15 +02:00
Henrik Rydgård e852771480 Integrate the voffset shuffle in ReadVector 2023-10-05 18:52:50 +02:00
fp64 49ac4c6774 Clarify 2023-10-02 14:05:49 -04:00
fp64 23e2d0f797 Add SSE2 version of vfpu_dot
See #18249. Speedup for this function ranges 10%..100%,
depending on system. Updated verification and speed measurements:
https://godbolt.org/z/W1z3sj6hz
2023-10-02 12:53:30 -04:00
fp64 dcaca7f111 Fix vrnd to the current understanding
Followup to #17506.
2023-06-04 16:44:27 -04:00
Henrik Rydgård 1ef1478cc8 Remove more impossibilities (GetMtxSize) 2023-06-04 11:48:43 +02:00
Henrik Rydgård a92cca2575 Don't check for impossibilities. Minor speedup for GetVecSize. 2023-06-04 11:28:39 +02:00
Henrik Rydgård 9db9fec898 VFPU: Some micro-optimizations. Don't fall back to interpreter path for vexp/vlog/vrexp. 2023-06-04 11:28:33 +02:00
fp64 a97c911d46 Address feedback 2023-05-25 17:28:38 -04:00
fp64 23ef21ba9b Fix a bug, and bump savestate version 2023-05-25 16:18:58 -04:00
fp64 71884d5843 Make vrnd match HW closer
See investigation starting
https://github.com/hrydgard/ppsspp/issues/16946#issuecomment-1467261209
for more details.
Still needs more testing.
2023-05-25 14:18:19 -04:00
Unknown W. Brackets 6da10463f9 Debugger: Make reg names safer, stop using v000.
Better to use S000, etc. as that's more clear throughout.
2023-04-29 09:48:33 -07:00
Henrik Rydgård 6945deec01 Replace a LOT of sprintf with snprintf, and a few strcpy with truncate_cpy 2023-04-28 21:04:05 +02:00
Henrik Rydgård aba026f7e9 Add back our older VFPU approximations, as fallbacks if files are missing.
PR #16984 added more accurate versions of these functions, but they require
large lookup tables stored in assets/.

If these files are missing, PPSSPP would simply crash, which isn't good.

We should probably try to warn the user somehow that these files are
missing, though...
2023-04-03 11:33:41 +02:00
Henrik Rydgård d996fb74d4 MSVC: Set language standard to c++17.
Noticed that we were getting some new warnings after merging the
constexpr stuff.
2023-04-02 17:55:15 +02:00
fp64 38fc21a2c0 Implement load-on-demand of vfpu tables 2023-03-12 08:21:15 -04:00
fp64 d3be0ee654 Fix tabs vs spaces, tweak comments, error on BE 2023-03-12 08:21:15 -04:00
fp64 67bb17eba3 Add more vfpu_*, move tables to assets 2023-03-12 08:21:15 -04:00
fp64 ee98603fe7 Fix the sign of cos(2*n+1)
Also fix the license text.
2023-03-12 08:21:14 -04:00
fp64 3661bb27ce Implement sin/cos as per #16946 2023-03-12 08:21:13 -04:00
Henrik Rydgård 9e736ca50c Workaround for sin/cos issue in GTA on Mac (and maybe others) 2023-02-07 17:43:12 +01:00
Unknown W. Brackets a7b7bf7826 Global: Set many read-only params as const.
This makes what they do and which args to use clearer, if nothing else.
2022-12-10 21:13:36 -08:00
Unknown W. Brackets 3f997518f3 irjit: Handle vrot overlap more correctly.
Sine ignores overlap, cosine does not.
2022-10-29 22:25:25 -07:00
Unknown W. Brackets 3df6cb704f Global: Fix some type conversion warnings.
Hidden by some warning disables.
2022-01-30 16:09:33 -08:00