A new FSinCos op writes both from one argument reduction. arm64 and x64 get
both back from a single call, packed in one double; RISC-V and LoongArch make
the two calls.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
vh2f always went to the interpreter in the IR. A new FHalfToFloat op
converts the lower or upper half of a word, and the native backends
call vfpu_h2f for it like FSin. The legacy arm64 JIT now makes the same
call instead of computing the conversion inline; vh2f is rare, and the
call is much less code.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A float made from a constant NaN bit pattern can come out quieted: MSVC
turned the 0x7F800001 that vcos, vexp2 and vlog2 return into 0x7FC00001,
which failed cpu/vfpu/exact on Windows. Each function now computes its
result's bits in integers and converts once at the end, like vrcp and
friends already did. Fixed-point results become floats by shifting,
which is exact since the VFPU keeps 22 significant bits. Identical to
the previous code over every 32-bit input.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Each of the three is now a range check, a fast path and a few special
cases. The fast path reads its segment from a table whose constant term
already includes the result's exponent bits, so what remains is two
multiplies, some shifts and one exponent adjustment, all in integers.
Identical to the previous code over every 32-bit input.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
They use the same quadratic interpolator as rcp and friends, with three
twists. sin indexes the quarter wave from the top, and asin and sin work
in a per-segment exponent whose 4-ulp truncation also applies to results
in a lower binade. log2 truncates exponent + log2(1.m) toward zero to 22
significant bits, and where that step is coarser than 2^-24 the datapath
drops coefficient bits to match; that also covers the region just below
1.0 that needed a special case.
vfpu_sincos now reduces the angle once. With every table gone, so are
the asset folder, the loader, InitVFPU and the fallbacks for tables that
failed to load. All seven functions are bit-exact with the table-based
code over every 32-bit input.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The four share one quadratic interpolator: 128 segments picked by the top
7 bits of the input, each with a constant, a linear and a squared-term
coefficient, and a squarer on the top 10 bits of the rest that rounds t^2
up to a multiple of 256. 128 small coefficient sets per function
replace the 1 MB of delta tables. Derived from the output of the
table-based code, and bit-exact with it over every 32-bit input.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Same shape as the NEON one, which now shares its rounding tail. SSE2 has
no per-lane shift, so the alignment shift is a multiply by a power of two
built from float bits and converted by truncation; the unsigned maxima
use the 16-bit instructions, since every value involved fits in 15 bits.
Nothing depends on the host rounding mode or flush-to-zero.
Checked against the reference by VFPUDot and on 300M more inputs offline,
also with MXCSR set to round toward zero with FTZ and DAZ.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The four lanes are computed together: exponents, 24x24-bit products with
round-to-odd, alignment by truncation and a signed horizontal sum. One
pairwise maximum finds both the alignment exponent and any inf or NaN,
which go to the reference. The final rounding is branch-free and in
integers, since the host rounding mode may be the game's.
About three times the throughput of the reference on Apple M-series
(5.2 vs 15.2 ns per call). VFPUDot checks it against the reference on
four million inputs picked to cover cancellation, ties, subnormals and
the overflow edges; a billion more matched offline.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
vf2h truncates the mantissa (no rounding), gives a signed zero below 2^-14
and a signed inf from 65536 up, and keeps the low ten mantissa bits of a
NaN, so one with those clear becomes inf. vh2f flushes a subnormal half to
a signed zero and keeps inf/NaN mantissa bits unshifted. The x86 JIT's own
vh2f gets the subnormal flush; the other backends go through the
interpreter. Recorded in cpu/vfpu/specials.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
mtvc to RCX0-7 keeps the low 20 bits of the value and forces the top
twelve to 0x3F8 (cpu/vfpu/vrnd), which is also the form vrnds and vrndi
leave them in. We masked with 0x3FFFFFFF. The generator itself matches
hardware for every seed and state in that test.
Interpreter, IR and the arm64 JIT. The x86 and ARM32 JITs don't mask
control register writes at all and are left as they were.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
WriteMMIO_U32 was missing the return after the kernel-mode check, so a
user-mode write raised the exception and then went through anyway. The
other five MMIO accessors already return here; this one lost it when the
GPIO/syscon branches were added.
Int_Vrot scanned all four entries of dregs, but GetVectorRegs only fills
the first n - so a vrot with vs == 0 (S000) matched a lane that isn't
there and took the cosine from it. The IR backend gets this right via
IsOverlapSafe, so the two disagreed.
The breakpoint checks in MIPSInterpret and RunUntilDowncountZeroWithChecks
read instr->flags without checking for null, which MIPSGetInstruction
returns for the eight primary opcodes (and many subops) that don't decode.
With a memcheck or register breakpoint active, landing on one of those
crashed instead of raising ILLEGAL.
Also: GetVectorOverlap decoded the second vector with size1 (no callers
today), and RegisterFunction left most of its AnalyzedFunction
uninitialized, including the size that ends up in knownfuncs.ini.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Was looking over VFPU stuff, and noticed this.
Presumably, on x86 the code was already doing exactly this, but still, ouch.
Did retest the code on all relevant (|x|>2^32) available inputs, all matches (except NaN payloads, as usual, but that is unrelated).
* Rename LogType to Log
* Explicitly use the Log:: enum when logging. Allows for autocomplete when editing.
* Mac/ARM64 buildfix
* Do the same with the hle result log macros
* Rename the log names to mixed case while at it.
* iOS buildfix
* Qt buildfix attempt, ARM32 buildfix
Drops these functions down the ranking of top functions by quite a bit in GTA,
speedup at most 0.5% though. But enough of these small ones and they
start adding up.
Not sure why GTA falls back to the interpreter for these so much though.
I guess some "uneaten" prefix..
See #18249. Speedup for this function ranges 10%..100%,
depending on system. Updated verification and speed measurements:
https://godbolt.org/z/W1z3sj6hz
PR #16984 added more accurate versions of these functions, but they require
large lookup tables stored in assets/.
If these files are missing, PPSSPP would simply crash, which isn't good.
We should probably try to warn the user somehow that these files are
missing, though...