They use the same quadratic interpolator as rcp and friends, with three
twists. sin indexes the quarter wave from the top, and asin and sin work
in a per-segment exponent whose 4-ulp truncation also applies to results
in a lower binade. log2 truncates exponent + log2(1.m) toward zero to 22
significant bits, and where that step is coarser than 2^-24 the datapath
drops coefficient bits to match; that also covers the region just below
1.0 that needed a special case.
vfpu_sincos now reduces the angle once. With every table gone, so are
the asset folder, the loader, InitVFPU and the fallbacks for tables that
failed to load. All seven functions are bit-exact with the table-based
code over every 32-bit input.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The four share one quadratic interpolator: 128 segments picked by the top
7 bits of the input, each with a constant, a linear and a squared-term
coefficient, and a squarer on the top 10 bits of the rest that rounds t^2
up to a multiple of 256. 128 small coefficient sets per function
replace the 1 MB of delta tables. Derived from the output of the
table-based code, and bit-exact with it over every 32-bit input.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The IO thread wrote results (file lists, sizes, loaded data) straight into
PSP memory while the game kept running, so they landed at an arbitrary
point in its code. utility/savedata/filelist caught it now and then: a
poll saw the entries written but the counts still zero. On a PSP it's all
there when Update returns. Host IO timing still uses the thread.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
sceCtrlPeek/ReadBufferPositive2/Negative2, which take a port before the
usual buffer and count. Not implemented.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
GameSharing InitStart used to return 0 without starting anything, so its
GetStatus stayed at NONE and a game waiting for the dialog to finish
would hang. It now goes through the normal lifecycle (INIT, RUNNING,
FINISHED, SHUTDOWN, NONE) and reports that the user cancelled.
PSPPlaceholderDialog was abstract and unused (and missing from CMake);
it's now that stand-in.
WRONG_TYPE from GameSharing GetStatus/Update/ShutdownStart is what a PSP
returns whenever another dialog type was the last one started, so log it
at debug like the other dialogs. Sega Rally polls it every frame.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
secureVersion and the key only matter for the full 1536-byte request; the
older sizes always save without a key. Otherwise SDK 2.07+ uses the new
keyed hash for versions 0 and 3, older SDKs the old keyed hash for 0 and
no key for 3, and version 2 is always the old keyed hash.
Versions 0, 2 and 3 all require a key, and are rejected with SAVE_PARAM
otherwise - including 0, which we used to save without a key. Matches the
SFO modes and save errors recorded in utility/savedata/secureversion; what
remains there is load strictness when secureVersion doesn't match the
file, which we keep lenient for older saves.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On a PSP, dialog animations advance by animSpeed frames per Update and the
fades take about 200ms. Ours took 500ms (1/30 s per animSpeed, over
FADE_TIME 1.0), and all input was ignored until it finished, which made
dialogs feel sluggish, noticeably so in 30fps games. Input is still
ignored while fading out, once a choice has been made.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
sceIoOpen on such a path returns an invalid-argument error on hardware,
not file-not-found (utility/savedata/idlist).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On a PSP, dialog init and shutdown happen partly at the accessThread
priority and partly at the graphicsThread priority, one phase after the
other. Model that with one helper thread that switches priority per phase,
and let starting it reschedule normally instead of disabling interrupts.
A caller with worse priority than both now sees shutdown complete inside
ShutdownStart, as on hardware. NFL Street 3 (graphics 17, access 19,
caller 111) calls the next InitStart right after ShutdownStart and used to
loop forever on 'A save request is already running' (#19957). The utility
pspautotests, where the caller has better priority, are unchanged.
Also logs the dialog thread priorities at debug level.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
An invalid value (e.g. --debugger-run swallowing the next flag) printed an
error but returned Exit, so the process exited 0.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Since 551e4cd0ab, headless without --graphics silently used the OpenGL
backend, since bSoftwareRendering defaulted to false. Under Mesa llvmpipe
on Linux/WSL, that hangs games early in boot (and test.py, which passes no
--graphics). The README already documents software as the default.
Also documents the traps hit while chasing this in docs/debugging.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A sum just below a power of two can round up into the next exponent, and
at the top of the range into inf; random inputs almost never land there.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Same shape as the NEON one, which now shares its rounding tail. SSE2 has
no per-lane shift, so the alignment shift is a multiply by a power of two
built from float bits and converted by truncation; the unsigned maxima
use the 16-bit instructions, since every value involved fits in 15 bits.
Nothing depends on the host rounding mode or flush-to-zero.
Checked against the reference by VFPUDot and on 300M more inputs offline,
also with MXCSR set to round toward zero with FTZ and DAZ.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The four lanes are computed together: exponents, 24x24-bit products with
round-to-odd, alignment by truncation and a signed horizontal sum. One
pairwise maximum finds both the alignment exponent and any inf or NaN,
which go to the reference. The final rounding is branch-free and in
integers, since the host rounding mode may be the game's.
About three times the throughput of the reference on Apple M-series
(5.2 vs 15.2 ns per call). VFPUDot checks it against the reference on
four million inputs picked to cover cancellation, ties, subnormals and
the overflow edges; a billion more matched offline.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Without --memstick, a library whose HLE the config has disabled has no firmware
to resolve against, and the game dies on unresolved imports rather than falling
back to our HLE. The run still logs happily for its whole timeout, so the only
tell is zero flips.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
videos_ only learns about the CSC output, so the texture we actually draw is
just an ordinary 512x512 8888 texture whose contents happen to be different
every frame: hash, miss, throw it in the secondary cache, rebuild, forever.
So track the copy. NotifyVideoCopy marks the destination as video when the
source is, and the copy funnels call it: sceDmacMemcpy, sceKernelMemcpy, and
the four replaced memcpy/memmove variants. It sits outside their "is either
side VRAM" gate, since a RAM-to-RAM copy of a frame is still a frame.
Being a video texture also gets it forced linear filtering and keeps it out of
texture upscaling, which is what you want for a video either way.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
A `video = true` in textures.ini opted a pack into replacing and dumping video
textures. Both halves are a bad deal. Dumping writes a file per decoded frame,
which fills a disk rather than producing anything a pack can use, and replacing
means a hash lookup on content that is different every frame and will never be
found twice.
It was also the only reason the texture cache still hashed video textures at
all, so it cost every game that has ever played a cutscene, not just the packs
that set it.
Video textures are now never replaced and never dumped, and skipHash is simply
isVideo. An existing ini keeping the key is harmless - unknown options are
ignored.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
NotifyWriteFormattedFromMemory appends to videos_ unconditionally. A game
blitting its decoded frame to the display buffer does that every displayed
frame while it waits for the next one to decode, so the same two or three
addresses come back over and over: Death Jr pushes 733 entries where there are
two distinct buffers, Tekken 6 around 53,000 where there are three. IsVideo()
walks that vector linearly on every texture. Refresh the matching entry instead
of appending a new one - Death Jr now holds 2 entries and Tekken 6 holds 18,
peak size 2 and 3.
The other half is the two TODOs that were already sitting there. A video
texture is new every frame by definition, so re-hashing it only confirms what
the VIDEO flag already said, and the secondary cache has nothing to offer a
frame that will never recur. Skip both, and with them the secondary lookup that
would otherwise key off a hash we no longer compute.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The ISA returns the canonical NaN from every operation, so a negative or
signaling NaN operand loses its sign and payload where the PSP keeps them.
Not worth a check per FP op in the JIT.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
fpu_branch_hazard (the compare-to-branch hazard), cacheop (the write-back
data cache seen through the uncached mirror) and fpu_nan (which NaN 0/0
makes, host dependent on x86) go to tests_next. cpu/fpu/fpu is re-recorded
from a binary built with the current toolchain, which prints -nan.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
x86 (SQRTSS and libm alike) returns 0xffc00000 for it, the PSP 0x7fc00000.
-0 stays -0 and a NaN input comes through as it is, so only a negative
input needs the sign cleared: a compare and an xor in the x64 JITs, a
branch in the interpreters. cpu/fpu/roundmode.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
The C cast is undefined past the int32 range, and x86 makes it INT_MIN, so
round/trunc/ceil/floor/cvt.w.s of anything from 2^31 up gave 0x80000000 on
x86 hosts while the PSP saturates to 0x7fffffff (cpu/fpu/roundmode). Route
all of them through SaturatedFloatToInt, which also covers NaN and inf, and
drop the special cases that did.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
specials stays in tests_next for vcmp on denormals and the NaN
canonicalization and denormal flush in vbfy/vocp/vavg/vfad/vsocp, which
overlap the USE_VFPU_DOT accuracy switch. overlap_vcrsp is vcrsp with an
overlapping destination, which the assembler refuses and the hardware
doesn't read-before-write for.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
The second pack read a lane the first one had just written
(vi2s.q C002, C000). Pack into temps when the outputs overlap the inputs.
Found by the corrected cpu/vfpu/overlap.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
EN/NN compared the source against an uninitialized temp, usually the
previous lane's all-ones mask (a NaN), so every lane came out as NaN. NI
used a less-than compare, which is false for a NaN. TR set the bit and
then replaced it with a garbage bit from the temp. Found by
cpu/vfpu/specials.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
In the interpreter, the IR interpreter, every backend's FSign and the x86
JIT's own vsgn. Recorded in cpu/vfpu/specials.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
The same ordering as vmin/vmax (sign and magnitude, denormals compare as
zero), and on a tie vsrt1/2 keep the lower lane of each pair, vsrt3/4 the
upper one. Recorded in cpu/vfpu/specials.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
vf2h truncates the mantissa (no rounding), gives a signed zero below 2^-14
and a signed inf from 65536 up, and keeps the low ten mantissa bits of a
NaN, so one with those clear becomes inf. vh2f flushes a subnormal half to
a signed zero and keeps inf/NaN mantissa bits unshifted. The x86 JIT's own
vh2f gets the subnormal flush; the other backends go through the
interpreter. Recorded in cpu/vfpu/specials.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
The other control registers were masked already; the CC path wasn't, so
a value with bits 6 or 7 set made bvt 6 and bvt 7 branch. Hardware never
sets those (cpu/vfpu/vbranch), and the IR masks them.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
The NaN check compared src1 with itself, so only a NaN in src1 reached
the slow path with the hardware's integer-compare rule. A NaN in src2 went
to MINSS/MAXSS, which just return it. Compare src1 with src2, as the arm64
backend does. Caught by cpu/vfpu/minmax.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
The same set of gates the arm64 JIT got: vabs/vneg, vsat1, the vrcp
group, vzero/vone/vidt, vdiv, vscl, vfad/vavg, vrot, vuc2i, and now also
vcrs/vcrsp/vqmul, vbfy and vtfm, which the IR frontend hands to the
interpreter as well. Swizzles naming a lane past the op's size go there
too. mtvc now keeps only the bits the hardware does (six for CC, the low
20 of a prefix, 0x3F8 on top of the RNG state). vcst's SIMD path skipped
the D prefix.
Fixes cpu/vfpu/prefix_ctrl on Windows CI. Verified on an x86_64 build
under Rosetta: prefix_ctrl and the full tests_good pass, prefix_consume
is down to the same three ops that differ on every backend.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
The riscv64/loongarch64 cross builds have no SDL, and Headless.cpp had a
matching "no window" path gated on those two architectures. Gate it on
the option instead (a HEADLESS_NO_SDL define), so any headless-only build
without a usable SDL gets it - an x86_64 build on an arm64 Mac, say, where
Homebrew's SDL3 is arm64 only. That's how the x86 JIT fixes here were
tested, under Rosetta.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
mtvc to RCX0-7 keeps the low 20 bits of the value and forces the top
twelve to 0x3F8 (cpu/vfpu/vrnd), which is also the form vrnds and vrndi
leave them in. We masked with 0x3FFFFFFF. The generator itself matches
hardware for every seed and state in that test.
Interpreter, IR and the arm64 JIT. The x86 and ARM32 JITs don't mask
control register writes at all and are left as they were.
Co-Authored-By: Claude Fable 5.1 <[email protected]>