Both were no-ops, so mfic left its destination unchanged. They read and
write the interrupt enable flag that sceKernelCpuSuspendIntr/ResumeIntr
use. Only bit 0 counts for mtic, which also goes for
sceKernelCpuResumeIntr, since on hardware it's just mtic.
Adds the intr/mfic test, recorded on hardware.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The third argument is a timeout pointer, as threadman.prx shows. A kernel
address from user mode is ILLEGAL_ADDR there; we used to write through it.
Also, no lookup by index: the syscall requires the exact uid.
Adds the threads/tls/allocate test, recorded on hardware.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This is the syscall usersystemlib's sceKernelGetTlsAddr makes when the
thread's cached TLS address is null, as (uid, &addr, 0). Code that has to
run before usersystemlib.prx is loaded (like plugins built with a Rust SDK)
inlines sceKernelGetTlsAddr and imports this directly.
Shares the allocation with sceKernelGetTlsAddr. A thread waiting on a full
pool stores its address pointer as the wait value, so no state changes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
OptimizeLoadsAfterStores only dropped a load right after a store of the same
reg. Now a load of anything the block stored or loaded before becomes a reg
move (with the extension for 8/16-bit loads), as long as nothing in between
may have changed the memory, the address reg, or the reg holding the value.
Only a store through the same base at a disjoint range is known not to alias,
and constant addresses outside RAM are left alone.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
When the cosine lane follows the sine lane, FSinCos can write both in place
instead of going through a temp and two FMovs.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The IR temps don't live past the block, but FlushAll stored them anyway at
every exit (the branch operands, lwl/lwr temps, VFPU temp lanes). At an exit,
discard the ones nothing later in the block reads.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- A constant stays known after it's written out for a read (by a store,
MovZ, a multiply...), so later uses still fold. Whatever an op writes is
forgotten after its inputs are written, and setting a reg to the value it
already holds isn't written twice.
- A conditional exit that isn't taken keeps the constants known.
- The saturating and min/max FP ops, FSign and the 31-bit Vec2 pack/unpack
no longer flush every GPR constant.
- A load through its own base is folded (lui v0, hi; lw v0, lo(v0)).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- An lwl/lwr pair was combined into one load even when the first half loads
into the base register, which changes the address of the second half.
- ApplyMemoryValidation shared one sp check across the block even past an
Interpret or CallReplacement, which may change sp.
- Drop a duplicate FSqrt meta entry, and name Load8Ext correctly.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Copy propagation through an FPR temp didn't stop when an instruction
rewrote the temp in place (it compared an FPR number against the +32
offset reg), so later reads lost that write.
- A read of the temp in both operands only had src1 replaced, yet the copy
into the temp was still removed.
- The replacement matched operands by number without checking their type,
so a StoreFloat whose GPR address had the temp's number got its address
replaced (IRVTEMP_PFX_S and IRTEMP_0 are both 192).
- A write to lanes 1-3 of a Vec4 temp wasn't noticed.
- IRReadsFromFPRs stopped after the F operands, missing Vec4Scale's vector.
- Exits and barriers didn't count as reading everything, so a write to a
real reg could be moved above an exit.
- Load32Linked and Store32Conditional were removed when their reg was
overwritten unread, losing LLBIT and the store.
Also fixes an off-by-one in the vec src3 read check. The unit test now
reports every failing case instead of stopping at the first.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
After the registers usable by compressed instructions, prefer s2-s7 and
fs2-fs11, so that fewer values have to be flushed around calls. The
dispatcher already saves them.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A new FSinCos op writes both from one argument reduction. arm64 and x64 get
both back from a single call, packed in one double; RISC-V and LoongArch make
the two calls.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The cosine is then taken of what vrot wrote to that lane: the sine, or zero.
The IR looked at the sine lane instead of the lane holding the angle, and the
legacy JITs ignored the overlap. The assembler refuses such a vrot, so those
now leave it to the interpreter, and don't pair one with the vrot before it.
Covered by the new cpu/vfpu/vrot test.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The ABI only preserves the low 64 bits of F24-F31, which are first in
the allocation order, so a four-lane vector there lost its upper half
across a call to a math helper. Flush those like arm64 does for
S8-S15.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The thunk saves a fixed set of registers and MXCSR on every call. The
math helpers (vrcp through vrexp2, and vrot's sincos) leave MXCSR alone,
so they now go through CallProtectedLeaf, which saves only the
caller-saved registers the caches are using, around a direct call. vrot
also no longer flushes everything first. x86-64 only; 32-bit x86 keeps
the thunk.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Values in S8-S15 survive a call, but the full flush wrote them back and
the following instructions loaded them again. The VFPU math callouts,
vrot and vh2f now flush only the caller-saved registers, plus the few
callee-saved ones they stage values in, and map the destinations
afterwards. In a normalize loop with sixteen VFPU registers live, that
takes a vrsq.s from about 15 to 11.5 ns.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With LSX a lane group can be mapped as one vector reg, and F() then returns
the same reg for every lane, so the per-lane code wrote only lane 0 (vs2i,
vus2i). Also drop the stale aliasing TODOs; each path reads its sources first.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On a PSP the Media Engine decodes, and the samples land in the output
buffer as the call returns, a couple of milliseconds in. We wrote them at
once and only then delayed the thread. Since sceAudio plays straight out
of game memory, that matters: Fired Up decodes each chunk to 0x40 bytes into
one of its two buffers, running over the first 16 samples of the other one,
which it has just queued, and relies on the mixer having read those first.
Writing early replaced them about 21 times a second, which is the constant
crackle in its music and intro (it showed up with the sceAudio buffering
rework, which stopped copying buffers at enqueue).
Now the decoder's output is set aside, the old contents put back, and a
CoreTiming event writes the samples just before the thread wakes. Pending
writes are kept in savestates.
Adds audio/blocking/parked, recorded on a PSP: a blocking output that had
to wait returns before any of its buffer has played, so the game really
does depend on the decode's latency.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Only 4x4 with source and destination transposed alike, and a scale
outside the destination, compiled; most vmscl in games are transposed
or 3x3. The rest now multiply element by element, with the scale
copied first, and only a partly overlapping source still falls back.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ll and sc always went to the interpreter; a few games use them
thousands of times. They now load and store directly with fast memory,
keeping llBit in MIPSState. vcmp's NaN and inf-or-NaN tests only look
at s, so they no longer need vt to be the same register: EN and NN
compare s with itself.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Without the Zbb extension, clz, rotr, wsbh, wsbw, bitrev, min and max
went to the IR interpreter, and wsbh always did. They're now base ISA
sequences: a branchless binary search for clz, paired shifts for the
rotates and byte swaps, mask-and-shift steps for bitrev, and a
compare-and-branch for min and max. Checked against a model of the
instructions over random inputs, since nothing here runs RISC-V.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
They always went to the interpreter. Each channel is now a shift, mask
and shift in the existing integer ops, so every backend handles them.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
vh2f always went to the interpreter in the IR. A new FHalfToFloat op
converts the lower or upper half of a word, and the native backends
call vfpu_h2f for it like FSin. The legacy arm64 JIT now makes the same
call instead of computing the conversion inline; vh2f is rare, and the
call is much less code.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
All went to the interpreter. vh2f works in integers to match vfpu_h2f,
since FCVTL neither flushes subnormal halves nor keeps inf/NaN mantissa
bits unshifted. The color conversions are bitfield extracts and inserts.
vbfy, vcrs and vdet follow the IR frontend, and vmscl scales element
by element. Results go through scratch registers, so a destination that
overlaps a source is fine.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Vec4ClampToZero and Vec2ClampToZero only ever fed Vec4Pack31To8 and
Vec2Pack31To16, for vi2uc and vi2us. The packs now clamp negative lanes
to zero themselves, which saves an op and a vector temp, and lets x64
clamp with PACKUSWB's saturation after an arithmetic shift.
While at it, RISC-V compiles Vec2Unpack16To31, Vec2Pack31To16 and
Vec4Pack32To8, and LoongArch Vec2Unpack16To31, Vec2Pack31To16 and the
non-LSX Vec4Pack32To8, all of which went to the IR interpreter.
LoongArch's Vec2Pack32To16 and Vec2Unpack16To32 now take their scalar
path with LSX too, instead of falling back.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
vi2uc, vi2c, vi2us, vi2s, vuc2i, vc2i, vus2i and vs2i lowered to IR ops
that x64 left to the IR interpreter. They're now SSE2: shifts into
place, then PACKSSDW/PACKUSWB for the packs and self-unpacks for the
unpacks. An output overlapping its input still goes the slow way.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
FCvtWS always went to the IR interpreter. Both now convert in the
current rounding mode, which ApplyRoundingMode keeps at the game's, as
x64 does. RISC-V's FCVT already saturates and gives INT_MAX for NaN;
LoongArch patches NaN to INT_MAX like its FRound.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
They always went to the IR interpreter. Now a signed compare-and-branch
on the normalized sources picks which one to move.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
These went to the interpreter. vf2in/vf2iz/vf2iu/vf2id use FCVT, which
saturates like the PSP, and patch NaN to 0x7FFFFFFF like the IR
backend. vsgn keeps the sign bit on 1.0 and gives 0 below the smallest
normal. vsge and vslt select 1 or 0 on a compare whose condition is
false when unordered. EI and NI test |s| against infinity in integers.
All match the interpreter on special values, ties and denormals.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
These were always interpreted. They're now IR ops that the native
backends compile to calls to vfpu_exp2 and vfpu_log2, like FSin and
FAsin. vrexp2 is FNeg followed by FExp2, which is how vfpu_rexp2
computes it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
vsin, vcos, vnsin, vasin, vexp2, vlog2 and vrexp2 went to the
interpreter. They now take the same direct call as vrcp and friends.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The vertex decoder's C++ steps may fuse, like the vertex JITs do; with
contraction off everywhere, RISC-V's VertexJit test found the JIT and
the steps disagreeing on morphed float UVs. The flag now applies to
Core/MIPS only, in CMake and the libretro Makefile. ndk-build has no
per-file flags, so the legacy Android.mk goes back to the default.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A float made from a constant NaN bit pattern can come out quieted: MSVC
turned the 0x7F800001 that vcos, vexp2 and vlog2 return into 0x7FC00001,
which failed cpu/vfpu/exact on Windows. Each function now computes its
result's bits in integers and converts once at the end, like vrcp and
friends already did. Fixed-point results become floats by shifting,
which is exact since the VFPU keeps 22 significant bits. Identical to
the previous code over every 32-bit input.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
They were added in a different order, so they rounded differently from
every other backend.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Each of the three is now a range check, a fast path and a few special
cases. The fast path reads its segment from a table whose constant term
already includes the result's exponent bits, so what remains is two
multiplies, some shifts and one exponent adjustment, all in integers.
Identical to the previous code over every 32-bit input.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
These went through the host's sqrt and division everywhere except the
interpreter's vrcp and vnrcp (vsqrt and vrsq there only behind
USE_VFPU_SQRT, now gone). They now always give the PSP's bits: the IR
gets FVSqrt (FSqrt stays the FPU's IEEE sqrt.s), and FRSqrt and FRecip,
which only the VFPU emits, become vfpu_rsqrt and vfpu_rcp; the IR
interpreter and the x64, arm64, RISC-V and LoongArch backends call them.
The old JITs call them directly, the ARM ones keeping the lanes in
callee-saved registers across the calls. cpu/vfpu/exact now passes on every core.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A request started since the last RequestManager::Update (headless never
calls it) sat in newDownloads_, which CancelAll skipped. It was then
destroyed along with the static g_DownloadManager at exit, and its
destructor removed its progress bar from the already destroyed g_OSD:
"mutex lock failed". Seen with a Netconf dialog still downloading the
infra DNS json when a test ended.
CancelAll now takes the new ones too, and runs at shutdown while g_OSD is
still there. The Netconf json request is also let go of when the emulator
shuts down, rather than living on into the next game.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
They use the same quadratic interpolator as rcp and friends, with three
twists. sin indexes the quarter wave from the top, and asin and sin work
in a per-segment exponent whose 4-ulp truncation also applies to results
in a lower binade. log2 truncates exponent + log2(1.m) toward zero to 22
significant bits, and where that step is coarser than 2^-24 the datapath
drops coefficient bits to match; that also covers the region just below
1.0 that needed a special case.
vfpu_sincos now reduces the angle once. With every table gone, so are
the asset folder, the loader, InitVFPU and the fallbacks for tables that
failed to load. All seven functions are bit-exact with the table-based
code over every 32-bit input.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The four share one quadratic interpolator: 128 segments picked by the top
7 bits of the input, each with a constant, a linear and a squared-term
coefficient, and a squarer on the top 10 bits of the rest that rounds t^2
up to a multiple of 256. 128 small coefficient sets per function
replace the 1 MB of delta tables. Derived from the output of the
table-based code, and bit-exact with it over every 32-bit input.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The IO thread wrote results (file lists, sizes, loaded data) straight into
PSP memory while the game kept running, so they landed at an arbitrary
point in its code. utility/savedata/filelist caught it now and then: a
poll saw the entries written but the counts still zero. On a PSP it's all
there when Update returns. Host IO timing still uses the thread.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
sceCtrlPeek/ReadBufferPositive2/Negative2, which take a port before the
usual buffer and count. Not implemented.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
GameSharing InitStart used to return 0 without starting anything, so its
GetStatus stayed at NONE and a game waiting for the dialog to finish
would hang. It now goes through the normal lifecycle (INIT, RUNNING,
FINISHED, SHUTDOWN, NONE) and reports that the user cancelled.
PSPPlaceholderDialog was abstract and unused (and missing from CMake);
it's now that stand-in.
WRONG_TYPE from GameSharing GetStatus/Update/ShutdownStart is what a PSP
returns whenever another dialog type was the last one started, so log it
at debug like the other dialogs. Sega Rally polls it every frame.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
secureVersion and the key only matter for the full 1536-byte request; the
older sizes always save without a key. Otherwise SDK 2.07+ uses the new
keyed hash for versions 0 and 3, older SDKs the old keyed hash for 0 and
no key for 3, and version 2 is always the old keyed hash.
Versions 0, 2 and 3 all require a key, and are rejected with SAVE_PARAM
otherwise - including 0, which we used to save without a key. Matches the
SFO modes and save errors recorded in utility/savedata/secureversion; what
remains there is load strictness when secureVersion doesn't match the
file, which we keep lenient for older saves.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On a PSP, dialog animations advance by animSpeed frames per Update and the
fades take about 200ms. Ours took 500ms (1/30 s per animSpeed, over
FADE_TIME 1.0), and all input was ignored until it finished, which made
dialogs feel sluggish, noticeably so in 30fps games. Input is still
ignored while fading out, once a choice has been made.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
sceIoOpen on such a path returns an invalid-argument error on hardware,
not file-not-found (utility/savedata/idlist).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>