Commit Graph
100 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 1f1c478c92 MsgDialog: An Abort takes effect from the 8th Update, like on a PSP
Until then the dialog runs normally (and writes result = 0, which an
immediate abort skipped). Measured with one Update per vblank; at one every
other vblank a PSP took 6, so it isn't purely a count.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 12:50:23 -06:00
Henrik RydgårdandClaude Opus 5.5 5f261bd7ff Utility: Stand in for the HtmlViewer, offering to open the page in a browser
Instead of the PSP's web browser, a dialog shows the URL the game wants
and opens it in the host's browser on X, or backs out on O. Either way the
game sees the browser closed normally. Platforms that can't open a URL
(the new SYSPROP_CAN_LAUNCH_URL) only offer to back out. Only plain
printable-ASCII http(s) addresses are handed over, and on Linux without a
shell.

What the firmware does (sceUtility_Driver, 6.61, plus
utility/dialog/htmlviewer): the HtmlViewer has its own state apart from the
other dialogs, so they don't block each other, and its calls return
WRONG_TYPE until one has started. The request size picks the 2.00 to 3.00
layout, and InitStart allocates 3.5MB of user memory (4.5MB with options
bit 0x400 from 2.70 on), failing with 800200d9 without it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 12:50:23 -06:00
Henrik RydgårdandClaude Opus 5.5 3c15b2b112 Utility: One dialog at a time whatever its type, like the firmware
On a PSP, every InitStart fails with INVALID_STATUS until the last dialog
started is back at NONE, including while it's shutting down, and before
its params are checked. A failed InitStart leaves the current type alone.
We returned WRONG_TYPE instead, and a failed InitStart (e.g. a bad size)
still switched the current type, so every later dialog was refused. (One of
ours that fails after already starting, as savedata can, still becomes the
current type, since the game may poll it.)

The busy check applies status changes that are due, but doesn't use up an
auto status dialog's one-time INITIALIZE/SHUTDOWN reports; one that only
waits to report SHUTDOWN is let finish. Auto status dialogs now release
volatile memory on the way to NONE, including when the game saw RUNNING
before the init thread was done, which used to leave it locked for the next
dialog. GamedataInstall no longer requires currentDialogActive, which its
ShutdownStart cleared even when it then failed, so it could never finish -
and would now have blocked every other dialog.

Also: MsgDialog accepts exactly the three sizes sceUtility_Driver does (we
memcpy'd whatever size was given), and HtmlViewer GetStatus answers
WRONG_TYPE.

Adds utility/dialog/status and utility/dialog/priority, recorded on a PSP.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 12:50:23 -06:00
Henrik Rydgård 74cbbc067c Merge pull request #22357 from hrydgard/ir-interpreter-opts
IR interpreter optimizations
2026-09-25 12:48:17 -06:00
Henrik RydgårdandClaude Opus 5.5 4fc858d6db arm64 IR JIT: Pair adjacent loads and stores into LDP/STP
When a Load32, Store32, LoadFloat or StoreFloat is followed by the same op
on the next or previous word through the same base, and the base can be
mapped as a pointer, emit one LDP/STP for both. A struct-copying loop runs
about 24% faster; code without such pairs is unaffected.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 11:47:17 -06:00
Henrik RydgårdandClaude Opus 5.5 d1c04d70a6 IR interpreter: Use threaded dispatch with GCC and Clang
Every op ended with a break back to one shared indirect jump, which the
CPU has to predict for every op in the program. With labels as values,
each op jumps through a table from its own site instead, which predicts
much better: an integer-heavy benchmark runs about 13% faster on an M1.
Other compilers keep the switch, and ops missing from the table fall back
to it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 11:47:17 -06:00
Henrik RydgårdandClaude Opus 5.5 7eb371b231 IR interpreter: Merge a conditional exit with the ExitToConst after it
Blocks ending in a branch dispatch a conditional exit and then the
fallthrough ExitToConst. One op now returns either target, reading the
second from the ExitToConst, which stays behind unexecuted.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 11:47:17 -06:00
Henrik RydgårdandClaude Opus 5.5 c6aea794f6 IR interpreter: Skip ReduceVec4Flush
It avoids flushes a native backend would need, at the cost of extra
instructions (a copy of each scalar before a Vec4Scale, for instance) that
the interpreter only has to dispatch.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 11:47:17 -06:00
Henrik Rydgård 0d93194d7f Merge pull request #22356 from hrydgard/mfic-mtic
Implement mfic and mtic CPU instructions
2026-09-25 11:40:38 -06:00
Henrik Rydgård d4544e3a42 Merge pull request #22355 from hrydgard/headless-gl-offscreen
Make headless work off-screen with OpenGL on the Mac
2026-09-25 11:14:46 -06:00
Henrik RydgårdandClaude Opus 5.5 feb6caa3c4 Implement mfic and mtic
Both were no-ops, so mfic left its destination unchanged. They read and
write the interrupt enable flag that sceKernelCpuSuspendIntr/ResumeIntr
use. Only bit 0 counts for mtic, which also goes for
sceKernelCpuResumeIntr, since on hardware it's just mtic.

Adds the intr/mfic test, recorded on hardware.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 11:08:13 -06:00
Henrik Rydgård a9d63d2ee1 Merge pull request #22354 from hrydgard/tlspl-allocate
Implement _sceKernelAllocateTlspl
2026-09-25 10:36:46 -06:00
Henrik RydgårdandClaude Opus 5.5 f47269864a _sceKernelAllocateTlspl: Check user pointers, support the timeout
The third argument is a timeout pointer, as threadman.prx shows. A kernel
address from user mode is ILLEGAL_ADDR there; we used to write through it.
Also, no lookup by index: the syscall requires the exact uid.

Adds the threads/tls/allocate test, recorded on hardware.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 10:05:07 -06:00
Henrik RydgårdandClaude Opus 5.5 ac39c55c36 Implement _sceKernelAllocateTlspl
This is the syscall usersystemlib's sceKernelGetTlsAddr makes when the
thread's cached TLS address is null, as (uid, &addr, 0). Code that has to
run before usersystemlib.prx is loaded (like plugins built with a Rust SDK)
inlines sceKernelGetTlsAddr and imports this directly.

Shares the allocation with sceKernelGetTlsAddr. A thread waiting on a full
pool stores its address pointer as the wait value, so no state changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 10:04:44 -06:00
Henrik RydgårdandClaude Opus 5.5 2b954b6a37 docs: Replace the headless hang history with what each GPU backend does
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 09:47:58 -06:00
Henrik RydgårdandClaude Opus 5.5 31dc2d47b5 docs: Headless OpenGL no longer hangs, and runs offscreen on macOS
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 09:47:58 -06:00
Henrik RydgårdandClaude Opus 5.5 ea101b714f Headless: Run OpenGL offscreen on macOS, without a window
A CGL context with no drawable, rendering into a framebuffer object of its
own that stands in for the backbuffer through g_defaultFBO. Core profile,
as the SDL app uses on macOS.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 09:47:58 -06:00
Henrik RydgårdandClaude Opus 5.5 c9cddb99e9 Headless: Present OpenGL frames too
OpenGL's render thread only finishes a frame when it's presented, so the
emu thread eventually waited forever in BeginFrame for a free frame, with
the render thread waiting for work. GPU tests and games hung, silently,
as a blocked host thread also defeats the timeouts.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 09:47:58 -06:00
Henrik Rydgård 47730e0eb0 Merge pull request #22353 from hrydgard/headless-vulkan-mac
Get Vulkan working on PPSSPPHeadless for Mac
2026-09-25 09:44:52 -06:00
Henrik Rydgård 48797bfe26 Merge pull request #22352 from hrydgard/ir-pass-fixes
Claude code review for IR optimization passes: Bug fixes and improvements
2026-09-25 09:26:22 -06:00
Henrik RydgårdandClaude Opus 5.5 bc8f0f29c4 docs: Headless Vulkan runs offscreen, and is the one to use for games
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 09:07:23 -06:00
Henrik RydgårdandClaude Opus 5.5 cc83084c94 Headless: Run Vulkan offscreen, without a window
Vulkan no longer needs a hidden window (which on macOS could never work, as
the Metal window description has no data2). It now uses the offscreen mode,
and presents each frame like the app does, as an unpresented frame would
wait forever for its next image. MoltenVK's console logging is limited to
errors, as it mixed into the test output.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 09:06:56 -06:00
Henrik RydgårdandClaude Opus 5.5 1a3addf0cb Mac: Look for MoltenVK outside the app bundle too
Executables outside a bundle (headless) found no Vulkan library. Also try
the app bundle built next to them, the Vulkan SDK's install and Homebrew's.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 09:06:56 -06:00
Henrik RydgårdandClaude Opus 5.5 23ba2f62f8 Vulkan: Don't wait for an image in frames that never acquired one
A frame that never draws to the backbuffer never acquires an image, but
its final submit still waited on the acquire semaphore, which nothing
signals, hanging the GPU. Skip the swap for such frames when finishing
them. This replaces the check for a frame with no steps at all, which
could also set it partway through a frame that acquires later.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 09:06:56 -06:00
Henrik RydgårdandClaude Opus 5.5 0b778d6b83 Vulkan: Add an offscreen mode to VulkanGraphicsContext
For when there's nothing to present to. Instead of a surface and swapchain,
it renders into images of its own through the VulkanPresentation interface
libretro uses, picking the graphics queue without a surface. Acquiring and
presenting signal and wait on the frame's semaphores with empty submits.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 09:06:56 -06:00
Henrik Rydgård a50fb6071f Merge pull request #22351 from hrydgard/jit-call-optimizations
JIT function-call optimizations
2026-09-24 16:59:39 -06:00
Henrik RydgårdandClaude Opus 5.5 dab5d82227 IR: Forward stores and loads to later loads in the block
OptimizeLoadsAfterStores only dropped a load right after a store of the same
reg. Now a load of anything the block stored or loaded before becomes a reg
move (with the extension for 8/16-bit loads), as long as nothing in between
may have changed the memory, the address reg, or the reg holding the value.
Only a store through the same base at a disjoint range is known not to alias,
and constant addresses outside RAM are left alone.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:56:03 -06:00
Henrik RydgårdandClaude Opus 5.5 1cdc432d6a IR: Let vrot's FSinCos write straight into an [s, c] pair
When the cosine lane follows the sine lane, FSinCos can write both in place
instead of going through a temp and two FMovs.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:42:03 -06:00
Henrik RydgårdandClaude Opus 5.5 498511fd51 IR: Drop everything after a conditional exit folded to always taken
Only the ExitToConst right after it was skipped before.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:39:24 -06:00
Henrik RydgårdandClaude Opus 5.5 00f5e12b47 IR JIT: Don't store dead temps at exits
The IR temps don't live past the block, but FlushAll stored them anyway at
every exit (the branch operands, lwl/lwr temps, VFPU temp lanes). At an exit,
discard the ones nothing later in the block reads.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:38:39 -06:00
Henrik RydgårdandClaude Opus 5.5 a99a82cc6b IR: Keep more constants known in PropagateConstants
- A constant stays known after it's written out for a read (by a store,
  MovZ, a multiply...), so later uses still fold. Whatever an op writes is
  forgotten after its inputs are written, and setting a reg to the value it
  already holds isn't written twice.
- A conditional exit that isn't taken keeps the constants known.
- The saturating and min/max FP ops, FSign and the 31-bit Vec2 pack/unpack
  no longer flush every GPR constant.
- A load through its own base is folded (lui v0, hi; lw v0, lo(v0)).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:35:11 -06:00
Henrik RydgårdandClaude Opus 5.5 25e8f1d57c IR: Fix lwl/lwr pairing into the base reg, and sp validation past barriers
- An lwl/lwr pair was combined into one load even when the first half loads
  into the base register, which changes the address of the second half.
- ApplyMemoryValidation shared one sp check across the block even past an
  Interpret or CallReplacement, which may change sp.
- Drop a duplicate FSqrt meta entry, and name Load8Ext correctly.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:35:11 -06:00
Henrik RydgårdandClaude Opus 5.5 1dbb62d570 IR: Fix PurgeTemps miscompiles
- Copy propagation through an FPR temp didn't stop when an instruction
  rewrote the temp in place (it compared an FPR number against the +32
  offset reg), so later reads lost that write.
- A read of the temp in both operands only had src1 replaced, yet the copy
  into the temp was still removed.
- The replacement matched operands by number without checking their type,
  so a StoreFloat whose GPR address had the temp's number got its address
  replaced (IRVTEMP_PFX_S and IRTEMP_0 are both 192).
- A write to lanes 1-3 of a Vec4 temp wasn't noticed.
- IRReadsFromFPRs stopped after the F operands, missing Vec4Scale's vector.
- Exits and barriers didn't count as reading everything, so a write to a
  real reg could be moved above an exit.
- Load32Linked and Store32Conditional were removed when their reg was
  overwritten unread, losing LLBIT and the store.

Also fixes an off-by-one in the vec src3 read check. The unit test now
reports every failing case instead of stopping at the first.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:35:11 -06:00
Henrik RydgårdandClaude Opus 5.5 f8004442c8 RISC-V IR: Allocate the saved registers before the temporaries
After the registers usable by compressed instructions, prefer s2-s7 and
fs2-fs11, so that fewer values have to be flushed around calls. The
dispatcher already saves them.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:22:41 -06:00
Henrik RydgårdandClaude Opus 5.5 c38d04a10c RISC-V/LoongArch IR: Get vrot's sine and cosine from one call
Like arm64 and x64, call vfpu_sincos_packed and split the double it returns.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:22:41 -06:00
Henrik RydgårdandClaude Opus 5.5 f92340c092 IR: Compute vrot's sine and cosine in one call
A new FSinCos op writes both from one argument reduction. arm64 and x64 get
both back from a single call, packed in one double; RISC-V and LoongArch make
the two calls.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:22:41 -06:00
Henrik RydgårdandClaude Opus 5.5 6fc4eb19df VFPU: Fix vrot with the angle in a destination lane
The cosine is then taken of what vrot wrote to that lane: the sine, or zero.
The IR looked at the sine lane instead of the lane holding the angle, and the
legacy JITs ignored the overlap. The assembler refuses such a vrot, so those
now leave it to the interpreter, and don't pair one with the vrot before it.

Covered by the new cpu/vfpu/vrot test.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:22:41 -06:00
Henrik RydgårdandClaude Opus 5.5 4ebf9ed037 LoongArch JIT: Flush LSX vectors in F24-F31 before calls
The ABI only preserves the low 64 bits of F24-F31, which are first in
the allocation order, so a four-lane vector there lost its upper half
across a call to a math helper. Flush those like arm64 does for
S8-S15.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:22:22 -06:00
Henrik RydgårdandClaude Opus 5.5 46ae865b95 x86 JIT: Call the VFPU math helpers without the full thunk
The thunk saves a fixed set of registers and MXCSR on every call. The
math helpers (vrcp through vrexp2, and vrot's sincos) leave MXCSR alone,
so they now go through CallProtectedLeaf, which saves only the
caller-saved registers the caches are using, around a direct call. vrot
also no longer flushes everything first. x86-64 only; 32-bit x86 keeps
the thunk.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:22:22 -06:00
Henrik RydgårdandClaude Opus 5.5 77ff1578c4 x86 JIT: Don't save callee-saved registers in the call thunk
ProtectFunction's thunk saved all of XMM2-15 and RBX around every call,
but the callee preserves XMM6-15 on Windows and RBX everywhere. On
Windows that drops ten 16-byte saves and loads from each protected call.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:22:22 -06:00
Henrik RydgårdandClaude Opus 5.5 f7ba67405d arm64 JIT: Keep S8-S15 mapped across VFPU math calls
Values in S8-S15 survive a call, but the full flush wrote them back and
the following instructions loaded them again. The VFPU math callouts,
vrot and vh2f now flush only the caller-saved registers, plus the few
callee-saved ones they stage values in, and map the destinations
afterwards. In a normalize loop with sixteen VFPU registers live, that
takes a vrsq.s from about 15 to 11.5 ns.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:22:22 -06:00
Henrik Rydgård bd8e321137 Merge pull request #22350 from hrydgard/jit-missing-ops
JIT (all of them): Fill out various missing ops
2026-09-24 16:05:35 -06:00
Henrik RydgårdandClaude Opus 5.5 8f804d6f0f Document the LoongArch LSX lane trap and testing under qemu
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 15:42:38 -06:00
Henrik RydgårdandClaude Opus 5.5 79e4568e50 LoongArch JIT: Leave the Vec2 packs to the interpreter with LSX
With LSX a lane group can be mapped as one vector reg, and F() then returns
the same reg for every lane, so the per-lane code wrote only lane 0 (vs2i,
vus2i). Also drop the stale aliasing TODOs; each path reads its sources first.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 15:39:39 -06:00
Henrik Rydgård be22f0270b Merge pull request #22348 from hrydgard/atrac-decode-latency
Atrac: Write decoded samples when sceAtracDecodeData returns, not when called
2026-09-24 15:17:52 -06:00
Henrik RydgårdandClaude Opus 5.5 ff7371dcc5 Atrac: Write decoded samples when sceAtracDecodeData returns, not when called
On a PSP the Media Engine decodes, and the samples land in the output
buffer as the call returns, a couple of milliseconds in. We wrote them at
once and only then delayed the thread. Since sceAudio plays straight out
of game memory, that matters: Fired Up decodes each chunk to 0x40 bytes into
one of its two buffers, running over the first 16 samples of the other one,
which it has just queued, and relies on the mixer having read those first.
Writing early replaced them about 21 times a second, which is the constant
crackle in its music and intro (it showed up with the sceAudio buffering
rework, which stopped copying buffers at enqueue).

Now the decoder's output is set aside, the old contents put back, and a
CoreTiming event writes the samples just before the thread wakes. Pending
writes are kept in savestates.

Adds audio/blocking/parked, recorded on a PSP: a blocking output that had
to wait returns before any of its buffer has played, so the game really
does depend on the decode's latency.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:49:09 -06:00
Henrik Rydgård 3d859bf7cc Merge pull request #22347 from ygordreyer/fix/umineko-texture-hash
Fix stale glyphs in swizzled CLUT4 atlases
2026-09-24 14:45:20 -06:00
Henrik RydgårdandClaude Opus 5.5 d1a8ce0dcc IR: Compile vmscl on any size and transposition
Only 4x4 with source and destination transposed alike, and a scale
outside the destination, compiled; most vmscl in games are transposed
or 3x3. The rest now multiply element by element, with the scale
copied first, and only a partly overlapping source still falls back.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 b87533a128 arm64 JIT: Compile ll/sc, and vcmp EN/NN/ES/NS with vs != vt
ll and sc always went to the interpreter; a few games use them
thousands of times. They now load and store directly with fast memory,
keeping llBit in MIPSState. vcmp's NaN and inf-or-NaN tests only look
at s, so they no longer need vt to be the same register: EN and NN
compare s with itself.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 d9eab0968a RISC-V JIT: Compile the bit ops and min/max without Zbb
Without the Zbb extension, clz, rotr, wsbh, wsbw, bitrev, min and max
went to the IR interpreter, and wsbh always did. They're now base ISA
sequences: a branchless binary search for clz, paired shifts for the
rotates and byte swaps, mask-and-shift steps for bitrev, and a
compare-and-branch for min and max. Checked against a model of the
instructions over random inputs, since nothing here runs RISC-V.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 9def690781 IR: Compile vt4444, vt5551 and vt5650
They always went to the interpreter. Each channel is now a shift, mask
and shift in the existing integer ops, so every backend handles them.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 6b78138c08 VFPU: Compile vh2f in the IR, and call vfpu_h2f for it everywhere
vh2f always went to the interpreter in the IR. A new FHalfToFloat op
converts the lower or upper half of a word, and the native backends
call vfpu_h2f for it like FSin. The legacy arm64 JIT now makes the same
call instead of computing the conversion inline; vh2f is rare, and the
call is much less code.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 ebd279fa70 arm64 JIT: Compile vh2f, vt4444/5551/5650, vbfy, vmscl, vcrs and vdet
All went to the interpreter. vh2f works in integers to match vfpu_h2f,
since FCVTL neither flushes subnormal halves nor keeps inf/NaN mantissa
bits unshifted. The color conversions are bitfield extracts and inserts.
vbfy, vcrs and vdet follow the IR frontend, and vmscl scales element
by element. Results go through scratch registers, so a destination that
overlaps a source is fine.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 eaf55c467b IR: Fold ClampToZero into the 31-bit packs
Vec4ClampToZero and Vec2ClampToZero only ever fed Vec4Pack31To8 and
Vec2Pack31To16, for vi2uc and vi2us. The packs now clamp negative lanes
to zero themselves, which saves an op and a vector temp, and lets x64
clamp with PACKUSWB's saturation after an arithmetic shift.

While at it, RISC-V compiles Vec2Unpack16To31, Vec2Pack31To16 and
Vec4Pack32To8, and LoongArch Vec2Unpack16To31, Vec2Pack31To16 and the
non-LSX Vec4Pack32To8, all of which went to the IR interpreter.
LoongArch's Vec2Pack32To16 and Vec2Unpack16To32 now take their scalar
path with LSX too, instead of falling back.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 3db7c65e86 x64 IR JIT: Compile the VFPU pack, unpack and clamp ops
vi2uc, vi2c, vi2us, vi2s, vuc2i, vc2i, vus2i and vs2i lowered to IR ops
that x64 left to the IR interpreter. They're now SSE2: shifts into
place, then PACKSSDW/PACKUSWB for the packs and self-unpacks for the
unpacks. An output overlapping its input still goes the slow way.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 dd416656a7 RISC-V/LoongArch JIT: Compile cvt.w.s
FCvtWS always went to the IR interpreter. Both now convert in the
current rounding mode, which ApplyRoundingMode keeps at the game's, as
x64 does. RISC-V's FCVT already saturates and gives INT_MAX for NaN;
LoongArch patches NaN to INT_MAX like its FRound.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 0c44852c9e LoongArch JIT: Compile min and max
They always went to the IR interpreter. Now a signed compare-and-branch
on the normalized sources picks which one to move.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik RydgårdandClaude Opus 5.5 8661fbf8ce arm64 JIT: Compile vf2i, vsgn, vslt, vsge and vcmp EI/NI
These went to the interpreter. vf2in/vf2iz/vf2iu/vf2id use FCVT, which
saturates like the PSP, and patch NaN to 0x7FFFFFFF like the IR
backend. vsgn keeps the sign bit on 1.0 and gives 0 below the smallest
normal. vsge and vslt select 1 or 0 on a compare whose condition is
false when unordered. EI and NI test |s| against infinity in integers.
All match the interpreter on special values, ties and denormals.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:31:22 -06:00
Henrik Rydgård 319fbedabf Merge pull request #22346 from hrydgard/macosx-fixes
MacOSX: Menubar fixes, compile warning fix
2026-09-24 14:27:24 -06:00
Henrik Rydgård 19f57ec438 Merge pull request #22345 from hrydgard/use-vfpu-interpolator
Use the new VFPU special-function implementations globally
2026-09-24 14:27:09 -06:00
Henrik RydgårdandClaude Opus 5.5 ecf9efcf84 macOS: Add menu items from the Windows menu
Stop, Switch UMD, savestate slots, Load/Save State (Cmd+L/Cmd+S) and state
files, shortcuts to the settings screens (plus Settings... with Cmd+, in the
app menu), Rendering Resolution, Texture Filtering, Enable Sound, Enable
Cheats and the website links.

Menu items now carry their action and how to tell whether they're checked
or enabled as blocks, checked each time a menu is shown, instead of tags
matched up in menuNeedsUpdate. The Graphics menu becomes Game Settings,
as on Windows.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:03:15 -06:00
Henrik RydgårdandClaude Opus 5.5 48b380a27e macOS: Fix menu bar items
- Recent files always booted the first one (no item had its tag set), and
  the list was never refreshed. Rebuild it when opened, with the path on each item.
- The Graphics and Debug menus were never updated when opened: they were
  matched by title, looked up under different translation keys than they
  were created with. Compare the menus themselves.
- The Fullscreen item toggled bFullScreen twice, doing nothing.
- The status counter items checked their tag instead of the flag.
- Remove Restart Graphics from the Debug menu (it shared tag 12 with Show
  Debug Statistics).
- Drop @available checks for versions below the deployment target.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 13:40:56 -06:00
Henrik RydgårdandClaude Opus 5.5 b832ceab83 IR: Add FExp2 and FLog2 for vexp2, vlog2 and vrexp2
These were always interpreted. They're now IR ops that the native
backends compile to calls to vfpu_exp2 and vfpu_log2, like FSin and
FAsin. vrexp2 is FNeg followed by FExp2, which is how vfpu_rexp2
computes it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 13:34:19 -06:00
Henrik RydgårdandClaude Opus 5.5 d313cf7626 arm64 JIT: Call the exact VFPU sin, cos, asin, exp2 and log2 too
vsin, vcos, vnsin, vasin, vexp2, vlog2 and vrexp2 went to the
interpreter. They now take the same direct call as vrcp and friends.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 13:34:19 -06:00
Henrik RydgårdandClaude Opus 5.5 1f5f2f5ebb macOS: Target 11.0, the oldest current libc++ supports
With a deployment target of 10.13 (10.14 for arm64 libretro), every file
that includes a libc++ header warned "The selected platform is no longer
supported by libc++." Apple Silicon Macs start at 11 anyway.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 13:21:36 -06:00
Henrik RydgårdandClaude Opus 5.5 e932d55f1c Build: Only turn off floating-point contraction for CPU emulation
The vertex decoder's C++ steps may fuse, like the vertex JITs do; with
contraction off everywhere, RISC-V's VertexJit test found the JIT and
the steps disagreeing on morphed float UVs. The flag now applies to
Core/MIPS only, in CMake and the libretro Makefile. ndk-build has no
per-file flags, so the legacy Android.mk goes back to the default.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 12:59:04 -06:00
Henrik RydgårdandClaude Opus 5.5 9b4d4ee35b VFPU: Build sin, cos, asin, exp2 and log2 results as integer bits
A float made from a constant NaN bit pattern can come out quieted: MSVC
turned the 0x7F800001 that vcos, vexp2 and vlog2 return into 0x7FC00001,
which failed cpu/vfpu/exact on Windows. Each function now computes its
result's bits in integers and converts once at the end, like vrcp and
friends already did. Fixed-point results become floats by shifting,
which is exact since the VFPU keeps 22 significant bits. Identical to
the previous code over every 32-bit input.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 12:59:04 -06:00
Henrik RydgårdandClaude Opus 5.5 7ecdbe7799 UnitTest: Remove the old asin and sin/cos approximation experiments
They only printed comparisons against libm and checked nothing. The VFPU
functions are exact now and have their own tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 12:24:43 -06:00
Henrik RydgårdandClaude Opus 5.5 eeb48b35c5 x86 JIT: Sum vqmul's y and w in the interpreter's order
They were added in a different order, so they rounded differently from
every other backend.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 12:24:43 -06:00
Henrik RydgårdandClaude Opus 5.5 1206528e45 Build: Turn off floating-point contraction for C++
Clang and GCC fuse a * b + c into one rounding on CPUs that have fused
multiply-add, which on ARM64 made the interpreter's dot products, vavg,
vdet and vqmul round differently from x86-64. Fused operations now only
happen where the code asks for them. MSVC doesn't contract by default.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 12:24:43 -06:00
Henrik RydgårdandClaude Opus 5.5 b9144d4fca VFPU: Fold the result exponent into the rcp, rsqrt and sqrt coefficients
Each of the three is now a range check, a fast path and a few special
cases. The fast path reads its segment from a table whose constant term
already includes the result's exponent bits, so what remains is two
multiplies, some shifts and one exponent adjustment, all in integers.
Identical to the previous code over every 32-bit input.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 12:24:43 -06:00
Henrik RydgårdandClaude Opus 5.5 ba77f2259d Add cpu/vfpu/exact to tests_good
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 12:24:42 -06:00
Henrik RydgårdandClaude Opus 5.5 a2b10778be VFPU: vsqrt, vrsq, vrcp and vnrcp are exact in every backend
These went through the host's sqrt and division everywhere except the
interpreter's vrcp and vnrcp (vsqrt and vrsq there only behind
USE_VFPU_SQRT, now gone). They now always give the PSP's bits: the IR
gets FVSqrt (FSqrt stays the FPU's IEEE sqrt.s), and FRSqrt and FRecip,
which only the VFPU emits, become vfpu_rsqrt and vfpu_rcp; the IR
interpreter and the x64, arm64, RISC-V and LoongArch backends call them.
The old JITs call them directly, the ARM ones keeping the lanes in
callee-saved registers across the calls. cpu/vfpu/exact now passes on every core.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 12:24:42 -06:00
Henrik Rydgård 87ad32d578 Merge pull request #22344 from hrydgard/download-cancel-at-exit
Don't crash at exit with a download still in flight
2026-09-24 12:10:21 -06:00
Henrik RydgårdandClaude Opus 5.5 f20dd7565b Don't crash at exit with a download still in flight
A request started since the last RequestManager::Update (headless never
calls it) sat in newDownloads_, which CancelAll skipped. It was then
destroyed along with the static g_DownloadManager at exit, and its
destructor removed its progress bar from the already destroyed g_OSD:
"mutex lock failed". Seen with a Netconf dialog still downloading the
infra DNS json when a test ended.

CancelAll now takes the new ones too, and runs at shutdown while g_OSD is
still there. The Netconf json request is also let go of when the emulator
shuts down, rather than living on into the next game.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 11:23:59 -06:00
Henrik Rydgård 428c1a6aa6 Merge pull request #22343 from hrydgard/vfpu-interpolator
VFPU: Replace fp64's correction tables with accurate implementation of special functions
2026-09-24 10:46:33 -06:00
Henrik Rydgård f7af4c0d4e Merge pull request #22339 from hrydgard/sceutility-lifecycle
sceUtility lifecycle fixes, GameSharing placeholder, headless fixes
2026-09-24 10:29:28 -06:00
Henrik RydgårdandClaude Opus 5.5 d6b3d590eb VFPU: Compute log2, sin/cos and asin without correction tables, too
They use the same quadratic interpolator as rcp and friends, with three
twists. sin indexes the quarter wave from the top, and asin and sin work
in a per-segment exponent whose 4-ulp truncation also applies to results
in a lower binade. log2 truncates exponent + log2(1.m) toward zero to 22
significant bits, and where that step is coarser than 2^-24 the datapath
drops coefficient bits to match; that also covers the region just below
1.0 that needed a special case.

vfpu_sincos now reduces the angle once. With every table gone, so are
the asset folder, the loader, InitVFPU and the fallbacks for tables that
failed to load. All seven functions are bit-exact with the table-based
code over every 32-bit input.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 10:07:01 -06:00
Henrik RydgårdandClaude Opus 5.5 dade876474 VFPU: Compute vrcp, vrsq, vsqrt and vexp2 without correction tables
The four share one quadratic interpolator: 128 segments picked by the top
7 bits of the input, each with a constant, a linear and a squared-term
coefficient, and a squarer on the top 10 bits of the rest that rounds t^2
up to a multiple of 256. 128 small coefficient sets per function
replace the 1 MB of delta tables. Derived from the output of the
table-based code, and bit-exact with it over every 32-bit input.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 10:06:53 -06:00
Henrik RydgårdandClaude Opus 5.5 157b233865 Savedata: TEMP TEST FIX: Do the hidden modes' IO inside Update, not on a host thread
The IO thread wrote results (file lists, sizes, loaded data) straight into
PSP memory while the game kept running, so they landed at an arbitrary
point in its code. utility/savedata/filelist caught it now and then: a
poll saw the entries written but the counts still zero. On a PSP it's all
there when Update returns. Host IO timing still uses the thread.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 19:41:41 -06:00
Henrik RydgårdandClaude Opus 5.5 d4de590e3f sceCtrl: Name the port-taking buffer functions
sceCtrlPeek/ReadBufferPositive2/Negative2, which take a port before the
usual buffer and count. Not implemented.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 19:30:31 -06:00
Henrik RydgårdandClaude Opus 5.5 55e9862b90 libretro: Build PSPPlaceholderDialog
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 19:30:31 -06:00
Henrik RydgårdandClaude Opus 5.5 30fada5833 Utility: Run GameSharing through a placeholder dialog, quiet its WRONG_TYPE logs
GameSharing InitStart used to return 0 without starting anything, so its
GetStatus stayed at NONE and a game waiting for the dialog to finish
would hang. It now goes through the normal lifecycle (INIT, RUNNING,
FINISHED, SHUTDOWN, NONE) and reports that the user cancelled.

PSPPlaceholderDialog was abstract and unused (and missing from CMake);
it's now that stand-in.

WRONG_TYPE from GameSharing GetStatus/Update/ShutdownStart is what a PSP
returns whenever another dialog type was the last one started, so log it
at debug like the other dialogs. Sega Rally polls it every frame.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 19:25:24 -06:00
Henrik RydgårdandClaude Opus 5.5 5a23bb2d94 Savedata: Pick the crypt mode the way the firmware does
secureVersion and the key only matter for the full 1536-byte request; the
older sizes always save without a key. Otherwise SDK 2.07+ uses the new
keyed hash for versions 0 and 3, older SDKs the old keyed hash for 0 and
no key for 3, and version 2 is always the old keyed hash.

Versions 0, 2 and 3 all require a key, and are rejected with SAVE_PARAM
otherwise - including 0, which we used to save without a key. Matches the
SFO modes and save errors recorded in utility/savedata/secureversion; what
remains there is load strictness when secureVersion doesn't match the
file, which we keep lenient for older saves.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 19:25:24 -06:00
Henrik RydgårdandClaude Opus 5.5 ec00700730 Dialog: Fade in 200ms like the firmware, and accept input while fading in
On a PSP, dialog animations advance by animSpeed frames per Update and the
fades take about 200ms. Ours took 500ms (1/30 s per animSpeed, over
FADE_TIME 1.0), and all input was ignored until it finished, which made
dialogs feel sluggish, noticeably so in 30fps games. Input is still
ignored while fading out, once a choice has been made.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 19:25:24 -06:00
Henrik RydgårdandClaude Opus 5.5 cafbab50d9 Io: Reject memory stick paths with characters FAT can't store
sceIoOpen on such a path returns an invalid-argument error on hardware,
not file-not-found (utility/savedata/idlist).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 19:25:24 -06:00
Henrik RydgårdandClaude Opus 5.5 74c5cbf503 Savedata: Leave GetSize's needed strings alone when nothing is needed
Matches hardware; moves utility/savedata/getsize to tests_good.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 19:25:24 -06:00
Henrik RydgårdandClaude Opus 5.5 20970d5fad Dialog: Follow the firmware's thread priorities in init/shutdown
On a PSP, dialog init and shutdown happen partly at the accessThread
priority and partly at the graphicsThread priority, one phase after the
other. Model that with one helper thread that switches priority per phase,
and let starting it reschedule normally instead of disabling interrupts.

A caller with worse priority than both now sees shutdown complete inside
ShutdownStart, as on hardware. NFL Street 3 (graphics 17, access 19,
caller 111) calls the next InitStart right after ShutdownStart and used to
loop forever on 'A save request is already running' (#19957). The utility
pspautotests, where the caller has better priority, are unchanged.

Also logs the dialog thread priorities at debug level.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 19:25:24 -06:00
Henrik RydgårdandClaude Opus 5.5 24e44cb5dd CmdLine: Exit with an error on a bad option value
An invalid value (e.g. --debugger-run swallowing the next flag) printed an
error but returned Exit, so the process exited 0.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 19:25:24 -06:00
Henrik RydgårdandClaude Opus 5.5 a293221fe2 Headless: Default to software rendering again
Since 551e4cd0ab, headless without --graphics silently used the OpenGL
backend, since bSoftwareRendering defaulted to false. Under Mesa llvmpipe
on Linux/WSL, that hangs games early in boot (and test.py, which passes no
--graphics). The README already documents software as the default.

Also documents the traps hit while chasing this in docs/debugging.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 19:25:24 -06:00
Henrik Rydgård 09cf031f2b Merge pull request #22342 from hrydgard/vdot-neon
Accurate vdot: NEON and SSE2 implementations
2026-09-23 19:24:16 -06:00
Henrik Rydgård f65598c60f Merge pull request #22341 from sum2012/gpu-readback
Add ForceEnableGPUReadback compat By Gemini
2026-09-23 09:02:45 -06:00
Henrik RydgårdandClaude Opus 5.5 0ac12b7bf3 VFPUDot: Cover sums that round into the next power of two
A sum just below a power of two can round up into the next exponent, and
at the top of the range into inf; random inputs almost never land there.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 08:18:25 -06:00
Henrik RydgårdandClaude Opus 5.5 cd005b9041 VFPU: SSE2 version of the exact vdot
Same shape as the NEON one, which now shares its rounding tail. SSE2 has
no per-lane shift, so the alignment shift is a multiply by a power of two
built from float bits and converted by truncation; the unsigned maxima
use the 16-bit instructions, since every value involved fits in 15 bits.
Nothing depends on the host rounding mode or flush-to-zero.

Checked against the reference by VFPUDot and on 300M more inputs offline,
also with MXCSR set to round toward zero with FTZ and DAZ.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 08:18:25 -06:00
Henrik RydgårdandClaude Opus 5.5 649561764d VFPU: NEON version of the exact vdot
The four lanes are computed together: exponents, 24x24-bit products with
round-to-odd, alignment by truncation and a signed horizontal sum. One
pairwise maximum finds both the alignment exponent and any inf or NaN,
which go to the reference. The final rounding is branch-free and in
integers, since the host rounding mode may be the game's.

About three times the throughput of the reference on Apple M-series
(5.2 vs 15.2 ns per call). VFPUDot checks it against the reference on
four million inputs picked to cover cancellation, ties, subnormals and
the overflow edges; a billion more matched offline.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 08:18:25 -06:00
Henrik Rydgård fa583d1d33 Merge pull request #22340 from sashwatpuri/fix/audio-file-duration-limit
Limit achievement sound effect duration
2026-09-22 21:39:19 -06:00
Henrik Rydgård 7c18ad3d1e Merge pull request #22338 from hrydgard/video-texture-hash
Video texture tracking and avoiding hashing
2026-09-22 16:50:34 -06:00
Henrik Rydgård 818c08be4c Merge pull request #22337 from hrydgard/cpu-test-holes
Claude with pspautotests: CPU test holes
2026-09-22 16:19:17 -06:00
Henrik RydgårdandClaude Opus 5 772f4f6e7a AGENTS.md: note that headless needs the app's memstick for firmware runs
Without --memstick, a library whose HLE the config has disabled has no firmware
to resolve against, and the game dies on unresolved imports rather than falling
back to our HLE. The run still logs happily for its whole timeout, so the only
tell is zero flips.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-22 16:10:52 -06:00
Henrik RydgårdandClaude Opus 5 4aef060293 Carry "this is video" across block copies
videos_ only learns about the CSC output, so the texture we actually draw is
just an ordinary 512x512 8888 texture whose contents happen to be different
every frame: hash, miss, throw it in the secondary cache, rebuild, forever.

So track the copy. NotifyVideoCopy marks the destination as video when the
source is, and the copy funnels call it: sceDmacMemcpy, sceKernelMemcpy, and
the four replaced memcpy/memmove variants. It sits outside their "is either
side VRAM" gate, since a RAM-to-RAM copy of a frame is still a frame.

Being a video texture also gets it forced linear filtering and keeps it out of
texture upscaling, which is what you want for a video either way.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-22 16:10:28 -06:00