sceUriParse's size query (no parsed-URI or work area) returns 0 on a PSP, not -1,
which made Wipeout Pure's embedded browser give up before connecting. And
sceHttpGetAllHeader hands out the header block as received, ending with the blank
line, NUL-terminated and with the NUL counted; without the blank line the browser
never displayed the image the page consists of. Both from the 6.60 firmware
modules (libparse_uri.prx, libhttp.prx).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
When a Load32, Store32, LoadFloat or StoreFloat is followed by the same op
on the next or previous word through the same base, and the base can be
mapped as a pointer, emit one LDP/STP for both. A struct-copying loop runs
about 24% faster; code without such pairs is unaffected.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Every op ended with a break back to one shared indirect jump, which the
CPU has to predict for every op in the program. With labels as values,
each op jumps through a table from its own site instead, which predicts
much better: an integer-heavy benchmark runs about 13% faster on an M1.
Other compilers keep the switch, and ops missing from the table fall back
to it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Blocks ending in a branch dispatch a conditional exit and then the
fallthrough ExitToConst. One op now returns either target, reading the
second from the ExitToConst, which stays behind unexecuted.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
It avoids flushes a native backend would need, at the cost of extra
instructions (a copy of each scalar before a Vec4Scale, for instance) that
the interpreter only has to dispatch.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Both were no-ops, so mfic left its destination unchanged. They read and
write the interrupt enable flag that sceKernelCpuSuspendIntr/ResumeIntr
use. Only bit 0 counts for mtic, which also goes for
sceKernelCpuResumeIntr, since on hardware it's just mtic.
Adds the intr/mfic test, recorded on hardware.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The third argument is a timeout pointer, as threadman.prx shows. A kernel
address from user mode is ILLEGAL_ADDR there; we used to write through it.
Also, no lookup by index: the syscall requires the exact uid.
Adds the threads/tls/allocate test, recorded on hardware.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This is the syscall usersystemlib's sceKernelGetTlsAddr makes when the
thread's cached TLS address is null, as (uid, &addr, 0). Code that has to
run before usersystemlib.prx is loaded (like plugins built with a Rust SDK)
inlines sceKernelGetTlsAddr and imports this directly.
Shares the allocation with sceKernelGetTlsAddr. A thread waiting on a full
pool stores its address pointer as the wait value, so no state changes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A CGL context with no drawable, rendering into a framebuffer object of its
own that stands in for the backbuffer through g_defaultFBO. Core profile,
as the SDL app uses on macOS.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
OpenGL's render thread only finishes a frame when it's presented, so the
emu thread eventually waited forever in BeginFrame for a free frame, with
the render thread waiting for work. GPU tests and games hung, silently,
as a blocked host thread also defeats the timeouts.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Vulkan no longer needs a hidden window (which on macOS could never work, as
the Metal window description has no data2). It now uses the offscreen mode,
and presents each frame like the app does, as an unpresented frame would
wait forever for its next image. MoltenVK's console logging is limited to
errors, as it mixed into the test output.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Executables outside a bundle (headless) found no Vulkan library. Also try
the app bundle built next to them, the Vulkan SDK's install and Homebrew's.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A frame that never draws to the backbuffer never acquires an image, but
its final submit still waited on the acquire semaphore, which nothing
signals, hanging the GPU. Skip the swap for such frames when finishing
them. This replaces the check for a frame with no steps at all, which
could also set it partway through a frame that acquires later.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
For when there's nothing to present to. Instead of a surface and swapchain,
it renders into images of its own through the VulkanPresentation interface
libretro uses, picking the graphics queue without a surface. Acquiring and
presenting signal and wait on the frame's semaphores with empty submits.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
OptimizeLoadsAfterStores only dropped a load right after a store of the same
reg. Now a load of anything the block stored or loaded before becomes a reg
move (with the extension for 8/16-bit loads), as long as nothing in between
may have changed the memory, the address reg, or the reg holding the value.
Only a store through the same base at a disjoint range is known not to alias,
and constant addresses outside RAM are left alone.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
When the cosine lane follows the sine lane, FSinCos can write both in place
instead of going through a temp and two FMovs.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The IR temps don't live past the block, but FlushAll stored them anyway at
every exit (the branch operands, lwl/lwr temps, VFPU temp lanes). At an exit,
discard the ones nothing later in the block reads.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- A constant stays known after it's written out for a read (by a store,
MovZ, a multiply...), so later uses still fold. Whatever an op writes is
forgotten after its inputs are written, and setting a reg to the value it
already holds isn't written twice.
- A conditional exit that isn't taken keeps the constants known.
- The saturating and min/max FP ops, FSign and the 31-bit Vec2 pack/unpack
no longer flush every GPR constant.
- A load through its own base is folded (lui v0, hi; lw v0, lo(v0)).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- An lwl/lwr pair was combined into one load even when the first half loads
into the base register, which changes the address of the second half.
- ApplyMemoryValidation shared one sp check across the block even past an
Interpret or CallReplacement, which may change sp.
- Drop a duplicate FSqrt meta entry, and name Load8Ext correctly.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Copy propagation through an FPR temp didn't stop when an instruction
rewrote the temp in place (it compared an FPR number against the +32
offset reg), so later reads lost that write.
- A read of the temp in both operands only had src1 replaced, yet the copy
into the temp was still removed.
- The replacement matched operands by number without checking their type,
so a StoreFloat whose GPR address had the temp's number got its address
replaced (IRVTEMP_PFX_S and IRTEMP_0 are both 192).
- A write to lanes 1-3 of a Vec4 temp wasn't noticed.
- IRReadsFromFPRs stopped after the F operands, missing Vec4Scale's vector.
- Exits and barriers didn't count as reading everything, so a write to a
real reg could be moved above an exit.
- Load32Linked and Store32Conditional were removed when their reg was
overwritten unread, losing LLBIT and the store.
Also fixes an off-by-one in the vec src3 read check. The unit test now
reports every failing case instead of stopping at the first.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
After the registers usable by compressed instructions, prefer s2-s7 and
fs2-fs11, so that fewer values have to be flushed around calls. The
dispatcher already saves them.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A new FSinCos op writes both from one argument reduction. arm64 and x64 get
both back from a single call, packed in one double; RISC-V and LoongArch make
the two calls.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The cosine is then taken of what vrot wrote to that lane: the sine, or zero.
The IR looked at the sine lane instead of the lane holding the angle, and the
legacy JITs ignored the overlap. The assembler refuses such a vrot, so those
now leave it to the interpreter, and don't pair one with the vrot before it.
Covered by the new cpu/vfpu/vrot test.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The ABI only preserves the low 64 bits of F24-F31, which are first in
the allocation order, so a four-lane vector there lost its upper half
across a call to a math helper. Flush those like arm64 does for
S8-S15.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The thunk saves a fixed set of registers and MXCSR on every call. The
math helpers (vrcp through vrexp2, and vrot's sincos) leave MXCSR alone,
so they now go through CallProtectedLeaf, which saves only the
caller-saved registers the caches are using, around a direct call. vrot
also no longer flushes everything first. x86-64 only; 32-bit x86 keeps
the thunk.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ProtectFunction's thunk saved all of XMM2-15 and RBX around every call,
but the callee preserves XMM6-15 on Windows and RBX everywhere. On
Windows that drops ten 16-byte saves and loads from each protected call.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Values in S8-S15 survive a call, but the full flush wrote them back and
the following instructions loaded them again. The VFPU math callouts,
vrot and vh2f now flush only the caller-saved registers, plus the few
callee-saved ones they stage values in, and map the destinations
afterwards. In a normalize loop with sixteen VFPU registers live, that
takes a vrsq.s from about 15 to 11.5 ns.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With LSX a lane group can be mapped as one vector reg, and F() then returns
the same reg for every lane, so the per-lane code wrote only lane 0 (vs2i,
vus2i). Also drop the stale aliasing TODOs; each path reads its sources first.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On a PSP the Media Engine decodes, and the samples land in the output
buffer as the call returns, a couple of milliseconds in. We wrote them at
once and only then delayed the thread. Since sceAudio plays straight out
of game memory, that matters: Fired Up decodes each chunk to 0x40 bytes into
one of its two buffers, running over the first 16 samples of the other one,
which it has just queued, and relies on the mixer having read those first.
Writing early replaced them about 21 times a second, which is the constant
crackle in its music and intro (it showed up with the sceAudio buffering
rework, which stopped copying buffers at enqueue).
Now the decoder's output is set aside, the old contents put back, and a
CoreTiming event writes the samples just before the thread wakes. Pending
writes are kept in savestates.
Adds audio/blocking/parked, recorded on a PSP: a blocking output that had
to wait returns before any of its buffer has played, so the game really
does depend on the decode's latency.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Only 4x4 with source and destination transposed alike, and a scale
outside the destination, compiled; most vmscl in games are transposed
or 3x3. The rest now multiply element by element, with the scale
copied first, and only a partly overlapping source still falls back.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ll and sc always went to the interpreter; a few games use them
thousands of times. They now load and store directly with fast memory,
keeping llBit in MIPSState. vcmp's NaN and inf-or-NaN tests only look
at s, so they no longer need vt to be the same register: EN and NN
compare s with itself.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Without the Zbb extension, clz, rotr, wsbh, wsbw, bitrev, min and max
went to the IR interpreter, and wsbh always did. They're now base ISA
sequences: a branchless binary search for clz, paired shifts for the
rotates and byte swaps, mask-and-shift steps for bitrev, and a
compare-and-branch for min and max. Checked against a model of the
instructions over random inputs, since nothing here runs RISC-V.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
They always went to the interpreter. Each channel is now a shift, mask
and shift in the existing integer ops, so every backend handles them.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
vh2f always went to the interpreter in the IR. A new FHalfToFloat op
converts the lower or upper half of a word, and the native backends
call vfpu_h2f for it like FSin. The legacy arm64 JIT now makes the same
call instead of computing the conversion inline; vh2f is rare, and the
call is much less code.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
All went to the interpreter. vh2f works in integers to match vfpu_h2f,
since FCVTL neither flushes subnormal halves nor keeps inf/NaN mantissa
bits unshifted. The color conversions are bitfield extracts and inserts.
vbfy, vcrs and vdet follow the IR frontend, and vmscl scales element
by element. Results go through scratch registers, so a destination that
overlaps a source is fine.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>