The blit rates were measured with nothing else running. In a game, threads
waking up and SAS mixing on the Media Engine compete with the GE for main RAM:
Star Wars: Lethal Alliance's movie blit takes 8.65ms alone and 10.3ms in the
game. We don't model that load, so RAM texture fetches get a fixed 1.17x for a
typical one. With it, that game's long movie plays at 30 fps as on hardware.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The SAS mix estimate was a guess capped at 1200us. Measured on a PSP
(pspautotests audio/timing/sastiming), a mix costs 110us plus 0.49us per grain
sample, plus per voice and sample 0.445us + 0.0675us per unit of pitch ratio
(VAG; PCM and noise slightly less), plus 0.64us per sample with a reverb type
set. Linear to 1% across 64-2048 samples and 0-32 voices: 32 VAG voices at 512
samples take 8.7ms, not 1.2.
The mix runs on the Media Engine, as do video and audio decoding, so they now
queue behind each other there (MEScheduleJob). A game that keeps SAS running
during a movie - Star Wars: Lethal Alliance - decodes more slowly for it, like
on hardware.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Full-screen clears measured on a PSP (pspautotests gpu/timing/blittiming):
0.49ms on a 16-bit framebuffer whatever is cleared, 0.69ms on 8888, 1.02ms on
8888 with depth. Charging them may help games that spin hard on an empty
screen, but it's off (chargeClearTime) until tried on some. The video blit
cost moves into the same function, now EstimateFillCycles.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The blit cost remembered only the last buffer a decoder wrote into, forever.
Move the texture cache's video list (with its ageing out a few flips after
the last write) into GPUCommon, so the texture cache, the blit cost and
SoftGPU all share one. That also counts both of a double-buffered player's
frames, which exposed that a clear drawn with texturing still enabled was
being charged as a blit - skip clears and draws without texture coordinates.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Measured on a PSP (pspautotests gpu/timing/blittiming), a full-screen blit
from an unswizzled texture costs what the texture fetch costs: 16-bit formats
half of 32-bit, VRAM a fifth of RAM, and rectangles wider than ~128 texels
~7.5x as much as narrow strips, from texture cache thrashing. Framebuffer
format, filtering and blending don't matter. Ys I & II draws its movie as one
full-width sprite from a 565 texture in RAM, which takes 33ms - that, not the
decode, is what holds it to 30 fps.
All of the ME and GE costs speed up with the clock (2/3 as long at 333/166),
since the whole system runs from the one PLL.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
sceMp4AacDecode of an AAC-LC stereo frame takes about 1.7ms on a PSP, measured
with pspautotests video/mp4/mp4timing.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Movie players like the one in Star Wars: Lethal Alliance present every decoded
frame after a single vblank wait, with no clock or timestamp check, so the
frame rate depends on decode, CSC, ATRAC decode and the GE blit adding up to
more than a vblank. We charged nearly nothing for any of them, so such movies
ran at 60 fps until the ringbuffer's slack ran out.
Costs measured on a PSP with a copy of that player (pspautotests
video/mpeg/playertiming), for a 480x272 frame:
- sceVideocodecDecode: 3.4ms (sceMpegAvcDecode 5.8ms less sceMpegAvcCsc 2.4ms)
- sceMpegBaseCscAvc: 2.4ms, was a flat 4ms
- sceAudiocodecDecode, ATRAC3+ only: 2.5ms per frame
- GE: 9.7ms for a through-mode rectangle blit from a decoded video frame,
charged by area, only for textures in the buffer a decoder last wrote.
GE time also now carries across stall address updates. Before, a list sent
in stalled chunks only had its last chunk's time counted, so sceGeDrawSync
returned 39us after a blit that takes 9.7ms. This affects every game that
builds its lists incrementally, so GE-timing-sensitive games need checking.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
While apctl was in JOINING, every Update restarted the fade-in, so the dialog
sat at its first, barely visible step until the state moved on to getting an
IP. It now keeps the fade-in it started with.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- The result was only ever written on cancel, so a connection that worked
gave the game back whatever was in the field, which some fill with -1.
- Infrastructure mode drew a Cancel button that did nothing, so if the
access point never gave an IP there was no way out. Cancelling now also
disconnects the connect it started, so the game isn't left connected
after being told the dialog was aborted.
- An unknown netAction went down neither path and never finished.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Loading a utility module without room for our (rough) memory block stored
(u32)-1 as its address, which av_atrac3plus then memset. It now loads
without the block, which HLE doesn't need.
- LoadNetModule's module id was unsigned, so a load error was passed on to
sceKernelStartModule as an id.
- The dialog helper threads put the game's priorities straight into ORI
immediates; ones that aren't a priority at all now fall back to 0x20.
- A fade never finished with an animSpeed of 0 or less.
- MsgDialog V3 button captions filling all 64 bytes ran on into the next
field.
- GamedataInstall: a file shorter than it claimed was retried forever, Abort
worked in any state and wrote through an unchecked pointer, and the game
and data names were read as C strings from fixed-size fields.
- NpSignin: a cancel was overwritten with SUCCESS in the same frame, so the
game saw a sign-in. The status is reset on start, so a reused struct no
longer hangs.
- Unloading av_atrac3plus never reached __AtracNotifyUnloadModule, leaving
the atrac state pointing at freed memory.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- InitStart checks the address and size the way sceUtility_Driver does
(0x40 and 0x44 are both fine, utility/dialog/sizes) and that the field
struct is in memory.
- The input text was read up to a terminator with no limit and no memory
check, and the output was written through an unchecked pointer every
Update. Both are bounded and checked now.
- Converting a string to UTF-8 checked for room before each character, then
wrote up to three bytes and a terminator, overflowing a 2048-byte stack
buffer on long non-ASCII text (both conversions did).
- An output buffer of length 0 made FieldMaxLength wrap around, and the
keyboard preview then indexed the text at -1.
- The native input box's callbacks captured the dialog and could write into
it after it was deleted, and its status was read and written without the
lock. They now share a small state object instead, a new one per start
and per state load (unless a box is still open, which answers into the
loaded state).
- Savestates keep the current keyboard and the Korean combining state. With
older ones, it picks a keyboard the field allows, as Init does.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- The Yugioh savedata workaround force-stopped the other dialog without
releasing the volatile memory it held, so the savedata helper then waited
for it forever.
- Screenshot's ShutdownStart and Update succeeded when no screenshot was
running, putting it back into SHUTDOWN.
- Savestates: Netconf didn't keep its request address, and a DNS json
download going on was gone after a load, so it could wait for it forever;
it now fetches the json again (it's cached). NpSignin didn't keep its
request address either. Both restart their timeouts instead of timing out
at once. With states from before, they keep the current address as they
used to.
- Loading a state from before NpSignin, GameSharing or HtmlViewer were saved
resets them rather than keeping this session's state (including the
HtmlViewer's memory block), and without Shutdown's side effects, which
would write to the loaded memory and release its volatile lock.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
InitStart sizes: Netconf and NpSignin accepted any size, then wrote
common.size bytes back from a 64-68 byte host struct, copying host memory
into PSP RAM. GamedataInstall looked for install files before checking the
size, and the HtmlViewer read options before checking the whole request was
in memory. All the dialogs now check the address, then the sizes
sceUtility_Driver accepts (utility/dialog/sizes), before anything else, as
the firmware does (a bad address is INVALID_ADDRESS), and write back no more
than the struct.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Replaces the temporary fix that did the hidden modes' IO inside Update.
The IO thread read and wrote the dialog's request, display state and save
list, all shared with the emulator thread, which kept using them to draw the
dialog and reload the request from the game.
Now the IO thread works on its own copy of the request, its own SavedataParam
and directory names resolved up front, and shares nothing else with the
emulator thread but the (locked) file system, MemoryStick_FreeSpace's cached
use and sceChnnlsv's scratch buffer and kirk state, the last two now under
locks too. It still reads and writes the game's buffers directly, like a PSP's
utility threads and sceIoReadAsync do, so a savestate waits for it before it
saves or loads memory. Save
bookkeeping, the save indicator and display changes happen on the emulator
thread when the results are taken, and only the request fields the IO changed
are copied back, so a game's own edits in the meantime survive.
Hidden modes take the results at the next Update (or, with Host IO timing,
the first Update that finds them done). The visible dialogs keep drawing and
take them once the IO is done; save and load used to stall the emulator
thread for the whole operation. Savestates keep results that haven't been
taken yet.
When the results land in PSP memory doesn't matter to games, so
utility/savedata/filelist now only prints them once the utility has finished.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Until then the dialog runs normally (and writes result = 0, which an
immediate abort skipped). Measured with one Update per vblank; at one every
other vblank a PSP took 6, so it isn't purely a count.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Instead of the PSP's web browser, a dialog shows the URL the game wants
and opens it in the host's browser on X, or backs out on O. Either way the
game sees the browser closed normally. Platforms that can't open a URL
(the new SYSPROP_CAN_LAUNCH_URL) only offer to back out. Only plain
printable-ASCII http(s) addresses are handed over, and on Linux without a
shell.
What the firmware does (sceUtility_Driver, 6.61, plus
utility/dialog/htmlviewer): the HtmlViewer has its own state apart from the
other dialogs, so they don't block each other, and its calls return
WRONG_TYPE until one has started. The request size picks the 2.00 to 3.00
layout, and InitStart allocates 3.5MB of user memory (4.5MB with options
bit 0x400 from 2.70 on), failing with 800200d9 without it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On a PSP, every InitStart fails with INVALID_STATUS until the last dialog
started is back at NONE, including while it's shutting down, and before
its params are checked. A failed InitStart leaves the current type alone.
We returned WRONG_TYPE instead, and a failed InitStart (e.g. a bad size)
still switched the current type, so every later dialog was refused. (One of
ours that fails after already starting, as savedata can, still becomes the
current type, since the game may poll it.)
The busy check applies status changes that are due, but doesn't use up an
auto status dialog's one-time INITIALIZE/SHUTDOWN reports; one that only
waits to report SHUTDOWN is let finish. Auto status dialogs now release
volatile memory on the way to NONE, including when the game saw RUNNING
before the init thread was done, which used to leave it locked for the next
dialog. GamedataInstall no longer requires currentDialogActive, which its
ShutdownStart cleared even when it then failed, so it could never finish -
and would now have blocked every other dialog.
Also: MsgDialog accepts exactly the three sizes sceUtility_Driver does (we
memcpy'd whatever size was given), and HtmlViewer GetStatus answers
WRONG_TYPE.
Adds utility/dialog/status and utility/dialog/priority, recorded on a PSP.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A stop ended the run unless the CPU had been told to break at start, which is
only --debugger. So under --debugger-run, pausing from the debugger (or any
breakpoint) exited the process. Key it on whether the debugger is on instead.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The apctl info copies the AutoDNS server when the connection gets its IP, which
normally happens before netconf has downloaded infra-dns.json, so games that read
the primary DNS server afterwards (to do their own lookups) got an empty string.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
sceUriParse's size query (no parsed-URI or work area) returns 0 on a PSP, not -1,
which made Wipeout Pure's embedded browser give up before connecting. And
sceHttpGetAllHeader hands out the header block as received, ending with the blank
line, NUL-terminated and with the NUL counted; without the blank line the browser
never displayed the image the page consists of. Both from the 6.60 firmware
modules (libparse_uri.prx, libhttp.prx).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
When a Load32, Store32, LoadFloat or StoreFloat is followed by the same op
on the next or previous word through the same base, and the base can be
mapped as a pointer, emit one LDP/STP for both. A struct-copying loop runs
about 24% faster; code without such pairs is unaffected.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Every op ended with a break back to one shared indirect jump, which the
CPU has to predict for every op in the program. With labels as values,
each op jumps through a table from its own site instead, which predicts
much better: an integer-heavy benchmark runs about 13% faster on an M1.
Other compilers keep the switch, and ops missing from the table fall back
to it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Blocks ending in a branch dispatch a conditional exit and then the
fallthrough ExitToConst. One op now returns either target, reading the
second from the ExitToConst, which stays behind unexecuted.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
It avoids flushes a native backend would need, at the cost of extra
instructions (a copy of each scalar before a Vec4Scale, for instance) that
the interpreter only has to dispatch.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Both were no-ops, so mfic left its destination unchanged. They read and
write the interrupt enable flag that sceKernelCpuSuspendIntr/ResumeIntr
use. Only bit 0 counts for mtic, which also goes for
sceKernelCpuResumeIntr, since on hardware it's just mtic.
Adds the intr/mfic test, recorded on hardware.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The third argument is a timeout pointer, as threadman.prx shows. A kernel
address from user mode is ILLEGAL_ADDR there; we used to write through it.
Also, no lookup by index: the syscall requires the exact uid.
Adds the threads/tls/allocate test, recorded on hardware.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This is the syscall usersystemlib's sceKernelGetTlsAddr makes when the
thread's cached TLS address is null, as (uid, &addr, 0). Code that has to
run before usersystemlib.prx is loaded (like plugins built with a Rust SDK)
inlines sceKernelGetTlsAddr and imports this directly.
Shares the allocation with sceKernelGetTlsAddr. A thread waiting on a full
pool stores its address pointer as the wait value, so no state changes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A CGL context with no drawable, rendering into a framebuffer object of its
own that stands in for the backbuffer through g_defaultFBO. Core profile,
as the SDL app uses on macOS.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
OpenGL's render thread only finishes a frame when it's presented, so the
emu thread eventually waited forever in BeginFrame for a free frame, with
the render thread waiting for work. GPU tests and games hung, silently,
as a blocked host thread also defeats the timeouts.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Vulkan no longer needs a hidden window (which on macOS could never work, as
the Metal window description has no data2). It now uses the offscreen mode,
and presents each frame like the app does, as an unpresented frame would
wait forever for its next image. MoltenVK's console logging is limited to
errors, as it mixed into the test output.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Executables outside a bundle (headless) found no Vulkan library. Also try
the app bundle built next to them, the Vulkan SDK's install and Homebrew's.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A frame that never draws to the backbuffer never acquires an image, but
its final submit still waited on the acquire semaphore, which nothing
signals, hanging the GPU. Skip the swap for such frames when finishing
them. This replaces the check for a frame with no steps at all, which
could also set it partway through a frame that acquires later.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
For when there's nothing to present to. Instead of a surface and swapchain,
it renders into images of its own through the VulkanPresentation interface
libretro uses, picking the graphics queue without a surface. Acquiring and
presenting signal and wait on the frame's semaphores with empty submits.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
OptimizeLoadsAfterStores only dropped a load right after a store of the same
reg. Now a load of anything the block stored or loaded before becomes a reg
move (with the extension for 8/16-bit loads), as long as nothing in between
may have changed the memory, the address reg, or the reg holding the value.
Only a store through the same base at a disjoint range is known not to alias,
and constant addresses outside RAM are left alone.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
When the cosine lane follows the sine lane, FSinCos can write both in place
instead of going through a temp and two FMovs.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The IR temps don't live past the block, but FlushAll stored them anyway at
every exit (the branch operands, lwl/lwr temps, VFPU temp lanes). At an exit,
discard the ones nothing later in the block reads.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- A constant stays known after it's written out for a read (by a store,
MovZ, a multiply...), so later uses still fold. Whatever an op writes is
forgotten after its inputs are written, and setting a reg to the value it
already holds isn't written twice.
- A conditional exit that isn't taken keeps the constants known.
- The saturating and min/max FP ops, FSign and the 31-bit Vec2 pack/unpack
no longer flush every GPR constant.
- A load through its own base is folded (lui v0, hi; lw v0, lo(v0)).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- An lwl/lwr pair was combined into one load even when the first half loads
into the base register, which changes the address of the second half.
- ApplyMemoryValidation shared one sp check across the block even past an
Interpret or CallReplacement, which may change sp.
- Drop a duplicate FSqrt meta entry, and name Load8Ext correctly.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Copy propagation through an FPR temp didn't stop when an instruction
rewrote the temp in place (it compared an FPR number against the +32
offset reg), so later reads lost that write.
- A read of the temp in both operands only had src1 replaced, yet the copy
into the temp was still removed.
- The replacement matched operands by number without checking their type,
so a StoreFloat whose GPR address had the temp's number got its address
replaced (IRVTEMP_PFX_S and IRTEMP_0 are both 192).
- A write to lanes 1-3 of a Vec4 temp wasn't noticed.
- IRReadsFromFPRs stopped after the F operands, missing Vec4Scale's vector.
- Exits and barriers didn't count as reading everything, so a write to a
real reg could be moved above an exit.
- Load32Linked and Store32Conditional were removed when their reg was
overwritten unread, losing LLBIT and the store.
Also fixes an off-by-one in the vec src3 read check. The unit test now
reports every failing case instead of stopping at the first.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
After the registers usable by compressed instructions, prefer s2-s7 and
fs2-fs11, so that fewer values have to be flushed around calls. The
dispatcher already saves them.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A new FSinCos op writes both from one argument reduction. arm64 and x64 get
both back from a single call, packed in one double; RISC-V and LoongArch make
the two calls.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The cosine is then taken of what vrot wrote to that lane: the sine, or zero.
The IR looked at the sine lane instead of the lane holding the angle, and the
legacy JITs ignored the overlap. The assembler refuses such a vrot, so those
now leave it to the interpreter, and don't pair one with the vrot before it.
Covered by the new cpu/vfpu/vrot test.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The ABI only preserves the low 64 bits of F24-F31, which are first in
the allocation order, so a four-lane vector there lost its upper half
across a call to a math helper. Flush those like arm64 does for
S8-S15.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The thunk saves a fixed set of registers and MXCSR on every call. The
math helpers (vrcp through vrexp2, and vrot's sincos) leave MXCSR alone,
so they now go through CallProtectedLeaf, which saves only the
caller-saved registers the caches are using, around a direct call. vrot
also no longer flushes everything first. x86-64 only; 32-bit x86 keeps
the thunk.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ProtectFunction's thunk saved all of XMM2-15 and RBX around every call,
but the callee preserves XMM6-15 on Windows and RBX everywhere. On
Windows that drops ten 16-byte saves and loads from each protected call.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Values in S8-S15 survive a call, but the full flush wrote them back and
the following instructions loaded them again. The VFPU math callouts,
vrot and vh2f now flush only the caller-saved registers, plus the few
callee-saved ones they stage values in, and map the destinations
afterwards. In a normalize loop with sixteen VFPU registers live, that
takes a vrsq.s from about 15 to 11.5 ns.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With LSX a lane group can be mapped as one vector reg, and F() then returns
the same reg for every lane, so the per-lane code wrote only lane 0 (vs2i,
vus2i). Also drop the stale aliasing TODOs; each path reads its sources first.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On a PSP the Media Engine decodes, and the samples land in the output
buffer as the call returns, a couple of milliseconds in. We wrote them at
once and only then delayed the thread. Since sceAudio plays straight out
of game memory, that matters: Fired Up decodes each chunk to 0x40 bytes into
one of its two buffers, running over the first 16 samples of the other one,
which it has just queued, and relies on the mixer having read those first.
Writing early replaced them about 21 times a second, which is the constant
crackle in its music and intro (it showed up with the sceAudio buffering
rework, which stopped copying buffers at enqueue).
Now the decoder's output is set aside, the old contents put back, and a
CoreTiming event writes the samples just before the thread wakes. Pending
writes are kept in savestates.
Adds audio/blocking/parked, recorded on a PSP: a blocking output that had
to wait returns before any of its buffer has played, so the game really
does depend on the decode's latency.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Only 4x4 with source and destination transposed alike, and a scale
outside the destination, compiled; most vmscl in games are transposed
or 3x3. The rest now multiply element by element, with the scale
copied first, and only a partly overlapping source still falls back.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ll and sc always went to the interpreter; a few games use them
thousands of times. They now load and store directly with fast memory,
keeping llBit in MIPSState. vcmp's NaN and inf-or-NaN tests only look
at s, so they no longer need vt to be the same register: EN and NN
compare s with itself.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Without the Zbb extension, clz, rotr, wsbh, wsbw, bitrev, min and max
went to the IR interpreter, and wsbh always did. They're now base ISA
sequences: a branchless binary search for clz, paired shifts for the
rotates and byte swaps, mask-and-shift steps for bitrev, and a
compare-and-branch for min and max. Checked against a model of the
instructions over random inputs, since nothing here runs RISC-V.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
They always went to the interpreter. Each channel is now a shift, mask
and shift in the existing integer ops, so every backend handles them.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
vh2f always went to the interpreter in the IR. A new FHalfToFloat op
converts the lower or upper half of a word, and the native backends
call vfpu_h2f for it like FSin. The legacy arm64 JIT now makes the same
call instead of computing the conversion inline; vh2f is rare, and the
call is much less code.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
All went to the interpreter. vh2f works in integers to match vfpu_h2f,
since FCVTL neither flushes subnormal halves nor keeps inf/NaN mantissa
bits unshifted. The color conversions are bitfield extracts and inserts.
vbfy, vcrs and vdet follow the IR frontend, and vmscl scales element
by element. Results go through scratch registers, so a destination that
overlaps a source is fine.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Vec4ClampToZero and Vec2ClampToZero only ever fed Vec4Pack31To8 and
Vec2Pack31To16, for vi2uc and vi2us. The packs now clamp negative lanes
to zero themselves, which saves an op and a vector temp, and lets x64
clamp with PACKUSWB's saturation after an arithmetic shift.
While at it, RISC-V compiles Vec2Unpack16To31, Vec2Pack31To16 and
Vec4Pack32To8, and LoongArch Vec2Unpack16To31, Vec2Pack31To16 and the
non-LSX Vec4Pack32To8, all of which went to the IR interpreter.
LoongArch's Vec2Pack32To16 and Vec2Unpack16To32 now take their scalar
path with LSX too, instead of falling back.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
vi2uc, vi2c, vi2us, vi2s, vuc2i, vc2i, vus2i and vs2i lowered to IR ops
that x64 left to the IR interpreter. They're now SSE2: shifts into
place, then PACKSSDW/PACKUSWB for the packs and self-unpacks for the
unpacks. An output overlapping its input still goes the slow way.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
FCvtWS always went to the IR interpreter. Both now convert in the
current rounding mode, which ApplyRoundingMode keeps at the game's, as
x64 does. RISC-V's FCVT already saturates and gives INT_MAX for NaN;
LoongArch patches NaN to INT_MAX like its FRound.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
They always went to the IR interpreter. Now a signed compare-and-branch
on the normalized sources picks which one to move.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
These went to the interpreter. vf2in/vf2iz/vf2iu/vf2id use FCVT, which
saturates like the PSP, and patch NaN to 0x7FFFFFFF like the IR
backend. vsgn keeps the sign bit on 1.0 and gives 0 below the smallest
normal. vsge and vslt select 1 or 0 on a compare whose condition is
false when unordered. EI and NI test |s| against infinity in integers.
All match the interpreter on special values, ties and denormals.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Stop, Switch UMD, savestate slots, Load/Save State (Cmd+L/Cmd+S) and state
files, shortcuts to the settings screens (plus Settings... with Cmd+, in the
app menu), Rendering Resolution, Texture Filtering, Enable Sound, Enable
Cheats and the website links.
Menu items now carry their action and how to tell whether they're checked
or enabled as blocks, checked each time a menu is shown, instead of tags
matched up in menuNeedsUpdate. The Graphics menu becomes Game Settings,
as on Windows.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Recent files always booted the first one (no item had its tag set), and
the list was never refreshed. Rebuild it when opened, with the path on each item.
- The Graphics and Debug menus were never updated when opened: they were
matched by title, looked up under different translation keys than they
were created with. Compare the menus themselves.
- The Fullscreen item toggled bFullScreen twice, doing nothing.
- The status counter items checked their tag instead of the flag.
- Remove Restart Graphics from the Debug menu (it shared tag 12 with Show
Debug Statistics).
- Drop @available checks for versions below the deployment target.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
These were always interpreted. They're now IR ops that the native
backends compile to calls to vfpu_exp2 and vfpu_log2, like FSin and
FAsin. vrexp2 is FNeg followed by FExp2, which is how vfpu_rexp2
computes it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
vsin, vcos, vnsin, vasin, vexp2, vlog2 and vrexp2 went to the
interpreter. They now take the same direct call as vrcp and friends.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With a deployment target of 10.13 (10.14 for arm64 libretro), every file
that includes a libc++ header warned "The selected platform is no longer
supported by libc++." Apple Silicon Macs start at 11 anyway.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The vertex decoder's C++ steps may fuse, like the vertex JITs do; with
contraction off everywhere, RISC-V's VertexJit test found the JIT and
the steps disagreeing on morphed float UVs. The flag now applies to
Core/MIPS only, in CMake and the libretro Makefile. ndk-build has no
per-file flags, so the legacy Android.mk goes back to the default.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A float made from a constant NaN bit pattern can come out quieted: MSVC
turned the 0x7F800001 that vcos, vexp2 and vlog2 return into 0x7FC00001,
which failed cpu/vfpu/exact on Windows. Each function now computes its
result's bits in integers and converts once at the end, like vrcp and
friends already did. Fixed-point results become floats by shifting,
which is exact since the VFPU keeps 22 significant bits. Identical to
the previous code over every 32-bit input.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
They only printed comparisons against libm and checked nothing. The VFPU
functions are exact now and have their own tests.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
They were added in a different order, so they rounded differently from
every other backend.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Clang and GCC fuse a * b + c into one rounding on CPUs that have fused
multiply-add, which on ARM64 made the interpreter's dot products, vavg,
vdet and vqmul round differently from x86-64. Fused operations now only
happen where the code asks for them. MSVC doesn't contract by default.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Each of the three is now a range check, a fast path and a few special
cases. The fast path reads its segment from a table whose constant term
already includes the result's exponent bits, so what remains is two
multiplies, some shifts and one exponent adjustment, all in integers.
Identical to the previous code over every 32-bit input.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
These went through the host's sqrt and division everywhere except the
interpreter's vrcp and vnrcp (vsqrt and vrsq there only behind
USE_VFPU_SQRT, now gone). They now always give the PSP's bits: the IR
gets FVSqrt (FSqrt stays the FPU's IEEE sqrt.s), and FRSqrt and FRecip,
which only the VFPU emits, become vfpu_rsqrt and vfpu_rcp; the IR
interpreter and the x64, arm64, RISC-V and LoongArch backends call them.
The old JITs call them directly, the ARM ones keeping the lanes in
callee-saved registers across the calls. cpu/vfpu/exact now passes on every core.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>