On Vulkan, reading back the display framebuffer after EndDrawFrame hit
the insideFrame_ assert in CopyFramebufferToMemory, so any run that
timed out with --screenshot-save crashed in debug builds.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The PSP blends save icons over a black background, so their transparent
parts come out black. 0a5fa27957 turned blending off instead, which is
wrong for icons that rely on it. Revert that, and draw a black rectangle
under each icon. Fixes#22280.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With xxhash 0.8.4, XXH3 is faster than StableQuickTexHash on ARM64
(about 37 vs 24 GB/s on a Snapdragon X), and its scalar path, which
RISC-V builds get, is about as fast as the quick hash's. It doesn't
collide the way the quick hash does (#8249).
Texture replacement still uses its own hash setting, so texture packs
are unaffected.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Keeps our two local changes (the ppsspp_config.h include and the ARM32
prefetch in the XXH32 loop). Hash values are unchanged.
XXH3 is much faster on ARM64 now: about 37 GB/s against 13.5 with v0.8.1
on a Snapdragon X, where our quick texture hash does 24.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Homebrew written for a PSP-2000+ under custom firmware can use the top
32MB of RAM directly without setting MEMSIZE, which leaves the user
partition at its normal size. NJEMU's slim builds do this (#8925).
Previously we didn't map that memory at all for PBPs; MEMSIZE=1 isn't a
workaround either, since it grows the partition and the heap and stacks
land where the program writes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
NJEMU's SystemButtons.prx kernel plugin imports it when the firmware
reports 3.71 or later, and polls it every frame to read HOME/volume.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The net tests need WLAN on, but that also made every game run log into a
real adhoc server on the internet, so results depended on the network.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
sceNetAdhocctlInit waits until the friend finder has logged into the adhoc
server, polling in emulated time but giving up only after 5s of wall time.
The friend finder makes one attempt per login request, and it too waited
the full 5s on a connection that had already been refused, because it
only looked for success. Now it checks the socket error and gives up at
once, and records the failure so the wait ends with it.
In headless, which runs far ahead of real time, Gods Eater Burst sat in
this wait for its whole run.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Start from the quality that fit the previous frame instead of the top, so
most frames encode once. Step down while a frame is too big, and step back
up when one comes out under half the limit. Windows and the recompression
fallback share the logic.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Android, iOS, macOS and Linux encode camera frames at a fixed quality, so a
detailed frame can exceed the game's framesize just as on Windows before.
pushCameraImage now decodes such a frame and re-encodes it at lower quality
until it fits.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The friend finder thread set friendFinderRunning itself, after a DNS
lookup of the adhoc server. A shutdown in that window cleared the flag
first; the thread then set it again, looped forever, and the join in
NetAdhocctl_Term() never returned. Gods Eater Burst hit this in about one
headless run in six. The flag is now set before the thread is created,
and a finished thread is joined before a new one replaces it, which
would otherwise call std::terminate.
The built-in adhoc server thread had the same race, behind its check for
an existing server, and gets the same fix.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The PSP camera keeps JPEG frames within the framesize from the video setup.
Go!Edit stores frames in 15KB slots and only takes frames that fit, so our
uncompressed-quality webcam frames were dropped and one frame got repeated
for the whole clip. Lower the JPEG quality until a frame fits.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With a host microphone present, a blocking read waited until the host had
delivered all the data. If it never did, the thread waited forever: Go!Edit's
sound thread stalled that way and its video recording never advanced. The
PSP mic streams in real time, so wake at the scheduled time and fill what
the host didn't deliver with silence.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
:screenshot saves gpu.buffer.screenshot as a PNG without dumping the data
URI into the output. The docs claimed nested parameters need a raw JSON
line, which gets no ticket; a single-quoted JSON value in key=value form
works and keeps it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
It returned at once. Go!Edit reads frames in a loop on a high-priority
thread (bhCameraGetJpeg) that only yields to its own priority level, so
after the "Loading complete" dialog it spun and starved the rest of the
game. Return at the camera's next frame tick, at the rate from the setup
params.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Only the rounding mode, flags, enables, cause, FCC and FS bits can be
written (0x0181FFFF, pspautotests cpu/fpu/fcr), as the interpreter, IR
and x86 already have it. Both ARM JITs stored the whole value.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware sceVaudioChReserve takes ~260us once it gets past the busy
check, succeed or fail, and worse threads can run meanwhile. Releasing
takes ~25us and doesn't wait. We returned at once, which is what
audio/sceaudio/reserve's [r] markers showed; it now passes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A delay's deadline is now + usec, and the clock is read again when the
alarm is set. If the deadline has passed by then, the call returns 0 at
once without giving up the CPU. On hardware that makes
sceKernelDelayThread(0) return at once about 60% of the time. On a thread's
first wait after it starts, a delay of 1 does so about two times in three
as well (pspautotests threads/scheduling/delayzero). We always waited at
least 210us.
The choice is pseudo-random off the tick count, not the tick phase, since
our cycle counts are regular enough for a polling loop to lock into never
yielding. Threads remember whether they've waited since starting (Thread
savestate section version 6).
Also moves threads/vpl/create into the passing tests: re-recorded on 6.61,
it agrees with what we do for partitions 8 and 9.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With dispatch suspended, IO fails in the driver when it tries to wait. The
memory stick driver returns SCE_KERNEL_ERROR_CAN_NOT_WAIT, while usbhostfs,
which serves host0: under PSPLink, returns -1. host0: is mostly what
homebrew developers run from, so it now does the same.
Also from threads/scheduling/dispatch, which now passes:
- sceIoRead reports an async operation still in progress before failing
on suspended dispatch.
- A write to stdout or stderr doesn't give up the CPU.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware a load costs an open and a read of the file, plus about 1ms
and 30us per KB of loader work (pspautotests threads/scheduling/callcosts),
with the caller waiting throughout. It now charges sceIoOpen's and
sceIoRead's estimates for the file plus that, instead of a flat 500us.
Also notes why sceKernelLoadModuleByID fails from a game's own fd on ms0:
or host0: on hardware, which we don't emulate.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
These own GPU objects, memory or refcounts in their destructors (or assert
there that they were torn down), so a copy would double-free. Nothing copies
them today; this keeps it that way. The manager base classes cover every
backend's subclass.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- The texture and fragment test caches dropped their GLRTexture objects on
DeviceLost without queueing them for deletion, leaking them on every
Android background/resume. The deleter already skips the GL calls when
the context is gone.
- GLRenderManager::ThreadEnd cleared unsubmitted init and render steps
without freeing the data they own. Run them through the dry run instead,
which now also frees stereo matrices and shader code.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Release the file reference when loading a level fails or finds nothing,
since only a loaded level takes ownership of it.
- Free the PNG image when the size changed since the header was read.
- Delete the VFS when a pack without an ini has no hash-named textures.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Put the anisotropy level in the sampler key, so changing it applies on
Vulkan and D3D11.
- Release CLUT textures at shutdown on GLES and D3D11.
- Fix the depth readback viewport, which squeezed the image whenever the
read rectangle was smaller than the fbo.
- Test the computed depth, not the unset result, in the equal-depth clear
check.
- Don't read back a CLUT from a framebuffer without an fbo.
- Tolerate null entries when releasing post-shader objects and CLUT
textures after a failed creation.
- ImGe: Don't crash on a framebuffer without an fbo, or on GetVFB under the
software renderer.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
DeviceLost saves the cache and clears everything, and the next save wrote
back only what was drawn since, so each Android background/resume cycle
shrank the cache.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
RecordNextFrame now refuses while a recording is active and during frame
dump playback (which asserted on the next replay). The callback handoff to
the CPU thread is locked, and a failed file open no longer crashes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- The block transfer overlap check passed the stride in pixels where bytes
are expected, so it only covered part of the rectangle.
- A selfrender/selfdepth flush in UpdateState dropped the current draw's
pending writes and reads, so later transfers didn't wait for it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- A list dropped for a bad pc or a GE error stayed RUNNING: its ID was never
freed and sceGeListSync on it never returned. Complete it like a finished
one.
- FlushImm switches to through mode and another vertex decoder, so flush the
queued draws first even when the immediate flags match.
- Clear leftover temporary GE breakpoints when setting or clearing the next
break, so a step that never got there doesn't trip later.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Flush before the queued draws would decode more than VERTEX_BUFFER_MAX
vertices. The batch was limited by index count, which doesn't bound a
sparse index range, and DecodeVerts silently stopped while DecodeInds
still emitted indices for the undecoded draws.
- Give TestBoundingBox its own scratch buffer. It used offsets in decoded_,
which can hold decoded vertices that aren't flushed yet.
- Read 32-bit indices the way the PSP does, ignoring the upper 16 bits.
IndexConverter and the fast bounding box test used all 32, so a game
setting them indexed far past the decoded vertices.
- D3D11: Flush in FinishDeferred like the other backends, since indices
are still read from PSP memory at flush time (#10095).
- Don't JIT new vertex decoders once the code space is full.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Setting data decodes and throws away the frames before the first sample,
so on hardware it costs a decoder setup plus that decode, with the caller
waiting: ~900us for mono Atrac3 and ~3.5ms for stereo Atrac3+, whatever the
buffer size (pspautotests threads/scheduling/callcosts). It was charged
100us.
Decoding a frame now costs what the same frame costs through
sceAudiocodecDecode, through the shared ME queue, instead of a flat
2300us. That's about the same for stereo Atrac3+ and less for Atrac3
(685us mono, ~1100us stereo). The first Atrac3+ frames after setup still
come out ~500us short.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
From pspautotests threads/scheduling/syscallkinds:
- sceKernelUnloadModule takes ~400us that better threads can preempt
and worse ones don't get in on, so it uses the busy delay rather than
a wait.
- A successful sceKernelVolatileMemTryLock takes under 100us. It ate
500000 cycles as a hack for Crash Tag Team Racing, which has since
moved to (and no longer needs) the DrawSyncEatCycles compat flag.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware the kernel fills a new thread's stack with interrupts on, so
a better thread that wakes during it runs before the call returns, the
time it takes doesn't count towards the call, and worse threads get
nothing (pspautotests threads/scheduling/preemptsyscall). PPSSPP ate the
whole cost at once and only rescheduled at the end.
__KernelBusyDelayResult() models such a syscall: the caller waits, an
idle thread stands in for it while nothing better wants the CPU, and its
remaining cycles only count down while that's the case. When done it
goes back ahead of threads of its own priority, having never given up
the CPU. sceKernelCreateThread uses it for the stack fill, unless a
thread event handler is about to run.
Booting 75 games against master shows no difference.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Wait on actionComplete instead of a bare condition variable wait, which
could miss the wakeup and hang until resume.
- Serialize requesters, so two debuggers can't overwrite each other's
action, and make SetCmdValue/FlushDrawing wait too.
- Give up and withdraw the request when stepping ends, instead of waiting
forever (this deadlocked game shutdown against the Win32 GE debugger).
- Run requests during CPU stepping, which already accepted them.
- Clear the stepping state on Core_Resume from GE stepping and on
GPU_Shutdown.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The OpenGL and Vulkan shader caches store raw shader IDs (and, for Vulkan,
pipeline keys) on disk. Add the rule to AGENTS.md and point to it from the
persisted types.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Skipping them is only a speed-up, but it changes how states get batched and
optimized, which changes the rendered output (NBA 2K13 and Virtua Tennis
frame dumps). Keep the old behaviour until that's understood.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Validating after every churn step was quadratic in the block count. Check
every 64 steps and at the end, and do fewer iterations.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A code space clear frees the functions the current state points to, but
the state was kept as long as the GE registers didn't change. Track the
clear generations and recompute, also when a compile during the state
computation clears the caches.
Also skip the binner flush for compiles on builds without the software
JIT, where Compile() does nothing.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>