- Flush() on an empty queue now trims the state and CLUT rings, since its
callers flush because one is full and push right after.
- BinQueue::Full() uses >=, so an overshoot can't go unnoticed.
- IsExactSelfRender compares against the target the queued draws were
binned for, not gstate, which already has the next one during a flush.
- A depth test without depth writes marks the depth buffer as read.
- The DarkStalkers untextured sprite recomputes the binner state around it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
These own GPU objects, memory or refcounts in their destructors (or assert
there that they were torn down), so a copy would double-free. Nothing copies
them today; this keeps it that way. The manager base classes cover every
backend's subclass.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- The texture and fragment test caches dropped their GLRTexture objects on
DeviceLost without queueing them for deletion, leaking them on every
Android background/resume. The deleter already skips the GL calls when
the context is gone.
- GLRenderManager::ThreadEnd cleared unsubmitted init and render steps
without freeing the data they own. Run them through the dry run instead,
which now also frees stereo matrices and shader code.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Release the file reference when loading a level fails or finds nothing,
since only a loaded level takes ownership of it.
- Free the PNG image when the size changed since the header was read.
- Delete the VFS when a pack without an ini has no hash-named textures.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Put the anisotropy level in the sampler key, so changing it applies on
Vulkan and D3D11.
- Release CLUT textures at shutdown on GLES and D3D11.
- Fix the depth readback viewport, which squeezed the image whenever the
read rectangle was smaller than the fbo.
- Test the computed depth, not the unset result, in the equal-depth clear
check.
- Don't read back a CLUT from a framebuffer without an fbo.
- Tolerate null entries when releasing post-shader objects and CLUT
textures after a failed creation.
- ImGe: Don't crash on a framebuffer without an fbo, or on GetVFB under the
software renderer.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
DeviceLost saves the cache and clears everything, and the next save wrote
back only what was drawn since, so each Android background/resume cycle
shrank the cache.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
RecordNextFrame now refuses while a recording is active and during frame
dump playback (which asserted on the next replay). The callback handoff to
the CPU thread is locked, and a failed file open no longer crashes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- The block transfer overlap check passed the stride in pixels where bytes
are expected, so it only covered part of the rectangle.
- A selfrender/selfdepth flush in UpdateState dropped the current draw's
pending writes and reads, so later transfers didn't wait for it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- A list dropped for a bad pc or a GE error stayed RUNNING: its ID was never
freed and sceGeListSync on it never returned. Complete it like a finished
one.
- FlushImm switches to through mode and another vertex decoder, so flush the
queued draws first even when the immediate flags match.
- Clear leftover temporary GE breakpoints when setting or clearing the next
break, so a step that never got there doesn't trip later.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Flush before the queued draws would decode more than VERTEX_BUFFER_MAX
vertices. The batch was limited by index count, which doesn't bound a
sparse index range, and DecodeVerts silently stopped while DecodeInds
still emitted indices for the undecoded draws.
- Give TestBoundingBox its own scratch buffer. It used offsets in decoded_,
which can hold decoded vertices that aren't flushed yet.
- Read 32-bit indices the way the PSP does, ignoring the upper 16 bits.
IndexConverter and the fast bounding box test used all 32, so a game
setting them indexed far past the decoded vertices.
- D3D11: Flush in FinishDeferred like the other backends, since indices
are still read from PSP memory at flush time (#10095).
- Don't JIT new vertex decoders once the code space is full.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Wait on actionComplete instead of a bare condition variable wait, which
could miss the wakeup and hang until resume.
- Serialize requesters, so two debuggers can't overwrite each other's
action, and make SetCmdValue/FlushDrawing wait too.
- Give up and withdraw the request when stepping ends, instead of waiting
forever (this deadlocked game shutdown against the Win32 GE debugger).
- Run requests during CPU stepping, which already accepted them.
- Clear the stepping state on Core_Resume from GE stepping and on
GPU_Shutdown.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The OpenGL and Vulkan shader caches store raw shader IDs (and, for Vulkan,
pipeline keys) on disk. Add the rule to AGENTS.md and point to it from the
persisted types.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Skipping them is only a speed-up, but it changes how states get batched and
optimized, which changes the rendered output (NBA 2K13 and Virtua Tennis
frame dumps). Keep the old behaviour until that's understood.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A code space clear frees the functions the current state points to, but
the state was kept as long as the GE registers didn't change. Track the
clear generations and recompute, also when a compile during the state
computation clears the caches.
Also skip the binner flush for compiles on builds without the software
JIT, where Compile() does nothing.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- DDS files with mips stored level 0's file reference in level 1 too, so it
was freed twice.
- The VFS was deleted on config changes (and ini reloads) while load tasks
still used it. Now each cached texture waits for its task and releases its
file references through the old VFS first, and reloads afterwards.
- A failed ini reload turns replacement off instead of leaving it on without
a VFS.
- Release file references on purge.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The pipeline deletion callbacks block on in-flight compiles, which use the
shader module promises that the shaders' deletion callbacks free. Queueing
the shaders first freed the promises under a pending compile.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Delete TexCacheEntry objects dropped on rehash instead of leaking them.
- Don't leave a released null entry in cache_ when the framebuffer match
returns before the slot is refilled.
- Reset clutRenderAddress_ in Clear(), which releases the dynamic CLUT FBOs.
- Don't cache a null texture in drawPixelsCache_ when creation fails.
- Fix the reversed subtraction in the failed-FBO retry check.
- Remove the never-taken buffered-rendering early-out in UpdateRenderSize.
Taking it would leave existing VFBs without an fbo.
- Include smoothedDepal in the depal shader cache key, and print/parse the
debug IDs as 64-bit.
- Release depal pipelines through Draw2DPipeline::Release so the shader
source isn't leaked, and make that null-safe.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
GLSLtoSPV takes an optional SPIRVCache, keyed on a 32-bit hash of the
source, stage and variant, plus the source length. A changed shader
simply misses. thin3d's shaders and the other fixed ones use a global
cache in PSP/SYSTEM/CACHE/vulkan_spirv.cache, loaded on first use and
saved after graphics init, when a game's cache is saved, and at
shutdown; it's flushed once it reaches 32 entries, about twice what a
session compiles, so outdated ones don't pile up. Game shaders keep
theirs in the .vkshadercache, ahead of the shader IDs so that the
compiles on load find it (version 60), and only what the session used
is saved.
A cold glslang costs about 40ms before its first shader here, and
0.3-0.9ms per shader after that.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Neither is serialized, and both went stale on load. The ME busy time was
measured against the pre-load clock, so loading an earlier state made the
next SAS/codec job wait until the old time came around, freezing the game.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Some DoState code meant for after a load ran on every save:
- scePower reset the bus frequency a game set (and with a locked CPU
speed, applied the current setting to the clock).
- sceDisplay reset the lag sync baseline, and could schedule lag sync in
the measuring pass only, which failed the save.
- GPUState dirtied the texture, sceUmd notified the UI, and sceMpeg
dropped a pending ringbuffer fix-up for an old state.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The blit rates were measured with nothing else running. In a game, threads
waking up and SAS mixing on the Media Engine compete with the GE for main RAM:
Star Wars: Lethal Alliance's movie blit takes 8.65ms alone and 10.3ms in the
game. We don't model that load, so RAM texture fetches get a fixed 1.17x for a
typical one. With it, that game's long movie plays at 30 fps as on hardware.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Full-screen clears measured on a PSP (pspautotests gpu/timing/blittiming):
0.49ms on a 16-bit framebuffer whatever is cleared, 0.69ms on 8888, 1.02ms on
8888 with depth. Charging them may help games that spin hard on an empty
screen, but it's off (chargeClearTime) until tried on some. The video blit
cost moves into the same function, now EstimateFillCycles.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The blit cost remembered only the last buffer a decoder wrote into, forever.
Move the texture cache's video list (with its ageing out a few flips after
the last write) into GPUCommon, so the texture cache, the blit cost and
SoftGPU all share one. That also counts both of a double-buffered player's
frames, which exposed that a clear drawn with texturing still enabled was
being charged as a blit - skip clears and draws without texture coordinates.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Measured on a PSP (pspautotests gpu/timing/blittiming), a full-screen blit
from an unswizzled texture costs what the texture fetch costs: 16-bit formats
half of 32-bit, VRAM a fifth of RAM, and rectangles wider than ~128 texels
~7.5x as much as narrow strips, from texture cache thrashing. Framebuffer
format, filtering and blending don't matter. Ys I & II draws its movie as one
full-width sprite from a 565 texture in RAM, which takes 33ms - that, not the
decode, is what holds it to 30 fps.
All of the ME and GE costs speed up with the clock (2/3 as long at 333/166),
since the whole system runs from the one PLL.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Movie players like the one in Star Wars: Lethal Alliance present every decoded
frame after a single vblank wait, with no clock or timestamp check, so the
frame rate depends on decode, CSC, ATRAC decode and the GE blit adding up to
more than a vblank. We charged nearly nothing for any of them, so such movies
ran at 60 fps until the ringbuffer's slack ran out.
Costs measured on a PSP with a copy of that player (pspautotests
video/mpeg/playertiming), for a 480x272 frame:
- sceVideocodecDecode: 3.4ms (sceMpegAvcDecode 5.8ms less sceMpegAvcCsc 2.4ms)
- sceMpegBaseCscAvc: 2.4ms, was a flat 4ms
- sceAudiocodecDecode, ATRAC3+ only: 2.5ms per frame
- GE: 9.7ms for a through-mode rectangle blit from a decoded video frame,
charged by area, only for textures in the buffer a decoder last wrote.
GE time also now carries across stall address updates. Before, a list sent
in stalled chunks only had its last chunk's time counted, so sceGeDrawSync
returned 39us after a blit that takes 9.7ms. This affects every game that
builds its lists incrementally, so GE-timing-sensitive games need checking.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
videos_ only learns about the CSC output, so the texture we actually draw is
just an ordinary 512x512 8888 texture whose contents happen to be different
every frame: hash, miss, throw it in the secondary cache, rebuild, forever.
So track the copy. NotifyVideoCopy marks the destination as video when the
source is, and the copy funnels call it: sceDmacMemcpy, sceKernelMemcpy, and
the four replaced memcpy/memmove variants. It sits outside their "is either
side VRAM" gate, since a RAM-to-RAM copy of a frame is still a frame.
Being a video texture also gets it forced linear filtering and keeps it out of
texture upscaling, which is what you want for a video either way.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
A `video = true` in textures.ini opted a pack into replacing and dumping video
textures. Both halves are a bad deal. Dumping writes a file per decoded frame,
which fills a disk rather than producing anything a pack can use, and replacing
means a hash lookup on content that is different every frame and will never be
found twice.
It was also the only reason the texture cache still hashed video textures at
all, so it cost every game that has ever played a cutscene, not just the packs
that set it.
Video textures are now never replaced and never dumped, and skipHash is simply
isVideo. An existing ini keeping the key is harmless - unknown options are
ignored.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
NotifyWriteFormattedFromMemory appends to videos_ unconditionally. A game
blitting its decoded frame to the display buffer does that every displayed
frame while it waits for the next one to decode, so the same two or three
addresses come back over and over: Death Jr pushes 733 entries where there are
two distinct buffers, Tekken 6 around 53,000 where there are three. IsVideo()
walks that vector linearly on every texture. Refresh the matching entry instead
of appending a new one - Death Jr now holds 2 entries and Tekken 6 holds 18,
peak size 2 and 3.
The other half is the two TODOs that were already sitting there. A video
texture is new every frame by definition, so re-hashing it only confirms what
the VIDEO flag already said, and the secondary cache has nothing to offer a
frame that will never recur. Skip both, and with them the secondary lookup that
would otherwise key off a hash we no longer compute.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The arm64 vertex JIT had no morph support (the table entries were commented
out), so every morphed format fell back to the step functions. Add the same
morph steps x86 has: texcoords (plain and prescaled), normals, positions and
the four color formats.
Each step sums its frames the way its step function does as compiled: fused
where that's plain C++, which the compiler turns into fused multiply-adds,
and separate multiply and add where it's CrossSIMD. The unit test confirms
they match bit for bit. Morph formats decode about 2.5-3x faster than with
the steps.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Enables TestVertexJitMatchesSteps on both. riscv64 now passes all 216000
jitted formats and loongarch64 all 204000, against 28906 on arm64.
riscv64:
- Jit_PosFloat didn't clean NaN or infinity at all, it just copied the
three words. Clamp to +-FLT_MAX like the x86 JIT does.
- Jit_PosFloatThrough was missing the truncation of Z to an integer.
- The morph helpers started the sum from the first product rather than
from +0.0, which rounds differently and lets a -0 term through, and
rounded that first product towards zero where the steps round to
nearest. The rest of the sum stays fused, since the compiler contracts
the steps into fused multiply-adds.
- The texcoord prescale and 5551 color morph paths read morph weights
from tables that GetMorphValueUsage never asked to be filled in, so
they used whatever an earlier vertex type had left there.
- The non-Zbb bounds update compared the wrong way around, so through
mode texcoord bounds came out inverted.
loongarch64:
- Jit_PosFloat had a TODO to clean NaN and infinity, and didn't.
- Jit_PosFloatThrough was missing the same Z truncation.
- Jit_WriteMorphColor narrowed with the logical saturating shifts, so a
negative channel became a huge unsigned value and saturated to 255
instead of clamping to 0, and it rounded where the steps truncate. It
also read the packed color back sign-extended, so any alpha above 0x7F
compared as larger than 0xFF000000 and claimed full alpha.
- The three packed color morph formats are rewritten. The LSX versions
built each channel with a chain of inserts, shifts and shuffles that
didn't survive being run; 4444 also broadcast its scale from the mask
register. They now follow the steps channel by channel. Note the
accumulator has to be an LSX scratch register - F4-F7 alias V4-V7,
which hold the skin matrix for the whole vertex.
- The vertex bounds were loaded with a signed halfword load, so the
0xFFFF they start at became -1 and no texcoord was ever below it.
PrescaleUV now fuses on these two as well - both JITs fuse it, and the
compiler would have contracted the plain expression there anyway.
The test tolerates a small relative difference on decoded floats, scaled
by the morph count since each term rounds once. Every bug above was
orders of magnitude larger than that.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
From another pass over ge.prx against our code, each checked on a PSP with
gpu/ge/callbackstate except the last:
- A list's context is restored after its finish callback, which sees the
state the list left. We restored at the FINISH, before it. Now that nothing
runs until InterruptEnd(), that's where it happens.
- sceGeSaveContext/RestoreContext only fail while the GE is executing. It's
stopped during a finish callback and a SUSPEND signal callback, however
much is queued, so they work there. We said busy whenever a list existed.
- sceGeListDeQueue emptying the queue doesn't turn completed lists into
nothing, only sceGeDrawSync does. CheckDrawSync() is gone.
- The "break in progress" flag that makes sceGeContinue only requeue the
list is cleared by an interrupt that follows the break at once, so it's
only seen from a callback or with interrupts off. Ours lasted until the
next GE interrupt of any kind.
- sceKernelLoadExec restarts the GE driver, which begins by zeroing every
register and matrix. Reinitialize() now does too, so a program started that
way finds the same GE as one booted directly, rather than its launcher's.
Not testable on hardware: nothing after the restart can report back.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Reinitialize() wiped the 64 display lists but kept the queue of their ids.
Whatever the old executable still had queued came back as lists with no state
and a pc of 0, behind the first list of the new executable, where they
blocked everything. Crazy Taxi: Fare Wars is a launcher for its two games,
and stopped at a black screen that way. Fixes#19894.
This removes the workaround for it, which dropped such a list but returned
before currentList was cleared, and only worked as long as something else
happened to clear it later. A list with a bad pc is now dropped like one that
ran into an error, instead of sitting at the head of the queue for good.
Also narrows what sceGeBreak(1) throws away to interrupts that have actually
been raised, which is what gpu/ge/intrsuspend shows. The ones we haven't
raised yet are only late because we execute lists ahead of time: a game that
breaks right after its last list, and then waits for what the finish callback
signals, got that callback long ago on hardware.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Resetting the GE also gets rid of an interrupt that was raised but not taken
yet, so a list that reached its FINISH just before never gets its finish
callback. We delivered one anyway, for a list that no longer existed. If the
break comes from inside a GE callback, the interrupt being handled is kept,
since its handler still has to return.
Found by gpu/ge/intrsuspend, which also confirms from a thread, with
interrupts suspended, that nothing moves along the queue until the FINISH
interrupt has been taken.
Savestates: bump GPUCommon to 7. We didn't use to mark a PAUSE signal as
delivered, which sceGeContinue now goes by, so a state saved with a list
paused that way would load into a game that could never continue it. Fixed
up on load.
gpu/signals/handlercalls goes in as known failing: with an old SDK version, a
stall address set from inside a SUSPEND callback doesn't reach the GE, which
we can't express with just the one stall address per list. See docs/sceGe.md.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
On hardware the GE stops at every SIGNAL and FINISH, and it's the interrupt
that gets it going again: on the same list after a signal, on the next one
after a FINISH - once the finish callback has run, with the finished list
still at the head of the queue. We ran the next list right away and dropped
the finished one at once, so a finish callback saw an empty queue. A list
enqueued from there was started instead of queued, and then couldn't be
dequeued, which hung the new gpu/ge/queue2 test.
ProcessDLQueue() now runs nothing while the head of the queue has an
interrupt pending, and InterruptEnd() is what takes a finished list off the
queue. This also keeps a stall update from restarting a list that's stopped
at a signal before the handler has run. drawCompleteTicks is still set when
the last list reaches its FINISH, so a sceGeDrawSync in between doesn't wait.
Other things gpu/ge/queue2 and gpu/ge/breakwait showed, all from a real PSP:
- sceGeListEnQueue compares against the address a list was enqueued with
(or stopped at by sceGeBreak), mirrors included, not against its current pc.
We had that the wrong way around.
- The stack-in-use check only applies to lists that have started executing.
This is probably what IgnoreEnqueue was added for (Metal Gear Acid 2,
#10906). The flag stays until someone has checked the game without it.
- A PAUSE signal makes the list PAUSED at once, before the FINISH delivers it.
In between, sceGeContinue and sceGeBreak say BUSY, and updating the stall
address does nothing, so a list that stalls there is stuck.
- A completed list can't be dequeued, with or without a context.
- sceGeDrawSync(1) looked at currentList instead of the list it had found.
- sceGeBreak(1) doesn't wake anyone, and a late interrupt for a list it reset
no longer marks that list completed. Threads in sceGeDrawSync are woken
before the ones waiting for the last list.
Also fixes currentList being lost when loading a state where it's list 0,
and makes ge_pending_cb a plain std::list - nothing else touches it, and the
GPU thread it was shared with is long gone. Same savestate format.
See docs/sceGe.md.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
The test didn't compile: VERTS is captured by reference into testFormat, and
MSVC won't use a captured constexpr as an array bound. Make it static.
On arm64 it then failed on the UV prescale steps, by one ULP. The arm64 JIT and
the NEON handwritten decoders fuse the multiply-add, and the steps only match
that when the compiler contracts a * b + c - which clang does and MSVC doesn't,
in Debug or Release. Spell out which one happens instead of relying on it.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The two handwritten SIMD decoders are used with or without the JIT, and had
drifted from the step functions:
- The God of War one ignored g_DoubleTextureCoordinates, giving HD Remaster
games the wrong UVs, and passed NaN and infinite positions through. Those
now come out finite like everywhere else.
- The GTA one expanded 5551 colors wrongly on NEON (a left shift where the
SSE version shifts right).
- On ARM64, both now fuse the UV scale and offset, like the JIT and the steps
as the compiler builds them.
Both are back in the unit test, which now also feeds NaN and infinity to plain
float positions and checks they come out finite, and runs only on x86-64 and
arm64 for now.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Add a unit test that decodes every vertex format through both the step
functions and the JIT and requires identical output, side effects included.
Only skinning may differ by rounding, since arm64 accumulates the bone
matrices with fused multiply-adds.
What it found, and fixed:
- x86 morph colors rounded to nearest where the steps truncate, and applied
the scale before the weight, which rounds differently.
- x86 through-mode u16 UV bounds compared signed, so texcoords above 32767
scrambled the bounds for everyone running the x86 JIT.
- x86 morph sums could produce -0 where the steps produce +0.
- Step_NormalS16Morph scaled by 1/32768 twice, giving near-zero normals.
- Step_PosFloatThrough lost the truncation of Z to an integer that the JITs
and the vertex reader did before it moved into the decoder.
- SetVertexType never reset skinInDecode.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The vertex decoder JIT was enabled only when g_Config.iCpuCore was one of the
JIT cores, which tied an unrelated GPU path to the CPU setting. Headless leaned
on that by forcing iCpuCore to the interpreter after ApplyToConfig(), to keep
the decoder JIT off in the tests.
Add CoreParameter::bUseVertexDecoderJit instead. Standalone and libretro set it
wherever the host can JIT, so IR interpreter users on such hosts now get the
decoder JIT too. Headless sets it false, which keeps the test results as they
were and lets --cpu go through ApplyToConfig() like every other option.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Our psmf and psmfPlayer HLE plays video by calling our sceMpeg HLE, so it has
nothing to talk to when the real mpeg.prx is running. The flag now gives way
when sceMpeg is LLE.
Moved below the force-enable and unavailable masks so it tests what sceMpeg
actually ended up as rather than what was asked for.
Also remove a bad assert.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>