688 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 759b494b6a GPU: Check post shaders once per host frame, after resizes
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:16:41 -06:00
Henrik RydgårdandClaude Opus 5.5 507bb0801b GE: Complete lists dropped on error, flush before immediate draws
- A list dropped for a bad pc or a GE error stayed RUNNING: its ID was never
  freed and sceGeListSync on it never returned. Complete it like a finished
  one.
- FlushImm switches to through mode and another vertex decoder, so flush the
  queued draws first even when the immediate flags match.
- Clear leftover temporary GE breakpoints when setting or clearing the next
  break, so a step that never got there doesn't trip later.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 40a70b004c Savestate: Reset video frame tracking and ME busy time on load
Neither is serialized, and both went stale on load. The ME busy time was
measured against the pre-load clock, so loading an earlier state made the
next SAS/codec job wait until the old time came around, freezing the game.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 09:36:48 -06:00
Henrik RydgårdandClaude Opus 5.5 4bfaab058b GPU: Reject savestates with display list ids out of range
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 09:34:05 -06:00
Henrik RydgårdandClaude Opus 5.5 ad25588651 GPU: Allow for main RAM contention in the movie blit cost
The blit rates were measured with nothing else running. In a game, threads
waking up and SAS mixing on the Media Engine compete with the GE for main RAM:
Star Wars: Lethal Alliance's movie blit takes 8.65ms alone and 10.3ms in the
game. We don't model that load, so RAM texture fetches get a fixed 1.17x for a
typical one. With it, that game's long movie plays at 30 fps as on hardware.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 17:37:53 -06:00
Henrik RydgårdandClaude Opus 5.5 8f66065f53 GPU: Add a clear cost, disabled for now
Full-screen clears measured on a PSP (pspautotests gpu/timing/blittiming):
0.49ms on a 16-bit framebuffer whatever is cleared, 0.69ms on 8888, 1.02ms on
8888 with depth. Charging them may help games that spin hard on an empty
screen, but it's off (chargeClearTime) until tried on some. The video blit
cost moves into the same function, now EstimateFillCycles.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 17:18:42 -06:00
Henrik RydgårdandClaude Opus 5.5 242ef0a984 GPU: Track video frames in GPUCommon, so they expire for the blit cost too
The blit cost remembered only the last buffer a decoder wrote into, forever.
Move the texture cache's video list (with its ageing out a few flips after
the last write) into GPUCommon, so the texture cache, the blit cost and
SoftGPU all share one. That also counts both of a double-buffered player's
frames, which exposed that a clear drawn with texturing still enabled was
being charged as a blit - skip clears and draws without texture coordinates.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 17:00:11 -06:00
Henrik RydgårdandClaude Opus 5.5 7a675b42d5 Model the movie blit by texture format, location and width; scale costs with the clock
Measured on a PSP (pspautotests gpu/timing/blittiming), a full-screen blit
from an unswizzled texture costs what the texture fetch costs: 16-bit formats
half of 32-bit, VRAM a fifth of RAM, and rectangles wider than ~128 texels
~7.5x as much as narrow strips, from texture cache thrashing. Framebuffer
format, filtering and blending don't matter. Ys I & II draws its movie as one
full-width sprite from a 565 texture in RAM, which takes 33ms - that, not the
decode, is what holds it to 30 fps.

All of the ME and GE costs speed up with the clock (2/3 as long at 333/166),
since the whole system runs from the one PLL.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 16:40:11 -06:00
Henrik RydgårdandClaude Opus 5.5 a0b812bc31 Charge hardware-measured time for movie decode, colour conversion and blit
Movie players like the one in Star Wars: Lethal Alliance present every decoded
frame after a single vblank wait, with no clock or timestamp check, so the
frame rate depends on decode, CSC, ATRAC decode and the GE blit adding up to
more than a vblank. We charged nearly nothing for any of them, so such movies
ran at 60 fps until the ringbuffer's slack ran out.

Costs measured on a PSP with a copy of that player (pspautotests
video/mpeg/playertiming), for a 480x272 frame:

- sceVideocodecDecode: 3.4ms (sceMpegAvcDecode 5.8ms less sceMpegAvcCsc 2.4ms)
- sceMpegBaseCscAvc: 2.4ms, was a flat 4ms
- sceAudiocodecDecode, ATRAC3+ only: 2.5ms per frame
- GE: 9.7ms for a through-mode rectangle blit from a decoded video frame,
  charged by area, only for textures in the buffer a decoder last wrote.

GE time also now carries across stall address updates. Before, a list sent
in stalled chunks only had its last chunk's time counted, so sceGeDrawSync
returned 39us after a blit that takes 9.7ms. This affects every game that
builds its lists incrementally, so GE-timing-sensitive games need checking.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 15:23:08 -06:00
Henrik RydgårdandClaude Opus 5 4aef060293 Carry "this is video" across block copies
videos_ only learns about the CSC output, so the texture we actually draw is
just an ordinary 512x512 8888 texture whose contents happen to be different
every frame: hash, miss, throw it in the secondary cache, rebuild, forever.

So track the copy. NotifyVideoCopy marks the destination as video when the
source is, and the copy funnels call it: sceDmacMemcpy, sceKernelMemcpy, and
the four replaced memcpy/memmove variants. It sits outside their "is either
side VRAM" gate, since a RAM-to-RAM copy of a frame is still a frame.

Being a video texture also gets it forced linear filtering and keeps it out of
texture upscaling, which is what you want for a video either way.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-22 16:10:28 -06:00
sum2012 f5b6b45b5b Remove IgnoreEnqueue hack 2026-09-22 20:19:57 +08:00
Henrik RydgårdandClaude Fable 5.1 a36d09daac sceGe: what callbacks see, and a full GE reset on sceKernelLoadExec
From another pass over ge.prx against our code, each checked on a PSP with
gpu/ge/callbackstate except the last:

- A list's context is restored after its finish callback, which sees the
  state the list left. We restored at the FINISH, before it. Now that nothing
  runs until InterruptEnd(), that's where it happens.
- sceGeSaveContext/RestoreContext only fail while the GE is executing. It's
  stopped during a finish callback and a SUSPEND signal callback, however
  much is queued, so they work there. We said busy whenever a list existed.
- sceGeListDeQueue emptying the queue doesn't turn completed lists into
  nothing, only sceGeDrawSync does. CheckDrawSync() is gone.
- The "break in progress" flag that makes sceGeContinue only requeue the
  list is cleared by an interrupt that follows the break at once, so it's
  only seen from a callback or with interrupts off. Ours lasted until the
  next GE interrupt of any kind.
- sceKernelLoadExec restarts the GE driver, which begins by zeroing every
  register and matrix. Reinitialize() now does too, so a program started that
  way finds the same GE as one booted directly, rather than its launcher's.
  Not testable on hardware: nothing after the restart can report back.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-21 13:21:03 -06:00
Henrik RydgårdandClaude Fable 5.1 cc41a23256 GPU: empty the display list queue on sceKernelLoadExec, fixes Crazy Taxi
Reinitialize() wiped the 64 display lists but kept the queue of their ids.
Whatever the old executable still had queued came back as lists with no state
and a pc of 0, behind the first list of the new executable, where they
blocked everything. Crazy Taxi: Fare Wars is a launcher for its two games,
and stopped at a black screen that way. Fixes #19894.

This removes the workaround for it, which dropped such a list but returned
before currentList was cleared, and only worked as long as something else
happened to clear it later. A list with a bad pc is now dropped like one that
ran into an error, instead of sitting at the head of the queue for good.

Also narrows what sceGeBreak(1) throws away to interrupts that have actually
been raised, which is what gpu/ge/intrsuspend shows. The ones we haven't
raised yet are only late because we execute lists ahead of time: a game that
breaks right after its last list, and then waits for what the finish callback
signals, got that callback long ago on hardware.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-21 13:10:02 -06:00
Henrik RydgårdandClaude Fable 5.1 91c9c4d14e sceGe: sceGeBreak(1) takes pending interrupts with it
Resetting the GE also gets rid of an interrupt that was raised but not taken
yet, so a list that reached its FINISH just before never gets its finish
callback. We delivered one anyway, for a list that no longer existed. If the
break comes from inside a GE callback, the interrupt being handled is kept,
since its handler still has to return.

Found by gpu/ge/intrsuspend, which also confirms from a thread, with
interrupts suspended, that nothing moves along the queue until the FINISH
interrupt has been taken.

Savestates: bump GPUCommon to 7. We didn't use to mark a PAUSE signal as
delivered, which sceGeContinue now goes by, so a state saved with a list
paused that way would load into a game that could never continue it. Fixed
up on load.

gpu/signals/handlercalls goes in as known failing: with an old SDK version, a
stall address set from inside a SUSPEND callback doesn't reach the GE, which
we can't express with just the one stall address per list. See docs/sceGe.md.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-21 12:48:03 -06:00
Henrik RydgårdandClaude Fable 5.1 8a23e633a1 sceGe: keep a finished list on the queue until its interrupt is done
On hardware the GE stops at every SIGNAL and FINISH, and it's the interrupt
that gets it going again: on the same list after a signal, on the next one
after a FINISH - once the finish callback has run, with the finished list
still at the head of the queue. We ran the next list right away and dropped
the finished one at once, so a finish callback saw an empty queue. A list
enqueued from there was started instead of queued, and then couldn't be
dequeued, which hung the new gpu/ge/queue2 test.

ProcessDLQueue() now runs nothing while the head of the queue has an
interrupt pending, and InterruptEnd() is what takes a finished list off the
queue. This also keeps a stall update from restarting a list that's stopped
at a signal before the handler has run. drawCompleteTicks is still set when
the last list reaches its FINISH, so a sceGeDrawSync in between doesn't wait.

Other things gpu/ge/queue2 and gpu/ge/breakwait showed, all from a real PSP:

- sceGeListEnQueue compares against the address a list was enqueued with
  (or stopped at by sceGeBreak), mirrors included, not against its current pc.
  We had that the wrong way around.
- The stack-in-use check only applies to lists that have started executing.
  This is probably what IgnoreEnqueue was added for (Metal Gear Acid 2,
  #10906). The flag stays until someone has checked the game without it.
- A PAUSE signal makes the list PAUSED at once, before the FINISH delivers it.
  In between, sceGeContinue and sceGeBreak say BUSY, and updating the stall
  address does nothing, so a list that stalls there is stuck.
- A completed list can't be dequeued, with or without a context.
- sceGeDrawSync(1) looked at currentList instead of the list it had found.
- sceGeBreak(1) doesn't wake anyone, and a late interrupt for a list it reset
  no longer marks that list completed. Threads in sceGeDrawSync are woken
  before the ones waiting for the last list.

Also fixes currentList being lost when loading a state where it's list 0,
and makes ge_pending_cb a plain std::list - nothing else touches it, and the
GPU thread it was shared with is long gone. Same savestate format.

See docs/sceGe.md.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-21 12:21:19 -06:00
Henrik Rydgård 2096179dce Drive-by code cleanup 2026-08-15 18:31:20 +02:00
Henrik Rydgård c5e4d0d90d Rename the get-memory-pointer functions to make it clear where CPU exceptions can happen. 2026-08-12 14:06:16 +02:00
Henrik Rydgård e9a3449ede More MIPSState * plumbing (manual) 2026-08-12 14:02:19 +02:00
Henrik Rydgård 4f8859d0ec Centralize a disassembly utility function between the debuggers 2026-08-11 09:08:52 +02:00
Henrik Rydgård 2be4d995f2 More Read_U32 cleanup 2026-08-10 11:23:23 +02:00
Henrik Rydgård 3eb056ad86 Move Common/GraphicsContext.h to Common/GPU/GraphicsContext.h 2026-07-26 13:58:17 +02:00
Henrik Rydgård cd40e2d4f0 Delete more code related to hardware skinning 2026-07-14 17:17:40 +02:00
Henrik Rydgård 61e1ef8f7a Remove the "Software skinning" option. Now always on. 2026-07-14 17:02:57 +02:00
Henrik Rydgård 18a9de3eca Move AdvanceVerts to GPUStateCache 2026-07-14 15:28:21 +02:00
Henrik Rydgård 6d2948a09b Remove the cached UVScale in gstate_c. Conversion is cheap enough to do directly from gstate, no point in caching. 2026-07-03 13:17:35 +02:00
Henrik Rydgård 1b2f87bf1a Rename GPUgstate to GEState 2026-07-02 20:34:09 +02:00
Henrik Rydgård 79a81186a6 Add some new stats (enqueues, stall address updates) 2026-06-09 23:49:54 +02:00
Henrik Rydgård d0f4e5c2d1 Visualize bounding boxes in ImGe debugger by drawing their corners 2026-06-03 13:46:30 +02:00
Henrik Rydgård ab108f7f0c BBOX: Avoid z clipping issues by culling in clip space. 2026-06-03 10:05:21 +02:00
Henrik Rydgård bb573e6e0c Well, it builds, but doesn't work yet. 2026-06-02 12:33:53 +02:00
Henrik Rydgård 2a5c2fa477 Check if the viewport transform matches the clip space. If so we can skip the near clip plane. 2026-05-30 19:07:59 +02:00
Henrik Rydgård 189dde8115 Cache the cull matrix. Minor bbox optimization. 2026-05-30 19:07:59 +02:00
Henrik Rydgård ae98055cec Delete a lot of legacy depth and viewport code, fix OpenGL 2026-05-30 19:07:58 +02:00
Henrik Rydgård f60e27a9b7 Just some refactoring of the GPUStatistics struct, and more use of StringWriter 2026-05-29 14:40:31 +02:00
Henrik Rydgård 5513fcb223 Merge pull request #21705 from GermanAizek/constexpr-cpp17
GPU: modernize use C++17 constexpr for precalculate compilation
2026-05-27 12:21:39 +02:00
Henrik Rydgård 5cfbf50111 Keep updated products of view*proj and world*view*proj matrices. Use to simplify bbox culling. 2026-05-21 19:18:32 +02:00
Herman Semenoff 400d136f51 GPU: modernize use C++17 constexpr for precalculate compilation 2026-05-19 21:57:45 +03:00
Herman Semenoff 5e2f51af8e gpu: elf: using reserve() for optimize inserts in for loop
From #21611
2026-04-28 10:53:32 +02:00
Henrik Rydgård d9ba20ce99 Add sanity checks in GPU::PerformMemoryCopy and GPU::PerformMemorySet
Part of #21055
2026-01-01 14:43:14 +01:00
Henrik Rydgård 5f7a937466 Rename ValidSize to ClampValidSizeAt 2025-12-30 20:31:07 +01:00
sum2012 736ee4890b Fix ReapplyGfxState for GE_CMD_LOADCLUT
Fix pink issue in dialog
Fix #21058
2025-12-07 16:24:43 +08:00
Henrik Rydgård 67010ff2af Split the display layout config between landscape and portrait orientations 2025-11-05 12:49:51 +01:00
Henrik Rydgård 8fb132b5a8 Inline a function that was only used in one place. 2025-11-04 13:35:45 +01:00
Henrik Rydgård e745ab16e4 Merge branch 'master' into IgnoreEnqueue 2025-08-15 11:30:10 +02:00
Henrik Rydgård 3e0fea509d Improve sanity checks for vertex ranges in GPU 2025-05-14 09:39:14 +02:00
Henrik Rydgård c1ae455ce8 Fix ImDebugger cleanup on exit 2025-04-07 13:03:33 +02:00
Henrik Rydgård ebfc467d5d Start removing bad coreState checks 2025-04-05 09:18:56 +02:00
Henrik Rydgård 2bfe327dbd Expose PSPThread in the same manner 2025-03-31 10:24:03 +02:00
Henrik Rydgård 0f840e6240 Move JPEG error codes to the big enum, some include cleanup 2025-03-21 20:44:46 +01:00
Henrik Rydgård 1f5cfe82ed Fix issue with hleLogDebugOrError where the return value argument got repeated.
Not good when the argument is a function call..
2025-03-05 11:24:44 +01:00