9978 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 7a675b42d5 Model the movie blit by texture format, location and width; scale costs with the clock
Measured on a PSP (pspautotests gpu/timing/blittiming), a full-screen blit
from an unswizzled texture costs what the texture fetch costs: 16-bit formats
half of 32-bit, VRAM a fifth of RAM, and rectangles wider than ~128 texels
~7.5x as much as narrow strips, from texture cache thrashing. Framebuffer
format, filtering and blending don't matter. Ys I & II draws its movie as one
full-width sprite from a 565 texture in RAM, which takes 33ms - that, not the
decode, is what holds it to 30 fps.

All of the ME and GE costs speed up with the clock (2/3 as long at 333/166),
since the whole system runs from the one PLL.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 16:40:11 -06:00
Henrik RydgårdandClaude Opus 5.5 a0b812bc31 Charge hardware-measured time for movie decode, colour conversion and blit
Movie players like the one in Star Wars: Lethal Alliance present every decoded
frame after a single vblank wait, with no clock or timestamp check, so the
frame rate depends on decode, CSC, ATRAC decode and the GE blit adding up to
more than a vblank. We charged nearly nothing for any of them, so such movies
ran at 60 fps until the ringbuffer's slack ran out.

Costs measured on a PSP with a copy of that player (pspautotests
video/mpeg/playertiming), for a 480x272 frame:

- sceVideocodecDecode: 3.4ms (sceMpegAvcDecode 5.8ms less sceMpegAvcCsc 2.4ms)
- sceMpegBaseCscAvc: 2.4ms, was a flat 4ms
- sceAudiocodecDecode, ATRAC3+ only: 2.5ms per frame
- GE: 9.7ms for a through-mode rectangle blit from a decoded video frame,
  charged by area, only for textures in the buffer a decoder last wrote.

GE time also now carries across stall address updates. Before, a list sent
in stalled chunks only had its last chunk's time counted, so sceGeDrawSync
returned 39us after a blit that takes 9.7ms. This affects every game that
builds its lists incrementally, so GE-timing-sensitive games need checking.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 15:23:08 -06:00
Ygor Dreyer a422b818b7 Hash full swizzled CLUT4 glyph atlases 2026-09-24 17:10:10 -03:00
sum2012 a8b85eec8b Add ForceEnableGPUReadback compat
To solve gpu readback issue
2026-09-23 22:06:28 +08:00
Henrik RydgårdandClaude Opus 5 4aef060293 Carry "this is video" across block copies
videos_ only learns about the CSC output, so the texture we actually draw is
just an ordinary 512x512 8888 texture whose contents happen to be different
every frame: hash, miss, throw it in the secondary cache, rebuild, forever.

So track the copy. NotifyVideoCopy marks the destination as video when the
source is, and the copy funnels call it: sceDmacMemcpy, sceKernelMemcpy, and
the four replaced memcpy/memmove variants. It sits outside their "is either
side VRAM" gate, since a RAM-to-RAM copy of a frame is still a frame.

Being a video texture also gets it forced linear filtering and keeps it out of
texture upscaling, which is what you want for a video either way.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-22 16:10:28 -06:00
Henrik RydgårdandClaude Opus 5 aa7b1de28c Drop the texture replacer's "video" option
A `video = true` in textures.ini opted a pack into replacing and dumping video
textures. Both halves are a bad deal. Dumping writes a file per decoded frame,
which fills a disk rather than producing anything a pack can use, and replacing
means a hash lookup on content that is different every frame and will never be
found twice.

It was also the only reason the texture cache still hashed video textures at
all, so it cost every game that has ever played a cutscene, not just the packs
that set it.

Video textures are now never replaced and never dumped, and skipHash is simply
isVideo. An existing ini keeping the key is harmless - unknown options are
ignored.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-22 16:10:16 -06:00
Henrik RydgårdandClaude Opus 5 14bfd332dc Stop stacking duplicate video entries, and stop hashing video textures
NotifyWriteFormattedFromMemory appends to videos_ unconditionally. A game
blitting its decoded frame to the display buffer does that every displayed
frame while it waits for the next one to decode, so the same two or three
addresses come back over and over: Death Jr pushes 733 entries where there are
two distinct buffers, Tekken 6 around 53,000 where there are three. IsVideo()
walks that vector linearly on every texture. Refresh the matching entry instead
of appending a new one - Death Jr now holds 2 entries and Tekken 6 holds 18,
peak size 2 and 3.

The other half is the two TODOs that were already sitting there. A video
texture is new every frame by definition, so re-hashing it only confirms what
the VIDEO flag already said, and the secondary cache has nothing to offer a
frame that will never recur. Skip both, and with them the secondary lookup that
would otherwise key off a hash we no longer compute.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-22 16:10:00 -06:00
sum2012 f5b6b45b5b Remove IgnoreEnqueue hack 2026-09-22 20:19:57 +08:00
Henrik RydgårdandClaude Opus 5 f28896e1ed Vertex decoder: JIT morph on arm64
The arm64 vertex JIT had no morph support (the table entries were commented
out), so every morphed format fell back to the step functions. Add the same
morph steps x86 has: texcoords (plain and prescaled), normals, positions and
the four color formats.

Each step sums its frames the way its step function does as compiled: fused
where that's plain C++, which the compiler turns into fused multiply-adds,
and separate multiply and add where it's CrossSIMD. The unit test confirms
they match bit for bit. Morph formats decode about 2.5-3x faster than with
the steps.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 16:51:00 -06:00
Henrik Rydgård 7b95ff2808 Merge pull request #22325 from hrydgard/vertex-decoder-jit-match
Vertex decoder: New test, make the JITs match the C++ decoder closely
2026-09-21 16:28:27 -06:00
Henrik RydgårdandClaude Opus 5 66b59c73a9 Vertex decoder: bring the riscv64 and loongarch64 JITs in line
Enables TestVertexJitMatchesSteps on both. riscv64 now passes all 216000
jitted formats and loongarch64 all 204000, against 28906 on arm64.

riscv64:
- Jit_PosFloat didn't clean NaN or infinity at all, it just copied the
  three words. Clamp to +-FLT_MAX like the x86 JIT does.
- Jit_PosFloatThrough was missing the truncation of Z to an integer.
- The morph helpers started the sum from the first product rather than
  from +0.0, which rounds differently and lets a -0 term through, and
  rounded that first product towards zero where the steps round to
  nearest. The rest of the sum stays fused, since the compiler contracts
  the steps into fused multiply-adds.
- The texcoord prescale and 5551 color morph paths read morph weights
  from tables that GetMorphValueUsage never asked to be filled in, so
  they used whatever an earlier vertex type had left there.
- The non-Zbb bounds update compared the wrong way around, so through
  mode texcoord bounds came out inverted.

loongarch64:
- Jit_PosFloat had a TODO to clean NaN and infinity, and didn't.
- Jit_PosFloatThrough was missing the same Z truncation.
- Jit_WriteMorphColor narrowed with the logical saturating shifts, so a
  negative channel became a huge unsigned value and saturated to 255
  instead of clamping to 0, and it rounded where the steps truncate. It
  also read the packed color back sign-extended, so any alpha above 0x7F
  compared as larger than 0xFF000000 and claimed full alpha.
- The three packed color morph formats are rewritten. The LSX versions
  built each channel with a chain of inserts, shifts and shuffles that
  didn't survive being run; 4444 also broadcast its scale from the mask
  register. They now follow the steps channel by channel. Note the
  accumulator has to be an LSX scratch register - F4-F7 alias V4-V7,
  which hold the skin matrix for the whole vertex.
- The vertex bounds were loaded with a signed halfword load, so the
  0xFFFF they start at became -1 and no texcoord was ever below it.

PrescaleUV now fuses on these two as well - both JITs fuse it, and the
compiler would have contracted the plain expression there anyway.

The test tolerates a small relative difference on decoded floats, scaled
by the morph count since each term rounds once. Every bug above was
orders of magnitude larger than that.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 13:48:56 -06:00
Henrik RydgårdandClaude Fable 5.1 a36d09daac sceGe: what callbacks see, and a full GE reset on sceKernelLoadExec
From another pass over ge.prx against our code, each checked on a PSP with
gpu/ge/callbackstate except the last:

- A list's context is restored after its finish callback, which sees the
  state the list left. We restored at the FINISH, before it. Now that nothing
  runs until InterruptEnd(), that's where it happens.
- sceGeSaveContext/RestoreContext only fail while the GE is executing. It's
  stopped during a finish callback and a SUSPEND signal callback, however
  much is queued, so they work there. We said busy whenever a list existed.
- sceGeListDeQueue emptying the queue doesn't turn completed lists into
  nothing, only sceGeDrawSync does. CheckDrawSync() is gone.
- The "break in progress" flag that makes sceGeContinue only requeue the
  list is cleared by an interrupt that follows the break at once, so it's
  only seen from a callback or with interrupts off. Ours lasted until the
  next GE interrupt of any kind.
- sceKernelLoadExec restarts the GE driver, which begins by zeroing every
  register and matrix. Reinitialize() now does too, so a program started that
  way finds the same GE as one booted directly, rather than its launcher's.
  Not testable on hardware: nothing after the restart can report back.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-21 13:21:03 -06:00
Henrik RydgårdandClaude Fable 5.1 cc41a23256 GPU: empty the display list queue on sceKernelLoadExec, fixes Crazy Taxi
Reinitialize() wiped the 64 display lists but kept the queue of their ids.
Whatever the old executable still had queued came back as lists with no state
and a pc of 0, behind the first list of the new executable, where they
blocked everything. Crazy Taxi: Fare Wars is a launcher for its two games,
and stopped at a black screen that way. Fixes #19894.

This removes the workaround for it, which dropped such a list but returned
before currentList was cleared, and only worked as long as something else
happened to clear it later. A list with a bad pc is now dropped like one that
ran into an error, instead of sitting at the head of the queue for good.

Also narrows what sceGeBreak(1) throws away to interrupts that have actually
been raised, which is what gpu/ge/intrsuspend shows. The ones we haven't
raised yet are only late because we execute lists ahead of time: a game that
breaks right after its last list, and then waits for what the finish callback
signals, got that callback long ago on hardware.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-21 13:10:02 -06:00
Henrik RydgårdandClaude Fable 5.1 91c9c4d14e sceGe: sceGeBreak(1) takes pending interrupts with it
Resetting the GE also gets rid of an interrupt that was raised but not taken
yet, so a list that reached its FINISH just before never gets its finish
callback. We delivered one anyway, for a list that no longer existed. If the
break comes from inside a GE callback, the interrupt being handled is kept,
since its handler still has to return.

Found by gpu/ge/intrsuspend, which also confirms from a thread, with
interrupts suspended, that nothing moves along the queue until the FINISH
interrupt has been taken.

Savestates: bump GPUCommon to 7. We didn't use to mark a PAUSE signal as
delivered, which sceGeContinue now goes by, so a state saved with a list
paused that way would load into a game that could never continue it. Fixed
up on load.

gpu/signals/handlercalls goes in as known failing: with an old SDK version, a
stall address set from inside a SUSPEND callback doesn't reach the GE, which
we can't express with just the one stall address per list. See docs/sceGe.md.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-21 12:48:03 -06:00
Henrik RydgårdandClaude Fable 5.1 8a23e633a1 sceGe: keep a finished list on the queue until its interrupt is done
On hardware the GE stops at every SIGNAL and FINISH, and it's the interrupt
that gets it going again: on the same list after a signal, on the next one
after a FINISH - once the finish callback has run, with the finished list
still at the head of the queue. We ran the next list right away and dropped
the finished one at once, so a finish callback saw an empty queue. A list
enqueued from there was started instead of queued, and then couldn't be
dequeued, which hung the new gpu/ge/queue2 test.

ProcessDLQueue() now runs nothing while the head of the queue has an
interrupt pending, and InterruptEnd() is what takes a finished list off the
queue. This also keeps a stall update from restarting a list that's stopped
at a signal before the handler has run. drawCompleteTicks is still set when
the last list reaches its FINISH, so a sceGeDrawSync in between doesn't wait.

Other things gpu/ge/queue2 and gpu/ge/breakwait showed, all from a real PSP:

- sceGeListEnQueue compares against the address a list was enqueued with
  (or stopped at by sceGeBreak), mirrors included, not against its current pc.
  We had that the wrong way around.
- The stack-in-use check only applies to lists that have started executing.
  This is probably what IgnoreEnqueue was added for (Metal Gear Acid 2,
  #10906). The flag stays until someone has checked the game without it.
- A PAUSE signal makes the list PAUSED at once, before the FINISH delivers it.
  In between, sceGeContinue and sceGeBreak say BUSY, and updating the stall
  address does nothing, so a list that stalls there is stuck.
- A completed list can't be dequeued, with or without a context.
- sceGeDrawSync(1) looked at currentList instead of the list it had found.
- sceGeBreak(1) doesn't wake anyone, and a late interrupt for a list it reset
  no longer marks that list completed. Threads in sceGeDrawSync are woken
  before the ones waiting for the last list.

Also fixes currentList being lost when loading a state where it's list 0,
and makes ge_pending_cb a plain std::list - nothing else touches it, and the
GPU thread it was shared with is long gone. Same savestate format.

See docs/sceGe.md.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-21 12:21:19 -06:00
Henrik RydgårdandClaude Opus 5 c58baedd75 Vertex decoder: fix the jit-match test on MSVC
The test didn't compile: VERTS is captured by reference into testFormat, and
MSVC won't use a captured constexpr as an array bound. Make it static.

On arm64 it then failed on the UV prescale steps, by one ULP. The arm64 JIT and
the NEON handwritten decoders fuse the multiply-add, and the steps only match
that when the compiler contracts a * b + c - which clang does and MSVC doesn't,
in Debug or Release. Spell out which one happens instead of relying on it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 12:13:09 -06:00
Henrik RydgårdandClaude Opus 5 a1ed24ae2d Vertex decoder: fix the handwritten decoders, and test them too
The two handwritten SIMD decoders are used with or without the JIT, and had
drifted from the step functions:
- The God of War one ignored g_DoubleTextureCoordinates, giving HD Remaster
  games the wrong UVs, and passed NaN and infinite positions through. Those
  now come out finite like everywhere else.
- The GTA one expanded 5551 colors wrongly on NEON (a left shift where the
  SSE version shifts right).
- On ARM64, both now fuse the UV scale and offset, like the JIT and the steps
  as the compiler builds them.

Both are back in the unit test, which now also feeds NaN and infinity to plain
float positions and checks they come out finite, and runs only on x86-64 and
arm64 for now.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 11:29:40 -06:00
Henrik RydgårdandClaude Opus 5 63a6c11214 Vertex decoder: make the JIT match the step functions
Add a unit test that decodes every vertex format through both the step
functions and the JIT and requires identical output, side effects included.
Only skinning may differ by rounding, since arm64 accumulates the bone
matrices with fused multiply-adds.

What it found, and fixed:
- x86 morph colors rounded to nearest where the steps truncate, and applied
  the scale before the weight, which rounds differently.
- x86 through-mode u16 UV bounds compared signed, so texcoords above 32767
  scrambled the bounds for everyone running the x86 JIT.
- x86 morph sums could produce -0 where the steps produce +0.
- Step_NormalS16Morph scaled by 1/32768 twice, giving near-zero normals.
- Step_PosFloatThrough lost the truncation of Z to an integer that the JITs
  and the vertex reader did before it moved into the decoder.
- SetVertexType never reset skinInDecode.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 11:29:40 -06:00
Henrik RydgårdandClaude Opus 5 0b5a8f537e Decide on the vertex decoder JIT with CoreParameter, not the CPU core
The vertex decoder JIT was enabled only when g_Config.iCpuCore was one of the
JIT cores, which tied an unrelated GPU path to the CPU setting. Headless leaned
on that by forcing iCpuCore to the interpreter after ApplyToConfig(), to keep
the decoder JIT off in the tests.

Add CoreParameter::bUseVertexDecoderJit instead. Standalone and libretro set it
wherever the host can JIT, so IR interpreter users on such hosts now get the
decoder JIT too. Headless sets it false, which keeps the test results as they
were and lets --cpu go through ApplyToConfig() like every other option.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 10:31:58 -06:00
Henrik RydgårdandClaude Opus 5 19bac4894c Let sceMpeg LLE override the ForceHLEPsmf compat flag
Our psmf and psmfPlayer HLE plays video by calling our sceMpeg HLE, so it has
nothing to talk to when the real mpeg.prx is running. The flag now gives way
when sceMpeg is LLE.

Moved below the force-enable and unavailable masks so it tests what sceMpeg
actually ended up as rather than what was asked for.

Also remove a bad assert.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 15:15:45 -06:00
KailashandClaude Opus 5 eb2f5dfc61 GLES: Pass known index range to glDrawRangeElements
DrawEngineGLES already knows that every index it generates is below the
decoded/transformed vertex count, but only passed the count to
glDrawElements. Drivers that need the vertex range before vertex
shading, like Mesa's Panfrost, then scan the index data on the CPU for
every draw, and their min/max caches can't help since the index push
buffer is rewritten every frame.

Only used for single-instance draws on desktop GL or GLES3, and only
when the caller provides the range, so other DrawIndexed callers are
unchanged.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-16 21:01:54 +05:30
Acts1631 90f7c6d195 Rely on PNG header dimension bounds
PNG replacement dimensions are validated by PNGHeaderPeek before the

decoded buffer is allocated, so the additional size_t overflow checks are

redundant.
2026-09-07 14:52:06 -04:00
Acts1631 9ad3b821d9 Bound replacement PNG dimensions safely
Reject non-positive or oversized dimensions independently in the PNG

header peeker, matching the existing 8192 pixel decode limit. Use

checked size_t arithmetic before allocating replacement RGBA data so

crafted dimensions cannot overflow the allocation size.
2026-09-05 16:24:24 -04:00
Henrik Rydgård 2de11efe4e TextureCache: stop clamping DXT decoding to bufw either
DecodeDXTBlocks limited its x loop to min(bufw, w), so when w > bufw everything
from bufw to w was simply never written - the destination is sized for w, so
those columns kept whatever the buffer held.

The software sampler doesn't do that. It addresses a block as
(v >> 2) * (texbufw >> 2) + (u >> 2), which for u past bufw runs on into the
next row's blocks, the same way the linear formats run into the next row. So the
two renderers disagreed on the same texture.

Follow the sampler: decode w texels per row and let the block index carry on,
with the range check widened to cover the blocks that reach past the last row's
stride. Shares the SourceExtent helper from the previous commit, counting 4x4
blocks rather than texels.

Rendering-visible where w > bufw for a DXT texture, which is the case that used
to leave stale contents behind.
2026-09-05 11:25:55 -06:00
Henrik Rydgård ca9d76fef3 TextureCache: bound the source by w as well as bufw, instead of clamping w
The non-DXT decode paths read w texels per row from a source whose stride, range
check and unswizzle buffer were all sized from bufw alone. w and bufw are
independent GE registers, so w > bufw is reachable, and the last row then runs
off the end of both the validated guest range and the temp buffer.

The previous commit clamped w down to bufw, which stops the overrun but is the
wrong shape twice over: the game asked for w texels and the destination is sized
for w, so the tail of every row is left holding whatever was in the buffer, and
it silently narrows a texture the hardware would have decoded in full.

Size the checks from both instead. The extent a decode touches is bufw per row
plus however far the last row reaches past its own stride - the same adjustment
TextureReplacer::ComputeHash and the GE recorder already make - so the range
check, the height it falls back to when the range is short, and the unswizzle
buffer are all computed that way now.

Pull the five copies of "resize tmpTexBuf32_, unswizzle into it" into a helper
while we're here, since they all need the same sizing. It zeroes the buffer when
w reaches past bufw, as UnswizzleFromMem only fills the stride and the tail would
otherwise be stale heap.

ComputeTextureHash has the same bufw-only assumption, but it's left alone
deliberately - changing what goes into a texture hash isn't worth the risk here.

DXT is left alone too - it clamps to minw, but there the limit is the block index
within a bufw/4 block row, so its range check already covers what it reads.
2026-09-05 10:26:14 -06:00
Henrik RydgårdandClaude Opus 5 ed5d9813bb TextureCache: fix two host-memory overruns from GE texture state
PrepareBuildTexture's mip scan stopped at the first level with a dimension of 1
*before* running the mip size check for that level, so a level like 1x256 under a
256x256 level 0 was accepted as a valid mip. The backend then sized the level from
the halved level-0 dimensions (128x128) while LoadTextureLevel re-read the real
height (256) from the GE state, writing past the allocation. Do the check first.

That check was also gated on GPU_USE_SAMPLER_LOD_CONTROL, so backends without it
did no mip dimension validation at all - there's no reason for the guard, the
result only feeds badMipSizes, so drop it.

Separately, the non-DXT decode paths read w texels per row from a source whose
stride, range check and unswizzle buffer are all computed from bufw. w and bufw
are independent GE registers, so w > bufw is reachable and read past the end of
both the validated guest range and the temp buffer. Clamp w to bufw in
DecodeTextureLevel, which is what the DXT paths have always done via minw.

Note: ungating the mip check means backends without SAMPLER_LOD_CONTROL can now
set badMipSizes where they previously didn't, which collapses such textures to a
single level. That's the intended behavior, but it is a rendering-visible change
on those backends.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01DCPmm7FoQUoqrbMdhfqhQ2
2026-09-05 10:25:47 -06:00
Henrik Rydgård 68fd30bbba Merge pull request #22217 from hrydgard/gpu-fixes
Claude code review: Vulkan
2026-09-04 18:03:47 -06:00
Henrik RydgårdandClaude Opus 5 859d09cc75 Vulkan: Two small fixes
The pipeline debug listing printed the color blend factors in the alpha slot,
which is doubly unhelpful since that branch is only taken when the alpha factors
differ from the defaults.

CompileShaderModuleAsync takes ownership of the tag but only deleted it on the
success path, leaking it whenever GLSLtoSPV failed.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01SGz54K3ZXa2Qzyc3aMEYyY
2026-09-04 12:10:03 -06:00
Henrik RydgårdandClaude Opus 5 9273764a3b Vulkan: Don't insert the same key twice when compute pipeline creation fails
The failure branch inserted a null pipeline and then fell through to the normal
insert of the same key, which trips DenseHashMap's duplicate-key assert - and
_assert_msg_ is live in release builds, so a logged error became a crash.

Also skip the null entries when deleting cached pipelines.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01SGz54K3ZXa2Qzyc3aMEYyY
2026-09-04 12:10:03 -06:00
Henrik RydgårdandClaude Opus 5 41d69e61e1 Vulkan: Set USES_DEPTH_STENCIL/USES_BLEND_CONSTANT before creating the pipeline
They were OR-ed into pipelineFlags just after the CreateGraphicsPipeline call that
consumes them, so the render manager's "don't compile a pipeline that requires
depth for a non-depth renderpass type" check could never fire for game pipelines.
thin3d_vulkan.cpp sets the flag before its call, which is the intended order.

Note this can now legitimately skip some variants when loading the shader cache -
those were invalid combinations that the check was written to reject.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01SGz54K3ZXa2Qzyc3aMEYyY
2026-09-04 12:10:03 -06:00
Henrik RydgårdandClaude Opus 5 38d5ab657e Vulkan: Disable hardware texture scaling if the constant buffer fails to load
If reading the shader's constant buffer file failed, we'd skip writing descriptor
binding 4 but still dispatch the compute shader, which declares it - a statically
used but unwritten descriptor. It also re-read the missing file on every single
texture upload. Now we drop the scaling shaders instead, so following textures
take the CPU scaling path.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01SGz54K3ZXa2Qzyc3aMEYyY
2026-09-04 12:10:03 -06:00
Henrik RydgårdandClaude Opus 5 79a48c7dbe Vulkan: Fix crash when a replaced texture fails to allocate
The out-of-VRAM retry path cleared plan.replaced but left plan.doReplace set.
GetMipSize() dereferences plan.replaced when doReplace is true, so the fallback
crashed instead of recovering. The common code sets both together.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01SGz54K3ZXa2Qzyc3aMEYyY
2026-09-04 12:10:03 -06:00
Henrik Rydgård 425f6e2c37 D3D11: Don't silently draw with a shader that failed to compile
The shader creation helpers returned S_FALSE when compilation produced no
bytecode - but S_FALSE is a success code, so the FAILED() checks in the
D3D11VertexShader/D3D11FragmentShader constructors never fired and failed_ was
never set. Return E_FAIL instead, and only hand out the bytecode when the shader
object was actually created.

Failed() had no callers at all, so a failed shader was handed to the draw as
usual: VSSetShader(nullptr), CreateInputLayout on empty bytecode, and the
ignored HRESULT from SetupDecFmtForDraw meant IASetInputLayout(nullptr). The
draw then did nothing, with no log line to explain it. Skip the draw and warn
instead, like the GL backend does.
2026-09-04 12:08:21 -06:00
Henrik Rydgård 15c702f203 D3D11: Fix texture upload buffer leaks
The cleanup loop was hardcoded to 12 entries while the array holds more and
levels can exceed that with a texture replacement pack - a 8192-pixel
replacement has 14 mip levels, so every level from 12 up leaked on each texture
build. Use ARRAY_SIZE, and shrink the array to 16 since D3D11 caps textures at
16384 pixels anyway.

The out-of-memory bail-out returned before that loop, leaking every level
decoded so far. Since the entry ends up without a texture it gets rebuilt, and
leaks again, every following frame.
2026-09-04 12:07:50 -06:00
Henrik Rydgård 050cacbbdf Merge pull request #22214 from hrydgard/gles-fixes
Claude code review: OpenGL backend
2026-09-04 11:45:42 -06:00
Henrik Rydgård 256984d561 Fix some Claude-isms 2026-09-04 10:41:31 -06:00
Henrik Rydgård ac016201dc GLES: Small cleanups
Scissor the stencil readback to the region actually being read back - latent,
every caller passes a zero origin today.

Remove a DecodeVerts call that can never do anything: both branches above it
have already advanced decodeVertsCounter_ to numDrawVerts_. Worse than useless,
since in the non-skinning branch the vertices went to the push buffer, so
decoded_ doesn't hold them.
2026-09-04 10:41:31 -06:00
Henrik Rydgård 78692deca2 GLES: Set IS_3D when creating the 3D texture, not after uploading it
The out-of-memory bail-out added in the previous commit returned before the
status flag was set, leaving a GL_TEXTURE_3D object bound while ApplyTexture
told the shader generator it was a 2D texture. The entry stays cached, so it
would repeat every frame, not just the one that failed to allocate.
2026-09-04 10:41:31 -06:00
Henrik Rydgård 89ebb8acc6 GLES: Actually apply anisotropic filtering
TextureCacheGLES passed a hardcoded 0.0f instead of key.aniso, so the
Anisotropic Filtering setting did nothing at all on the OpenGL backend, even
though GPU_USE_ANISOTROPY was advertised and D3D11/Vulkan both honor it. Looks
like it was left behind by the 2017 render manager refactor.

The queue runner now clamps to the device maximum it already queried into
maxAnisotropyLevel_ (until now unused), and only touches the parameter when the
extension is actually supported - the anisotropy branch there has been dead
since every caller passed 0.0f, so this is the first time it runs.

0.0f keeps its meaning of "don't care" for the CLUT/fragment-test/thin3d
callers; the texture cache now passes 1.0f when the setting is off, so turning
it off takes effect on already-uploaded textures instead of only new ones.

TexCache: Never use anisotropic filtering for CLUT8-indexed textures

What gets sampled for those is palette indices, depalettized by the shader
afterwards - averaging indices across an anisotropic footprint produces garbage
colors. Affects all backends, not just the GL one that just started honoring
key.aniso.

TexCache: Clear key.aniso wherever filtering is forced to nearest

It was only cleared in the two places inside the AUTO_MAX_QUALITY branch, so the
TEX_FILTER_AUTO path (pixel-mapped textures, the ugly color test heuristic), the
FORCE_NEAREST setting and the replacement-texture override could all end up
requesting nearest filtering with anisotropy still on.

Doing it in the switch that applies forceFiltering covers every path, so it
can't drift apart again.

GLES: Only record the applied anisotropy, and log skipped draws

The queue runner updated tex->anisotropy even when it skipped the call because
the value was 0.0f ("don't care") - harmless while nothing ever set anisotropy,
but now it would make the tracked state disagree with GL, so a later request for
the value it thinks is set would be wrongly skipped.

Also log when a draw is skipped for a missing vertex shader. The failure is
cached per shader ID, so without it geometry silently disappears for the rest of
the session after the one-shot OSD message.
2026-09-04 10:41:15 -06:00
Henrik Rydgård 388eef9d88 GLES: Harden the shader disk cache loader and the 3D texture upload path
The cache loader indexed &vec[0] on vectors that can legitimately be empty (a
header-sized file with zero counts passes both sanity checks), and the counts
are signed ints where only the upper bound was checked - a negative count would
reach resize() as a huge size_t.

The 3D texture branch had the out-of-memory assert but not the bail-out the 2D
branch has, so an ignored assert fell straight into memset(nullptr).
2026-09-03 20:43:06 -06:00
Henrik Rydgård ad131e522f GLES: Handle shader compilation failure in the hardware transform path
ApplyVertexShader can return null - if the requested shader fails to compile it
retries with a software transform ID, and if that fails too it returns (and
caches) null. We then called UseHWTransform() on it.

ApplyFragmentShader can likewise return null, and the hardware path ignored it,
unlike the software path. Without a linked shader nothing binds a program for
this render pass, so the draw would have gone through with whatever program a
previous pass left bound.
2026-09-03 20:43:06 -06:00
Henrik Rydgård 245ed61c0f GLES: Actually apply the stencil write mask in ApplyDrawStateLate
53aa2cc596 changed the first argument from "true" to stencilState_.writeMask,
but that slot is "bool enabled" - the writeMask argument stayed hardcoded to
0xFF, so the mask still never reached GL. The clear-mode call just above gets
the slots right.

Reachable because SoftwareTransformCommon refuses the fast clear path when the
stencil write mask is partial, so exactly those clears end up here.
2026-09-03 20:43:06 -06:00
Henrik Rydgård de0dc2d7d4 Remove the leftover geometry shader scaffolding
Nothing has generated or used a geometry shader since the GS paths were removed
- GeometryShaderGenerator is gone, and ShaderWriter's BeginGSMain/EndGSMain had
no callers at all. Removes ShaderStage::Geometry and everything hanging off it:
the GS preambles and GSMain helpers in ShaderWriter, the stage mappings in all
three thin3d backends, the D3D11 geometry shader plumbing (curGS_, the pipeline
and module members, gs_4_0 compilation), CreateGeometryShaderD3D11, the unused
PipelineFlags::USES_GEOMETRY_SHADER and PipelineManagerVulkan's
UsesGeometryShader().

Also stop enabling the Vulkan geometryShader device feature, since we no longer
have any use for it.

Kept on purpose: the device feature is still listed in the feature dumps (like
other capabilities we don't use), and the Vulkan shader cache header keeps its
now-always-zero geometry shader count so the on-disk format stays compatible.
2026-09-03 11:24:03 -06:00
Henrik Rydgård 293cf0c78c Revert the logic changes from "Add new TexCache logging channel"
This partially reverts commit ee1314f803.
2026-09-02 15:56:10 -06:00
Henrik Rydgård d4b6965db0 TestBoundingBox: bail out when the indices reach past the scratch space
corners and verts are carved out of decoded_ at fixed offsets 6*65536 bytes apart,
and NormalizeVertices fills corners with indexUpperBound - indexLowerBound + 1
SimpleVertex. The vertexCount > 1024 guard doesn't bound that: the index values come
from the game, so 1024 indices can span the full 16-bit range and run corners into
verts, making the cull decision from overwritten data. Bail on an index over 1024
and report visible - a bbox test that large isn't worth doing anyway.
2026-09-02 18:06:47 +02:00
Henrik Rydgård 506cfb6701 Bound two values that texture packs and shader inis supply
A hashrange of 'addr,w,h = 0,0' passed validation (0 isn't bigger than the source),
became desc_.newW/newH, and ReplacedTexture::Prepare divides by them. A post-shader
SSAA level multiplies the render resolution with no upper bound, while the
texture-shader Scale sitting a few lines away is checked against 2..8.
2026-09-02 18:06:47 +02:00
Henrik Rydgård 56bba5f6f5 Merge pull request #22189 from hrydgard/debug-input-rewind-overflows
Fix three more out-of-bounds writes (minor)
2026-08-31 17:53:10 +02:00
Henrik Rydgård 9cb50459d6 Merge pull request #22186 from hrydgard/medium-correctness-fixes
Medium-severity correctness fixes
2026-08-31 16:30:53 +02:00
Henrik Rydgård 1ee9168711 Merge pull request #22184 from hrydgard/fix-adreno-workaround-regression
Fix fragment shader logic error (Adreno stencil driver bug workaround)
2026-08-31 13:30:19 +02:00
Henrik Rydgård 496eddb7fb Merge pull request #22185 from hrydgard/framebuffer-and-null-deref-fixes
Framebuffer and null deref fixes
2026-08-31 13:29:40 +02:00