3634 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 333035df03 Lighting: Compute the GE's pow from the float's bits
Read as an integer, a float's bits are its log2 with the mantissa a
straight line between powers of two, scaled by 2^23, and writing an
integer back is the matching exp2. That's exactly the GE's
approximation, without log2/exp2/floor. pspPow now also returns 1 for
e <= 0 itself, so the callers drop their checks. Shader languages
without integers fall back to a true pow.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 15:47:16 -06:00
Henrik RydgårdandClaude Opus 5.5 a6849661ba Shade mapping: Use the light vector as lighting sees it
Environment map S and T are (N.L + 1) / 2 with L the light's vector as
lighting uses it: from the vertex to the light for point and spot lights,
a zero vector staying zero, and the half vector for a light that does
specular. Whether lighting or the light is enabled still doesn't matter
(gpu/lighting/shademap).

The vertex shader ID now carries the type and computation of the shade
mapping lights (the ubershader reads them from u_lightControl), so both
shader caches get a new version.

Fixes the hair shine in iDOLM@STER SP (#12376).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 15:47:16 -06:00
Henrik RydgårdandClaude Opus 5.5 6f6d3e2a81 Lighting: Take the specular viewer direction from the view matrix
The viewer is at infinity along view space +z, so in world space, where
lighting happens, it's the view matrix's third column rather than
(0,0,1): turning the camera moves the highlights (gpu/lighting/specular).
The shaders read it from u_view.

In a Need for Speed Carbon frame replayed on a PSP, this and the pow
bring the error of the cars from MSE 957 to 120 (Vulkan).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 15:47:16 -06:00
Henrik RydgårdandClaude Opus 5.5 48f3d5c907 Lighter: No specular from a light behind the surface
The shaders and the software renderer only add specular when N.L >= 0,
as the GE does (gpu/lighting/specular); the CPU lighter, used for points,
lines and rectangles, didn't check.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 15:47:16 -06:00
Henrik RydgårdandClaude Opus 5.5 d7a96875fb Lighting: Use the GE's approximate pow for specular, diffuse and spot
The GE computes these powers as exp2(e * log2(x)), with log2 and exp2
each a straight line between powers of two (Mitchell's approximation),
and only uses the top 4 bits of the specular coefficient's mantissa.
Through a highlight's falloff a true pow is 10-30 steps of 255 brighter
at the exponents games use. Measured in gpu/lighting/specular.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 15:47:16 -06:00
Henrik Rydgård a7ec706a4c Merge pull request #22400 from hrydgard/pacman-spline-lighting
Spline lighting fixes
2026-09-30 15:46:20 -06:00
Henrik RydgårdandClaude Opus 5.5 0715f43ab8 Splines: Detect poles relative to the other derivative
With animated control points the pole is only nearly degenerate, so the
vanishing derivative is rounding noise rather than exactly zero, and the
absolute threshold missed it. The resulting random normals still showed
as dark patches on Pac-Man Arrangement's ghosts (#12354).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 13:20:01 -06:00
Henrik RydgårdandClaude Opus 5.5 aa38fafd77 DrawEngine: Don't drop large batches of points, lines and rectangles
Software transform expands each point, line and rectangle to four
vertices, and gave up on the whole draw when that didn't fit
VERTEX_BUFFER_MAX. Batching only counted input vertices, so a batch over
16384 points (or 32768 line or rectangle vertices) vanished silently,
whether it came from one PRIM or several merged ones. Count the expanded
vertices when batching, and submit a PRIM too big on its own in parts.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 12:04:53 -06:00
Henrik RydgårdandClaude Opus 5.5 22c2cea923 Splines: Fix a single patch with both edges open
With only one patch, the open-last-edge adjustment assumed the first edge
was closed, so one knot interval came out as 2 instead of 1, and the
patch wasn't the Bezier patch a fully clamped cubic is. Positions were off
by up to about 7% of the patch.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 11:11:23 -06:00
Henrik RydgårdandClaude Opus 5.5 7a06e25aa0 GPU: Remove leftovers from hardware tessellation
The GLES sampler uniforms and texture slots for the control points and
weights, the Vulkan storage buffer bindings, and the u_spline_counts
uniform, which becomes padding (the C++ side already was).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 11:00:02 -06:00
Henrik RydgårdandClaude Opus 5.5 f87a07b1a3 Splines: Use the limit normal at a pole instead of NaN
Where all the control points along a patch edge meet at one point, like
the top of a dome, one derivative is zero and so is the cross product,
and normalizing it gave NaN. Use the limit instead, built from the mixed
second derivative. Fixes the dark spots on the ghosts' heads in Pac-Man
Arrangement (#12354).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 10:57:05 -06:00
Henrik Rydgård ae460c9e1a Merge pull request #22390 from hrydgard/gpu-review-leftovers
Claude code review of GPU: Fix leftover findings
2026-09-29 15:43:31 -06:00
Henrik RydgårdandClaude Opus 5.5 b9501df529 TextureReplacer: Fix stale and shared lookup results
- Reloading the ini clears the per-key lookup caches, which could keep
  saying "no replacement" for textures the new ini replaces.
- With ignoreAddress, hash ranges were skipped when sizing the
  replacement, though ComputeHash applies them. cache_ is now keyed by
  the full key, so its lookups hit too.
- Textures sharing files but differing in size, hash range or filtering
  no longer share one ReplacedTexture (the first one's settings won).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:28:21 -06:00
Henrik RydgårdandClaude Opus 5.5 b8789af997 GPU: Draw frames displayed from RAM in non-buffered mode
They were uploaded but never drawn, and the block transfer hack drew
whatever source was left over. Draw them straight into the backbuffer
pass, without post shaders, which would need to bind their own targets.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:17:01 -06:00
Henrik RydgårdandClaude Opus 5.5 bc581349fd GPU: Install the draw engines' invalidation callback from BeginFrame
The draw engine is created on the loader thread while the UI thread may
already be rendering and invoking the callback.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:16:42 -06:00
Henrik RydgårdandClaude Opus 5.5 4c2086aac5 GPU: Keep spline/bezier tessellation from reaching zero
Reduce only the larger factor when over the vertex limit, and stop at 1.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:16:40 -06:00
Henrik RydgårdandClaude Opus 5.5 7de7b2d76b TextureCache: Use XXH3 instead of the quick hash for texture data
With xxhash 0.8.4, XXH3 is faster than StableQuickTexHash on ARM64
(about 37 vs 24 GB/s on a Snapdragon X), and its scalar path, which
RISC-V builds get, is about as fast as the quick hash's. It doesn't
collide the way the quick hash does (#8249).

Texture replacement still uses its own hash setting, so texture packs
are unaffected.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:10:28 -06:00
Henrik RydgårdandClaude Opus 5.5 93f57a4f38 GPU: Delete copy operations on classes that own resources
These own GPU objects, memory or refcounts in their destructors (or assert
there that they were torn down), so a copy would double-free. Nothing copies
them today; this keeps it that way. The manager base classes cover every
backend's subclass.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 b72927bbeb Texture replacement: Plug leaks on error paths
- Release the file reference when loading a level fails or finds nothing,
  since only a loaded level takes ownership of it.
- Free the PNG image when the size changed since the header was read.
- Delete the VFS when a pack without an ini has no hash-named textures.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 e108e41675 GPU: Fix assorted leaks, null derefs and small rendering bugs
- Put the anisotropy level in the sampler key, so changing it applies on
  Vulkan and D3D11.
- Release CLUT textures at shutdown on GLES and D3D11.
- Fix the depth readback viewport, which squeezed the image whenever the
  read rectangle was smaller than the fbo.
- Test the computed depth, not the unset result, in the equal-depth clear
  check.
- Don't read back a CLUT from a framebuffer without an fbo.
- Tolerate null entries when releasing post-shader objects and CLUT
  textures after a failed creation.
- ImGe: Don't crash on a framebuffer without an fbo, or on GetVFB under the
  software renderer.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 16ea2cbf88 GPU: Fix draw engine buffer overruns and stale vertex data
- Flush before the queued draws would decode more than VERTEX_BUFFER_MAX
  vertices. The batch was limited by index count, which doesn't bound a
  sparse index range, and DecodeVerts silently stopped while DecodeInds
  still emitted indices for the undecoded draws.
- Give TestBoundingBox its own scratch buffer. It used offsets in decoded_,
  which can hold decoded vertices that aren't flushed yet.
- Read 32-bit indices the way the PSP does, ignoring the upper 16 bits.
  IndexConverter and the fast bounding box test used all 32, so a game
  setting them indexed far past the decoded vertices.
- D3D11: Flush in FinishDeferred like the other backends, since indices
  are still read from PSP memory at flush time (#10095).
- Don't JIT new vertex decoders once the code space is full.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 24b71386ff Document that shader cache key changes need a CACHE_VERSION bump
The OpenGL and Vulkan shader caches store raw shader IDs (and, for Vulkan,
pipeline keys) on disk. Add the rule to AGENTS.md and point to it from the
persisted types.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 45d997c3ea Texture replacement: Fix DDS mip double free and VFS lifetime
- DDS files with mips stored level 0's file reference in level 1 too, so it
  was freed twice.
- The VFS was deleted on config changes (and ini reloads) while load tasks
  still used it. Now each cached texture waits for its task and releases its
  file references through the old VFS first, and reloads afterwards.
- A failed ini reload turns replacement off instead of leaving it on without
  a VFS.
- Release file references on purge.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 b912b5823b GPU: Fix texture cache and framebuffer cache lifetime bugs
- Delete TexCacheEntry objects dropped on rehash instead of leaking them.
- Don't leave a released null entry in cache_ when the framebuffer match
  returns before the slot is refilled.
- Reset clutRenderAddress_ in Clear(), which releases the dynamic CLUT FBOs.
- Don't cache a null texture in drawPixelsCache_ when creation fails.
- Fix the reversed subtraction in the failed-FBO retry check.
- Remove the never-taken buffered-rendering early-out in UpdateRenderSize.
  Taking it would leave existing VFBs without an fbo.
- Include smoothedDepal in the depal shader cache key, and print/parse the
  debug IDs as 64-bit.
- Release depal pipelines through Draw2DPipeline::Release so the shader
  source isn't leaked, and make that null-safe.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 3f44da709b Clamp sampling of video textures and direct-displayed video to the 480x272 frame
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:22:01 -06:00
Henrik RydgårdandClaude Opus 5.5 242ef0a984 GPU: Track video frames in GPUCommon, so they expire for the blit cost too
The blit cost remembered only the last buffer a decoder wrote into, forever.
Move the texture cache's video list (with its ageing out a few flips after
the last write) into GPUCommon, so the texture cache, the blit cost and
SoftGPU all share one. That also counts both of a double-buffered player's
frames, which exposed that a clear drawn with texturing still enabled was
being charged as a blit - skip clears and draws without texture coordinates.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 17:00:11 -06:00
Ygor Dreyer a422b818b7 Hash full swizzled CLUT4 glyph atlases 2026-09-24 17:10:10 -03:00
sum2012 a8b85eec8b Add ForceEnableGPUReadback compat
To solve gpu readback issue
2026-09-23 22:06:28 +08:00
Henrik RydgårdandClaude Opus 5 4aef060293 Carry "this is video" across block copies
videos_ only learns about the CSC output, so the texture we actually draw is
just an ordinary 512x512 8888 texture whose contents happen to be different
every frame: hash, miss, throw it in the secondary cache, rebuild, forever.

So track the copy. NotifyVideoCopy marks the destination as video when the
source is, and the copy funnels call it: sceDmacMemcpy, sceKernelMemcpy, and
the four replaced memcpy/memmove variants. It sits outside their "is either
side VRAM" gate, since a RAM-to-RAM copy of a frame is still a frame.

Being a video texture also gets it forced linear filtering and keeps it out of
texture upscaling, which is what you want for a video either way.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-22 16:10:28 -06:00
Henrik RydgårdandClaude Opus 5 aa7b1de28c Drop the texture replacer's "video" option
A `video = true` in textures.ini opted a pack into replacing and dumping video
textures. Both halves are a bad deal. Dumping writes a file per decoded frame,
which fills a disk rather than producing anything a pack can use, and replacing
means a hash lookup on content that is different every frame and will never be
found twice.

It was also the only reason the texture cache still hashed video textures at
all, so it cost every game that has ever played a cutscene, not just the packs
that set it.

Video textures are now never replaced and never dumped, and skipHash is simply
isVideo. An existing ini keeping the key is harmless - unknown options are
ignored.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-22 16:10:16 -06:00
Henrik RydgårdandClaude Opus 5 14bfd332dc Stop stacking duplicate video entries, and stop hashing video textures
NotifyWriteFormattedFromMemory appends to videos_ unconditionally. A game
blitting its decoded frame to the display buffer does that every displayed
frame while it waits for the next one to decode, so the same two or three
addresses come back over and over: Death Jr pushes 733 entries where there are
two distinct buffers, Tekken 6 around 53,000 where there are three. IsVideo()
walks that vector linearly on every texture. Refresh the matching entry instead
of appending a new one - Death Jr now holds 2 entries and Tekken 6 holds 18,
peak size 2 and 3.

The other half is the two TODOs that were already sitting there. A video
texture is new every frame by definition, so re-hashing it only confirms what
the VIDEO flag already said, and the secondary cache has nothing to offer a
frame that will never recur. Skip both, and with them the secondary lookup that
would otherwise key off a hash we no longer compute.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-22 16:10:00 -06:00
Henrik RydgårdandClaude Opus 5 f28896e1ed Vertex decoder: JIT morph on arm64
The arm64 vertex JIT had no morph support (the table entries were commented
out), so every morphed format fell back to the step functions. Add the same
morph steps x86 has: texcoords (plain and prescaled), normals, positions and
the four color formats.

Each step sums its frames the way its step function does as compiled: fused
where that's plain C++, which the compiler turns into fused multiply-adds,
and separate multiply and add where it's CrossSIMD. The unit test confirms
they match bit for bit. Morph formats decode about 2.5-3x faster than with
the steps.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 16:51:00 -06:00
Henrik RydgårdandClaude Opus 5 66b59c73a9 Vertex decoder: bring the riscv64 and loongarch64 JITs in line
Enables TestVertexJitMatchesSteps on both. riscv64 now passes all 216000
jitted formats and loongarch64 all 204000, against 28906 on arm64.

riscv64:
- Jit_PosFloat didn't clean NaN or infinity at all, it just copied the
  three words. Clamp to +-FLT_MAX like the x86 JIT does.
- Jit_PosFloatThrough was missing the truncation of Z to an integer.
- The morph helpers started the sum from the first product rather than
  from +0.0, which rounds differently and lets a -0 term through, and
  rounded that first product towards zero where the steps round to
  nearest. The rest of the sum stays fused, since the compiler contracts
  the steps into fused multiply-adds.
- The texcoord prescale and 5551 color morph paths read morph weights
  from tables that GetMorphValueUsage never asked to be filled in, so
  they used whatever an earlier vertex type had left there.
- The non-Zbb bounds update compared the wrong way around, so through
  mode texcoord bounds came out inverted.

loongarch64:
- Jit_PosFloat had a TODO to clean NaN and infinity, and didn't.
- Jit_PosFloatThrough was missing the same Z truncation.
- Jit_WriteMorphColor narrowed with the logical saturating shifts, so a
  negative channel became a huge unsigned value and saturated to 255
  instead of clamping to 0, and it rounded where the steps truncate. It
  also read the packed color back sign-extended, so any alpha above 0x7F
  compared as larger than 0xFF000000 and claimed full alpha.
- The three packed color morph formats are rewritten. The LSX versions
  built each channel with a chain of inserts, shifts and shuffles that
  didn't survive being run; 4444 also broadcast its scale from the mask
  register. They now follow the steps channel by channel. Note the
  accumulator has to be an LSX scratch register - F4-F7 alias V4-V7,
  which hold the skin matrix for the whole vertex.
- The vertex bounds were loaded with a signed halfword load, so the
  0xFFFF they start at became -1 and no texcoord was ever below it.

PrescaleUV now fuses on these two as well - both JITs fuse it, and the
compiler would have contracted the plain expression there anyway.

The test tolerates a small relative difference on decoded floats, scaled
by the morph count since each term rounds once. Every bug above was
orders of magnitude larger than that.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 13:48:56 -06:00
Henrik RydgårdandClaude Opus 5 c58baedd75 Vertex decoder: fix the jit-match test on MSVC
The test didn't compile: VERTS is captured by reference into testFormat, and
MSVC won't use a captured constexpr as an array bound. Make it static.

On arm64 it then failed on the UV prescale steps, by one ULP. The arm64 JIT and
the NEON handwritten decoders fuse the multiply-add, and the steps only match
that when the compiler contracts a * b + c - which clang does and MSVC doesn't,
in Debug or Release. Spell out which one happens instead of relying on it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 12:13:09 -06:00
Henrik RydgårdandClaude Opus 5 a1ed24ae2d Vertex decoder: fix the handwritten decoders, and test them too
The two handwritten SIMD decoders are used with or without the JIT, and had
drifted from the step functions:
- The God of War one ignored g_DoubleTextureCoordinates, giving HD Remaster
  games the wrong UVs, and passed NaN and infinite positions through. Those
  now come out finite like everywhere else.
- The GTA one expanded 5551 colors wrongly on NEON (a left shift where the
  SSE version shifts right).
- On ARM64, both now fuse the UV scale and offset, like the JIT and the steps
  as the compiler builds them.

Both are back in the unit test, which now also feeds NaN and infinity to plain
float positions and checks they come out finite, and runs only on x86-64 and
arm64 for now.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 11:29:40 -06:00
Henrik RydgårdandClaude Opus 5 63a6c11214 Vertex decoder: make the JIT match the step functions
Add a unit test that decodes every vertex format through both the step
functions and the JIT and requires identical output, side effects included.
Only skinning may differ by rounding, since arm64 accumulates the bone
matrices with fused multiply-adds.

What it found, and fixed:
- x86 morph colors rounded to nearest where the steps truncate, and applied
  the scale before the weight, which rounds differently.
- x86 through-mode u16 UV bounds compared signed, so texcoords above 32767
  scrambled the bounds for everyone running the x86 JIT.
- x86 morph sums could produce -0 where the steps produce +0.
- Step_NormalS16Morph scaled by 1/32768 twice, giving near-zero normals.
- Step_PosFloatThrough lost the truncation of Z to an integer that the JITs
  and the vertex reader did before it moved into the decoder.
- SetVertexType never reset skinInDecode.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 11:29:40 -06:00
Henrik RydgårdandClaude Opus 5 0b5a8f537e Decide on the vertex decoder JIT with CoreParameter, not the CPU core
The vertex decoder JIT was enabled only when g_Config.iCpuCore was one of the
JIT cores, which tied an unrelated GPU path to the CPU setting. Headless leaned
on that by forcing iCpuCore to the interpreter after ApplyToConfig(), to keep
the decoder JIT off in the tests.

Add CoreParameter::bUseVertexDecoderJit instead. Standalone and libretro set it
wherever the host can JIT, so IR interpreter users on such hosts now get the
decoder JIT too. Headless sets it false, which keeps the test results as they
were and lets --cpu go through ApplyToConfig() like every other option.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 10:31:58 -06:00
Henrik RydgårdandClaude Opus 5 19bac4894c Let sceMpeg LLE override the ForceHLEPsmf compat flag
Our psmf and psmfPlayer HLE plays video by calling our sceMpeg HLE, so it has
nothing to talk to when the real mpeg.prx is running. The flag now gives way
when sceMpeg is LLE.

Moved below the force-enable and unavailable masks so it tests what sceMpeg
actually ended up as rather than what was asked for.

Also remove a bad assert.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 15:15:45 -06:00
Acts1631 90f7c6d195 Rely on PNG header dimension bounds
PNG replacement dimensions are validated by PNGHeaderPeek before the

decoded buffer is allocated, so the additional size_t overflow checks are

redundant.
2026-09-07 14:52:06 -04:00
Acts1631 9ad3b821d9 Bound replacement PNG dimensions safely
Reject non-positive or oversized dimensions independently in the PNG

header peeker, matching the existing 8192 pixel decode limit. Use

checked size_t arithmetic before allocating replacement RGBA data so

crafted dimensions cannot overflow the allocation size.
2026-09-05 16:24:24 -04:00
Henrik Rydgård 2de11efe4e TextureCache: stop clamping DXT decoding to bufw either
DecodeDXTBlocks limited its x loop to min(bufw, w), so when w > bufw everything
from bufw to w was simply never written - the destination is sized for w, so
those columns kept whatever the buffer held.

The software sampler doesn't do that. It addresses a block as
(v >> 2) * (texbufw >> 2) + (u >> 2), which for u past bufw runs on into the
next row's blocks, the same way the linear formats run into the next row. So the
two renderers disagreed on the same texture.

Follow the sampler: decode w texels per row and let the block index carry on,
with the range check widened to cover the blocks that reach past the last row's
stride. Shares the SourceExtent helper from the previous commit, counting 4x4
blocks rather than texels.

Rendering-visible where w > bufw for a DXT texture, which is the case that used
to leave stale contents behind.
2026-09-05 11:25:55 -06:00
Henrik Rydgård ca9d76fef3 TextureCache: bound the source by w as well as bufw, instead of clamping w
The non-DXT decode paths read w texels per row from a source whose stride, range
check and unswizzle buffer were all sized from bufw alone. w and bufw are
independent GE registers, so w > bufw is reachable, and the last row then runs
off the end of both the validated guest range and the temp buffer.

The previous commit clamped w down to bufw, which stops the overrun but is the
wrong shape twice over: the game asked for w texels and the destination is sized
for w, so the tail of every row is left holding whatever was in the buffer, and
it silently narrows a texture the hardware would have decoded in full.

Size the checks from both instead. The extent a decode touches is bufw per row
plus however far the last row reaches past its own stride - the same adjustment
TextureReplacer::ComputeHash and the GE recorder already make - so the range
check, the height it falls back to when the range is short, and the unswizzle
buffer are all computed that way now.

Pull the five copies of "resize tmpTexBuf32_, unswizzle into it" into a helper
while we're here, since they all need the same sizing. It zeroes the buffer when
w reaches past bufw, as UnswizzleFromMem only fills the stride and the tail would
otherwise be stale heap.

ComputeTextureHash has the same bufw-only assumption, but it's left alone
deliberately - changing what goes into a texture hash isn't worth the risk here.

DXT is left alone too - it clamps to minw, but there the limit is the block index
within a bufw/4 block row, so its range check already covers what it reads.
2026-09-05 10:26:14 -06:00
Henrik RydgårdandClaude Opus 5 ed5d9813bb TextureCache: fix two host-memory overruns from GE texture state
PrepareBuildTexture's mip scan stopped at the first level with a dimension of 1
*before* running the mip size check for that level, so a level like 1x256 under a
256x256 level 0 was accepted as a valid mip. The backend then sized the level from
the halved level-0 dimensions (128x128) while LoadTextureLevel re-read the real
height (256) from the GE state, writing past the allocation. Do the check first.

That check was also gated on GPU_USE_SAMPLER_LOD_CONTROL, so backends without it
did no mip dimension validation at all - there's no reason for the guard, the
result only feeds badMipSizes, so drop it.

Separately, the non-DXT decode paths read w texels per row from a source whose
stride, range check and unswizzle buffer are all computed from bufw. w and bufw
are independent GE registers, so w > bufw is reachable and read past the end of
both the validated guest range and the temp buffer. Clamp w to bufw in
DecodeTextureLevel, which is what the DXT paths have always done via minw.

Note: ungating the mip check means backends without SAMPLER_LOD_CONTROL can now
set badMipSizes where they previously didn't, which collapses such textures to a
single level. That's the intended behavior, but it is a rendering-visible change
on those backends.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01DCPmm7FoQUoqrbMdhfqhQ2
2026-09-05 10:25:47 -06:00
Henrik Rydgård 89ebb8acc6 GLES: Actually apply anisotropic filtering
TextureCacheGLES passed a hardcoded 0.0f instead of key.aniso, so the
Anisotropic Filtering setting did nothing at all on the OpenGL backend, even
though GPU_USE_ANISOTROPY was advertised and D3D11/Vulkan both honor it. Looks
like it was left behind by the 2017 render manager refactor.

The queue runner now clamps to the device maximum it already queried into
maxAnisotropyLevel_ (until now unused), and only touches the parameter when the
extension is actually supported - the anisotropy branch there has been dead
since every caller passed 0.0f, so this is the first time it runs.

0.0f keeps its meaning of "don't care" for the CLUT/fragment-test/thin3d
callers; the texture cache now passes 1.0f when the setting is off, so turning
it off takes effect on already-uploaded textures instead of only new ones.

TexCache: Never use anisotropic filtering for CLUT8-indexed textures

What gets sampled for those is palette indices, depalettized by the shader
afterwards - averaging indices across an anisotropic footprint produces garbage
colors. Affects all backends, not just the GL one that just started honoring
key.aniso.

TexCache: Clear key.aniso wherever filtering is forced to nearest

It was only cleared in the two places inside the AUTO_MAX_QUALITY branch, so the
TEX_FILTER_AUTO path (pixel-mapped textures, the ugly color test heuristic), the
FORCE_NEAREST setting and the replacement-texture override could all end up
requesting nearest filtering with anisotropy still on.

Doing it in the switch that applies forceFiltering covers every path, so it
can't drift apart again.

GLES: Only record the applied anisotropy, and log skipped draws

The queue runner updated tex->anisotropy even when it skipped the call because
the value was 0.0f ("don't care") - harmless while nothing ever set anisotropy,
but now it would make the tracked state disagree with GL, so a later request for
the value it thinks is set would be wrongly skipped.

Also log when a draw is skipped for a missing vertex shader. The failure is
cached per shader ID, so without it geometry silently disappears for the rest of
the session after the one-shot OSD message.
2026-09-04 10:41:15 -06:00
Henrik Rydgård 293cf0c78c Revert the logic changes from "Add new TexCache logging channel"
This partially reverts commit ee1314f803.
2026-09-02 15:56:10 -06:00
Henrik Rydgård d4b6965db0 TestBoundingBox: bail out when the indices reach past the scratch space
corners and verts are carved out of decoded_ at fixed offsets 6*65536 bytes apart,
and NormalizeVertices fills corners with indexUpperBound - indexLowerBound + 1
SimpleVertex. The vertexCount > 1024 guard doesn't bound that: the index values come
from the game, so 1024 indices can span the full 16-bit range and run corners into
verts, making the cull decision from overwritten data. Bail on an index over 1024
and report visible - a bbox test that large isn't worth doing anyway.
2026-09-02 18:06:47 +02:00
Henrik Rydgård 506cfb6701 Bound two values that texture packs and shader inis supply
A hashrange of 'addr,w,h = 0,0' passed validation (0 isn't bigger than the source),
became desc_.newW/newH, and ReplacedTexture::Prepare divides by them. A post-shader
SSAA level multiplies the render resolution with no upper bound, while the
texture-shader Scale sitting a few lines away is checked against 2..8.
2026-09-02 18:06:47 +02:00
Henrik Rydgård 56bba5f6f5 Merge pull request #22189 from hrydgard/debug-input-rewind-overflows
Fix three more out-of-bounds writes (minor)
2026-08-31 17:53:10 +02:00
Henrik Rydgård 9cb50459d6 Merge pull request #22186 from hrydgard/medium-correctness-fixes
Medium-severity correctness fixes
2026-08-31 16:30:53 +02:00
Henrik Rydgård 1ee9168711 Merge pull request #22184 from hrydgard/fix-adreno-workaround-regression
Fix fragment shader logic error (Adreno stencil driver bug workaround)
2026-08-31 13:30:19 +02:00