Read as an integer, a float's bits are its log2 with the mantissa a
straight line between powers of two, scaled by 2^23, and writing an
integer back is the matching exp2. That's exactly the GE's
approximation, without log2/exp2/floor. pspPow now also returns 1 for
e <= 0 itself, so the callers drop their checks. Shader languages
without integers fall back to a true pow.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Environment map S and T are (N.L + 1) / 2 with L the light's vector as
lighting uses it: from the vertex to the light for point and spot lights,
a zero vector staying zero, and the half vector for a light that does
specular. Whether lighting or the light is enabled still doesn't matter
(gpu/lighting/shademap).
The vertex shader ID now carries the type and computation of the shade
mapping lights (the ubershader reads them from u_lightControl), so both
shader caches get a new version.
Fixes the hair shine in iDOLM@STER SP (#12376).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The viewer is at infinity along view space +z, so in world space, where
lighting happens, it's the view matrix's third column rather than
(0,0,1): turning the camera moves the highlights (gpu/lighting/specular).
The shaders read it from u_view.
In a Need for Speed Carbon frame replayed on a PSP, this and the pow
bring the error of the cars from MSE 957 to 120 (Vulkan).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The shaders and the software renderer only add specular when N.L >= 0,
as the GE does (gpu/lighting/specular); the CPU lighter, used for points,
lines and rectangles, didn't check.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The GE computes these powers as exp2(e * log2(x)), with log2 and exp2
each a straight line between powers of two (Mitchell's approximation),
and only uses the top 4 bits of the specular coefficient's mantissa.
Through a highlight's falloff a true pow is 10-30 steps of 255 brighter
at the exponents games use. Measured in gpu/lighting/specular.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With animated control points the pole is only nearly degenerate, so the
vanishing derivative is rounding noise rather than exactly zero, and the
absolute threshold missed it. The resulting random normals still showed
as dark patches on Pac-Man Arrangement's ghosts (#12354).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Software transform expands each point, line and rectangle to four
vertices, and gave up on the whole draw when that didn't fit
VERTEX_BUFFER_MAX. Batching only counted input vertices, so a batch over
16384 points (or 32768 line or rectangle vertices) vanished silently,
whether it came from one PRIM or several merged ones. Count the expanded
vertices when batching, and submit a PRIM too big on its own in parts.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With only one patch, the open-last-edge adjustment assumed the first edge
was closed, so one knot interval came out as 2 instead of 1, and the
patch wasn't the Bezier patch a fully clamped cubic is. Positions were off
by up to about 7% of the patch.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The GLES sampler uniforms and texture slots for the control points and
weights, the Vulkan storage buffer bindings, and the u_spline_counts
uniform, which becomes padding (the C++ side already was).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Where all the control points along a patch edge meet at one point, like
the top of a dome, one derivative is zero and so is the cross product,
and normalizing it gave NaN. Use the limit instead, built from the mixed
second derivative. Fixes the dark spots on the ghosts' heads in Pac-Man
Arrangement (#12354).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Reloading the ini clears the per-key lookup caches, which could keep
saying "no replacement" for textures the new ini replaces.
- With ignoreAddress, hash ranges were skipped when sizing the
replacement, though ComputeHash applies them. cache_ is now keyed by
the full key, so its lookups hit too.
- Textures sharing files but differing in size, hash range or filtering
no longer share one ReplacedTexture (the first one's settings won).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
They were uploaded but never drawn, and the block transfer hack drew
whatever source was left over. Draw them straight into the backbuffer
pass, without post shaders, which would need to bind their own targets.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The draw engine is created on the loader thread while the UI thread may
already be rendering and invoking the callback.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With xxhash 0.8.4, XXH3 is faster than StableQuickTexHash on ARM64
(about 37 vs 24 GB/s on a Snapdragon X), and its scalar path, which
RISC-V builds get, is about as fast as the quick hash's. It doesn't
collide the way the quick hash does (#8249).
Texture replacement still uses its own hash setting, so texture packs
are unaffected.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
These own GPU objects, memory or refcounts in their destructors (or assert
there that they were torn down), so a copy would double-free. Nothing copies
them today; this keeps it that way. The manager base classes cover every
backend's subclass.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Release the file reference when loading a level fails or finds nothing,
since only a loaded level takes ownership of it.
- Free the PNG image when the size changed since the header was read.
- Delete the VFS when a pack without an ini has no hash-named textures.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Put the anisotropy level in the sampler key, so changing it applies on
Vulkan and D3D11.
- Release CLUT textures at shutdown on GLES and D3D11.
- Fix the depth readback viewport, which squeezed the image whenever the
read rectangle was smaller than the fbo.
- Test the computed depth, not the unset result, in the equal-depth clear
check.
- Don't read back a CLUT from a framebuffer without an fbo.
- Tolerate null entries when releasing post-shader objects and CLUT
textures after a failed creation.
- ImGe: Don't crash on a framebuffer without an fbo, or on GetVFB under the
software renderer.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Flush before the queued draws would decode more than VERTEX_BUFFER_MAX
vertices. The batch was limited by index count, which doesn't bound a
sparse index range, and DecodeVerts silently stopped while DecodeInds
still emitted indices for the undecoded draws.
- Give TestBoundingBox its own scratch buffer. It used offsets in decoded_,
which can hold decoded vertices that aren't flushed yet.
- Read 32-bit indices the way the PSP does, ignoring the upper 16 bits.
IndexConverter and the fast bounding box test used all 32, so a game
setting them indexed far past the decoded vertices.
- D3D11: Flush in FinishDeferred like the other backends, since indices
are still read from PSP memory at flush time (#10095).
- Don't JIT new vertex decoders once the code space is full.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The OpenGL and Vulkan shader caches store raw shader IDs (and, for Vulkan,
pipeline keys) on disk. Add the rule to AGENTS.md and point to it from the
persisted types.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- DDS files with mips stored level 0's file reference in level 1 too, so it
was freed twice.
- The VFS was deleted on config changes (and ini reloads) while load tasks
still used it. Now each cached texture waits for its task and releases its
file references through the old VFS first, and reloads afterwards.
- A failed ini reload turns replacement off instead of leaving it on without
a VFS.
- Release file references on purge.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Delete TexCacheEntry objects dropped on rehash instead of leaking them.
- Don't leave a released null entry in cache_ when the framebuffer match
returns before the slot is refilled.
- Reset clutRenderAddress_ in Clear(), which releases the dynamic CLUT FBOs.
- Don't cache a null texture in drawPixelsCache_ when creation fails.
- Fix the reversed subtraction in the failed-FBO retry check.
- Remove the never-taken buffered-rendering early-out in UpdateRenderSize.
Taking it would leave existing VFBs without an fbo.
- Include smoothedDepal in the depal shader cache key, and print/parse the
debug IDs as 64-bit.
- Release depal pipelines through Draw2DPipeline::Release so the shader
source isn't leaked, and make that null-safe.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The blit cost remembered only the last buffer a decoder wrote into, forever.
Move the texture cache's video list (with its ageing out a few flips after
the last write) into GPUCommon, so the texture cache, the blit cost and
SoftGPU all share one. That also counts both of a double-buffered player's
frames, which exposed that a clear drawn with texturing still enabled was
being charged as a blit - skip clears and draws without texture coordinates.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
videos_ only learns about the CSC output, so the texture we actually draw is
just an ordinary 512x512 8888 texture whose contents happen to be different
every frame: hash, miss, throw it in the secondary cache, rebuild, forever.
So track the copy. NotifyVideoCopy marks the destination as video when the
source is, and the copy funnels call it: sceDmacMemcpy, sceKernelMemcpy, and
the four replaced memcpy/memmove variants. It sits outside their "is either
side VRAM" gate, since a RAM-to-RAM copy of a frame is still a frame.
Being a video texture also gets it forced linear filtering and keeps it out of
texture upscaling, which is what you want for a video either way.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
A `video = true` in textures.ini opted a pack into replacing and dumping video
textures. Both halves are a bad deal. Dumping writes a file per decoded frame,
which fills a disk rather than producing anything a pack can use, and replacing
means a hash lookup on content that is different every frame and will never be
found twice.
It was also the only reason the texture cache still hashed video textures at
all, so it cost every game that has ever played a cutscene, not just the packs
that set it.
Video textures are now never replaced and never dumped, and skipHash is simply
isVideo. An existing ini keeping the key is harmless - unknown options are
ignored.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
NotifyWriteFormattedFromMemory appends to videos_ unconditionally. A game
blitting its decoded frame to the display buffer does that every displayed
frame while it waits for the next one to decode, so the same two or three
addresses come back over and over: Death Jr pushes 733 entries where there are
two distinct buffers, Tekken 6 around 53,000 where there are three. IsVideo()
walks that vector linearly on every texture. Refresh the matching entry instead
of appending a new one - Death Jr now holds 2 entries and Tekken 6 holds 18,
peak size 2 and 3.
The other half is the two TODOs that were already sitting there. A video
texture is new every frame by definition, so re-hashing it only confirms what
the VIDEO flag already said, and the secondary cache has nothing to offer a
frame that will never recur. Skip both, and with them the secondary lookup that
would otherwise key off a hash we no longer compute.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The arm64 vertex JIT had no morph support (the table entries were commented
out), so every morphed format fell back to the step functions. Add the same
morph steps x86 has: texcoords (plain and prescaled), normals, positions and
the four color formats.
Each step sums its frames the way its step function does as compiled: fused
where that's plain C++, which the compiler turns into fused multiply-adds,
and separate multiply and add where it's CrossSIMD. The unit test confirms
they match bit for bit. Morph formats decode about 2.5-3x faster than with
the steps.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Enables TestVertexJitMatchesSteps on both. riscv64 now passes all 216000
jitted formats and loongarch64 all 204000, against 28906 on arm64.
riscv64:
- Jit_PosFloat didn't clean NaN or infinity at all, it just copied the
three words. Clamp to +-FLT_MAX like the x86 JIT does.
- Jit_PosFloatThrough was missing the truncation of Z to an integer.
- The morph helpers started the sum from the first product rather than
from +0.0, which rounds differently and lets a -0 term through, and
rounded that first product towards zero where the steps round to
nearest. The rest of the sum stays fused, since the compiler contracts
the steps into fused multiply-adds.
- The texcoord prescale and 5551 color morph paths read morph weights
from tables that GetMorphValueUsage never asked to be filled in, so
they used whatever an earlier vertex type had left there.
- The non-Zbb bounds update compared the wrong way around, so through
mode texcoord bounds came out inverted.
loongarch64:
- Jit_PosFloat had a TODO to clean NaN and infinity, and didn't.
- Jit_PosFloatThrough was missing the same Z truncation.
- Jit_WriteMorphColor narrowed with the logical saturating shifts, so a
negative channel became a huge unsigned value and saturated to 255
instead of clamping to 0, and it rounded where the steps truncate. It
also read the packed color back sign-extended, so any alpha above 0x7F
compared as larger than 0xFF000000 and claimed full alpha.
- The three packed color morph formats are rewritten. The LSX versions
built each channel with a chain of inserts, shifts and shuffles that
didn't survive being run; 4444 also broadcast its scale from the mask
register. They now follow the steps channel by channel. Note the
accumulator has to be an LSX scratch register - F4-F7 alias V4-V7,
which hold the skin matrix for the whole vertex.
- The vertex bounds were loaded with a signed halfword load, so the
0xFFFF they start at became -1 and no texcoord was ever below it.
PrescaleUV now fuses on these two as well - both JITs fuse it, and the
compiler would have contracted the plain expression there anyway.
The test tolerates a small relative difference on decoded floats, scaled
by the morph count since each term rounds once. Every bug above was
orders of magnitude larger than that.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The test didn't compile: VERTS is captured by reference into testFormat, and
MSVC won't use a captured constexpr as an array bound. Make it static.
On arm64 it then failed on the UV prescale steps, by one ULP. The arm64 JIT and
the NEON handwritten decoders fuse the multiply-add, and the steps only match
that when the compiler contracts a * b + c - which clang does and MSVC doesn't,
in Debug or Release. Spell out which one happens instead of relying on it.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The two handwritten SIMD decoders are used with or without the JIT, and had
drifted from the step functions:
- The God of War one ignored g_DoubleTextureCoordinates, giving HD Remaster
games the wrong UVs, and passed NaN and infinite positions through. Those
now come out finite like everywhere else.
- The GTA one expanded 5551 colors wrongly on NEON (a left shift where the
SSE version shifts right).
- On ARM64, both now fuse the UV scale and offset, like the JIT and the steps
as the compiler builds them.
Both are back in the unit test, which now also feeds NaN and infinity to plain
float positions and checks they come out finite, and runs only on x86-64 and
arm64 for now.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Add a unit test that decodes every vertex format through both the step
functions and the JIT and requires identical output, side effects included.
Only skinning may differ by rounding, since arm64 accumulates the bone
matrices with fused multiply-adds.
What it found, and fixed:
- x86 morph colors rounded to nearest where the steps truncate, and applied
the scale before the weight, which rounds differently.
- x86 through-mode u16 UV bounds compared signed, so texcoords above 32767
scrambled the bounds for everyone running the x86 JIT.
- x86 morph sums could produce -0 where the steps produce +0.
- Step_NormalS16Morph scaled by 1/32768 twice, giving near-zero normals.
- Step_PosFloatThrough lost the truncation of Z to an integer that the JITs
and the vertex reader did before it moved into the decoder.
- SetVertexType never reset skinInDecode.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The vertex decoder JIT was enabled only when g_Config.iCpuCore was one of the
JIT cores, which tied an unrelated GPU path to the CPU setting. Headless leaned
on that by forcing iCpuCore to the interpreter after ApplyToConfig(), to keep
the decoder JIT off in the tests.
Add CoreParameter::bUseVertexDecoderJit instead. Standalone and libretro set it
wherever the host can JIT, so IR interpreter users on such hosts now get the
decoder JIT too. Headless sets it false, which keeps the test results as they
were and lets --cpu go through ApplyToConfig() like every other option.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Our psmf and psmfPlayer HLE plays video by calling our sceMpeg HLE, so it has
nothing to talk to when the real mpeg.prx is running. The flag now gives way
when sceMpeg is LLE.
Moved below the force-enable and unavailable masks so it tests what sceMpeg
actually ended up as rather than what was asked for.
Also remove a bad assert.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
PNG replacement dimensions are validated by PNGHeaderPeek before the
decoded buffer is allocated, so the additional size_t overflow checks are
redundant.
Reject non-positive or oversized dimensions independently in the PNG
header peeker, matching the existing 8192 pixel decode limit. Use
checked size_t arithmetic before allocating replacement RGBA data so
crafted dimensions cannot overflow the allocation size.
DecodeDXTBlocks limited its x loop to min(bufw, w), so when w > bufw everything
from bufw to w was simply never written - the destination is sized for w, so
those columns kept whatever the buffer held.
The software sampler doesn't do that. It addresses a block as
(v >> 2) * (texbufw >> 2) + (u >> 2), which for u past bufw runs on into the
next row's blocks, the same way the linear formats run into the next row. So the
two renderers disagreed on the same texture.
Follow the sampler: decode w texels per row and let the block index carry on,
with the range check widened to cover the blocks that reach past the last row's
stride. Shares the SourceExtent helper from the previous commit, counting 4x4
blocks rather than texels.
Rendering-visible where w > bufw for a DXT texture, which is the case that used
to leave stale contents behind.
The non-DXT decode paths read w texels per row from a source whose stride, range
check and unswizzle buffer were all sized from bufw alone. w and bufw are
independent GE registers, so w > bufw is reachable, and the last row then runs
off the end of both the validated guest range and the temp buffer.
The previous commit clamped w down to bufw, which stops the overrun but is the
wrong shape twice over: the game asked for w texels and the destination is sized
for w, so the tail of every row is left holding whatever was in the buffer, and
it silently narrows a texture the hardware would have decoded in full.
Size the checks from both instead. The extent a decode touches is bufw per row
plus however far the last row reaches past its own stride - the same adjustment
TextureReplacer::ComputeHash and the GE recorder already make - so the range
check, the height it falls back to when the range is short, and the unswizzle
buffer are all computed that way now.
Pull the five copies of "resize tmpTexBuf32_, unswizzle into it" into a helper
while we're here, since they all need the same sizing. It zeroes the buffer when
w reaches past bufw, as UnswizzleFromMem only fills the stride and the tail would
otherwise be stale heap.
ComputeTextureHash has the same bufw-only assumption, but it's left alone
deliberately - changing what goes into a texture hash isn't worth the risk here.
DXT is left alone too - it clamps to minw, but there the limit is the block index
within a bufw/4 block row, so its range check already covers what it reads.
PrepareBuildTexture's mip scan stopped at the first level with a dimension of 1
*before* running the mip size check for that level, so a level like 1x256 under a
256x256 level 0 was accepted as a valid mip. The backend then sized the level from
the halved level-0 dimensions (128x128) while LoadTextureLevel re-read the real
height (256) from the GE state, writing past the allocation. Do the check first.
That check was also gated on GPU_USE_SAMPLER_LOD_CONTROL, so backends without it
did no mip dimension validation at all - there's no reason for the guard, the
result only feeds badMipSizes, so drop it.
Separately, the non-DXT decode paths read w texels per row from a source whose
stride, range check and unswizzle buffer are all computed from bufw. w and bufw
are independent GE registers, so w > bufw is reachable and read past the end of
both the validated guest range and the temp buffer. Clamp w to bufw in
DecodeTextureLevel, which is what the DXT paths have always done via minw.
Note: ungating the mip check means backends without SAMPLER_LOD_CONTROL can now
set badMipSizes where they previously didn't, which collapses such textures to a
single level. That's the intended behavior, but it is a rendering-visible change
on those backends.
Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01DCPmm7FoQUoqrbMdhfqhQ2
TextureCacheGLES passed a hardcoded 0.0f instead of key.aniso, so the
Anisotropic Filtering setting did nothing at all on the OpenGL backend, even
though GPU_USE_ANISOTROPY was advertised and D3D11/Vulkan both honor it. Looks
like it was left behind by the 2017 render manager refactor.
The queue runner now clamps to the device maximum it already queried into
maxAnisotropyLevel_ (until now unused), and only touches the parameter when the
extension is actually supported - the anisotropy branch there has been dead
since every caller passed 0.0f, so this is the first time it runs.
0.0f keeps its meaning of "don't care" for the CLUT/fragment-test/thin3d
callers; the texture cache now passes 1.0f when the setting is off, so turning
it off takes effect on already-uploaded textures instead of only new ones.
TexCache: Never use anisotropic filtering for CLUT8-indexed textures
What gets sampled for those is palette indices, depalettized by the shader
afterwards - averaging indices across an anisotropic footprint produces garbage
colors. Affects all backends, not just the GL one that just started honoring
key.aniso.
TexCache: Clear key.aniso wherever filtering is forced to nearest
It was only cleared in the two places inside the AUTO_MAX_QUALITY branch, so the
TEX_FILTER_AUTO path (pixel-mapped textures, the ugly color test heuristic), the
FORCE_NEAREST setting and the replacement-texture override could all end up
requesting nearest filtering with anisotropy still on.
Doing it in the switch that applies forceFiltering covers every path, so it
can't drift apart again.
GLES: Only record the applied anisotropy, and log skipped draws
The queue runner updated tex->anisotropy even when it skipped the call because
the value was 0.0f ("don't care") - harmless while nothing ever set anisotropy,
but now it would make the tracked state disagree with GL, so a later request for
the value it thinks is set would be wrongly skipped.
Also log when a draw is skipped for a missing vertex shader. The failure is
cached per shader ID, so without it geometry silently disappears for the rest of
the session after the one-shot OSD message.
corners and verts are carved out of decoded_ at fixed offsets 6*65536 bytes apart,
and NormalizeVertices fills corners with indexUpperBound - indexLowerBound + 1
SimpleVertex. The vertexCount > 1024 guard doesn't bound that: the index values come
from the game, so 1024 indices can span the full 16-bit range and run corners into
verts, making the cull decision from overwritten data. Bail on an index over 1024
and report visible - a bbox test that large isn't worth doing anyway.
A hashrange of 'addr,w,h = 0,0' passed validation (0 isn't bigger than the source),
became desc_.newW/newH, and ReplacedTexture::Prepare divides by them. A post-shader
SSAA level multiplies the render resolution with no upper bound, while the
texture-shader Scale sitting a few lines away is checked against 2..8.