- Delete TexCacheEntry objects dropped on rehash instead of leaking them.
- Don't leave a released null entry in cache_ when the framebuffer match
returns before the slot is refilled.
- Reset clutRenderAddress_ in Clear(), which releases the dynamic CLUT FBOs.
- Don't cache a null texture in drawPixelsCache_ when creation fails.
- Fix the reversed subtraction in the failed-FBO retry check.
- Remove the never-taken buffered-rendering early-out in UpdateRenderSize.
Taking it would leave existing VFBs without an fbo.
- Include smoothedDepal in the depal shader cache key, and print/parse the
debug IDs as 64-bit.
- Release depal pipelines through Draw2DPipeline::Release so the shader
source isn't leaked, and make that null-safe.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
GLSLtoSPV takes an optional SPIRVCache, keyed on a 32-bit hash of the
source, stage and variant, plus the source length. A changed shader
simply misses. thin3d's shaders and the other fixed ones use a global
cache in PSP/SYSTEM/CACHE/vulkan_spirv.cache, loaded on first use and
saved after graphics init, when a game's cache is saved, and at
shutdown; it's flushed once it reaches 32 entries, about twice what a
session compiles, so outdated ones don't pile up. Game shaders keep
theirs in the .vkshadercache, ahead of the shader IDs so that the
compiles on load find it (version 60), and only what the session used
is saved.
A cold glslang costs about 40ms before its first shader here, and
0.3-0.9ms per shader after that.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Neither is serialized, and both went stale on load. The ME busy time was
measured against the pre-load clock, so loading an earlier state made the
next SAS/codec job wait until the old time came around, freezing the game.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Some DoState code meant for after a load ran on every save:
- scePower reset the bus frequency a game set (and with a locked CPU
speed, applied the current setting to the clock).
- sceDisplay reset the lag sync baseline, and could schedule lag sync in
the measuring pass only, which failed the save.
- GPUState dirtied the texture, sceUmd notified the UI, and sceMpeg
dropped a pending ringbuffer fix-up for an old state.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The blit rates were measured with nothing else running. In a game, threads
waking up and SAS mixing on the Media Engine compete with the GE for main RAM:
Star Wars: Lethal Alliance's movie blit takes 8.65ms alone and 10.3ms in the
game. We don't model that load, so RAM texture fetches get a fixed 1.17x for a
typical one. With it, that game's long movie plays at 30 fps as on hardware.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Full-screen clears measured on a PSP (pspautotests gpu/timing/blittiming):
0.49ms on a 16-bit framebuffer whatever is cleared, 0.69ms on 8888, 1.02ms on
8888 with depth. Charging them may help games that spin hard on an empty
screen, but it's off (chargeClearTime) until tried on some. The video blit
cost moves into the same function, now EstimateFillCycles.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The blit cost remembered only the last buffer a decoder wrote into, forever.
Move the texture cache's video list (with its ageing out a few flips after
the last write) into GPUCommon, so the texture cache, the blit cost and
SoftGPU all share one. That also counts both of a double-buffered player's
frames, which exposed that a clear drawn with texturing still enabled was
being charged as a blit - skip clears and draws without texture coordinates.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Measured on a PSP (pspautotests gpu/timing/blittiming), a full-screen blit
from an unswizzled texture costs what the texture fetch costs: 16-bit formats
half of 32-bit, VRAM a fifth of RAM, and rectangles wider than ~128 texels
~7.5x as much as narrow strips, from texture cache thrashing. Framebuffer
format, filtering and blending don't matter. Ys I & II draws its movie as one
full-width sprite from a 565 texture in RAM, which takes 33ms - that, not the
decode, is what holds it to 30 fps.
All of the ME and GE costs speed up with the clock (2/3 as long at 333/166),
since the whole system runs from the one PLL.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Movie players like the one in Star Wars: Lethal Alliance present every decoded
frame after a single vblank wait, with no clock or timestamp check, so the
frame rate depends on decode, CSC, ATRAC decode and the GE blit adding up to
more than a vblank. We charged nearly nothing for any of them, so such movies
ran at 60 fps until the ringbuffer's slack ran out.
Costs measured on a PSP with a copy of that player (pspautotests
video/mpeg/playertiming), for a 480x272 frame:
- sceVideocodecDecode: 3.4ms (sceMpegAvcDecode 5.8ms less sceMpegAvcCsc 2.4ms)
- sceMpegBaseCscAvc: 2.4ms, was a flat 4ms
- sceAudiocodecDecode, ATRAC3+ only: 2.5ms per frame
- GE: 9.7ms for a through-mode rectangle blit from a decoded video frame,
charged by area, only for textures in the buffer a decoder last wrote.
GE time also now carries across stall address updates. Before, a list sent
in stalled chunks only had its last chunk's time counted, so sceGeDrawSync
returned 39us after a blit that takes 9.7ms. This affects every game that
builds its lists incrementally, so GE-timing-sensitive games need checking.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
videos_ only learns about the CSC output, so the texture we actually draw is
just an ordinary 512x512 8888 texture whose contents happen to be different
every frame: hash, miss, throw it in the secondary cache, rebuild, forever.
So track the copy. NotifyVideoCopy marks the destination as video when the
source is, and the copy funnels call it: sceDmacMemcpy, sceKernelMemcpy, and
the four replaced memcpy/memmove variants. It sits outside their "is either
side VRAM" gate, since a RAM-to-RAM copy of a frame is still a frame.
Being a video texture also gets it forced linear filtering and keeps it out of
texture upscaling, which is what you want for a video either way.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
A `video = true` in textures.ini opted a pack into replacing and dumping video
textures. Both halves are a bad deal. Dumping writes a file per decoded frame,
which fills a disk rather than producing anything a pack can use, and replacing
means a hash lookup on content that is different every frame and will never be
found twice.
It was also the only reason the texture cache still hashed video textures at
all, so it cost every game that has ever played a cutscene, not just the packs
that set it.
Video textures are now never replaced and never dumped, and skipHash is simply
isVideo. An existing ini keeping the key is harmless - unknown options are
ignored.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
NotifyWriteFormattedFromMemory appends to videos_ unconditionally. A game
blitting its decoded frame to the display buffer does that every displayed
frame while it waits for the next one to decode, so the same two or three
addresses come back over and over: Death Jr pushes 733 entries where there are
two distinct buffers, Tekken 6 around 53,000 where there are three. IsVideo()
walks that vector linearly on every texture. Refresh the matching entry instead
of appending a new one - Death Jr now holds 2 entries and Tekken 6 holds 18,
peak size 2 and 3.
The other half is the two TODOs that were already sitting there. A video
texture is new every frame by definition, so re-hashing it only confirms what
the VIDEO flag already said, and the secondary cache has nothing to offer a
frame that will never recur. Skip both, and with them the secondary lookup that
would otherwise key off a hash we no longer compute.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The arm64 vertex JIT had no morph support (the table entries were commented
out), so every morphed format fell back to the step functions. Add the same
morph steps x86 has: texcoords (plain and prescaled), normals, positions and
the four color formats.
Each step sums its frames the way its step function does as compiled: fused
where that's plain C++, which the compiler turns into fused multiply-adds,
and separate multiply and add where it's CrossSIMD. The unit test confirms
they match bit for bit. Morph formats decode about 2.5-3x faster than with
the steps.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Enables TestVertexJitMatchesSteps on both. riscv64 now passes all 216000
jitted formats and loongarch64 all 204000, against 28906 on arm64.
riscv64:
- Jit_PosFloat didn't clean NaN or infinity at all, it just copied the
three words. Clamp to +-FLT_MAX like the x86 JIT does.
- Jit_PosFloatThrough was missing the truncation of Z to an integer.
- The morph helpers started the sum from the first product rather than
from +0.0, which rounds differently and lets a -0 term through, and
rounded that first product towards zero where the steps round to
nearest. The rest of the sum stays fused, since the compiler contracts
the steps into fused multiply-adds.
- The texcoord prescale and 5551 color morph paths read morph weights
from tables that GetMorphValueUsage never asked to be filled in, so
they used whatever an earlier vertex type had left there.
- The non-Zbb bounds update compared the wrong way around, so through
mode texcoord bounds came out inverted.
loongarch64:
- Jit_PosFloat had a TODO to clean NaN and infinity, and didn't.
- Jit_PosFloatThrough was missing the same Z truncation.
- Jit_WriteMorphColor narrowed with the logical saturating shifts, so a
negative channel became a huge unsigned value and saturated to 255
instead of clamping to 0, and it rounded where the steps truncate. It
also read the packed color back sign-extended, so any alpha above 0x7F
compared as larger than 0xFF000000 and claimed full alpha.
- The three packed color morph formats are rewritten. The LSX versions
built each channel with a chain of inserts, shifts and shuffles that
didn't survive being run; 4444 also broadcast its scale from the mask
register. They now follow the steps channel by channel. Note the
accumulator has to be an LSX scratch register - F4-F7 alias V4-V7,
which hold the skin matrix for the whole vertex.
- The vertex bounds were loaded with a signed halfword load, so the
0xFFFF they start at became -1 and no texcoord was ever below it.
PrescaleUV now fuses on these two as well - both JITs fuse it, and the
compiler would have contracted the plain expression there anyway.
The test tolerates a small relative difference on decoded floats, scaled
by the morph count since each term rounds once. Every bug above was
orders of magnitude larger than that.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
From another pass over ge.prx against our code, each checked on a PSP with
gpu/ge/callbackstate except the last:
- A list's context is restored after its finish callback, which sees the
state the list left. We restored at the FINISH, before it. Now that nothing
runs until InterruptEnd(), that's where it happens.
- sceGeSaveContext/RestoreContext only fail while the GE is executing. It's
stopped during a finish callback and a SUSPEND signal callback, however
much is queued, so they work there. We said busy whenever a list existed.
- sceGeListDeQueue emptying the queue doesn't turn completed lists into
nothing, only sceGeDrawSync does. CheckDrawSync() is gone.
- The "break in progress" flag that makes sceGeContinue only requeue the
list is cleared by an interrupt that follows the break at once, so it's
only seen from a callback or with interrupts off. Ours lasted until the
next GE interrupt of any kind.
- sceKernelLoadExec restarts the GE driver, which begins by zeroing every
register and matrix. Reinitialize() now does too, so a program started that
way finds the same GE as one booted directly, rather than its launcher's.
Not testable on hardware: nothing after the restart can report back.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Reinitialize() wiped the 64 display lists but kept the queue of their ids.
Whatever the old executable still had queued came back as lists with no state
and a pc of 0, behind the first list of the new executable, where they
blocked everything. Crazy Taxi: Fare Wars is a launcher for its two games,
and stopped at a black screen that way. Fixes#19894.
This removes the workaround for it, which dropped such a list but returned
before currentList was cleared, and only worked as long as something else
happened to clear it later. A list with a bad pc is now dropped like one that
ran into an error, instead of sitting at the head of the queue for good.
Also narrows what sceGeBreak(1) throws away to interrupts that have actually
been raised, which is what gpu/ge/intrsuspend shows. The ones we haven't
raised yet are only late because we execute lists ahead of time: a game that
breaks right after its last list, and then waits for what the finish callback
signals, got that callback long ago on hardware.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Resetting the GE also gets rid of an interrupt that was raised but not taken
yet, so a list that reached its FINISH just before never gets its finish
callback. We delivered one anyway, for a list that no longer existed. If the
break comes from inside a GE callback, the interrupt being handled is kept,
since its handler still has to return.
Found by gpu/ge/intrsuspend, which also confirms from a thread, with
interrupts suspended, that nothing moves along the queue until the FINISH
interrupt has been taken.
Savestates: bump GPUCommon to 7. We didn't use to mark a PAUSE signal as
delivered, which sceGeContinue now goes by, so a state saved with a list
paused that way would load into a game that could never continue it. Fixed
up on load.
gpu/signals/handlercalls goes in as known failing: with an old SDK version, a
stall address set from inside a SUSPEND callback doesn't reach the GE, which
we can't express with just the one stall address per list. See docs/sceGe.md.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
On hardware the GE stops at every SIGNAL and FINISH, and it's the interrupt
that gets it going again: on the same list after a signal, on the next one
after a FINISH - once the finish callback has run, with the finished list
still at the head of the queue. We ran the next list right away and dropped
the finished one at once, so a finish callback saw an empty queue. A list
enqueued from there was started instead of queued, and then couldn't be
dequeued, which hung the new gpu/ge/queue2 test.
ProcessDLQueue() now runs nothing while the head of the queue has an
interrupt pending, and InterruptEnd() is what takes a finished list off the
queue. This also keeps a stall update from restarting a list that's stopped
at a signal before the handler has run. drawCompleteTicks is still set when
the last list reaches its FINISH, so a sceGeDrawSync in between doesn't wait.
Other things gpu/ge/queue2 and gpu/ge/breakwait showed, all from a real PSP:
- sceGeListEnQueue compares against the address a list was enqueued with
(or stopped at by sceGeBreak), mirrors included, not against its current pc.
We had that the wrong way around.
- The stack-in-use check only applies to lists that have started executing.
This is probably what IgnoreEnqueue was added for (Metal Gear Acid 2,
#10906). The flag stays until someone has checked the game without it.
- A PAUSE signal makes the list PAUSED at once, before the FINISH delivers it.
In between, sceGeContinue and sceGeBreak say BUSY, and updating the stall
address does nothing, so a list that stalls there is stuck.
- A completed list can't be dequeued, with or without a context.
- sceGeDrawSync(1) looked at currentList instead of the list it had found.
- sceGeBreak(1) doesn't wake anyone, and a late interrupt for a list it reset
no longer marks that list completed. Threads in sceGeDrawSync are woken
before the ones waiting for the last list.
Also fixes currentList being lost when loading a state where it's list 0,
and makes ge_pending_cb a plain std::list - nothing else touches it, and the
GPU thread it was shared with is long gone. Same savestate format.
See docs/sceGe.md.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
The test didn't compile: VERTS is captured by reference into testFormat, and
MSVC won't use a captured constexpr as an array bound. Make it static.
On arm64 it then failed on the UV prescale steps, by one ULP. The arm64 JIT and
the NEON handwritten decoders fuse the multiply-add, and the steps only match
that when the compiler contracts a * b + c - which clang does and MSVC doesn't,
in Debug or Release. Spell out which one happens instead of relying on it.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The two handwritten SIMD decoders are used with or without the JIT, and had
drifted from the step functions:
- The God of War one ignored g_DoubleTextureCoordinates, giving HD Remaster
games the wrong UVs, and passed NaN and infinite positions through. Those
now come out finite like everywhere else.
- The GTA one expanded 5551 colors wrongly on NEON (a left shift where the
SSE version shifts right).
- On ARM64, both now fuse the UV scale and offset, like the JIT and the steps
as the compiler builds them.
Both are back in the unit test, which now also feeds NaN and infinity to plain
float positions and checks they come out finite, and runs only on x86-64 and
arm64 for now.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Add a unit test that decodes every vertex format through both the step
functions and the JIT and requires identical output, side effects included.
Only skinning may differ by rounding, since arm64 accumulates the bone
matrices with fused multiply-adds.
What it found, and fixed:
- x86 morph colors rounded to nearest where the steps truncate, and applied
the scale before the weight, which rounds differently.
- x86 through-mode u16 UV bounds compared signed, so texcoords above 32767
scrambled the bounds for everyone running the x86 JIT.
- x86 morph sums could produce -0 where the steps produce +0.
- Step_NormalS16Morph scaled by 1/32768 twice, giving near-zero normals.
- Step_PosFloatThrough lost the truncation of Z to an integer that the JITs
and the vertex reader did before it moved into the decoder.
- SetVertexType never reset skinInDecode.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The vertex decoder JIT was enabled only when g_Config.iCpuCore was one of the
JIT cores, which tied an unrelated GPU path to the CPU setting. Headless leaned
on that by forcing iCpuCore to the interpreter after ApplyToConfig(), to keep
the decoder JIT off in the tests.
Add CoreParameter::bUseVertexDecoderJit instead. Standalone and libretro set it
wherever the host can JIT, so IR interpreter users on such hosts now get the
decoder JIT too. Headless sets it false, which keeps the test results as they
were and lets --cpu go through ApplyToConfig() like every other option.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Our psmf and psmfPlayer HLE plays video by calling our sceMpeg HLE, so it has
nothing to talk to when the real mpeg.prx is running. The flag now gives way
when sceMpeg is LLE.
Moved below the force-enable and unavailable masks so it tests what sceMpeg
actually ended up as rather than what was asked for.
Also remove a bad assert.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
DrawEngineGLES already knows that every index it generates is below the
decoded/transformed vertex count, but only passed the count to
glDrawElements. Drivers that need the vertex range before vertex
shading, like Mesa's Panfrost, then scan the index data on the CPU for
every draw, and their min/max caches can't help since the index push
buffer is rewritten every frame.
Only used for single-instance draws on desktop GL or GLES3, and only
when the caller provides the range, so other DrawIndexed callers are
unchanged.
Co-Authored-By: Claude Opus 5 <[email protected]>
PNG replacement dimensions are validated by PNGHeaderPeek before the
decoded buffer is allocated, so the additional size_t overflow checks are
redundant.
Reject non-positive or oversized dimensions independently in the PNG
header peeker, matching the existing 8192 pixel decode limit. Use
checked size_t arithmetic before allocating replacement RGBA data so
crafted dimensions cannot overflow the allocation size.
DecodeDXTBlocks limited its x loop to min(bufw, w), so when w > bufw everything
from bufw to w was simply never written - the destination is sized for w, so
those columns kept whatever the buffer held.
The software sampler doesn't do that. It addresses a block as
(v >> 2) * (texbufw >> 2) + (u >> 2), which for u past bufw runs on into the
next row's blocks, the same way the linear formats run into the next row. So the
two renderers disagreed on the same texture.
Follow the sampler: decode w texels per row and let the block index carry on,
with the range check widened to cover the blocks that reach past the last row's
stride. Shares the SourceExtent helper from the previous commit, counting 4x4
blocks rather than texels.
Rendering-visible where w > bufw for a DXT texture, which is the case that used
to leave stale contents behind.
The non-DXT decode paths read w texels per row from a source whose stride, range
check and unswizzle buffer were all sized from bufw alone. w and bufw are
independent GE registers, so w > bufw is reachable, and the last row then runs
off the end of both the validated guest range and the temp buffer.
The previous commit clamped w down to bufw, which stops the overrun but is the
wrong shape twice over: the game asked for w texels and the destination is sized
for w, so the tail of every row is left holding whatever was in the buffer, and
it silently narrows a texture the hardware would have decoded in full.
Size the checks from both instead. The extent a decode touches is bufw per row
plus however far the last row reaches past its own stride - the same adjustment
TextureReplacer::ComputeHash and the GE recorder already make - so the range
check, the height it falls back to when the range is short, and the unswizzle
buffer are all computed that way now.
Pull the five copies of "resize tmpTexBuf32_, unswizzle into it" into a helper
while we're here, since they all need the same sizing. It zeroes the buffer when
w reaches past bufw, as UnswizzleFromMem only fills the stride and the tail would
otherwise be stale heap.
ComputeTextureHash has the same bufw-only assumption, but it's left alone
deliberately - changing what goes into a texture hash isn't worth the risk here.
DXT is left alone too - it clamps to minw, but there the limit is the block index
within a bufw/4 block row, so its range check already covers what it reads.
PrepareBuildTexture's mip scan stopped at the first level with a dimension of 1
*before* running the mip size check for that level, so a level like 1x256 under a
256x256 level 0 was accepted as a valid mip. The backend then sized the level from
the halved level-0 dimensions (128x128) while LoadTextureLevel re-read the real
height (256) from the GE state, writing past the allocation. Do the check first.
That check was also gated on GPU_USE_SAMPLER_LOD_CONTROL, so backends without it
did no mip dimension validation at all - there's no reason for the guard, the
result only feeds badMipSizes, so drop it.
Separately, the non-DXT decode paths read w texels per row from a source whose
stride, range check and unswizzle buffer are all computed from bufw. w and bufw
are independent GE registers, so w > bufw is reachable and read past the end of
both the validated guest range and the temp buffer. Clamp w to bufw in
DecodeTextureLevel, which is what the DXT paths have always done via minw.
Note: ungating the mip check means backends without SAMPLER_LOD_CONTROL can now
set badMipSizes where they previously didn't, which collapses such textures to a
single level. That's the intended behavior, but it is a rendering-visible change
on those backends.
Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01DCPmm7FoQUoqrbMdhfqhQ2
The pipeline debug listing printed the color blend factors in the alpha slot,
which is doubly unhelpful since that branch is only taken when the alpha factors
differ from the defaults.
CompileShaderModuleAsync takes ownership of the tag but only deleted it on the
success path, leaking it whenever GLSLtoSPV failed.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01SGz54K3ZXa2Qzyc3aMEYyY
The failure branch inserted a null pipeline and then fell through to the normal
insert of the same key, which trips DenseHashMap's duplicate-key assert - and
_assert_msg_ is live in release builds, so a logged error became a crash.
Also skip the null entries when deleting cached pipelines.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01SGz54K3ZXa2Qzyc3aMEYyY
They were OR-ed into pipelineFlags just after the CreateGraphicsPipeline call that
consumes them, so the render manager's "don't compile a pipeline that requires
depth for a non-depth renderpass type" check could never fire for game pipelines.
thin3d_vulkan.cpp sets the flag before its call, which is the intended order.
Note this can now legitimately skip some variants when loading the shader cache -
those were invalid combinations that the check was written to reject.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01SGz54K3ZXa2Qzyc3aMEYyY
If reading the shader's constant buffer file failed, we'd skip writing descriptor
binding 4 but still dispatch the compute shader, which declares it - a statically
used but unwritten descriptor. It also re-read the missing file on every single
texture upload. Now we drop the scaling shaders instead, so following textures
take the CPU scaling path.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01SGz54K3ZXa2Qzyc3aMEYyY
The out-of-VRAM retry path cleared plan.replaced but left plan.doReplace set.
GetMipSize() dereferences plan.replaced when doReplace is true, so the fallback
crashed instead of recovering. The common code sets both together.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01SGz54K3ZXa2Qzyc3aMEYyY
The shader creation helpers returned S_FALSE when compilation produced no
bytecode - but S_FALSE is a success code, so the FAILED() checks in the
D3D11VertexShader/D3D11FragmentShader constructors never fired and failed_ was
never set. Return E_FAIL instead, and only hand out the bytecode when the shader
object was actually created.
Failed() had no callers at all, so a failed shader was handed to the draw as
usual: VSSetShader(nullptr), CreateInputLayout on empty bytecode, and the
ignored HRESULT from SetupDecFmtForDraw meant IASetInputLayout(nullptr). The
draw then did nothing, with no log line to explain it. Skip the draw and warn
instead, like the GL backend does.
The cleanup loop was hardcoded to 12 entries while the array holds more and
levels can exceed that with a texture replacement pack - a 8192-pixel
replacement has 14 mip levels, so every level from 12 up leaked on each texture
build. Use ARRAY_SIZE, and shrink the array to 16 since D3D11 caps textures at
16384 pixels anyway.
The out-of-memory bail-out returned before that loop, leaking every level
decoded so far. Since the entry ends up without a texture it gets rebuilt, and
leaks again, every following frame.
Scissor the stencil readback to the region actually being read back - latent,
every caller passes a zero origin today.
Remove a DecodeVerts call that can never do anything: both branches above it
have already advanced decodeVertsCounter_ to numDrawVerts_. Worse than useless,
since in the non-skinning branch the vertices went to the push buffer, so
decoded_ doesn't hold them.
The out-of-memory bail-out added in the previous commit returned before the
status flag was set, leaving a GL_TEXTURE_3D object bound while ApplyTexture
told the shader generator it was a 2D texture. The entry stays cached, so it
would repeat every frame, not just the one that failed to allocate.
TextureCacheGLES passed a hardcoded 0.0f instead of key.aniso, so the
Anisotropic Filtering setting did nothing at all on the OpenGL backend, even
though GPU_USE_ANISOTROPY was advertised and D3D11/Vulkan both honor it. Looks
like it was left behind by the 2017 render manager refactor.
The queue runner now clamps to the device maximum it already queried into
maxAnisotropyLevel_ (until now unused), and only touches the parameter when the
extension is actually supported - the anisotropy branch there has been dead
since every caller passed 0.0f, so this is the first time it runs.
0.0f keeps its meaning of "don't care" for the CLUT/fragment-test/thin3d
callers; the texture cache now passes 1.0f when the setting is off, so turning
it off takes effect on already-uploaded textures instead of only new ones.
TexCache: Never use anisotropic filtering for CLUT8-indexed textures
What gets sampled for those is palette indices, depalettized by the shader
afterwards - averaging indices across an anisotropic footprint produces garbage
colors. Affects all backends, not just the GL one that just started honoring
key.aniso.
TexCache: Clear key.aniso wherever filtering is forced to nearest
It was only cleared in the two places inside the AUTO_MAX_QUALITY branch, so the
TEX_FILTER_AUTO path (pixel-mapped textures, the ugly color test heuristic), the
FORCE_NEAREST setting and the replacement-texture override could all end up
requesting nearest filtering with anisotropy still on.
Doing it in the switch that applies forceFiltering covers every path, so it
can't drift apart again.
GLES: Only record the applied anisotropy, and log skipped draws
The queue runner updated tex->anisotropy even when it skipped the call because
the value was 0.0f ("don't care") - harmless while nothing ever set anisotropy,
but now it would make the tracked state disagree with GL, so a later request for
the value it thinks is set would be wrongly skipped.
Also log when a draw is skipped for a missing vertex shader. The failure is
cached per shader ID, so without it geometry silently disappears for the rest of
the session after the one-shot OSD message.
The cache loader indexed &vec[0] on vectors that can legitimately be empty (a
header-sized file with zero counts passes both sanity checks), and the counts
are signed ints where only the upper bound was checked - a negative count would
reach resize() as a huge size_t.
The 3D texture branch had the out-of-memory assert but not the bail-out the 2D
branch has, so an ignored assert fell straight into memset(nullptr).