Commit Graph
505 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 056fd05231 UnitTest: Make BlockAllocator run in about a second
Validating after every churn step was quadratic in the block count. Check
every 64 steps and at the end, and do fewer iterations.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Artem Lytkin a3e84adab9 CachingFileLoader: keep the partial last block of the file
since 7e6405291 only full 64 KB reads get cached, so any read touching the
shorter last block of a file came back short. count blocks with
blocks_.size() too, so failed reads can't inflate the count until
MakeCacheSpaceFor loops forever. also stop read-ahead from asking the
backend for blocks past the end of the file.
2026-09-27 15:31:04 +03:00
Henrik RydgårdandClaude Opus 5.5 7eb371b231 IR interpreter: Merge a conditional exit with the ExitToConst after it
Blocks ending in a branch dispatch a conditional exit and then the
fallthrough ExitToConst. One op now returns either target, reading the
second from the ExitToConst, which stays behind unexecuted.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 11:47:17 -06:00
Henrik RydgårdandClaude Opus 5.5 dab5d82227 IR: Forward stores and loads to later loads in the block
OptimizeLoadsAfterStores only dropped a load right after a store of the same
reg. Now a load of anything the block stored or loaded before becomes a reg
move (with the extension for 8/16-bit loads), as long as nothing in between
may have changed the memory, the address reg, or the reg holding the value.
Only a store through the same base at a disjoint range is known not to alias,
and constant addresses outside RAM are left alone.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:56:03 -06:00
Henrik RydgårdandClaude Opus 5.5 498511fd51 IR: Drop everything after a conditional exit folded to always taken
Only the ExitToConst right after it was skipped before.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:39:24 -06:00
Henrik RydgårdandClaude Opus 5.5 a99a82cc6b IR: Keep more constants known in PropagateConstants
- A constant stays known after it's written out for a read (by a store,
  MovZ, a multiply...), so later uses still fold. Whatever an op writes is
  forgotten after its inputs are written, and setting a reg to the value it
  already holds isn't written twice.
- A conditional exit that isn't taken keeps the constants known.
- The saturating and min/max FP ops, FSign and the 31-bit Vec2 pack/unpack
  no longer flush every GPR constant.
- A load through its own base is folded (lui v0, hi; lw v0, lo(v0)).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:35:11 -06:00
Henrik RydgårdandClaude Opus 5.5 25e8f1d57c IR: Fix lwl/lwr pairing into the base reg, and sp validation past barriers
- An lwl/lwr pair was combined into one load even when the first half loads
  into the base register, which changes the address of the second half.
- ApplyMemoryValidation shared one sp check across the block even past an
  Interpret or CallReplacement, which may change sp.
- Drop a duplicate FSqrt meta entry, and name Load8Ext correctly.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:35:11 -06:00
Henrik RydgårdandClaude Opus 5.5 1dbb62d570 IR: Fix PurgeTemps miscompiles
- Copy propagation through an FPR temp didn't stop when an instruction
  rewrote the temp in place (it compared an FPR number against the +32
  offset reg), so later reads lost that write.
- A read of the temp in both operands only had src1 replaced, yet the copy
  into the temp was still removed.
- The replacement matched operands by number without checking their type,
  so a StoreFloat whose GPR address had the temp's number got its address
  replaced (IRVTEMP_PFX_S and IRTEMP_0 are both 192).
- A write to lanes 1-3 of a Vec4 temp wasn't noticed.
- IRReadsFromFPRs stopped after the F operands, missing Vec4Scale's vector.
- Exits and barriers didn't count as reading everything, so a write to a
  real reg could be moved above an exit.
- Load32Linked and Store32Conditional were removed when their reg was
  overwritten unread, losing LLBIT and the store.

Also fixes an off-by-one in the vec src3 read check. The unit test now
reports every failing case instead of stopping at the first.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:35:11 -06:00
Henrik RydgårdandClaude Opus 5.5 7ecdbe7799 UnitTest: Remove the old asin and sin/cos approximation experiments
They only printed comparisons against libm and checked nothing. The VFPU
functions are exact now and have their own tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 12:24:43 -06:00
Henrik RydgårdandClaude Opus 5.5 d6b3d590eb VFPU: Compute log2, sin/cos and asin without correction tables, too
They use the same quadratic interpolator as rcp and friends, with three
twists. sin indexes the quarter wave from the top, and asin and sin work
in a per-segment exponent whose 4-ulp truncation also applies to results
in a lower binade. log2 truncates exponent + log2(1.m) toward zero to 22
significant bits, and where that step is coarser than 2^-24 the datapath
drops coefficient bits to match; that also covers the region just below
1.0 that needed a special case.

vfpu_sincos now reduces the angle once. With every table gone, so are
the asset folder, the loader, InitVFPU and the fallbacks for tables that
failed to load. All seven functions are bit-exact with the table-based
code over every 32-bit input.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 10:07:01 -06:00
Henrik RydgårdandClaude Opus 5.5 0ac12b7bf3 VFPUDot: Cover sums that round into the next power of two
A sum just below a power of two can round up into the next exponent, and
at the top of the range into inf; random inputs almost never land there.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 08:18:25 -06:00
Henrik RydgårdandClaude Opus 5.5 cd005b9041 VFPU: SSE2 version of the exact vdot
Same shape as the NEON one, which now shares its rounding tail. SSE2 has
no per-lane shift, so the alignment shift is a multiply by a power of two
built from float bits and converted by truncation; the unsigned maxima
use the 16-bit instructions, since every value involved fits in 15 bits.
Nothing depends on the host rounding mode or flush-to-zero.

Checked against the reference by VFPUDot and on 300M more inputs offline,
also with MXCSR set to round toward zero with FTZ and DAZ.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 08:18:25 -06:00
Henrik RydgårdandClaude Opus 5.5 649561764d VFPU: NEON version of the exact vdot
The four lanes are computed together: exponents, 24x24-bit products with
round-to-odd, alignment by truncation and a signed horizontal sum. One
pairwise maximum finds both the alignment exponent and any inf or NaN,
which go to the reference. The final rounding is branch-free and in
integers, since the host rounding mode may be the game's.

About three times the throughput of the reference on Apple M-series
(5.2 vs 15.2 ns per call). VFPUDot checks it against the reference on
four million inputs picked to cover cancellation, ties, subnormals and
the overflow edges; a billion more matched offline.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 08:18:25 -06:00
Henrik Rydgård 7a9a354e50 Add a MIPSState context pointer to the JitAt call 2026-09-21 17:34:10 -06:00
Henrik RydgårdandClaude Opus 5 66b59c73a9 Vertex decoder: bring the riscv64 and loongarch64 JITs in line
Enables TestVertexJitMatchesSteps on both. riscv64 now passes all 216000
jitted formats and loongarch64 all 204000, against 28906 on arm64.

riscv64:
- Jit_PosFloat didn't clean NaN or infinity at all, it just copied the
  three words. Clamp to +-FLT_MAX like the x86 JIT does.
- Jit_PosFloatThrough was missing the truncation of Z to an integer.
- The morph helpers started the sum from the first product rather than
  from +0.0, which rounds differently and lets a -0 term through, and
  rounded that first product towards zero where the steps round to
  nearest. The rest of the sum stays fused, since the compiler contracts
  the steps into fused multiply-adds.
- The texcoord prescale and 5551 color morph paths read morph weights
  from tables that GetMorphValueUsage never asked to be filled in, so
  they used whatever an earlier vertex type had left there.
- The non-Zbb bounds update compared the wrong way around, so through
  mode texcoord bounds came out inverted.

loongarch64:
- Jit_PosFloat had a TODO to clean NaN and infinity, and didn't.
- Jit_PosFloatThrough was missing the same Z truncation.
- Jit_WriteMorphColor narrowed with the logical saturating shifts, so a
  negative channel became a huge unsigned value and saturated to 255
  instead of clamping to 0, and it rounded where the steps truncate. It
  also read the packed color back sign-extended, so any alpha above 0x7F
  compared as larger than 0xFF000000 and claimed full alpha.
- The three packed color morph formats are rewritten. The LSX versions
  built each channel with a chain of inserts, shifts and shuffles that
  didn't survive being run; 4444 also broadcast its scale from the mask
  register. They now follow the steps channel by channel. Note the
  accumulator has to be an LSX scratch register - F4-F7 alias V4-V7,
  which hold the skin matrix for the whole vertex.
- The vertex bounds were loaded with a signed halfword load, so the
  0xFFFF they start at became -1 and no texcoord was ever below it.

PrescaleUV now fuses on these two as well - both JITs fuse it, and the
compiler would have contracted the plain expression there anyway.

The test tolerates a small relative difference on decoded floats, scaled
by the morph count since each term rounds once. Every bug above was
orders of magnitude larger than that.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 13:48:56 -06:00
Henrik RydgårdandClaude Opus 5 c58baedd75 Vertex decoder: fix the jit-match test on MSVC
The test didn't compile: VERTS is captured by reference into testFormat, and
MSVC won't use a captured constexpr as an array bound. Make it static.

On arm64 it then failed on the UV prescale steps, by one ULP. The arm64 JIT and
the NEON handwritten decoders fuse the multiply-add, and the steps only match
that when the compiler contracts a * b + c - which clang does and MSVC doesn't,
in Debug or Release. Spell out which one happens instead of relying on it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 12:13:09 -06:00
Henrik RydgårdandClaude Opus 5 a1ed24ae2d Vertex decoder: fix the handwritten decoders, and test them too
The two handwritten SIMD decoders are used with or without the JIT, and had
drifted from the step functions:
- The God of War one ignored g_DoubleTextureCoordinates, giving HD Remaster
  games the wrong UVs, and passed NaN and infinite positions through. Those
  now come out finite like everywhere else.
- The GTA one expanded 5551 colors wrongly on NEON (a left shift where the
  SSE version shifts right).
- On ARM64, both now fuse the UV scale and offset, like the JIT and the steps
  as the compiler builds them.

Both are back in the unit test, which now also feeds NaN and infinity to plain
float positions and checks they come out finite, and runs only on x86-64 and
arm64 for now.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 11:29:40 -06:00
Henrik RydgårdandClaude Opus 5 63a6c11214 Vertex decoder: make the JIT match the step functions
Add a unit test that decodes every vertex format through both the step
functions and the JIT and requires identical output, side effects included.
Only skinning may differ by rounding, since arm64 accumulates the bone
matrices with fused multiply-adds.

What it found, and fixed:
- x86 morph colors rounded to nearest where the steps truncate, and applied
  the scale before the weight, which rounds differently.
- x86 through-mode u16 UV bounds compared signed, so texcoords above 32767
  scrambled the bounds for everyone running the x86 JIT.
- x86 morph sums could produce -0 where the steps produce +0.
- Step_NormalS16Morph scaled by 1/32768 twice, giving near-zero normals.
- Step_PosFloatThrough lost the truncation of Z to an integer that the JITs
  and the vertex reader did before it moved into the decoder.
- SetVertexType never reset skinInDecode.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 11:29:40 -06:00
Henrik RydgårdandClaude Opus 5 cd9bb569a3 unittest: don't pin CleanNaNInfs to one implementation's output
The contract is only that a bad lane becomes something that yields zero
when multiplied by zero, by whatever route is cheapest. SSE2 clamps to
+-FLT_MAX, NEON, LSX and the scalar fallback zero the lane - both fine.
Check that property, and that good lanes are untouched.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-19 13:21:36 -06:00
Henrik RydgårdandClaude Opus 5 5d703b014e unittest: move the CrossSIMD test to its own file and expand it
CrossSIMD has four independent implementations and only one is compiled
per machine, so each architecture has to check its own copy against
hand-worked results. Running this under qemu gives us coverage of the
targets we have no hardware for - it's already caught two LSX bugs.

Covers Vec4S32 and Vec4F32 arithmetic, comparisons and the mask helpers,
the lane accessors and shuffles, transpose, the various loads (including
the 24-bit and normalizing ones), NaN/Inf handling, and the matrix
routines the old test already had.

Verified on three of the four implementations: NEON natively, the scalar
fallback via TEST_FALLBACK, and LSX under qemu-loongarch64. SSE2 is left
to CI.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-19 12:35:09 -06:00
Henrik Rydgård 8f3da1a84a Merge pull request #22233 from hrydgard/combo-suppress-singles
Don't fire single-button mappings while a combo using them is held
2026-09-18 16:58:46 -06:00
Henrik RydgårdandClaude Opus 5 5c5c7dec86 Include what these files use
UnitTest.h's EXPECT_ macros all call printf and EXPECT_EQ_MEM calls memcmp, but
it included neither <cstdio> nor <cstring> - it has been relying on whatever the
including file happened to pull in first, and TestMpegCsc was the first not to.
The same shape in sceMpegbase.cpp and sceVideocodec.cpp, which use std::min,
std::move and memcpy without saying where they come from.

Also drop an abs() from TestMpegCsc rather than include <cstdlib> for one
subtraction.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 11:59:48 -06:00
Henrik RydgårdandClaude Opus 5 4cb01fea67 sceMpegbase: convert with swscale, keeping the scalar path as the fallback
The planes the de-tiling produces are already the YUV420P swscale wants, and our
sceMpeg HLE converts the same frames the same way, so the pixel formats and the
studio-range setup come straight from MediaEngine::getSwsFormat. It is 3-4x
quicker than going a pixel at a time: 0.37-0.44ms a frame becomes 0.09-0.12ms,
which is the whole reason sceMpegBaseCscAvc was at the top of a profile.

Chroma is upsampled with SWS_POINT rather than the HLE's SWS_BILINEAR, since
replicating is what the scalar path does and, being a fixed-function block,
almost certainly what the hardware does.

It is not bit-identical - swscale rounds its own way. TestMpegCsc measures the
gap per channel rather than per byte, so the number means something for a packed
16-bit pixel: worst 1 step of 31 for 5650 and 5551, 2 of 15 for 4444, 3 of 255
for 8888, with means around a fifth of a step. The scalar path stays as what the
longhand reference is checked against, and takes anything swscale won't - an odd
range origin, or a build without ffmpeg.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 11:38:49 -06:00
Henrik RydgårdandClaude Opus 5 02906f0510 sceMpegbase: write alpha as zero, and stop rebuilding the planes every frame
The colour conversion was writing alpha fully set - 0xFF000000, or the top bit
for 5551 - where the hardware writes zero. Our sceMpeg HLE already masks it off
and names Sword Art Online as a game that depends on it: it doesn't clear the
alpha in the buffer it hands over, and expects the video not to set it. The two
paths now agree.

The de-tiling ahead of it becomes UntileYCbCr, taking the eight buffers already
resolved, so it can be measured and compared against the original longhand
version in TestMpegCsc. Its planes move to scratch that persists between calls -
a movie converts one frame per displayed frame, and this was allocating and
clearing about 200KB every time - and the per-pixel bounds checks in the chroma
loop, which only depend on the group of eight, are hoisted out of it.

That last part is worth 2529 -> 3201 MPix/s, but the point of measuring was to
find out whether it mattered, and it doesn't much: de-tiling is 0.04ms of a
frame against the conversion's 0.4ms. The conversion is where the time is.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 11:38:49 -06:00
Henrik RydgårdandClaude Opus 5 a6996b2c3f sceMpegbase: pull the colour conversion out, and measure it
sceMpegBaseCscAvc is the top of a profile during video playback, so the loop
that does the work becomes MpegCscRange - a pure function with the HLE plumbing
left behind - and TestMpegCsc measures and checks it.

The measuring half reports megapixels per second for a 480x272 frame in each of
the four pixel formats. The checking half compares against the conversion
written out longhand, over whole frames and over partial ranges with odd offsets
and sizes, plus one-pixel, one-row and one-column ranges and one that reaches
the far edge of the frame. Those are the cases an optimized version gets wrong:
chroma is half resolution, so an odd left edge starts mid-sample, and anything
handling two pixels at a time has to deal with the leftover. The destination is
padded and prefilled, so writing outside the range fails too.

This is only the move - the loop is the same one, so the numbers it gives are
the baseline to improve on. On a Snapdragon X Elite it runs at about 300 MPix/s,
0.44ms for a frame.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 11:38:49 -06:00
Henrik RydgårdandClaude Opus 5 e8fa4e3f56 headless: split --timeout into --timeout-wall and --timeout-emulated
--timeout was wall-clock seconds, which is what CI wants but not what you want
when the question is whether the game has had long enough to get somewhere: a
heavy scene runs many times slower than real time and a near-idle one much
faster, so the same budget means very different amounts of game time. Booting a
firmware VSH is a good example - 10 emulated seconds is about 25 real ones on
6.61 and about 7 on 2.00, and judging those two by the same wall-clock number
makes a working shell look stuck.

Both limits can be set at once and whichever is reached first ends the run,
which also says which one it was. --timeout still works as the old name for
--timeout-wall. The IsDebuggerPresent() exemption stays on the wall-clock check
only; the emulated one doesn't need it, since sitting at a native breakpoint
burns no emulated time.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-17 16:01:47 -06:00
Henrik RydgårdandClaude Opus 5 16bf3a6519 BlockAllocator: let an allocator with no blocks save and load
An allocator that nothing has Init'd yet is a real state, not a broken one -
sceVideocodec keeps one for Media Engine memory that stays empty until a game
plays a video - but DoState asserted on bottom_ when writing, and on reading
treated a block count of zero as corrupt. Both ends handle it now, and the
write loop no longer special-cases the first block.

The stream layout is unchanged, so states written before this still load: they
always had at least one block, and take the same path they always did.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-15 11:06:56 -06:00
Henrik RydgårdandClaude Opus 5 4ae682283c Path: make WithReplacedExtension(old, new) report a mismatch instead of hiding it
It used to return the path unchanged when the path didn't end in oldExtension,
so a caller that guessed wrong silently went on using the original file - and
"the screenshot next to this savestate" quietly becomes "this savestate".
Every caller had to know to check the extension first, and most didn't.

Now it's [[nodiscard]] bool with an out-param, in the style of ComputePathTo
next door, so the mismatch has to be handled. Changing the signature rather
than the behaviour means no call site can keep the old assumption by accident.
All four callers wanted "skip it" or "fall back", which they now say out loud.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-12 11:54:11 -06:00
Henrik Rydgård c71fa5e32b sceIo: stop clobbering st_private, report FAT permissions, implement sceIoChstat
Cashing in the io/stat and io/shortname recordings.

__IoGetStat began with memset(stat, 0xfe, sizeof(SceIoStat)), which destroyed 24 bytes of the
caller's buffer that a real PSP never touches - it writes only as far as the timestamps and
leaves all six st_private words exactly as it found them. It also wrote a made-up sector number
into st_private[0] on the memory stick. That word carries the LBN on a UMD, which games read to
build disc0:/sce_lbn paths, so it stays for non-FAT and is left alone otherwise.

FAT has no permissions of its own and everything reads back as 0777. We were passing the host's
idea of the file through instead. The existing "all files look executable on FAT" hack for Beats
(issue #14812) was right in substance but lived only in sceIoDread, so sceIoGetstat and
sceIoDread disagreed about the same file where hardware has them agree. Both now go through one
path, which also gets the read-only case right: no write bits means mode 0555 and attr 0x21.

sceIoGetstat on the root of a volume is refused, as on hardware.

sceIoChstat was a logging stub. It now applies the read-only flag, which is what st_mode's write
bits and st_attr's 0x01 both mean on FAT - setting either produces both, and it's reversible.
That needs a new IFileSystem::SetFileWritable, defaulting to "can't" so read-only filesystems and
hosts that can't express it (Android content URIs) are unaffected; the call still succeeds there,
since hardware would have.

GenerateFatShortNames now accounts for capitalisation. FAT keeps a lowercase flag for the base and
another for the extension, but the PSP only honours the base one, so "shrt" becomes SHRT while
"readme.txt" becomes README~1.TXT. We were only adding a counter on collision. The unit test
carries the whole recorded set, including the corrected README~1.MD.

io/shortname stays in tests_next: its d_name column can't match while SimulateVFATBug is
uppercasing lowercase 8.3 names, which is deliberate and load-bearing for homebrew.
2026-09-10 09:52:01 -06:00
Henrik RydgårdandClaude Opus 5 fb8c99ad49 sceIo: generate and resolve FAT 8.3 short names
sceIoDread hands back a dirent whose d_private holds the 8.3 short name
ahead of the long name, and we never wrote the short name at all - the game
got whatever was on the stack there. Crazy Taxi: Fare Wars reads it rather
than d_name, so it rejected every file in ms0:/MUSIC, ended up with an empty
playlist and never even reserved an mp3 handle: custom soundtracks were
silently dead, with the game spinning on InitResource/SetLoopNum forever.

Generate the names from the directory listing, and resolve them back in
DirectoryFileSystem so a game can open a file by the short name it was given.
Both sides come from the same function, so they agree.

We can't lean on the host for any of this. Linux, macOS and Android have no
8.3 names at all, and while Windows does keep aliases it generates them by a
different rule - it counts to ~4 and then switches to a hash - so resolution
runs before the literal path is tried rather than as a fallback, or on
Windows we'd quietly open a different file than the one we handed the game.

The exact names a real PSP produces are still unverified - no pspautotest
covers d_private - so this implements the ordinary FAT rule and the new
FatShortNames unit test pins that down until hardware can settle it.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-07 10:20:49 -06:00
Henrik RydgårdandClaude Opus 5 50616825da Don't fire single-button mappings while a combo using them is held
If you map something to L2+R2, the mappings for L2 and R2 on their own
would fire as well. Now, while a combo mapping is fully held, the
shorter mappings that share an input with it are suppressed - longest
match wins. Releasing part of the combo brings the shorter mappings
back, for the inputs that are still held.

Adds a ControlMapper unit test covering the sequence.

Fixes #20621

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01JvJR8oJNSCimCM9KXVLjfq
2026-09-03 12:37:31 -06:00
Henrik Rydgård de0dc2d7d4 Remove the leftover geometry shader scaffolding
Nothing has generated or used a geometry shader since the GS paths were removed
- GeometryShaderGenerator is gone, and ShaderWriter's BeginGSMain/EndGSMain had
no callers at all. Removes ShaderStage::Geometry and everything hanging off it:
the GS preambles and GSMain helpers in ShaderWriter, the stage mappings in all
three thin3d backends, the D3D11 geometry shader plumbing (curGS_, the pipeline
and module members, gs_4_0 compilation), CreateGeometryShaderD3D11, the unused
PipelineFlags::USES_GEOMETRY_SHADER and PipelineManagerVulkan's
UsesGeometryShader().

Also stop enabling the Vulkan geometryShader device feature, since we no longer
have any use for it.

Kept on purpose: the device feature is still listed in the feature dumps (like
other capabilities we don't use), and the Vulkan shader cache header keeps its
now-always-zero geometry shader count so the on-disk format stays compatible.
2026-09-03 11:24:03 -06:00
Henrik Rydgård 2d5d2b98a7 Merge pull request #22188 from hrydgard/misc-crash-and-state-fixes
Claude fixing more assorted bugs
2026-09-02 18:03:30 +02:00
Henrik Rydgård cf8ef68f6b Merge pull request #22182 from hrydgard/remote-iso-empty-share-dir
Remote ISO: don't serve the whole filesystem when no folder is set
2026-08-31 12:47:23 +02:00
Henrik Rydgård e2ab84087e Fix five ways to end up stuck, crashed, or silently degraded
System.cpp stamped BootState::Complete unconditionally after InitGPU(), overwriting
the Failed that InitGPU sets when GPU_Init() fails - after it has already run
CPU_Shutdown(). PSP_InitUpdate then took the success path on a core that no longer
existed, down to a null Memory::base, and the first guest access dereferenced it.
InitGPU now reports failure and both callers honor it. (The libretro path had the
same problem from the other direction: it calls InitGPU after the Failed check.)

HandleAssert called g_assertCancelCallback directly on the IDCANCEL path, without
the null check its own BreakIntoPSPDebugger() helper does - and EmuScreen clears the
callback when a game is unloaded. So any assert after returning to the menu turned
"Cancel: skip and break into PPSSPP debugger" into a null jump, from the one button
whose entire purpose is surviving the assert.

__CheatDoState registered the cheat event type when the savestate had no CwCheat
section, but never scheduled it. CoreTiming::DoState has already swapped in the
state's event queue by then, which doesn't contain one either - so loading an old
savestate silently killed cheats, and the enable/disable polling with them, for the
rest of the session.

Achievements::ChangeUMD set g_isIdentifying and returned without clearing it when
hashing failed, leaving IsBlockingExecution() true forever - EmuScreen stops running
the CPU and the game is frozen until restart. Reachable from a disc swap on any ISO
whose PARAM.SFO or EBOOT.BIN can't be read.

x64Analyzer routed opcode 0x88 into the write path but had no case for it, so it hit
the default, logged from inside the crash handler, and failed. 0x88 is exactly what
the x64 JIT emits for a guest sb, so MemFault could never skip or ignore a bad byte
store the way it can a word one. Handle the 8-bit forms, and drop the 0x8a/0x8b cases
in the read path that the same 0xF0 mask made unreachable. Covered by a new
CheckAnalyze case, which fails without this change.

314 pspautotests pass, all unit tests pass.
2026-08-31 12:27:24 +02:00
Henrik Rydgård bdce0e2e44 Buildfix: include <cfloat> for FLT_MAX in TestArm64Emitter
MSVC pulls it in transitively, gcc and clang don't - broke the gcc-normal,
clang-normal, macos and test-headless-alpine CI jobs.
2026-08-31 00:29:22 +02:00
Henrik RydgårdandClaude Opus 5 b35f28e200 Remote ISO: don't serve the whole filesystem when no folder is set
In LOCAL_FOLDER share mode, LocalFromRemotePath ended with

    return Path(g_Config.sRemoteISOSharedDir) / decoded;

sRemoteISOSharedDir defaults to empty and nothing requires the user to pick a
folder before pressing "Share Games (Server)". Path::operator/ doesn't insert a
separator when the component already starts with one, so with an empty base it
returns the component verbatim - "GET /etc/passwd" resolved to Path("/etc/passwd"),
which is non-empty and went straight to DiscHandler. The backslash, "/.." and "//"
filters never fired, because no traversal is needed to get there. That is an
unauthenticated arbitrary file read for anything that can reach the port.

Refuse to resolve anything when no shared directory is configured, and check that
the joined path actually stays inside it. HandleListing needs the same guard: it
called GetFilesInDir on the empty path, which on Windows becomes
FindFirstFile("\*") - a listing of the root of the current drive.

Also log a warning when the server starts in this state, so "nothing is shared"
doesn't look like a mysterious failure.

The empty-base behavior of Path::operator/ is surprising enough to be worth
pinning down, so TestPath now asserts it.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01DCPmm7FoQUoqrbMdhfqhQ2
2026-08-30 13:51:02 +02:00
Henrik RydgårdandClaude Opus 5 5f9131ed2b Emitters: fix a batch of encoding bugs
Found by a review pass over Common/. Three of these affect code the JITs
actually emit today:

* ARM64 TryMOVI(8) returned true unconditionally ("can always do 8"), but MOVI
  with an 8-bit element replicates imm8 into every byte, so it can only encode a
  byte-uniform value. TryAnyMOVI always tries size 8 first, so it succeeded for
  every constant. MOVI2FDUP(FLT_MAX) - VertexDecoderArm64's Jit_PosFloat - came
  out as "movi v0.16b, #0xff", a quiet NaN, and since FMINNM/FMAXNM return the
  other operand for a quiet NaN, the infinity clamp silently did nothing.
  TryAnyMOVI's replication loop was also shifting by every bit position instead
  of by multiples of the element size, and it now only tries an element size the
  value actually repeats at. Regression test added.

* RISC-V SW()'s stack-pointer compression path called C_LWSP instead of C_SWSP,
  turning a store into a load that clobbers rs2 whenever autocompress is on
  (which RiscVJit and VertexDecoderRiscV both enable).

* LoongArch64 EncodeDFj passed the raw register enum instead of DecodeReg(fj),
  so bit 10 was always set and MOVFR2GR_S emitted movfr2gr.d - live in the
  LoongArch JIT's mfc1 and its FPU/vector compilers.

The rest have no callers today, but are wrong as written:

* ARM64: MOVI/MVNI computed the MSL cmode one too high (MSL #8 is 1100, not
  1101); TryMOVI's MVNI-with-MSL branch passed the value instead of its
  complement; TBZ/TBNZ put the register size in bit 31 where b5 belongs and
  didn't mask the bit index to 5 bits; the LDR/LDRSW/PRFM literal form checked
  the wrong mask for imm19 and wrote it unmasked; FCVTZS/FCVTZU's GPR-
  destination branch skipped DecodeReg and derived the type field from the GPR
  rather than from the float source.
* LoongArch64: LDPTR_D/STPTR_W/STPTR_D all passed Opcode32::LDPTR_W;
  AMCAS_DB_D duplicated AMSWAP_DB_D's opcode; EncodeJK shifted rk by 5 instead
  of 10; BYTEPICK_D masked its shift to 2 bits instead of 3.
* x64: VGATHERDPD/VGATHERQPS/VGATHERQPD used the wrong opcode/W combinations
  (only VGATHERDPS was right).

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01DCPmm7FoQUoqrbMdhfqhQ2
2026-08-30 13:44:34 +02:00
Henrik RydgårdandClaude Opus 5 2f5bb829f3 Demangle: rewrite the SN Systems demangler
The old one was reverse engineered from a handful of symbols and got the
shape of the format wrong - it required a digit right after the kind
character, which most real symbols don't have. Measured against a PSP
executable that shipped with its symbol table intact, it decoded 238 of
4662 mangled symbols, most of those incorrectly.

Worked out properly from that binary, the format turns out to be:

  __0 <kind> <name...> <params> [_ <return type>] [<qualifier>]

where the kind character (member function, free function, operator, data)
is the only thing that says how many name components follow, since nothing
separates the last one from the first parameter. Lengths are letters
(A = 0, a = 26); "5" marks an enclosing namespace; "7...._" is a template
argument list, with "4" plus a compact integer for a non-type argument and
"9<index>A" for a back-reference to one; "T<index>" and "N<count><index>"
repeat an earlier parameter; a trailing "K" is const and a trailing "T" is
a static member function. Also handles __TID_/__T_ (the two halves of a
class's RTTI) and __sti__ (a translation unit's static initializers).

That decodes 4661 of the 4662. The one holdout is an STL symbol whose
template argument is a reference to a member of another template.

Declarator wrapping is shared with the CodeWarrior demangler now, so
pointers to arrays come out as "short (**)[64]" in both.

docs/SNSystemsMangling.md describes the format, marking what's inferred
rather than attested.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01SF5eS5QDNexLksRDeDZvwY
2026-08-29 23:44:35 +02:00
Henrik RydgårdandClaude Opus 5 dfc04f3578 Demangle: handle CodeWarrior templates, function pointers and @-symbols
Checked against two PSP binaries that shipped with intact symbol tables,
which turned up several constructs the format's usual description doesn't
mention:

- Template arguments are written literally inside the length-prefixed name
  ("39CList<Q38hlScreen5Brwsr13CContentsUnit>"), not with a "__PT" prefix,
  and they nest. Function templates put theirs in the base name instead,
  followed by the return type.
- A family of "@"-decorated symbols for things with no C++ name: thunks
  ("@12@__dt__3SonFv"), string literals, function-local statics and their
  guard variables. Plus __vt__/__RTTI__/__sinit_, printed in the same style
  as the Itanium special names.
- Types are now built as a split declarator, so a pointer to a function
  comes out as "int (*)(int)" rather than "int (int) *".

Also stop the lenient pass from turning plain C names with a "__" in them
into nonsense - "I3dClut__FlushCache" became "I3dClut(long, ...)". It now
requires a class qualifier, which costs nothing: over ~10000 symbols the
lenient pass rescued none and only produced those false positives.

Symbol map names go from 128 to 256 characters, since a demangled name
keeps its parameters and templates make short work of 128.

docs/CodeWarriorMangling.md describes the format, marking the parts that
are inferred from cfront rather than attested in a real binary.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01SF5eS5QDNexLksRDeDZvwY
2026-08-29 23:23:57 +02:00
Henrik RydgårdandClaude Opus 5 27869abb6a Demangle: add CodeWarrior and SN Systems symbol demanglers
Older PSP binaries weren't built with GCC, so the Itanium demangler doesn't
help with them. Add two more, tried in turn by DemangleSymbolName():

- Metrowerks CodeWarrior, a descendant of the AT&T cfront scheme
  ("getDistance__6KzUtilFP7st_unitP7st_unit"). Handles Q<n> qualified names,
  the cfront type codes including T/N back-references, cv-qualifiers, and the
  operator/ctor/dtor name codes.
- SN Systems SNC/ProDG ("__0f5DstdIbad_castEwhatvK"), which encodes name
  component lengths as letters. Reverse engineered from a small sample, so
  the parts that are guesses are marked as such - they don't affect the name.

Both are rougher than the Itanium one: they aim for a correctly qualified name
plus a plausible parameter list, and print "..." for a parameter they can't
decode rather than throwing the name away. Results come back as a
DemangledSymbol with the name, parameters, return type and qualifiers kept
separate, in case a caller wants more than the printed string.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EFV5DUTc9ZYAKgsCMZGwX8
2026-08-29 12:18:27 +02:00
Henrik Rydgård fad7b93776 Merge pull request #22151 from hrydgard/ui-tab-navigation
UI: Add tab navigation
2026-08-27 22:17:47 +02:00
Henrik RydgårdandClaude Opus 5 ea1ad8ffed UI: Tab and Shift+Tab move focus through the view hierarchy
Unlike the directional moves, this doesn't look at where anything ended up on
screen - it walks the hierarchy in the order views were added, flattening nested
groups in place. That's what makes it predictable in the layouts where "what's
to the right of this" has no good answer.

A view is a stop if it's focusable and enabled, the same test the directional
moves apply, so the two agree on what's reachable. Hidden subtrees are skipped
whole, which is what keeps a TabHolder's inactive tabs - V_GONE rather than
removed - out of the order without any special casing. Containers are gated on
visibility only, not enabled, matching Key/Touch/Axis: disabling a container
doesn't stop its children being interactive anywhere else either.

Ctrl+Tab stays with ChoiceStrip, which uses it to switch tabs.

focusMoves now holds FocusMove rather than raw keycodes, so the direction is
decided in one place while the modifiers are still around, and a held key
repeats in the direction it was originally pressed with - the synthesized repeat
has no modifiers of its own. That also retires the keycode switch in
UpdateViewHierarchy and IsScrollKey, which had no other callers.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01SfY7iFJEjmRXf1XGrTs4MF
2026-08-27 20:52:20 +02:00
Henrik Rydgård 04bf3e56ef Merge pull request #22149 from hrydgard/symbol-map-cache-fix
ImDebugger: fix stale symbol list after a game is reloaded
2026-08-27 20:19:03 +02:00
Henrik RydgårdandClaude Opus 5 daa18fc25a ImDebugger: fix stale symbol list after a game is reloaded
The disasm window cached the flattened symbol list and only rebuilt it when one
of three menu items said so. Nothing marked it dirty when a game booted or
exited, and a new SymbolMap is allocated per boot, so the list kept showing the
previous game's functions.

Give SymbolMap a version counter that every mutator bumps, and let the window
compare against it instead. The counter is process-wide rather than per-map, so
a fresh map can't hand out a version a cached copy already holds.

Also re-find the selected symbol by address after a rebuild (the index means
something else afterwards), and drop the unused symbol cache members in
ImMemWindow.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01SfY7iFJEjmRXf1XGrTs4MF
2026-08-27 10:43:08 +02:00
Henrik RydgårdandClaude Opus 5 049bcd5483 Demangle C++ symbol names when loading ELF symbols
C++ homebrew has an unreadable symbol table -
everything is _ZN10PxRenderer7DrawImmE... - which makes the disassembly and
symbol list nearly useless. Add an Itanium C++ ABI demangler and run ELF
symbols through it on load, in both ElfReader::LoadSymbols (unstripped EXECs,
which is what a CMake pspdev EBOOT actually contains) and the companion-ELF
path.

The demangling standard is called Itanium for historical reasons - it
was defined for Itanium but ended up being almost universally
applicable.

Written from scratch rather than using __cxa_demangle, which doesn't exist on
MSVC/UWP, or vendoring LLVM's demangler, whose license doesn't fit. Anything
unrecognized (arbitrary constant expressions, decltype) aborts the parse and
the caller gets the original mangled name back, so a caller never sees a
half-parsed result. Recursion is depth-capped since the input comes from a
file we didn't write.

Checked against c++filt as an oracle: of 1089 mangled symbols in a real C++
homebrew EBOOT, one differs; of 55189 from libstdc++/libLLVM/cc1plus, 22
differ and 413 are declined. Fuzzed with 220k mutated and random inputs under
ASan/UBSan.

Also adds a right-click menu to the ImDebugger symbol list.

Note that SymbolMap stores names in char[128], so the longest STL names get
truncated in the UI. Still far more readable than the mangled form.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01X3DbkJ8ShYiXU7q5Tv1LZu
2026-08-26 08:12:45 +02:00
Henrik Rydgård 4fb5d5965c MMIO: Simplest possible kernel-mode permission check, start work on supporting in JIT 2026-08-21 10:32:22 +02:00
Henrik RydgårdandClaude Opus 5 b42a03e095 Add savestate serializer tests, fix three bounds-check bugs
PointerWrap and the Do() overloads around it are how every savestate is
written and read, and had no direct coverage. Everything read back came off
disk, so the corrupt-input paths matter as much as the round trips.

Three bugs, all in the bounds checking added in 58d4759ceb:

1. sizeof(T) is not a lower bound on how many bytes an element serializes to.
   It only holds for the types DoHelper_ writes out raw. A std::string is 32-40
   bytes in memory and serializes to as few as five; a T* serializes to whatever
   T::DoState() writes. So DoVector/DoList/DoSet/DoMap could reject a perfectly
   valid savestate whenever count * sizeof(element) exceeded the bytes left in
   the buffer. That is not hypothetical: pspFileSystem is serialized dead last
   in SaveStart::DoState, and MetaFileSystem::DoState does Do(p, currentDir) on
   a std::map<int, std::string>, so the check runs with only a few hundred bytes
   remaining and claims 44 bytes per entry against roughly 22 actual. Added
   SerializeMinElemSize<T>(), mirroring DoHelper_'s own condition, and used it
   in all five containers. The bound is only loosened, so nothing that loaded
   before can stop loading.

2. Do(p, std::map<K, T *> &) deletes every value before reading the new ones,
   and DoMap then returned on a bad count without clearing - leaving the map
   full of freed pointers to be used or deleted again. Six live maps go through
   this (sceMpeg, sceMp3, sceAac, sceFont, sceHeap, sceKernelThread's pending
   calls), so a corrupt savestate meant a use-after-free. Clear before the guard
   can bail out, in DoMap, DoMultimap and DoSet.

3. The wstring and u16string overloads validated stringLen < 0 but not 0, and
   didn't require a whole number of characters. read() computes
   stringLen / sizeof(char) - 1, so a length of 0 resized to SIZE_MAX and
   memcpy'd with a wrapped-around size. PSPOskDialog::DoState serializes both
   (inputChars at v2, a legacy wstring below that), so this was reachable: the
   test aborts the process without the fix.

The test covers round trips of PODs, strings (empty, embedded NUL), vector,
map, set, list and map-of-pointers, section titles and version gating in both
directions, marker mismatches, measure-vs-write checkpoint disagreement, the
error latch dropping to MODE_NOOP, every truncation of a valid buffer, and
hand-corrupted counts and lengths.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
2026-08-17 14:50:57 +02:00
Henrik RydgårdandClaude Opus 5 8265044a79 Add Hashmaps unit tests, stop tombstones from filling the table
DenseHashMap and PrehashMap are the open-addressed, linear-probing maps behind
the texture cache, the shader managers and the software renderer's
sampler/drawpixel caches, and had no coverage.

Writing the tests turned up a latent hang. Removal leaves tombstones, which
occupy probe slots exactly like live entries, but the load factor check only
looked at count_. So a workload that inserts and removes distinct keys keeps
count_ low forever while REMOVED fills the table, and no Grow is ever triggered.
Once there is no FREE bucket left, a lookup for a missing key has nothing to
terminate on - and the probe loops don't break out after their "Hit full"
assert, which is compiled out in release builds. The test reproduced it as a
hard hang in about a second.

Two fixes: count tombstones towards the load factor (rebuilding in place when
the load is mostly tombstones, growing otherwise), and make the probe loops
return instead of spinning if they ever do wrap all the way around.

Not reachable today - nothing in GPU/ calls Remove() on these maps, and
Maintain(), which exists to rebuild when tombstones pile up, is never called
anywhere. But Remove() is public API and the first caller to use it in a loop
would have hit an unexplained freeze.

Tests cover insert/get/miss/remove/size, tombstones not cutting a probe chain,
Iterate visiting exactly the live entries, Clear, growth past the initial
capacity, Rebuild compacting, a 20000-operation differential test against
std::unordered_map, and the tombstone churn above. PrehashMap gets the same
treatment.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
2026-08-17 14:50:57 +02:00
Henrik RydgårdandClaude Opus 5 ee8aaff6d2 Add SymbolMap unit tests, fix absolute symbols vanishing with no module loaded
SymbolMap had no coverage. It stores symbols relative to a module so they
survive that module being unloaded and reloaded elsewhere, and only symbols
belonging to a loaded module count as active - that indirection is where the
surprises are, so the tests concentrate on module lifetime and the shared label
table.

Writing them turned up a real bug. UpdateActiveSymbols() bailed out early when
activeModuleEnds was empty, as a "tiny optimization" for startup and shutdown,
having already cleared the active maps. But symbols with module index 0 are
absolute by design - they belong to no module, which is how you label a heap or
stack address - and the loops it skipped are exactly what keeps those alive.
So an absolute symbol disappeared as soon as the last module was unloaded, and
didn't exist at all before the first one was loaded. Dropping activeModuleEnds
from the early-out condition fixes it; the symbol-count half still gives the
intended fast path when there's nothing to do.

Tests cover function and data lookup by containing address, SetFunctionSize,
RemoveFunction/RemoveData, symbols surviving an unload/reload at a different
address, absolute (module 0) symbols, GetSymbolInfo/GetDescription, and Clear.
Two of them pin down behaviour that catches people out rather than asserting
it's right: AddLabel deliberately won't overwrite an existing label, and because
functions and data share one label table, renaming or removing via one affects
the other.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GZq8ZtJmFY7bkX5FVkr3P9
2026-08-17 14:50:57 +02:00