GLSLtoSPV takes an optional SPIRVCache, keyed on a 32-bit hash of the
source, stage and variant, plus the source length. A changed shader
simply misses. thin3d's shaders and the other fixed ones use a global
cache in PSP/SYSTEM/CACHE/vulkan_spirv.cache, loaded on first use and
saved after graphics init, when a game's cache is saved, and at
shutdown; it's flushed once it reaches 32 entries, about twice what a
session compiles, so outdated ones don't pile up. Game shaders keep
theirs in the .vkshadercache, ahead of the shader IDs so that the
compiles on load find it (version 60), and only what the session used
is saved.
A cold glslang costs about 40ms before its first shader here, and
0.3-0.9ms per shader after that.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
MP3 emulation already went through FFmpeg, leaving MiniMp3Audio dead.
The one live user was loading MP3 UI sound effects (custom achievement
sounds), which now splits the file into frames and decodes them with
the FFmpeg MP3 decoder.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
SetError overwrote the first bad section with whichever section a later
error came from, and an error before any section read an uninitialized
curTitle_.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
DoMap and DoSet cleared them, but only once the count had been read. A
state truncated right there left the deleted pointers in place, to be
freed again when the failed load reset the game.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Reject sizes past the end of the state before allocating (FPL, PGF,
achievements, SAS grain, savedata list, the memory fast path), fail
instead of desyncing on a SAS voice count mismatch, and free what old
states' paths and shrinking pointer containers dropped.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Containers of pointers are filled with nullptr and then DoClass'd, and
once an error switches the load to MODE_NOOP, every remaining element
called DoState on null.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
'PPSSPPDebug64.exe' (Win32): Unloaded 'C:\Windows\System32\mfreadwrite.dll'
The thread 'RecentISOThreadFunc' (9920) has exited with code 0 (0x0).
The thread 'Console' (10488) has exited with code 0 (0x0).
Detected memory leaks!
Dumping objects ->
{17571338} normal block at 0x00000214CC4A3FE0, 16 bytes long.
Data: < Cf > E8 43 66 CD 14 02 00 00 00 00 00 00 00 00 00 00
{17571337} normal block at 0x00000214CC4A3B80, 16 bytes long.
Data: < Cf > C8 43 66 CD 14 02 00 00 00 00 00 00 00 00 00 00
D:\project\memory_leak\ppsspp\Windows\Debugger\Debugger_Disasm.cpp(174) : {17571336} normal block at 0x00000214CD664190, 640 bytes long.
Data: < C > 90 43 F2 C0 F6 7F 00 00 F6 0C 04 00 00 00 00 00
D:\project\memory_leak\ppsspp\UI\NativeApp.cpp(879) : {84582} normal block at 0x00000214CE8C5DB0, 4224 bytes long.
Data: < > 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
{2432} normal block at 0x00000214D6255C20, 16 bytes long.
Data: <0 $ > 30 E7 24 D6 14 02 00 00 00 00 00 00 00 00 00 00
D:\project\memory_leak\ppsspp\Common\Net\HTTPNaettRequest.cpp(43) : {2431} normal block at 0x00000214D624E730, 32 bytes long.
Data: < \% > 20 5C 25 D6 14 02 00 00 00 00 00 00 00 00 00 00
Object dump complete.
The thread 22000 has exited with code 0 (0x0).
The thread 21248 has exited with code 0 (0x0).
The thread 38104 has exited with code 0 (0x0).
The thread 27860 has exited with code 0 (0x0).
The thread 33240 has exited with code 0 (0x0).
The thread 29352 has exited with code 0 (0x0).
The thread 13964 has exited with code 0 (0x0).
The thread 5640 has exited with code 0 (0x0).
The thread 10528 has exited with code 0 (0x0).
The thread 15252 has exited with code 0 (0x0).
The thread 23016 has exited with code 0 (0x0).
The thread 31424 has exited with code 0 (0x0).
The thread 7000 has exited with code 0 (0x0).
The thread 11620 has exited with code 0 (0x0).
The thread 36264 has exited with code 0 (0x0).
The thread 27248 has exited with code 0 (0x0).
The thread 24780 has exited with code 0 (0x0).
The thread 26228 has exited with code 0 (0x0).
The thread 13676 has exited with code 0 (0x0).
The thread 16828 has exited with code 0 (0x0).
The thread 27388 has exited with code 0 (0x0).
The thread 1668 has exited with code 0 (0x0).
The thread 37272 has exited with code 0 (0x0).
The program '[23456] PPSSPPDebug64.exe' has exited with code 0 (0x0).
Instead of the PSP's web browser, a dialog shows the URL the game wants
and opens it in the host's browser on X, or backs out on O. Either way the
game sees the browser closed normally. Platforms that can't open a URL
(the new SYSPROP_CAN_LAUNCH_URL) only offer to back out. Only plain
printable-ASCII http(s) addresses are handed over, and on Linux without a
shell.
What the firmware does (sceUtility_Driver, 6.61, plus
utility/dialog/htmlviewer): the HtmlViewer has its own state apart from the
other dialogs, so they don't block each other, and its calls return
WRONG_TYPE until one has started. The request size picks the 2.00 to 3.00
layout, and InitStart allocates 3.5MB of user memory (4.5MB with options
bit 0x400 from 2.70 on), failing with 800200d9 without it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Executables outside a bundle (headless) found no Vulkan library. Also try
the app bundle built next to them, the Vulkan SDK's install and Homebrew's.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A frame that never draws to the backbuffer never acquires an image, but
its final submit still waited on the acquire semaphore, which nothing
signals, hanging the GPU. Skip the swap for such frames when finishing
them. This replaces the check for a frame with no steps at all, which
could also set it partway through a frame that acquires later.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
For when there's nothing to present to. Instead of a surface and swapchain,
it renders into images of its own through the VulkanPresentation interface
libretro uses, picking the graphics queue without a surface. Acquiring and
presenting signal and wait on the frame's semaphores with empty submits.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ProtectFunction's thunk saved all of XMM2-15 and RBX around every call,
but the callee preserves XMM6-15 on Windows and RBX everywhere. On
Windows that drops ten 16-byte saves and loads from each protected call.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A request started since the last RequestManager::Update (headless never
calls it) sat in newDownloads_, which CancelAll skipped. It was then
destroyed along with the static g_DownloadManager at exit, and its
destructor removed its progress bar from the already destroyed g_OSD:
"mutex lock failed". Seen with a Netconf dialog still downloading the
infra DNS json when a test ended.
CancelAll now takes the new ones too, and runs at shutdown while g_OSD is
still there. The Netconf json request is also let go of when the emulator
shuts down, rather than living on into the next game.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The C cast is undefined past the int32 range, and x86 makes it INT_MIN, so
round/trunc/ceil/floor/cvt.w.s of anything from 2^31 up gave 0x80000000 on
x86 hosts while the PSP saturates to 0x7fffffff (cpu/fpu/roundmode). Route
all of them through SaturatedFloatToInt, which also covers NaN and inf, and
drop the special cases that did.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Two bugs stacked on each other, and between them every vector path in
this backend was either dead or miscompiled.
The register cache declared mapFPUSIMD unconditionally, but every
compiler here picks its path from cpu_info.LOONGARCH_LSX at runtime.
Without LSX the scalar fallbacks ran against a SIMD mapping and reached
lanes with F(reg + n), which addresses nothing there - ApplyMapping only
allocates per-lane registers when mapFPUSIMD is false. Vec4Unpack8To32
and Vec4DuplicateUpperBitsAndShift1 came out with only their first lane,
Vec2Unpack16To32 put both halves in one register, and the pack ops read
back whatever was next door. riscv64 sets the flag false and has never
had the problem. Tie the two together and the fallbacks are correct as
they stand.
That path was reachable because of the second one: the USE_CPU_FEATURES
block assigned over the hwcaps, and GetLoongArchInfo reads /proc/cpuinfo
and nothing else. Under qemu-user that file belongs to the host, so it
found no features and turned LSX off - leaving the LSX paths dead and the
untested scalar ones running, which is not what any Loongson 3A5000 or
later does. Let it only ever add to what the hwcaps found.
cpu/vfpu/convert, gum and matrix pass now, both with LSX and with it
masked off, putting loongarch64 at 338/342 - the same four as riscv64.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
EncodeCR, EncodeCI and EncodeCSS shifted the RiscVReg enum value straight
into the instruction without DecodeReg(), which every 32-bit encoder in
the file uses. FPRs are 0x20..0x3F in that enum, so bit 5 is set for all
of them and spilled into a neighbouring field - for the CI format, into
bit 12, which holds imm[5].
The visible effect: the dispatcher's epilogue restores the callee-saved
FP registers with c.fldsp, and every offset whose bit 5 was clear came
out 32 bytes too high. fs3-fs6 were restored from the wrong slots and
fs11 from 224(sp), past the 208-byte frame. So the native JIT corrupted
whatever the C++ caller had in those registers - which, in the headless
test runner, was the wall-clock deadline, making every test report an
instant TIMEOUT.
pspautotests on riscv64 under qemu goes from 0/342 to 338/342 with this,
matching the IR interpreter exactly.
Found with the disassembly the enableDisasm flag in RiscVAsm.cpp emits -
the save offsets and the restore offsets simply didn't match.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
LoadConvertU8, StoreConvertToU8 and LoadTranspose existed in the SSE2,
NEON and LSX implementations but not in the scalar one, so anything using
them wouldn't build on a target without SIMD.
The two LSX bugs were found by the new CrossSIMD unit test, run under
qemu-loongarch64:
- Vec4F32::operator[] had a switch with no breaks, so every index fell
through to the default and returned lane 3.
- StoreConvertToU8 narrowed with the logical (unsigned) saturating shifts,
which turn a negative value into a huge unsigned one and saturate it to
255. It should clamp to 0, as the packs/packus pair in the SSE version
does. Narrow signed->signed and then signed->unsigned instead.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
MOVE(LoongArch64Reg::X4, arg) targets X4, which is an LASX 256-bit vector
register (0x64), not a0. The LP64D ABI puts the first integer argument in
a0, which is R4.
This is not a corner case: GenerateFixedCode uses QuickCallFunctionR for
CoreTiming::Advance, so the emitter asserted ("DJK instruction rd must be
GPR") while building the dispatcher, before a single block was compiled.
The LoongArch native JIT could not start at all.
Found by running the loongarch64 cross build under qemu-user.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The riscv64 cross build is the first target to take the non-SIMD path,
and it didn't compile:
- LoadF24x4 called LoadR24x3_One, which doesn't exist. It should shift
all four lanes, like the SSE/NEON/LSX versions do.
- isnan/isinf were unqualified, and <cmath> wasn't included.
- WithLane3From and AnyCompareBitsSet were missing entirely.
LoadF24x3_One also left lane 3 as zero, where all three SIMD versions set
it to 1.0f - a behavioural bug that would only have shown up once someone
ran this path.
Verified by flipping TEST_FALLBACK in the header: the whole tree builds,
and with the scalar path in use the unit tests and pspautotests both pass
in full (342/342, software renderer, which leans on this code heavily).
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Headless is the only thing that uses it, and it had its own shorter format with
no thread, file or line, so a pattern that matched a log from the app quietly
matched nothing in one from headless. Costs a wrong conclusion the first time
you do it.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
DrawEngineGLES already knows that every index it generates is below the
decoded/transformed vertex count, but only passed the count to
glDrawElements. Drivers that need the vertex range before vertex
shading, like Mesa's Panfrost, then scan the index data on the CPU for
every draw, and their min/max caches can't help since the index push
buffer is rewritten every frame.
Only used for single-instance draws on desktop GL or GLES3, and only
when the caller provides the range, so other DrawIndexed callers are
unchanged.
Co-Authored-By: Claude Opus 5 <[email protected]>
The name and value columns were hardcoded at x=17 and x=77. The control has
a fixed width in the dialog and can not be widened, so the 8 hex digits of a
GPR only just fit - and once the register list got a vertical scrollbar, the
~17px it takes off the client width pushed the last digits off the edge.
Derive the value column from the client width each paint instead, pulling it
in far enough for a whole value to fit and clipping anything longer rather
than letting it spill. Same for the category tab labels.
Float registers print with %g rather than %f: in a column that narrow, a
clipped %f of a large value is worse than useless (1e20 would read as
100000000), while %g keeps six significant digits and stays short for the
values you normally see.
imgui 1.92 hands the backend a list of dirty rectangles when its font atlas
grows, but thin3d could only replace a whole mip level, so every new glyph
re-uploaded the entire atlas.
Adds DrawContext::UpdateTextureRegions, taking a batch of regions so each
backend can submit them together:
- Vulkan: packs the regions into the push pool and issues a single
vkCmdCopyBufferToImage from the init command buffer, transitioning only
the level being written.
- OpenGL: new TEXTURE_SUBIMAGE init step. The existing sub-image path is a
render command needing an active render pass and a texture slot, which
doesn't fit here. Rows are packed caller-side since GLES2 lacks
GL_UNPACK_ROW_LENGTH.
- D3D11: UpdateSubresource with a box, no staging texture needed.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
It used to return the path unchanged when the path didn't end in oldExtension,
so a caller that guessed wrong silently went on using the original file - and
"the screenshot next to this savestate" quietly becomes "this savestate".
Every caller had to know to check the extension first, and most didn't.
Now it's [[nodiscard]] bool with an out-param, in the style of ComputePathTo
next door, so the mismatch has to be handled. Changing the signature rather
than the behaviour means no call site can keep the old assumption by accident.
All four callers wanted "skip it" or "fall back", which they now say out loud.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Cashing in the io/stat and io/shortname recordings.
__IoGetStat began with memset(stat, 0xfe, sizeof(SceIoStat)), which destroyed 24 bytes of the
caller's buffer that a real PSP never touches - it writes only as far as the timestamps and
leaves all six st_private words exactly as it found them. It also wrote a made-up sector number
into st_private[0] on the memory stick. That word carries the LBN on a UMD, which games read to
build disc0:/sce_lbn paths, so it stays for non-FAT and is left alone otherwise.
FAT has no permissions of its own and everything reads back as 0777. We were passing the host's
idea of the file through instead. The existing "all files look executable on FAT" hack for Beats
(issue #14812) was right in substance but lived only in sceIoDread, so sceIoGetstat and
sceIoDread disagreed about the same file where hardware has them agree. Both now go through one
path, which also gets the read-only case right: no write bits means mode 0555 and attr 0x21.
sceIoGetstat on the root of a volume is refused, as on hardware.
sceIoChstat was a logging stub. It now applies the read-only flag, which is what st_mode's write
bits and st_attr's 0x01 both mean on FAT - setting either produces both, and it's reversible.
That needs a new IFileSystem::SetFileWritable, defaulting to "can't" so read-only filesystems and
hosts that can't express it (Android content URIs) are unaffected; the call still succeeds there,
since hardware would have.
GenerateFatShortNames now accounts for capitalisation. FAT keeps a lowercase flag for the base and
another for the extension, but the PSP only honours the base one, so "shrt" becomes SHRT while
"readme.txt" becomes README~1.TXT. We were only adding a counter on collision. The unit test
carries the whole recorded set, including the corrected README~1.MD.
io/shortname stays in tests_next: its d_name column can't match while SimulateVFATBug is
uppercasing lowercase 8.3 names, which is deliberate and load-bearing for homebrew.
Three fixes, all of them things the new hardware tests turned up.
sceMd5Block* and sceKernelUtilsMd5Block* shared one static md5_context and ignored the context
pointer the caller passed in, with a TODO saying it would do "unless games do several MD5
concurrently". hash/md5ctx shows a real PSP keeps everything in the caller's 96 bytes and happily
runs two digests at once, so do that instead: the state, the counters and the block buffer now
live at ctxAddr in the game's own memory, in the layout the test pins down. Two interleaved
digests come out right, and a context that gets copied mid-digest carries on correctly. As a
side effect the state is now covered by savestates, which a file-static never was.
MersenneTwister masked both halves with 0x80000000 where the low half needs 0x7FFFFFFF, so
sceMt19937UInt and sceKernelUtilsMt19937UInt were returning a sequence that isn't MT19937 at
all - every number differed from hardware from the first draw. hash/mt19937ctx computes the
reference sequence itself and confirms the PSP is plain MT19937; with the mask fixed we match it
for both seeds tested. Init also twists the array immediately, as hardware does, so a context
that has been seeded but not drawn from now holds what a real one would.
sceKernelCreateTlspl accepted partitions up to 9 before falling through to the permission check.
Hardware draws the line at 6 - threads/tls/create records 7, 8, 9 and 10 all returning
ILLEGAL_ARGUMENT - so 8 and 9 were coming back ILLEGAL_PERM. Note this is genuinely different
from sceKernelCreateVpl right above it, which does let 8 and 9 through to ILLEGAL_PERM; the two
had been sharing a check that was only ever right for Vpl.
Risk worth naming: the MT19937 change alters the numbers any game gets from these calls. That's
the point - they were wrong - but a savestate taken mid-sequence will resume with a generator
that behaves differently from the one that made it.
The log manager had a single external-callback slot, but the WebSocket
debugger registers one per *connection* - LogBroadcaster is a local in the
per-connection handler. So with two clients attached (the bundled JS debugger
in a browser and Tools/wsdbg, say) the second one to connect silently took the
log stream away from the first, and then whichever disconnected first cleared
the slot and stopped delivery to the other as well. A one-shot wsdbg command is
enough to do it: connect, take the stream, exit, and the long-lived listener
that was watching the log goes quiet with nothing to say why.
Make it a list with add/remove by handle. The dispatch loop holds the lock
across the callbacks so a listener can't be freed while one is running - which
is what lets LogBroadcaster delete its listener straight after removing it.
Enabling and disabling LogOutput::ExternalCallback belongs to the list now, and
disabling only happens when the last callback goes away.
libretro registers one of these too, and never removes it; it just moves to the
new call. It can't actually collide with the debugger - the libretro build
doesn't compile Core/Debugger/WebSocket at all - but there's no reason for it
to keep using an API that only has room for one caller.
Co-Authored-By: Claude Opus 5 <[email protected]>
RAIntegration keeps its cache and local achievement data next to the
executable. If PPSSPP is installed somewhere that needs elevation to write -
Program Files being the obvious case - that write fails and takes the emulator
down as soon as a set or code notes are loaded, with no log to show for it since
the log can't be written either.
Check whether the exe directory is writable before handing the DLL to rcheevos,
and if it isn't, say so and point at the portable .zip instead. Achievements
themselves still work, so carry on to the normal login.
Fixes#21260
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01GgACRqkQNpfJQ4fjwyoEup
Make missing JSON dictionaries null-safe and reject infra DNS data
without the required default object or games array. Fetch the metadata
over HTTPS so a network attacker cannot replace a valid response with a
crash-triggering document in transit.
Reject non-positive or oversized dimensions independently in the PNG
header peeker, matching the existing 8192 pixel decode limit. Use
checked size_t arithmetic before allocating replacement RGBA data so
crafted dimensions cannot overflow the allocation size.
unique_ptr, not shared_ptr. Nothing ever shares it: naett holds a raw pointer
and takes no part in the counting, and the abandonment path hands the sink over
and lets go in the same breath. What's really going on is a single owner that
moves, which is what unique_ptr says and shared_ptr left you to work out from
reading all the uses.
naett never frees the pointer it's given for the writer - it only reads
bodyWriterData and hands it back to the callback - so deleting the sink is
ours to do. The things naett does allocate, the request and response, stay raw
pointers with explicit naettFree/naettClose.
The length field was just buffer.size() written down twice.
It asserted that it was the first call, so net::Init - which is documented as
safe to call repeatedly, and is - needed a bool of its own purely to guard this
one line. That's bookkeeping the library may as well do itself, and the WSA half
of the same function already says as much: it doesn't track anything because
WSAStartup counts its own references.
So a repeat call does nothing, and net::Init loses g_naettInitialized. Asking
HTTPSAvailable() twice costs nothing - on Linux it's a cached dlopen result and
everywhere else it's a constant.
Found reviewing the commits before this one. On Windows a short write marked the
request complete and then queued another read regardless, which is worse than it
sounds: the caller is free to close a response the moment it sees complete, so
WinHTTP carried on writing into memory that was being freed. It also meant
cancelling didn't actually stop anything there. Stop at that point, and ignore
anything raised after the request is complete.
The Apple callbacks do the same check, so the cancellation we ask for after a
failed write doesn't overwrite the error that caused it.
naettMake only asserted that its request was non-NULL, and the request
constructors do return NULL when a platform can't set one up - an unparseable
URL is enough on Windows. Asserts are compiled out in release, so that was a
null dereference. It returns NULL now, and HTTPSRequest checks for it and fails
the request instead of handing it on.
Also: a request cancelled between construction and Start() didn't carry that
into the sink, and -1000 - our own cancellation code, not naett's - was logging
"Unhandled naett error" on a perfectly normal cancel.
Cancel() worked on the plain HTTP path - cancelled_ is threaded down as
progress->cancelled and checked while connecting and reading - but on the naett
path it only set a flag that nothing looked at. The transfer ran to completion
and cancelling just relabelled the result afterwards. Cancelling a store icon
that scrolled away, or a homebrew download the user gave up on, kept using the
bandwidth either way.
naett has one hook for this: a body writer that takes less than it was given
fails the request. So HTTPSRequest installs its own writer, which refuses
everything once cancelled. That works on all four backends and needs nothing
from naettClose, which is only really safe on Android.
The catch is lifetime. The writer runs on naett's transfer thread, and a request
that's still going when we're torn down would then be writing into a destroyed
HTTPSRequest - which is why the buffer it writes into is a separate refcounted
sink rather than a member. Join() on an unfinished request parks the sink where
it won't be freed and lets the request go, so a chunk that lands afterwards
writes somewhere that still exists.
That path is shutdown-only: RequestManager only cancels from its destructor, and
Update() waits for Done() before joining. Join now also notices a request that
finished while nobody was polling, and closes it properly instead of abandoning
it. What's left leaking at that point is a request still in flight as the
process exits, which is what already happened, just deliberate now and logged as
such rather than as an error.