Commit Graph
48061 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 21ee60a8ab Build: Update xxhash to v0.8.4.
Keeps our two local changes (the ppsspp_config.h include and the ARM32
prefetch in the XXH32 loop). Hash values are unchanged.

XXH3 is much faster on ARM64 now: about 37 GB/s against 13.5 with v0.8.1
on a Snapdragon X, where our quick texture hash does 24.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:04:08 -06:00
Henrik Rydgård 1b8553956f Merge pull request #22383 from hrydgard/preemptible-syscalls
Fixes after hardware scheduling tests by Claude
2026-09-29 12:35:27 -06:00
Henrik RydgårdandClaude Opus 5.5 5865a560eb ARM JITs: Mask what ctc1 writes to fcr31
Only the rounding mode, flags, enables, cause, FCC and FS bits can be
written (0x0181FFFF, pspautotests cpu/fpu/fcr), as the interpreter, IR
and x86 already have it. Both ARM JITs stored the whole value.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 11:09:47 -06:00
Henrik RydgårdandClaude Opus 5.5 178186ef4e sceVaudio: Reserving the channel waits about 250us
On hardware sceVaudioChReserve takes ~260us once it gets past the busy
check, succeed or fail, and worse threads can run meanwhile. Releasing
takes ~25us and doesn't wait. We returned at once, which is what
audio/sceaudio/reserve's [r] markers showed; it now passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 11:09:47 -06:00
Henrik RydgårdandClaude Opus 5.5 69e10728c9 test.py: Add gpu/primitives/indices32 to the to-do list
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 11:09:47 -06:00
Henrik RydgårdandClaude Opus 5.5 af8088b790 Threads: A delay can end before it starts waiting
A delay's deadline is now + usec, and the clock is read again when the
alarm is set. If the deadline has passed by then, the call returns 0 at
once without giving up the CPU. On hardware that makes
sceKernelDelayThread(0) return at once about 60% of the time. On a thread's
first wait after it starts, a delay of 1 does so about two times in three
as well (pspautotests threads/scheduling/delayzero). We always waited at
least 210us.

The choice is pseudo-random off the tick count, not the tick phase, since
our cycle counts are regular enough for a polling loop to lock into never
yielding. Threads remember whether they've waited since starting (Thread
savestate section version 6).

Also moves threads/vpl/create into the passing tests: re-recorded on 6.61,
it agrees with what we do for partitions 8 and 9.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 11:09:47 -06:00
Henrik RydgårdandClaude Opus 5.5 13dee85375 sceIo: host0: behaves like usbhostfs with dispatch suspended
With dispatch suspended, IO fails in the driver when it tries to wait. The
memory stick driver returns SCE_KERNEL_ERROR_CAN_NOT_WAIT, while usbhostfs,
which serves host0: under PSPLink, returns -1. host0: is mostly what
homebrew developers run from, so it now does the same.

Also from threads/scheduling/dispatch, which now passes:
- sceIoRead reports an async operation still in progress before failing
  on suspended dispatch.
- A write to stdout or stderr doesn't give up the CPU.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 11:09:47 -06:00
Henrik RydgårdandClaude Opus 5.5 f2eba8f711 sceKernelLoadModule: Charge for reading the file and loading it
On hardware a load costs an open and a read of the file, plus about 1ms
and 30us per KB of loader work (pspautotests threads/scheduling/callcosts),
with the caller waiting throughout. It now charges sceIoOpen's and
sceIoRead's estimates for the file plus that, instead of a flat 500us.

Also notes why sceKernelLoadModuleByID fails from a game's own fd on ms0:
or host0: on hardware, which we don't emulate.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 11:09:47 -06:00
Henrik Rydgård 806618f9a3 Merge pull request #22382 from hrydgard/gpu-lifecycle-fixes-2
Claude code review: GPU fixes 2
2026-09-29 11:04:38 -06:00
Henrik RydgårdandClaude Opus 5.5 93f57a4f38 GPU: Delete copy operations on classes that own resources
These own GPU objects, memory or refcounts in their destructors (or assert
there that they were torn down), so a copy would double-free. Nothing copies
them today; this keeps it that way. The manager base classes cover every
backend's subclass.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 39049a67fd GLES: Free textures on device lost, and unsubmitted step data at exit
- The texture and fragment test caches dropped their GLRTexture objects on
  DeviceLost without queueing them for deletion, leaking them on every
  Android background/resume. The deleter already skips the GL calls when
  the context is gone.
- GLRenderManager::ThreadEnd cleared unsubmitted init and render steps
  without freeing the data they own. Run them through the dry run instead,
  which now also frees stereo matrices and shader code.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 b72927bbeb Texture replacement: Plug leaks on error paths
- Release the file reference when loading a level fails or finds nothing,
  since only a loaded level takes ownership of it.
- Free the PNG image when the size changed since the header was read.
- Delete the VFS when a pack without an ini has no hash-named textures.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 e108e41675 GPU: Fix assorted leaks, null derefs and small rendering bugs
- Put the anisotropy level in the sampler key, so changing it applies on
  Vulkan and D3D11.
- Release CLUT textures at shutdown on GLES and D3D11.
- Fix the depth readback viewport, which squeezed the image whenever the
  read rectangle was smaller than the fbo.
- Test the computed depth, not the unset result, in the equal-depth clear
  check.
- Don't read back a CLUT from a framebuffer without an fbo.
- Tolerate null entries when releasing post-shader objects and CLUT
  textures after a failed creation.
- ImGe: Don't crash on a framebuffer without an fbo, or on GetVFB under the
  software renderer.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 71bc3187db Vulkan: Reload the shader cache after a device restore
DeviceLost saves the cache and clears everything, and the next save wrote
back only what was drawn since, so each Android background/resume cycle
shrank the cache.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 71210f3fa2 GE recorder: Don't let a second request hijack a recording
RecordNextFrame now refuses while a recording is active and during frame
dump playback (which asserted on the next replay). The callback handoff to
the CPU thread is locked, and a failed file open no longer crashes.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 d6f8615d2a SoftGPU: Fix block transfer and self-render overlap tracking
- The block transfer overlap check passed the stride in pixels where bytes
  are expected, so it only covered part of the rectangle.
- A selfrender/selfdepth flush in UpdateState dropped the current draw's
  pending writes and reads, so later transfers didn't wait for it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 507bb0801b GE: Complete lists dropped on error, flush before immediate draws
- A list dropped for a bad pc or a GE error stayed RUNNING: its ID was never
  freed and sceGeListSync on it never returned. Complete it like a finished
  one.
- FlushImm switches to through mode and another vertex decoder, so flush the
  queued draws first even when the immediate flags match.
- Clear leftover temporary GE breakpoints when setting or clearing the next
  break, so a step that never got there doesn't trip later.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 16ea2cbf88 GPU: Fix draw engine buffer overruns and stale vertex data
- Flush before the queued draws would decode more than VERTEX_BUFFER_MAX
  vertices. The batch was limited by index count, which doesn't bound a
  sparse index range, and DecodeVerts silently stopped while DecodeInds
  still emitted indices for the undecoded draws.
- Give TestBoundingBox its own scratch buffer. It used offsets in decoded_,
  which can hold decoded vertices that aren't flushed yet.
- Read 32-bit indices the way the PSP does, ignoring the upper 16 bits.
  IndexConverter and the fast bounding box test used all 32, so a game
  setting them indexed far past the decoded vertices.
- D3D11: Flush in FinishDeferred like the other backends, since indices
  are still read from PSP memory at flush time (#10095).
- Don't JIT new vertex decoders once the code space is full.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik Rydgård c40cce5173 Merge pull request #22381 from hrydgard/gpu-lifecycle-fixes
Claude code review: GPU lifecycle fixes
2026-09-29 10:24:27 -06:00
Henrik RydgårdandClaude Opus 5.5 4ba981b5b3 sceAtrac: Charge setting data and decoding what the ME takes
Setting data decodes and throws away the frames before the first sample,
so on hardware it costs a decoder setup plus that decode, with the caller
waiting: ~900us for mono Atrac3 and ~3.5ms for stereo Atrac3+, whatever the
buffer size (pspautotests threads/scheduling/callcosts). It was charged
100us.

Decoding a frame now costs what the same frame costs through
sceAudiocodecDecode, through the shared ME queue, instead of a flat
2300us. That's about the same for stereo Atrac3+ and less for Atrac3
(685us mono, ~1100us stereo). The first Atrac3+ frames after setup still
come out ~500us short.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:17:46 -06:00
Henrik Rydgård 52a16ea365 Make TempImage act a bit more robustly. 2026-09-29 10:16:08 -06:00
Henrik RydgårdandClaude Opus 5.5 c1e70c9686 Unloading a module is busy work, sceKernelVolatileMemTryLock is quick
From pspautotests threads/scheduling/syscallkinds:

- sceKernelUnloadModule takes ~400us that better threads can preempt
  and worse ones don't get in on, so it uses the busy delay rather than
  a wait.
- A successful sceKernelVolatileMemTryLock takes under 100us. It ate
  500000 cycles as a hack for Crash Tag Team Racing, which has since
  moved to (and no longer needs) the DrawSyncEatCycles compat flag.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:56:28 -06:00
Henrik RydgårdandClaude Opus 5.5 8301c35d7d Threads: Let better threads preempt a long sceKernelCreateThread
On hardware the kernel fills a new thread's stack with interrupts on, so
a better thread that wakes during it runs before the call returns, the
time it takes doesn't count towards the call, and worse threads get
nothing (pspautotests threads/scheduling/preemptsyscall). PPSSPP ate the
whole cost at once and only rescheduled at the end.

__KernelBusyDelayResult() models such a syscall: the caller waits, an
idle thread stands in for it while nothing better wants the CPU, and its
remaining cycles only count down while that's the case. When done it
goes back ahead of threads of its own priority, having never given up
the CPU. sceKernelCreateThread uses it for the stack fill, unless a
thread event handler is about to run.

Booting 75 games against master shows no difference.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:56:28 -06:00
Henrik Rydgård 16135e4e2f Merge pull request #22378 from hrydgard/intr-waits
Claude hardware testing: Interrupt waits
2026-09-29 09:56:08 -06:00
Henrik RydgårdandClaude Opus 5.5 24017a1ba4 GE debugger: Fix the stepping request handshake
- Wait on actionComplete instead of a bare condition variable wait, which
  could miss the wakeup and hang until resume.
- Serialize requesters, so two debuggers can't overwrite each other's
  action, and make SetCmdValue/FlushDrawing wait too.
- Give up and withdraw the request when stepping ends, instead of waiting
  forever (this deadlocked game shutdown against the Win32 GE debugger).
- Run requests during CPU stepping, which already accepted them.
- Clear the stepping state on Core_Resume from GE stepping and on
  GPU_Shutdown.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:49:10 -06:00
Henrik RydgårdandClaude Opus 5.5 24b71386ff Document that shader cache key changes need a CACHE_VERSION bump
The OpenGL and Vulkan shader caches store raw shader IDs (and, for Vulkan,
pipeline keys) on disk. Add the rule to AGENTS.md and point to it from the
persisted types.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 a9eacff1e1 SoftGPU: Keep the compile flushes on builds without the JIT
Skipping them is only a speed-up, but it changes how states get batched and
optimized, which changes the rendered output (NBA 2K13 and Virtua Tennis
frame dumps). Keep the old behaviour until that's understood.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 056fd05231 UnitTest: Make BlockAllocator run in about a second
Validating after every churn step was quadratic in the block count. Check
every 64 steps and at the end, and do fewer iterations.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 a9c782bfcf SoftGPU: Recompute the raster state after a JIT cache clear
A code space clear frees the functions the current state points to, but
the state was kept as long as the GE registers didn't change. Track the
clear generations and recompute, also when a compile during the state
computation clears the caches.

Also skip the binner flush for compiles on builds without the software
JIT, where Compile() does nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 45d997c3ea Texture replacement: Fix DDS mip double free and VFS lifetime
- DDS files with mips stored level 0's file reference in level 1 too, so it
  was freed twice.
- The VFS was deleted on config changes (and ini reloads) while load tasks
  still used it. Now each cached texture waits for its task and releases its
  file references through the old VFS first, and reloads afterwards.
- A failed ini reload turns replacement off instead of leaving it on without
  a VFS.
- Release file references on purge.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 4df1b9c2c3 Vulkan: Clear pipelines before shaders when use flags change
The pipeline deletion callbacks block on in-flight compiles, which use the
shader module promises that the shaders' deletion callbacks free. Queueing
the shaders first freed the promises under a pending compile.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 b912b5823b GPU: Fix texture cache and framebuffer cache lifetime bugs
- Delete TexCacheEntry objects dropped on rehash instead of leaking them.
- Don't leave a released null entry in cache_ when the framebuffer match
  returns before the slot is refilled.
- Reset clutRenderAddress_ in Clear(), which releases the dynamic CLUT FBOs.
- Don't cache a null texture in drawPixelsCache_ when creation fails.
- Fix the reversed subtraction in the failed-FBO retry check.
- Remove the never-taken buffered-rendering early-out in UpdateRenderSize.
  Taking it would leave existing VFBs without an fbo.
- Include smoothedDepal in the depal shader cache key, and print/parse the
  debug IDs as 64-bit.
- Release depal pipelines through Draw2DPipeline::Release so the shader
  source isn't leaked, and make that null-safe.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:35:17 -06:00
Henrik RydgårdandClaude Opus 5.5 5be18675e0 docs: Rebuild and re-record a whole pspautotests directory at a time
Rebuilding one test with the current SDK while its neighbours keep old
binaries hides toolchain differences until someone happens to touch
them, like the tests/intr module manager stubs that did nothing. So the
policy is now to rebuild every .prx in the directory and re-record them
all, diffing each .expected against the old one.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 8e291b8e12 Interrupts: Interrupt 8 has no handler on 6.61
intr/registersub, re-recorded on a 6.61 PSP with every test in the
directory rebuilt, finds no handler on interrupt 8 where the old
recording found one that didn't take user sub-interrupts. The old one was
probably made on an earlier firmware, whose drivers hooked it. Follow
6.61, the firmware PPSSPP models.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 8ed2d170f5 Audio: Don't emulate a failed blocking wait leaving the channel busy for good
When sceAudioOutputBlocking has to wait and the wait fails at once
(interrupts or dispatch disabled, inside an interrupt), the firmware
returns the error but leaves the channel's waiting flag set, and the
channel can never be used or released again. Keep the error, drop the
rest: whether the channel was busy at that moment is timing, and a
small difference in ours could lose a channel for the rest of a game
where hardware wouldn't. No game can depend on losing one.

Savestates made while this was emulated have the flag cleared on load.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 68dbbb253d sceDisplay: Vblank lasts 770us, not 731.5us
Measured with pspautotests display/vblanklen: 730-770us from
sceDisplayWaitVblankStart returning to the end of vblank, with an hcount
of up to 14 inside it. The old value dated from the first source drop
and left the highest hcount at 13. display/hcount now passes (with the
test fixed not to depend on where a line boundary falls).

Booting 75 games against master shows no difference.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 0928202310 scePower: CPU clock can't exceed the PLL, float frequency to the bit
scePowerSetCpuClockFrequency refuses a CPU clock above the PLL's, and
scePowerGetCpuClockFrequencyFloat computes pll * n / 511 in single
precision like the firmware, instead of converting whole Hz, which was
off in the last digit. power/freq now passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 960da1596c Threads: Charge for filling the stack on create, and for delete
From pspautotests threads/scheduling/costs: sceKernelCreateThread takes
about 150us plus roughly a cycle per byte of stack, which the kernel fills
with 0xFF (1.3ms for 256KB), and sceKernelDeleteThread 50-100us whatever
the stack size. Brings threads/scheduling/scheduling a good deal closer;
what's left needs a thread that wakes during a long syscall to preempt it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 436db85cf5 Reschedule on IO completion, drop the reschedule in time queries
The time queries rescheduled since 2013, so that a game spinning on the
clock would let a thread that a timing event had woken run (it fixed
audio in Crimson Gem Saga and Where Is My Heart?). In 2014 audio and
delay wakeups started rescheduling themselves, but IO completion never
did, and a movie reader thread in Driver 76 was only getting in through
the time queries. Now IO completion dispatches like any other wakeup.

The PSP doesn't dispatch in a time query, and doing so let a thread
that a terminate woke run too early. threads/threads/terminate now
passes.

Checked by booting 75 games against master: the same in all of them,
with Asphalt Urban GT2 getting further in the same time.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 638da935ea Threads: sceKernelStartThread hands the CPU straight to a better thread
When the new thread outranks the caller, the firmware switches to it
directly, even if a thread of still better priority is ready but hasn't
been dispatched (one that a sceKernelTerminateThread woke, say). Verified
against the new pspautotests threads/threads/termsuspended.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 23fc0cd422 Interrupts: Refuse sub-interrupt handlers where the firmware does
Only the GE and vblank interrupts take user sub-interrupt handlers, and
vblank only in slots 0-15, with some of the rest already held by the
kernel. The errors follow interruptman.prx's checks, and which interrupts
have handlers at all is read back from pspautotests intr/registersub and
intr/releasesub, which now pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 a74882b013 Audio/VolatileMem: Match hardware when a blocking call can't wait
From pspautotests intr/waits:

- sceAudioOutputBlocking sets the channel's waiting flag before its
  event flag wait, and when that wait fails at once (interrupts or
  dispatch disabled, or inside an interrupt) it returns the error
  without clearing the flag. The channel stays busy from then on, and
  can't be released.
- The SRC blocking output fails the same way even when a completion is
  already there, leaving the buffer armed.
- After a block that had samples in it, the mixer DMA is still playing
  it out, so a buffer arriving then isn't read early or restarts it.
- sceKernelVolatileMemLock only writes the fake address and size
  through pointers that are there, instead of faulting on NULL.

intr/waits now runs to the end; one scheduling marker still differs, from
async IO timing.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 5674c789ef sceUmd: Match hardware's parameter checks and wait timeouts
- A timeout of 0 to sceUmdWaitDriveStatWithTimer/CB means no timeout,
  not a tiny one (or 8ms for the CB version).
- Timeouts round like the event flag wait does.
- A wait with no timeout no longer times out right after a callback.
- sceUmdRegisterUMDCallBack only accepts callbacks.
- sceUmdActivate requires the name to be exactly "disc0:", and it and
  sceUmdDeactivate/sceUmdGetDiscInfo reject kernel pointers.
- sceUmdDeactivate needs a name in mode 2.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik Rydgård b04d03615b Merge pull request #22375 from hrydgard/startup-caching
Startup caching
2026-09-28 21:29:06 -06:00
Henrik Rydgård eab0b53a7a Merge pull request #22371 from hrydgard/audiocodec-fixes
sceAudiocodec and sceVideocodec timing and Atrac3+ fixes
2026-09-28 17:26:04 -06:00
Henrik Rydgård d3ab52b4b3 Buildfix 2026-09-28 17:05:44 -06:00
Henrik RydgårdandClaude Opus 5.5 023ad93ed3 sceVideocodec: Don't hold the ME for Init and Delete
They take tens of milliseconds for the caller, but queueing that time on
the shared ME timeline made the SAS mix wait behind them. In Jak and
Daxter that held up the sound threads at the end of the first clip, so
video_sound_thread got its last wake only after the game had deleted it
(NOT_DORMANT), and the orphaned thread then read a freed context.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 16:57:41 -06:00
Henrik RydgårdandClaude Opus 5.5 c2bc2d9308 ME: Charge measured times for sceVideocodec calls and the rest of sceAudiocodec
Measured on a PSP (pspautotests video/mp4/mp4timing, audio/audiocodec/timing):

- sceVideocodec Open, GetEDRAM, GetVersion and ReleaseEDRAM take ~70-150us,
  Init ~26.6ms (sceMpegCreate is 27-28ms), Delete ~21ms (was 2ms), and
  Stop 132us with nothing held back. All go through the ME queue now.
- Decodes that return no picture take as long as those that do; they
  were free.
- Open reports the EDRAM the decoder needs (0x3c2c) at ctx+0x18, which
  mpeg.prx passes on to GetEDRAM.
- sceAudiocodec: failed decodes (214/142/169us) and mono Atrac3+ init (524us).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 16:57:41 -06:00
Henrik RydgårdandClaude Opus 5.5 c162eb3d74 sceAudiocodec: Match hardware setup, framing, errors and timing; fix Atrac3 polarity
Checked against pspautotests audio/audiocodec, recorded on a PSP.

- Atrac3+: at3Related selects headered (mpeg.prx) or raw (libatrac3plus)
  frames, instead of sniffing for the sync word. The header's size field
  is 10 bits, as the context's. Header errors 0x211/0x213, bad frames
  0x20a, all returning SCE_AVCODEC_ERROR_INVALID_DATA with nothing read.
- The first successfully decoded Atrac3+ frame, and the first two AAC
  frames, produce no output. Checked sample-for-sample against hardware.
- Atrac3: the parameter at 0x28 selects the frame layout, as
  libatrac3plus.prx's table maps it. We used to read its low bit as a
  joint-stereo flag, which decoded mono (0x0F) streams as stereo garbage.
  AtracCtx2 had the table's fields swapped the same way.
- at3_standalone's Atrac3 output was inverted relative to the PSP's
  (sceAtrac too). Negate the IMDCT scale.
- CheckNeedMem sizes (AAC is 0x658c), codec 0x1004/0x1005, Init
  validation (AAC sample rate, Atrac3 parameter, Atrac3+ channels), and
  ReleaseEDRAM clearing edramAddr.
- Every call that reaches the ME now blocks for its measured time, and
  decode time is modelled per codec and frame size.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 16:57:41 -06:00
Henrik Rydgård 0f4bb62fa7 Merge pull request #22377 from hrydgard/callback-test-fixes
PSP kernel: Callback fixes
2026-09-28 16:24:06 -06:00