6739 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 a974440c79 Debugger: Time input.buttons.press in emulated vblanks
The release was counted down on the WebSocket thread, one step per poll
of host time however many vblanks had passed, so how long a scripted
press lasted depended on how fast the emulator ran, and scripted runs
went different ways. sceCtrl now releases it after that many vblank
samples, on the emulator thread; the debugger only reports when it's done.

Also: wsdbg's :screenshot works in headless with Vulkan.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:58:37 -06:00
Henrik RydgårdandClaude Opus 5.5 00c9a2d389 Utility: A savedata shutdown ends at priority 0x20
On hardware the last part of a savedata shutdown runs at priority 0x20,
whatever the dialog's own thread priorities, so a caller at 0x20 gets the
CPU back first and sees SHUTDOWN, and one at 0x21 or worse only sees NONE
(pspautotests utility/savedata/shutdownstatus). We ended it at the access
thread's priority, so Freak Out, which calls ShutdownStart from 0x20 and
waits for SHUTDOWN, sat at 'Please press START' forever. NFL Street 3,
which calls it from 111 and then InitStart straight away, still gets NONE.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:58:37 -06:00
Henrik RydgårdandClaude Opus 5.5 d86896bed9 Tlspl: Time out at once like other waits
Recorded on hardware (pspautotests threads/tls/timeout), a Tlspl
allocation follows the same timeout rule as the other waits, including
failing at once for 0 and 1us without writing the timeout back, which the
shared rule it moved to in the last commits didn't give it yet. Before
that it waited the raw timeout, ~30us short.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 80938529a9 Kernel waits: Share waiter ordering and clearing between objects
Priority-ordered waiting lists were sorted with a comparator wrapper per
object (msgpipe, fpl, vpl), or searched with a copy of the same function
(mutex, mbx). HLEKernel::SortWaitingThreadsByPriority() and
FindBestPriorityWaiter() now do both for any waiting list, of thread ids
or of structs with a threadID.

HLEKernel::ClearWaitingThreads() replaces the identical cancel/delete
loops in semaphores, event flags, fpl and vpl.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 cc5ee42dda Kernel waits: One timeout event for every kind of object
Semaphores, event flags, mutexes, lwmutexes, mbx, msgpipes, fpl, vpl,
tlspl and WaitThreadEnd each had their own CoreTiming event, handler
registration and savestate entry for wait timeouts, and their own function
to schedule one. Now one event (WaitThreadEnd's, renamed) times out all of
them, keyed by thread, and dispatches on the thread's wait type to a
timeoutFunc registered alongside the begin/end callback functions.
__KernelWaitCurThreadWithTimeout() starts such a wait, and the HLEKernel
helpers have overloads that use the shared event.

Old savestates still load: each object's section reads its old event id
and points it at the shared handler, so a timeout pending in the state
goes off as before. Checked with a state saved mid-wait by the previous
build, and with four games.

The one behaviour change: tlspl timeouts now follow the same hardware
rule as the others, where they used the raw timeout.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 804dd59df0 Interrupts: Charge for alarm handlers, and stop parking threads on idle
An interrupt with no handler to run, a vblank with none registered for
example, switched the running thread off to idle and left it there until
some later event rescheduled: ~775us of every frame in a game that spins
without a vblank handler. It now reschedules at once. Taking an interrupt
also clears the ll bit directly, which that switch had been doing.

Interrupt handlers can now carry a cost before they run and after the
last queued one returns. Alarms use it: on hardware a thread that keeps
running loses ~70us to an alarm handler, and a thread the handler wakes
runs ~50us after it (pspautotests threads/scheduling/alarmcosts), so
17us in and 40us out. sceKernelSetAlarm's 40us is split evenly around the
deadline, keeping the handler ~1040us after a 1000us alarm.

A handler's return value re-arms its alarm counting from the previous
deadline, so a repeating alarm doesn't drift by those costs, unless
that's already past, as after interrupts were suspended for a while.

The vblank's own cost (~62us of CPU on hardware) isn't charged yet: with
it, a waiter ~90us after the vblank still reads hcount 1 on hardware, but
line 2 here. Hardware evidently raises the interrupt ~40us before the
line count wraps. That's noted where the waiters are released.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 02b2b7b699 sceDisplay: Release vblank waiters ~50us after the vblank
On hardware a thread waiting for vblank returns ~53us after it, where we
had it back in ~5us, and the first of four waiters runs after ~85us: each
waiter beyond the first adds ~9us (pspautotests threads/scheduling/
vblankwake). The waiters are now released by a separate event 48us + 9us
per extra waiter after the vblank. Which vblank a wait is for is still
decided at the vblank, so a thread that starts waiting in between still
waits a whole frame (sceDisplay section version 8).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 f26ca37dad Threads: Context switches cost what they do on hardware
Timed on a PSP, each way of handing the CPU to another thread (the call
and the switch together):

                                     hardware   before   now
  rotate to an equal thread              7        14       7
  signal, better thread runs            10        17      10
  it waits again, back to caller        10        19      12
  wakeup, better thread runs             8        13       6
  it sleeps again, back to caller        7        12       6
  start a better thread, entry          30        28      30
  thread ends, back to its waiter       21        13      20
  notify, better thread's callback      14        13      14

A switch between two threads now costs 1150 cycles instead of 2700.
Starting a better thread costs 2000 cycles more, ending a thread 3300,
and setting up a callback 1800.

Also splits a wait timeout's ~30us into the deadline being taken 12us
into the call and the timeout going off 18us after it. That only changes
the time left written back, which threads/semaphores/wait and
threads/fpl/cancel pin between them. intr/vblank is re-recorded so it
no longer depends on the phase of the frame.

threads/callbacks/combos now passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 a17a9eb147 Callbacks: A notify takes the thread out of its CB wait right away
On hardware, notifying a callback of a thread in a CB wait takes it out of
the wait at once, even though the callback only runs when the thread would
get the CPU. A semaphore signalled in between doesn't end the wait: the
callback runs first, then the wait resumes and takes it (pspautotests
threads/callbacks/combos). We left the thread on the wait list until the
callback started, so the signal ended the wait and the callback didn't
run.

The notify now pauses the wait, as starting a callback used to. If the
callbacks are canceled before the thread's turn comes, the wait just
resumes (Thread savestate section version 7).

threads/callbacks/combos goes in the to-do list: a callback returning to
the thread that notified it still takes ~13us where hardware takes ~9,
part of the context switch cost.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 fb9ac8397e Threads: Wait timeouts work like hardware's, one rule for all of them
Every wait with a timeout behaves the same on hardware (pspautotests
threads/scheduling/waittimeouts). The deadline is taken, and the alarm set
up a moment later. If the deadline has passed by then, the wait fails with
WAIT_TIMEOUT at once, without yielding or writing the timeout back. That's
usual for 0us, half the time for 1us, and rare after; AllocateVpl does more
first. Otherwise it ends max(t, 205us) + ~35us after the call. Each object
had its own guess (24/245, 25/250, 20/250 and so on), and only MsgPipe had
the immediate case.

__KernelWaitTimesOutAtOnce() and __KernelWaitTimeoutUs() now do it for
semaphores, event flags, mutexes, lwmutexes, mbx, msgpipes, fpl, vpl and
WaitThreadEnd. The latency past the deadline isn't counted in the time
left written back.

Outcomes that hardware decides by the clock's phase (these, and
sceKernelDelayThread returning at once) go with the likelier one. Ones
between 50% and certain are instead spread evenly over calls, so a polling
loop can't lock into never yielding (sceKernelThread section version 7).
This replaces the pseudo-random choice for delays.

Also adds threads/scheduling/readyqueue, which already passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 f731e45f40 Alarms: Setting one takes ~40us, and it can't go off within ~215us
On hardware sceKernelSetAlarm takes about 40us, and the handler never runs
sooner than about 215us after the alarm is set, however short it asked
for (pspautotests threads/scheduling/alarmcosts). Also clamps huge
sysclock alarms before converting to cycles; LONG_LONG_MAX used to
overflow and go off at once, which the late-firing events had hidden.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 9ba40416da sceAtrac: Charge ME time when SetData's first frame doesn't decode
That failure comes after the codec is set up and the frame has been tried
on the ME, so the thread waits for both, unlike the other SetData errors.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 17:14:03 -06:00
Henrik RydgårdandClaude Opus 5.5 b9ea98d2c5 Atrac3: Set up the decoder the way libatrac3plus does
libatrac3plus picks the codec parameter from the frame size and the
header's joint stereo flag, and the channel count plays no part. We
guessed joint stereo from the frame size and channel count instead, so
LocoRoco 2's MuiMui house music never started: the game writes a 2-channel
normal-stereo header for every track it streams, and that 0xC0 track holds
one mono sound unit per frame. Taken as joint stereo, its first frame failed
during setup, and the game retried forever.

Now the joint stereo flag comes from the track header, and the decoder's
channel count from the parameter it maps to, so that track decodes as mono
into both output channels, as on hardware. Low-level decoding, which has
no header, still goes by the frame size. Atrac2 saves the flag; older
states fall back to the guess.

Fixes #8647.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 17:14:00 -06:00
Henrik Rydgård b9c5b28b8f Merge pull request #22386 from hrydgard/adhoc-shutdown-race
Adhoc: Fix shutdown hanging when it comes right after adhoc init
2026-09-29 13:29:20 -06:00
Henrik Rydgård 0c6a28fe11 Merge pull request #22387 from hrydgard/homebrew-slim-extra-ram
Expose extra ram to homebrew apps if Slim model is emulated
2026-09-29 13:13:10 -06:00
Henrik Rydgård 36cbb32e77 Merge pull request #22385 from hrydgard/goedit-camera-hang
Fix Go!Edit camera hang, and some other related issues
2026-09-29 13:00:49 -06:00
Henrik RydgårdandClaude Opus 5.5 3062ffd695 Map the PSP-2000 extra RAM for homebrew PBPs, outside the user partition
Homebrew written for a PSP-2000+ under custom firmware can use the top
32MB of RAM directly without setting MEMSIZE, which leaves the user
partition at its normal size. NJEMU's slim builds do this (#8925).
Previously we didn't map that memory at all for PBPs; MEMSIZE=1 isn't a
workaround either, since it grows the partition and the heap and stacks
land where the program writes.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 12:40:59 -06:00
Henrik RydgårdandClaude Opus 5.5 c9a631c336 sceCtrl: Add the 3.71 sceCtrl_driver NID for sceCtrlPeekBufferPositive
NJEMU's SystemButtons.prx kernel plugin imports it when the firmware
reports 3.71 or later, and polls it every frame to read HOME/volume.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 12:40:59 -06:00
Henrik RydgårdandClaude Opus 5.5 7ec860c24f Adhoc: Stop waiting for a login that has already failed
sceNetAdhocctlInit waits until the friend finder has logged into the adhoc
server, polling in emulated time but giving up only after 5s of wall time.
The friend finder makes one attempt per login request, and it too waited
the full 5s on a connection that had already been refused, because it
only looked for success. Now it checks the socket error and gives up at
once, and records the failure so the wait ends with it.

In headless, which runs far ahead of real time, Gods Eater Burst sat in
this wait for its whole run.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 12:37:48 -06:00
Henrik RydgårdandClaude Opus 5.5 9e54cebcc2 Camera: Remember the JPEG quality that fit the last frame
Start from the quality that fit the previous frame instead of the top, so
most frames encode once. Step down while a frame is too big, and step back
up when one comes out under half the limit. Windows and the recompression
fallback share the logic.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 12:29:04 -06:00
Henrik RydgårdandClaude Opus 5.5 07556d064e Camera: Recompress oversized frames on every platform
Android, iOS, macOS and Linux encode camera frames at a fixed quality, so a
detailed frame can exceed the game's framesize just as on Windows before.
pushCameraImage now decodes such a frame and re-encodes it at lower quality
until it fits.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 12:24:40 -06:00
Henrik RydgårdandClaude Opus 5.5 2cd84012d0 Adhoc: Fix shutdown hanging when it comes right after adhoc init
The friend finder thread set friendFinderRunning itself, after a DNS
lookup of the adhoc server. A shutdown in that window cleared the flag
first; the thread then set it again, looped forever, and the join in
NetAdhocctl_Term() never returned. Gods Eater Burst hit this in about one
headless run in six. The flag is now set before the thread is created,
and a finished thread is joined before a new one replaces it, which
would otherwise call std::terminate.

The built-in adhoc server thread had the same race, behind its check for
an existing server, and gets the same fix.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 11:58:42 -06:00
Henrik RydgårdandClaude Opus 5.5 9d33e828ab Camera: Compress webcam frames to fit the game's framesize
The PSP camera keeps JPEG frames within the framesize from the video setup.
Go!Edit stores frames in 15KB slots and only takes frames that fit, so our
uncompressed-quality webcam frames were dropped and one frame got repeated
for the whole clip. Lower the JPEG quality until a frame fits.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 11:44:10 -06:00
Henrik RydgårdandClaude Opus 5.5 948062907b sceUsbMic: Complete blocking reads when the samples are due
With a host microphone present, a blocking read waited until the host had
delivered all the data. If it never did, the thread waited forever: Go!Edit's
sound thread stalled that way and its video recording never advanced. The
PSP mic streams in real time, so wake at the scheduled time and fill what
the host didn't deliver with silence.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 11:44:10 -06:00
Henrik RydgårdandClaude Opus 5.5 4504092d3d sceUsbMic: Don't crash when there's no Windows capture device object
Only the app creates winMic, so a blocking mic read crashed headless.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 11:19:18 -06:00
Henrik RydgårdandClaude Opus 5.5 280ae09856 sceUsbCam: Make sceUsbCamReadVideoFrameBlocking wait for the next frame
It returned at once. Go!Edit reads frames in a loop on a high-priority
thread (bhCameraGetJpeg) that only yields to its own priority level, so
after the "Loading complete" dialog it spun and starved the rest of the
game. Return at the camera's next frame tick, at the rate from the setup
params.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 11:19:18 -06:00
Henrik RydgårdandClaude Opus 5.5 178186ef4e sceVaudio: Reserving the channel waits about 250us
On hardware sceVaudioChReserve takes ~260us once it gets past the busy
check, succeed or fail, and worse threads can run meanwhile. Releasing
takes ~25us and doesn't wait. We returned at once, which is what
audio/sceaudio/reserve's [r] markers showed; it now passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 11:09:47 -06:00
Henrik RydgårdandClaude Opus 5.5 af8088b790 Threads: A delay can end before it starts waiting
A delay's deadline is now + usec, and the clock is read again when the
alarm is set. If the deadline has passed by then, the call returns 0 at
once without giving up the CPU. On hardware that makes
sceKernelDelayThread(0) return at once about 60% of the time. On a thread's
first wait after it starts, a delay of 1 does so about two times in three
as well (pspautotests threads/scheduling/delayzero). We always waited at
least 210us.

The choice is pseudo-random off the tick count, not the tick phase, since
our cycle counts are regular enough for a polling loop to lock into never
yielding. Threads remember whether they've waited since starting (Thread
savestate section version 6).

Also moves threads/vpl/create into the passing tests: re-recorded on 6.61,
it agrees with what we do for partitions 8 and 9.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 11:09:47 -06:00
Henrik RydgårdandClaude Opus 5.5 13dee85375 sceIo: host0: behaves like usbhostfs with dispatch suspended
With dispatch suspended, IO fails in the driver when it tries to wait. The
memory stick driver returns SCE_KERNEL_ERROR_CAN_NOT_WAIT, while usbhostfs,
which serves host0: under PSPLink, returns -1. host0: is mostly what
homebrew developers run from, so it now does the same.

Also from threads/scheduling/dispatch, which now passes:
- sceIoRead reports an async operation still in progress before failing
  on suspended dispatch.
- A write to stdout or stderr doesn't give up the CPU.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 11:09:47 -06:00
Henrik RydgårdandClaude Opus 5.5 f2eba8f711 sceKernelLoadModule: Charge for reading the file and loading it
On hardware a load costs an open and a read of the file, plus about 1ms
and 30us per KB of loader work (pspautotests threads/scheduling/callcosts),
with the caller waiting throughout. It now charges sceIoOpen's and
sceIoRead's estimates for the file plus that, instead of a flat 500us.

Also notes why sceKernelLoadModuleByID fails from a game's own fd on ms0:
or host0: on hardware, which we don't emulate.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 11:09:47 -06:00
Henrik RydgårdandClaude Opus 5.5 4ba981b5b3 sceAtrac: Charge setting data and decoding what the ME takes
Setting data decodes and throws away the frames before the first sample,
so on hardware it costs a decoder setup plus that decode, with the caller
waiting: ~900us for mono Atrac3 and ~3.5ms for stereo Atrac3+, whatever the
buffer size (pspautotests threads/scheduling/callcosts). It was charged
100us.

Decoding a frame now costs what the same frame costs through
sceAudiocodecDecode, through the shared ME queue, instead of a flat
2300us. That's about the same for stereo Atrac3+ and less for Atrac3
(685us mono, ~1100us stereo). The first Atrac3+ frames after setup still
come out ~500us short.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:17:46 -06:00
Henrik RydgårdandClaude Opus 5.5 c1e70c9686 Unloading a module is busy work, sceKernelVolatileMemTryLock is quick
From pspautotests threads/scheduling/syscallkinds:

- sceKernelUnloadModule takes ~400us that better threads can preempt
  and worse ones don't get in on, so it uses the busy delay rather than
  a wait.
- A successful sceKernelVolatileMemTryLock takes under 100us. It ate
  500000 cycles as a hack for Crash Tag Team Racing, which has since
  moved to (and no longer needs) the DrawSyncEatCycles compat flag.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:56:28 -06:00
Henrik RydgårdandClaude Opus 5.5 8301c35d7d Threads: Let better threads preempt a long sceKernelCreateThread
On hardware the kernel fills a new thread's stack with interrupts on, so
a better thread that wakes during it runs before the call returns, the
time it takes doesn't count towards the call, and worse threads get
nothing (pspautotests threads/scheduling/preemptsyscall). PPSSPP ate the
whole cost at once and only rescheduled at the end.

__KernelBusyDelayResult() models such a syscall: the caller waits, an
idle thread stands in for it while nothing better wants the CPU, and its
remaining cycles only count down while that's the case. When done it
goes back ahead of threads of its own priority, having never given up
the CPU. sceKernelCreateThread uses it for the stack fill, unless a
thread event handler is about to run.

Booting 75 games against master shows no difference.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:56:28 -06:00
Henrik RydgårdandClaude Opus 5.5 8e291b8e12 Interrupts: Interrupt 8 has no handler on 6.61
intr/registersub, re-recorded on a 6.61 PSP with every test in the
directory rebuilt, finds no handler on interrupt 8 where the old
recording found one that didn't take user sub-interrupts. The old one was
probably made on an earlier firmware, whose drivers hooked it. Follow
6.61, the firmware PPSSPP models.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 8ed2d170f5 Audio: Don't emulate a failed blocking wait leaving the channel busy for good
When sceAudioOutputBlocking has to wait and the wait fails at once
(interrupts or dispatch disabled, inside an interrupt), the firmware
returns the error but leaves the channel's waiting flag set, and the
channel can never be used or released again. Keep the error, drop the
rest: whether the channel was busy at that moment is timing, and a
small difference in ours could lose a channel for the rest of a game
where hardware wouldn't. No game can depend on losing one.

Savestates made while this was emulated have the flag cleared on load.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 68dbbb253d sceDisplay: Vblank lasts 770us, not 731.5us
Measured with pspautotests display/vblanklen: 730-770us from
sceDisplayWaitVblankStart returning to the end of vblank, with an hcount
of up to 14 inside it. The old value dated from the first source drop
and left the highest hcount at 13. display/hcount now passes (with the
test fixed not to depend on where a line boundary falls).

Booting 75 games against master shows no difference.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 0928202310 scePower: CPU clock can't exceed the PLL, float frequency to the bit
scePowerSetCpuClockFrequency refuses a CPU clock above the PLL's, and
scePowerGetCpuClockFrequencyFloat computes pll * n / 511 in single
precision like the firmware, instead of converting whole Hz, which was
off in the last digit. power/freq now passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 960da1596c Threads: Charge for filling the stack on create, and for delete
From pspautotests threads/scheduling/costs: sceKernelCreateThread takes
about 150us plus roughly a cycle per byte of stack, which the kernel fills
with 0xFF (1.3ms for 256KB), and sceKernelDeleteThread 50-100us whatever
the stack size. Brings threads/scheduling/scheduling a good deal closer;
what's left needs a thread that wakes during a long syscall to preempt it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 436db85cf5 Reschedule on IO completion, drop the reschedule in time queries
The time queries rescheduled since 2013, so that a game spinning on the
clock would let a thread that a timing event had woken run (it fixed
audio in Crimson Gem Saga and Where Is My Heart?). In 2014 audio and
delay wakeups started rescheduling themselves, but IO completion never
did, and a movie reader thread in Driver 76 was only getting in through
the time queries. Now IO completion dispatches like any other wakeup.

The PSP doesn't dispatch in a time query, and doing so let a thread
that a terminate woke run too early. threads/threads/terminate now
passes.

Checked by booting 75 games against master: the same in all of them,
with Asphalt Urban GT2 getting further in the same time.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 638da935ea Threads: sceKernelStartThread hands the CPU straight to a better thread
When the new thread outranks the caller, the firmware switches to it
directly, even if a thread of still better priority is ready but hasn't
been dispatched (one that a sceKernelTerminateThread woke, say). Verified
against the new pspautotests threads/threads/termsuspended.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 23fc0cd422 Interrupts: Refuse sub-interrupt handlers where the firmware does
Only the GE and vblank interrupts take user sub-interrupt handlers, and
vblank only in slots 0-15, with some of the rest already held by the
kernel. The errors follow interruptman.prx's checks, and which interrupts
have handlers at all is read back from pspautotests intr/registersub and
intr/releasesub, which now pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 a74882b013 Audio/VolatileMem: Match hardware when a blocking call can't wait
From pspautotests intr/waits:

- sceAudioOutputBlocking sets the channel's waiting flag before its
  event flag wait, and when that wait fails at once (interrupts or
  dispatch disabled, or inside an interrupt) it returns the error
  without clearing the flag. The channel stays busy from then on, and
  can't be released.
- The SRC blocking output fails the same way even when a completion is
  already there, leaving the buffer armed.
- After a block that had samples in it, the mixer DMA is still playing
  it out, so a buffer arriving then isn't read early or restarts it.
- sceKernelVolatileMemLock only writes the fake address and size
  through pointers that are there, instead of faulting on NULL.

intr/waits now runs to the end; one scheduling marker still differs, from
async IO timing.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 5674c789ef sceUmd: Match hardware's parameter checks and wait timeouts
- A timeout of 0 to sceUmdWaitDriveStatWithTimer/CB means no timeout,
  not a tiny one (or 8ms for the CB version).
- Timeouts round like the event flag wait does.
- A wait with no timeout no longer times out right after a callback.
- sceUmdRegisterUMDCallBack only accepts callbacks.
- sceUmdActivate requires the name to be exactly "disc0:", and it and
  sceUmdDeactivate/sceUmdGetDiscInfo reject kernel pointers.
- sceUmdDeactivate needs a name in mode 2.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik Rydgård eab0b53a7a Merge pull request #22371 from hrydgard/audiocodec-fixes
sceAudiocodec and sceVideocodec timing and Atrac3+ fixes
2026-09-28 17:26:04 -06:00
Henrik RydgårdandClaude Opus 5.5 023ad93ed3 sceVideocodec: Don't hold the ME for Init and Delete
They take tens of milliseconds for the caller, but queueing that time on
the shared ME timeline made the SAS mix wait behind them. In Jak and
Daxter that held up the sound threads at the end of the first clip, so
video_sound_thread got its last wake only after the game had deleted it
(NOT_DORMANT), and the orphaned thread then read a freed context.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 16:57:41 -06:00
Henrik RydgårdandClaude Opus 5.5 c2bc2d9308 ME: Charge measured times for sceVideocodec calls and the rest of sceAudiocodec
Measured on a PSP (pspautotests video/mp4/mp4timing, audio/audiocodec/timing):

- sceVideocodec Open, GetEDRAM, GetVersion and ReleaseEDRAM take ~70-150us,
  Init ~26.6ms (sceMpegCreate is 27-28ms), Delete ~21ms (was 2ms), and
  Stop 132us with nothing held back. All go through the ME queue now.
- Decodes that return no picture take as long as those that do; they
  were free.
- Open reports the EDRAM the decoder needs (0x3c2c) at ctx+0x18, which
  mpeg.prx passes on to GetEDRAM.
- sceAudiocodec: failed decodes (214/142/169us) and mono Atrac3+ init (524us).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 16:57:41 -06:00
Henrik RydgårdandClaude Opus 5.5 c162eb3d74 sceAudiocodec: Match hardware setup, framing, errors and timing; fix Atrac3 polarity
Checked against pspautotests audio/audiocodec, recorded on a PSP.

- Atrac3+: at3Related selects headered (mpeg.prx) or raw (libatrac3plus)
  frames, instead of sniffing for the sync word. The header's size field
  is 10 bits, as the context's. Header errors 0x211/0x213, bad frames
  0x20a, all returning SCE_AVCODEC_ERROR_INVALID_DATA with nothing read.
- The first successfully decoded Atrac3+ frame, and the first two AAC
  frames, produce no output. Checked sample-for-sample against hardware.
- Atrac3: the parameter at 0x28 selects the frame layout, as
  libatrac3plus.prx's table maps it. We used to read its low bit as a
  joint-stereo flag, which decoded mono (0x0F) streams as stereo garbage.
  AtracCtx2 had the table's fields swapped the same way.
- at3_standalone's Atrac3 output was inverted relative to the PSP's
  (sceAtrac too). Negate the IMDCT scale.
- CheckNeedMem sizes (AAC is 0x658c), codec 0x1004/0x1005, Init
  validation (AAC sample rate, Atrac3 parameter, Atrac3+ channels), and
  ReleaseEDRAM clearing edramAddr.
- Every call that reaches the ME now blocks for its measured time, and
  decode time is modelled per codec and frame size.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 16:57:41 -06:00
Henrik RydgårdandClaude Opus 5.5 c11e46aeea LwMutex: Take the lock after a callback if it was released during it
The end callback checked the kernel object's lockThread, which for an
lwmutex is only refreshed by sceKernelReferLwMutexStatus. The lock state
lives in the workarea, so an unlock during the callback left the waiter
waiting forever. Verified against pspautotests threads/lwmutex/callbacks.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 15:59:25 -06:00
Henrik RydgårdandClaude Opus 5.5 eda92c4ef2 Callbacks: Track callback nesting per thread, one level deep
Replaces the global in-callback counter with each thread's own mipscall
chain, so several threads can be inside callbacks at once, and other
threads' callbacks (better priority ones right away) run while one is.
Verified against pspautotests threads/callbacks/otherthread, recursion
and intrnotify:

- A callback nests only one level: a CB wait that would go deeper never
  returns on hardware, so the callback is left pending instead.
- A non-CB wait inside a callback no longer runs callbacks because of
  the CB wait the callback interrupted.
- Callbacks for a waiting thread are only taken when it beats both the
  running thread and every ready one. After an interrupt (which runs on
  the idle thread) that's the thread about to resume, not idle.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 15:59:25 -06:00
Henrik RydgårdandClaude Opus 5.5 24b41d3f23 Kernel waits: Fix VPL/FPL and msgpipe waits around callbacks, report timeout left
Verified against pspautotests threads/callbacks/waittypes:

- __KernelThreadingInit() cleared the wait type callback table after
  __KernelMemoryInit() had registered VPL and FPL in it, so a VPL or FPL
  wait interrupted by a callback was never paused or resumed, and could
  hang forever.
- A msgpipe deleted during a callback left its waiter waiting, instead
  of waking it with WAIT_DELETE.
- A wait that got its object during a callback reported no time left;
  put the timer back before trying to unlock, so the unlock writes what
  remains.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 15:59:25 -06:00