libatrac3plus picks the codec parameter from the frame size and the
header's joint stereo flag, and the channel count plays no part. We
guessed joint stereo from the frame size and channel count instead, so
LocoRoco 2's MuiMui house music never started: the game writes a 2-channel
normal-stereo header for every track it streams, and that 0xC0 track holds
one mono sound unit per frame. Taken as joint stereo, its first frame failed
during setup, and the game retried forever.
Now the joint stereo flag comes from the track header, and the decoder's
channel count from the parameter it maps to, so that track decodes as mono
into both output channels, as on hardware. Low-level decoding, which has
no header, still goes by the frame size. Atrac2 saves the flag; older
states fall back to the guess.
Fixes#8647.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The PSP blends save icons over a black background, so their transparent
parts come out black. 0a5fa27957 turned blending off instead, which is
wrong for icons that rely on it. Revert that, and draw a black rectangle
under each icon. Fixes#22280.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
There's no swapchain, so reading the backbuffer asserted. Fail the
readback, and have gpu.buffer.screenshot fall back to the displayed
framebuffer.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Homebrew written for a PSP-2000+ under custom firmware can use the top
32MB of RAM directly without setting MEMSIZE, which leaves the user
partition at its normal size. NJEMU's slim builds do this (#8925).
Previously we didn't map that memory at all for PBPs; MEMSIZE=1 isn't a
workaround either, since it grows the partition and the heap and stacks
land where the program writes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
NJEMU's SystemButtons.prx kernel plugin imports it when the firmware
reports 3.71 or later, and polls it every frame to read HOME/volume.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
sceNetAdhocctlInit waits until the friend finder has logged into the adhoc
server, polling in emulated time but giving up only after 5s of wall time.
The friend finder makes one attempt per login request, and it too waited
the full 5s on a connection that had already been refused, because it
only looked for success. Now it checks the socket error and gives up at
once, and records the failure so the wait ends with it.
In headless, which runs far ahead of real time, Gods Eater Burst sat in
this wait for its whole run.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Start from the quality that fit the previous frame instead of the top, so
most frames encode once. Step down while a frame is too big, and step back
up when one comes out under half the limit. Windows and the recompression
fallback share the logic.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Android, iOS, macOS and Linux encode camera frames at a fixed quality, so a
detailed frame can exceed the game's framesize just as on Windows before.
pushCameraImage now decodes such a frame and re-encodes it at lower quality
until it fits.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The friend finder thread set friendFinderRunning itself, after a DNS
lookup of the adhoc server. A shutdown in that window cleared the flag
first; the thread then set it again, looped forever, and the join in
NetAdhocctl_Term() never returned. Gods Eater Burst hit this in about one
headless run in six. The flag is now set before the thread is created,
and a finished thread is joined before a new one replaces it, which
would otherwise call std::terminate.
The built-in adhoc server thread had the same race, behind its check for
an existing server, and gets the same fix.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The PSP camera keeps JPEG frames within the framesize from the video setup.
Go!Edit stores frames in 15KB slots and only takes frames that fit, so our
uncompressed-quality webcam frames were dropped and one frame got repeated
for the whole clip. Lower the JPEG quality until a frame fits.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With a host microphone present, a blocking read waited until the host had
delivered all the data. If it never did, the thread waited forever: Go!Edit's
sound thread stalled that way and its video recording never advanced. The
PSP mic streams in real time, so wake at the scheduled time and fill what
the host didn't deliver with silence.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
It returned at once. Go!Edit reads frames in a loop on a high-priority
thread (bhCameraGetJpeg) that only yields to its own priority level, so
after the "Loading complete" dialog it spun and starved the rest of the
game. Return at the camera's next frame tick, at the rate from the setup
params.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Only the rounding mode, flags, enables, cause, FCC and FS bits can be
written (0x0181FFFF, pspautotests cpu/fpu/fcr), as the interpreter, IR
and x86 already have it. Both ARM JITs stored the whole value.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware sceVaudioChReserve takes ~260us once it gets past the busy
check, succeed or fail, and worse threads can run meanwhile. Releasing
takes ~25us and doesn't wait. We returned at once, which is what
audio/sceaudio/reserve's [r] markers showed; it now passes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A delay's deadline is now + usec, and the clock is read again when the
alarm is set. If the deadline has passed by then, the call returns 0 at
once without giving up the CPU. On hardware that makes
sceKernelDelayThread(0) return at once about 60% of the time. On a thread's
first wait after it starts, a delay of 1 does so about two times in three
as well (pspautotests threads/scheduling/delayzero). We always waited at
least 210us.
The choice is pseudo-random off the tick count, not the tick phase, since
our cycle counts are regular enough for a polling loop to lock into never
yielding. Threads remember whether they've waited since starting (Thread
savestate section version 6).
Also moves threads/vpl/create into the passing tests: re-recorded on 6.61,
it agrees with what we do for partitions 8 and 9.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With dispatch suspended, IO fails in the driver when it tries to wait. The
memory stick driver returns SCE_KERNEL_ERROR_CAN_NOT_WAIT, while usbhostfs,
which serves host0: under PSPLink, returns -1. host0: is mostly what
homebrew developers run from, so it now does the same.
Also from threads/scheduling/dispatch, which now passes:
- sceIoRead reports an async operation still in progress before failing
on suspended dispatch.
- A write to stdout or stderr doesn't give up the CPU.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware a load costs an open and a read of the file, plus about 1ms
and 30us per KB of loader work (pspautotests threads/scheduling/callcosts),
with the caller waiting throughout. It now charges sceIoOpen's and
sceIoRead's estimates for the file plus that, instead of a flat 500us.
Also notes why sceKernelLoadModuleByID fails from a game's own fd on ms0:
or host0: on hardware, which we don't emulate.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Setting data decodes and throws away the frames before the first sample,
so on hardware it costs a decoder setup plus that decode, with the caller
waiting: ~900us for mono Atrac3 and ~3.5ms for stereo Atrac3+, whatever the
buffer size (pspautotests threads/scheduling/callcosts). It was charged
100us.
Decoding a frame now costs what the same frame costs through
sceAudiocodecDecode, through the shared ME queue, instead of a flat
2300us. That's about the same for stereo Atrac3+ and less for Atrac3
(685us mono, ~1100us stereo). The first Atrac3+ frames after setup still
come out ~500us short.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
From pspautotests threads/scheduling/syscallkinds:
- sceKernelUnloadModule takes ~400us that better threads can preempt
and worse ones don't get in on, so it uses the busy delay rather than
a wait.
- A successful sceKernelVolatileMemTryLock takes under 100us. It ate
500000 cycles as a hack for Crash Tag Team Racing, which has since
moved to (and no longer needs) the DrawSyncEatCycles compat flag.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware the kernel fills a new thread's stack with interrupts on, so
a better thread that wakes during it runs before the call returns, the
time it takes doesn't count towards the call, and worse threads get
nothing (pspautotests threads/scheduling/preemptsyscall). PPSSPP ate the
whole cost at once and only rescheduled at the end.
__KernelBusyDelayResult() models such a syscall: the caller waits, an
idle thread stands in for it while nothing better wants the CPU, and its
remaining cycles only count down while that's the case. When done it
goes back ahead of threads of its own priority, having never given up
the CPU. sceKernelCreateThread uses it for the stack fill, unless a
thread event handler is about to run.
Booting 75 games against master shows no difference.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Wait on actionComplete instead of a bare condition variable wait, which
could miss the wakeup and hang until resume.
- Serialize requesters, so two debuggers can't overwrite each other's
action, and make SetCmdValue/FlushDrawing wait too.
- Give up and withdraw the request when stepping ends, instead of waiting
forever (this deadlocked game shutdown against the Win32 GE debugger).
- Run requests during CPU stepping, which already accepted them.
- Clear the stepping state on Core_Resume from GE stepping and on
GPU_Shutdown.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
intr/registersub, re-recorded on a 6.61 PSP with every test in the
directory rebuilt, finds no handler on interrupt 8 where the old
recording found one that didn't take user sub-interrupts. The old one was
probably made on an earlier firmware, whose drivers hooked it. Follow
6.61, the firmware PPSSPP models.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
When sceAudioOutputBlocking has to wait and the wait fails at once
(interrupts or dispatch disabled, inside an interrupt), the firmware
returns the error but leaves the channel's waiting flag set, and the
channel can never be used or released again. Keep the error, drop the
rest: whether the channel was busy at that moment is timing, and a
small difference in ours could lose a channel for the rest of a game
where hardware wouldn't. No game can depend on losing one.
Savestates made while this was emulated have the flag cleared on load.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Measured with pspautotests display/vblanklen: 730-770us from
sceDisplayWaitVblankStart returning to the end of vblank, with an hcount
of up to 14 inside it. The old value dated from the first source drop
and left the highest hcount at 13. display/hcount now passes (with the
test fixed not to depend on where a line boundary falls).
Booting 75 games against master shows no difference.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
scePowerSetCpuClockFrequency refuses a CPU clock above the PLL's, and
scePowerGetCpuClockFrequencyFloat computes pll * n / 511 in single
precision like the firmware, instead of converting whole Hz, which was
off in the last digit. power/freq now passes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
From pspautotests threads/scheduling/costs: sceKernelCreateThread takes
about 150us plus roughly a cycle per byte of stack, which the kernel fills
with 0xFF (1.3ms for 256KB), and sceKernelDeleteThread 50-100us whatever
the stack size. Brings threads/scheduling/scheduling a good deal closer;
what's left needs a thread that wakes during a long syscall to preempt it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The time queries rescheduled since 2013, so that a game spinning on the
clock would let a thread that a timing event had woken run (it fixed
audio in Crimson Gem Saga and Where Is My Heart?). In 2014 audio and
delay wakeups started rescheduling themselves, but IO completion never
did, and a movie reader thread in Driver 76 was only getting in through
the time queries. Now IO completion dispatches like any other wakeup.
The PSP doesn't dispatch in a time query, and doing so let a thread
that a terminate woke run too early. threads/threads/terminate now
passes.
Checked by booting 75 games against master: the same in all of them,
with Asphalt Urban GT2 getting further in the same time.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
When the new thread outranks the caller, the firmware switches to it
directly, even if a thread of still better priority is ready but hasn't
been dispatched (one that a sceKernelTerminateThread woke, say). Verified
against the new pspautotests threads/threads/termsuspended.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Only the GE and vblank interrupts take user sub-interrupt handlers, and
vblank only in slots 0-15, with some of the rest already held by the
kernel. The errors follow interruptman.prx's checks, and which interrupts
have handlers at all is read back from pspautotests intr/registersub and
intr/releasesub, which now pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
From pspautotests intr/waits:
- sceAudioOutputBlocking sets the channel's waiting flag before its
event flag wait, and when that wait fails at once (interrupts or
dispatch disabled, or inside an interrupt) it returns the error
without clearing the flag. The channel stays busy from then on, and
can't be released.
- The SRC blocking output fails the same way even when a completion is
already there, leaving the buffer armed.
- After a block that had samples in it, the mixer DMA is still playing
it out, so a buffer arriving then isn't read early or restarts it.
- sceKernelVolatileMemLock only writes the fake address and size
through pointers that are there, instead of faulting on NULL.
intr/waits now runs to the end; one scheduling marker still differs, from
async IO timing.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- A timeout of 0 to sceUmdWaitDriveStatWithTimer/CB means no timeout,
not a tiny one (or 8ms for the CB version).
- Timeouts round like the event flag wait does.
- A wait with no timeout no longer times out right after a callback.
- sceUmdRegisterUMDCallBack only accepts callbacks.
- sceUmdActivate requires the name to be exactly "disc0:", and it and
sceUmdDeactivate/sceUmdGetDiscInfo reject kernel pointers.
- sceUmdDeactivate needs a name in mode 2.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
They take tens of milliseconds for the caller, but queueing that time on
the shared ME timeline made the SAS mix wait behind them. In Jak and
Daxter that held up the sound threads at the end of the first clip, so
video_sound_thread got its last wake only after the game had deleted it
(NOT_DORMANT), and the orphaned thread then read a freed context.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Measured on a PSP (pspautotests video/mp4/mp4timing, audio/audiocodec/timing):
- sceVideocodec Open, GetEDRAM, GetVersion and ReleaseEDRAM take ~70-150us,
Init ~26.6ms (sceMpegCreate is 27-28ms), Delete ~21ms (was 2ms), and
Stop 132us with nothing held back. All go through the ME queue now.
- Decodes that return no picture take as long as those that do; they
were free.
- Open reports the EDRAM the decoder needs (0x3c2c) at ctx+0x18, which
mpeg.prx passes on to GetEDRAM.
- sceAudiocodec: failed decodes (214/142/169us) and mono Atrac3+ init (524us).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Checked against pspautotests audio/audiocodec, recorded on a PSP.
- Atrac3+: at3Related selects headered (mpeg.prx) or raw (libatrac3plus)
frames, instead of sniffing for the sync word. The header's size field
is 10 bits, as the context's. Header errors 0x211/0x213, bad frames
0x20a, all returning SCE_AVCODEC_ERROR_INVALID_DATA with nothing read.
- The first successfully decoded Atrac3+ frame, and the first two AAC
frames, produce no output. Checked sample-for-sample against hardware.
- Atrac3: the parameter at 0x28 selects the frame layout, as
libatrac3plus.prx's table maps it. We used to read its low bit as a
joint-stereo flag, which decoded mono (0x0F) streams as stereo garbage.
AtracCtx2 had the table's fields swapped the same way.
- at3_standalone's Atrac3 output was inverted relative to the PSP's
(sceAtrac too). Negate the IMDCT scale.
- CheckNeedMem sizes (AAC is 0x658c), codec 0x1004/0x1005, Init
validation (AAC sample rate, Atrac3 parameter, Atrac3+ channels), and
ReleaseEDRAM clearing edramAddr.
- Every call that reaches the ME now blocks for its measured time, and
decode time is modelled per codec and frame size.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The end callback checked the kernel object's lockThread, which for an
lwmutex is only refreshed by sceKernelReferLwMutexStatus. The lock state
lives in the workarea, so an unlock during the callback left the waiter
waiting forever. Verified against pspautotests threads/lwmutex/callbacks.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Replaces the global in-callback counter with each thread's own mipscall
chain, so several threads can be inside callbacks at once, and other
threads' callbacks (better priority ones right away) run while one is.
Verified against pspautotests threads/callbacks/otherthread, recursion
and intrnotify:
- A callback nests only one level: a CB wait that would go deeper never
returns on hardware, so the callback is left pending instead.
- A non-CB wait inside a callback no longer runs callbacks because of
the CB wait the callback interrupted.
- Callbacks for a waiting thread are only taken when it beats both the
running thread and every ready one. After an interrupt (which runs on
the idle thread) that's the thread about to resume, not idle.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Verified against pspautotests threads/callbacks/waittypes:
- __KernelThreadingInit() cleared the wait type callback table after
__KernelMemoryInit() had registered VPL and FPL in it, so a VPL or FPL
wait interrupted by a callback was never paused or resumed, and could
hang forever.
- A msgpipe deleted during a callback left its waiter waiting, instead
of waking it with WAIT_DELETE.
- A wait that got its object during a callback reported no time left;
put the timer back before trying to unlock, so the unlock writes what
remains.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Verified against pspautotests threads/callbacks/delivery:
- Notifying the callback of a better priority thread in a CB wait runs
it right away. Callbacks of other waiting threads stay pending until
those threads would get to run, rather than being taken at any
reschedule, so they can still be counted or canceled.
- sceKernelCancelCallback clears the notify count, not just the arg.
threads/callbacks/cancel, count and umd/wait/wait now pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Verified against new pspautotests threads/callbacks/afterwait and nested:
- A thread inside a callback runs its own pending callbacks (even the
same one again) nested, when it enters a CB wait. Waits paused by a
nested callback are keyed by the outer callback's id.
- sceKernelCheckCallback inside a callback returns ILLEGAL_CONTEXT
without running anything.
- sceKernelSleepThreadCB with a queued wakeup runs pending callbacks
before consuming it.
- A thread whose wait ended during a callback keeps the CPU, instead of
queueing behind threads of the same priority.
threads/callbacks/notify now passes too.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>