On hardware a load costs an open and a read of the file, plus about 1ms
and 30us per KB of loader work (pspautotests threads/scheduling/callcosts),
with the caller waiting throughout. It now charges sceIoOpen's and
sceIoRead's estimates for the file plus that, instead of a flat 500us.
Also notes why sceKernelLoadModuleByID fails from a game's own fd on ms0:
or host0: on hardware, which we don't emulate.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Setting data decodes and throws away the frames before the first sample,
so on hardware it costs a decoder setup plus that decode, with the caller
waiting: ~900us for mono Atrac3 and ~3.5ms for stereo Atrac3+, whatever the
buffer size (pspautotests threads/scheduling/callcosts). It was charged
100us.
Decoding a frame now costs what the same frame costs through
sceAudiocodecDecode, through the shared ME queue, instead of a flat
2300us. That's about the same for stereo Atrac3+ and less for Atrac3
(685us mono, ~1100us stereo). The first Atrac3+ frames after setup still
come out ~500us short.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
From pspautotests threads/scheduling/syscallkinds:
- sceKernelUnloadModule takes ~400us that better threads can preempt
and worse ones don't get in on, so it uses the busy delay rather than
a wait.
- A successful sceKernelVolatileMemTryLock takes under 100us. It ate
500000 cycles as a hack for Crash Tag Team Racing, which has since
moved to (and no longer needs) the DrawSyncEatCycles compat flag.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware the kernel fills a new thread's stack with interrupts on, so
a better thread that wakes during it runs before the call returns, the
time it takes doesn't count towards the call, and worse threads get
nothing (pspautotests threads/scheduling/preemptsyscall). PPSSPP ate the
whole cost at once and only rescheduled at the end.
__KernelBusyDelayResult() models such a syscall: the caller waits, an
idle thread stands in for it while nothing better wants the CPU, and its
remaining cycles only count down while that's the case. When done it
goes back ahead of threads of its own priority, having never given up
the CPU. sceKernelCreateThread uses it for the stack fill, unless a
thread event handler is about to run.
Booting 75 games against master shows no difference.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Rebuilding one test with the current SDK while its neighbours keep old
binaries hides toolchain differences until someone happens to touch
them, like the tests/intr module manager stubs that did nothing. So the
policy is now to rebuild every .prx in the directory and re-record them
all, diffing each .expected against the old one.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
intr/registersub, re-recorded on a 6.61 PSP with every test in the
directory rebuilt, finds no handler on interrupt 8 where the old
recording found one that didn't take user sub-interrupts. The old one was
probably made on an earlier firmware, whose drivers hooked it. Follow
6.61, the firmware PPSSPP models.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
When sceAudioOutputBlocking has to wait and the wait fails at once
(interrupts or dispatch disabled, inside an interrupt), the firmware
returns the error but leaves the channel's waiting flag set, and the
channel can never be used or released again. Keep the error, drop the
rest: whether the channel was busy at that moment is timing, and a
small difference in ours could lose a channel for the rest of a game
where hardware wouldn't. No game can depend on losing one.
Savestates made while this was emulated have the flag cleared on load.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Measured with pspautotests display/vblanklen: 730-770us from
sceDisplayWaitVblankStart returning to the end of vblank, with an hcount
of up to 14 inside it. The old value dated from the first source drop
and left the highest hcount at 13. display/hcount now passes (with the
test fixed not to depend on where a line boundary falls).
Booting 75 games against master shows no difference.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
scePowerSetCpuClockFrequency refuses a CPU clock above the PLL's, and
scePowerGetCpuClockFrequencyFloat computes pll * n / 511 in single
precision like the firmware, instead of converting whole Hz, which was
off in the last digit. power/freq now passes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
From pspautotests threads/scheduling/costs: sceKernelCreateThread takes
about 150us plus roughly a cycle per byte of stack, which the kernel fills
with 0xFF (1.3ms for 256KB), and sceKernelDeleteThread 50-100us whatever
the stack size. Brings threads/scheduling/scheduling a good deal closer;
what's left needs a thread that wakes during a long syscall to preempt it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The time queries rescheduled since 2013, so that a game spinning on the
clock would let a thread that a timing event had woken run (it fixed
audio in Crimson Gem Saga and Where Is My Heart?). In 2014 audio and
delay wakeups started rescheduling themselves, but IO completion never
did, and a movie reader thread in Driver 76 was only getting in through
the time queries. Now IO completion dispatches like any other wakeup.
The PSP doesn't dispatch in a time query, and doing so let a thread
that a terminate woke run too early. threads/threads/terminate now
passes.
Checked by booting 75 games against master: the same in all of them,
with Asphalt Urban GT2 getting further in the same time.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
When the new thread outranks the caller, the firmware switches to it
directly, even if a thread of still better priority is ready but hasn't
been dispatched (one that a sceKernelTerminateThread woke, say). Verified
against the new pspautotests threads/threads/termsuspended.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Only the GE and vblank interrupts take user sub-interrupt handlers, and
vblank only in slots 0-15, with some of the rest already held by the
kernel. The errors follow interruptman.prx's checks, and which interrupts
have handlers at all is read back from pspautotests intr/registersub and
intr/releasesub, which now pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
From pspautotests intr/waits:
- sceAudioOutputBlocking sets the channel's waiting flag before its
event flag wait, and when that wait fails at once (interrupts or
dispatch disabled, or inside an interrupt) it returns the error
without clearing the flag. The channel stays busy from then on, and
can't be released.
- The SRC blocking output fails the same way even when a completion is
already there, leaving the buffer armed.
- After a block that had samples in it, the mixer DMA is still playing
it out, so a buffer arriving then isn't read early or restarts it.
- sceKernelVolatileMemLock only writes the fake address and size
through pointers that are there, instead of faulting on NULL.
intr/waits now runs to the end; one scheduling marker still differs, from
async IO timing.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- A timeout of 0 to sceUmdWaitDriveStatWithTimer/CB means no timeout,
not a tiny one (or 8ms for the CB version).
- Timeouts round like the event flag wait does.
- A wait with no timeout no longer times out right after a callback.
- sceUmdRegisterUMDCallBack only accepts callbacks.
- sceUmdActivate requires the name to be exactly "disc0:", and it and
sceUmdDeactivate/sceUmdGetDiscInfo reject kernel pointers.
- sceUmdDeactivate needs a name in mode 2.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
They take tens of milliseconds for the caller, but queueing that time on
the shared ME timeline made the SAS mix wait behind them. In Jak and
Daxter that held up the sound threads at the end of the first clip, so
video_sound_thread got its last wake only after the game had deleted it
(NOT_DORMANT), and the orphaned thread then read a freed context.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Measured on a PSP (pspautotests video/mp4/mp4timing, audio/audiocodec/timing):
- sceVideocodec Open, GetEDRAM, GetVersion and ReleaseEDRAM take ~70-150us,
Init ~26.6ms (sceMpegCreate is 27-28ms), Delete ~21ms (was 2ms), and
Stop 132us with nothing held back. All go through the ME queue now.
- Decodes that return no picture take as long as those that do; they
were free.
- Open reports the EDRAM the decoder needs (0x3c2c) at ctx+0x18, which
mpeg.prx passes on to GetEDRAM.
- sceAudiocodec: failed decodes (214/142/169us) and mono Atrac3+ init (524us).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Checked against pspautotests audio/audiocodec, recorded on a PSP.
- Atrac3+: at3Related selects headered (mpeg.prx) or raw (libatrac3plus)
frames, instead of sniffing for the sync word. The header's size field
is 10 bits, as the context's. Header errors 0x211/0x213, bad frames
0x20a, all returning SCE_AVCODEC_ERROR_INVALID_DATA with nothing read.
- The first successfully decoded Atrac3+ frame, and the first two AAC
frames, produce no output. Checked sample-for-sample against hardware.
- Atrac3: the parameter at 0x28 selects the frame layout, as
libatrac3plus.prx's table maps it. We used to read its low bit as a
joint-stereo flag, which decoded mono (0x0F) streams as stereo garbage.
AtracCtx2 had the table's fields swapped the same way.
- at3_standalone's Atrac3 output was inverted relative to the PSP's
(sceAtrac too). Negate the IMDCT scale.
- CheckNeedMem sizes (AAC is 0x658c), codec 0x1004/0x1005, Init
validation (AAC sample rate, Atrac3 parameter, Atrac3+ channels), and
ReleaseEDRAM clearing edramAddr.
- Every call that reaches the ME now blocks for its measured time, and
decode time is modelled per codec and frame size.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
GLSLtoSPV takes an optional SPIRVCache, keyed on a 32-bit hash of the
source, stage and variant, plus the source length. A changed shader
simply misses. thin3d's shaders and the other fixed ones use a global
cache in PSP/SYSTEM/CACHE/vulkan_spirv.cache, loaded on first use and
saved after graphics init, when a game's cache is saved, and at
shutdown; it's flushed once it reaches 32 entries, about twice what a
session compiles, so outdated ones don't pile up. Game shaders keep
theirs in the .vkshadercache, ahead of the shader IDs so that the
compiles on load find it (version 60), and only what the session used
is saved.
A cold glslang costs about 40ms before its first shader here, and
0.3-0.9ms per shader after that.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Like the UI atlas, keep the decoded image around and only recreate the
texture when a new UIContext asks for it, rather than reading the
metadata and decoding the ZIM from disk each time.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The end callback checked the kernel object's lockThread, which for an
lwmutex is only refreshed by sceKernelReferLwMutexStatus. The lock state
lives in the workarea, so an unlock during the callback left the waiter
waiting forever. Verified against pspautotests threads/lwmutex/callbacks.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Replaces the global in-callback counter with each thread's own mipscall
chain, so several threads can be inside callbacks at once, and other
threads' callbacks (better priority ones right away) run while one is.
Verified against pspautotests threads/callbacks/otherthread, recursion
and intrnotify:
- A callback nests only one level: a CB wait that would go deeper never
returns on hardware, so the callback is left pending instead.
- A non-CB wait inside a callback no longer runs callbacks because of
the CB wait the callback interrupted.
- Callbacks for a waiting thread are only taken when it beats both the
running thread and every ready one. After an interrupt (which runs on
the idle thread) that's the thread about to resume, not idle.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Verified against pspautotests threads/callbacks/waittypes:
- __KernelThreadingInit() cleared the wait type callback table after
__KernelMemoryInit() had registered VPL and FPL in it, so a VPL or FPL
wait interrupted by a callback was never paused or resumed, and could
hang forever.
- A msgpipe deleted during a callback left its waiter waiting, instead
of waking it with WAIT_DELETE.
- A wait that got its object during a callback reported no time left;
put the timer back before trying to unlock, so the unlock writes what
remains.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Verified against pspautotests threads/callbacks/delivery:
- Notifying the callback of a better priority thread in a CB wait runs
it right away. Callbacks of other waiting threads stay pending until
those threads would get to run, rather than being taken at any
reschedule, so they can still be counted or canceled.
- sceKernelCancelCallback clears the notify count, not just the arg.
threads/callbacks/cancel, count and umd/wait/wait now pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Verified against new pspautotests threads/callbacks/afterwait and nested:
- A thread inside a callback runs its own pending callbacks (even the
same one again) nested, when it enters a CB wait. Waits paused by a
nested callback are keyed by the outer callback's id.
- sceKernelCheckCallback inside a callback returns ILLEGAL_CONTEXT
without running anything.
- sceKernelSleepThreadCB with a queued wakeup runs pending callbacks
before consuming it.
- A thread whose wait ended during a callback keeps the CPU, instead of
queueing behind threads of the same priority.
threads/callbacks/notify now passes too.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A thread whose wait ends without a context switch (for example when a
callback run from the wait satisfies it) keeps its old waitType. If it
later ran a callback, from sceUmdWaitDriveStatCB with the drive already
ready for instance, the stale wait's begin/end hooks ran, couldn't find
the paused wait, and resumed the thread with SCE_KERNEL_ERROR_WAIT_DELETE,
overwriting the HLE call's return value.
Should fix "sceUmdWaitDriveStatCB: error 0x800201b5" dialog in
Maru Goukaku TOEIC Test Portable (#7576).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
How to keep lldb and gdb from stopping on the faults the handler is meant
to catch, and the crash_*.prx tests as a quick check that it works.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Apple platforms now use the same SIGSEGV/SIGBUS handler as Linux instead of
a Mach exception port. The Mach port was only set on the installing thread,
which is the loader thread, not the one running JIT code. ARM64 had no
context definition at all.
Also pass the thread state (not the mcontext) to the handler, accept SIGBUS
fault codes, and fix restoring a disabled altstack on Darwin.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
An interpreter state always claimed an uneaten VFPU prefix, which made a
JIT loading it run in unknown-prefix mode for the rest of the session.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Drop an ISO handle whose file isn't in the loaded image instead of keeping
a null file, and don't reopen a directory file with exclusive create.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Missing sections (videocodec, audiocodec, aac, mp3) and old-version
branches (impose, io, umd, gps, mic, display, font) left the session
before the load in place.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
sceMpeg's AVC resource flag, whether the VSH is running, sceReg's handle
counter, sceNet's pending apctl events and product code block, and the
save dialog's copy of the original request (without which the first
Update after a load reloaded the request and lost its results).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
__DmacDoState and __UsbCamDoState existed but weren't in the module list,
so the memcpy deadline and camera state carried over from before a load.
The NpDrm licensee key wasn't saved or reset at all, so EDATA opened after
a load in a fresh session couldn't be decrypted.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Without saved flags, those states were resolved against today's defaults,
and everything graduated since (sceMpeg, sceFont, the leaf libraries) ended
up on unresolved stubs.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The first one got n-1, which is the state's last event, so that one moved
to a new id while its queued occurrences fired the newcomer. n-1 dates
from before RestoreRegisterEvent could fall back when out of range. Also
recompute a debugger run-until deadline against the loaded clock.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
gpu/clut/offset and gpu/commands/material take over 1s each in a Debug
build, and every so often pushed past the 5s limit mid-run, which reads
as a failure with truncated output.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Same as the exit callback: states from before them numbered the action
types without them, so the boot-time ids can belong to other types.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A state from before the event kept its boot-time id, which Atrac restores
first, so the event the state had under that id moved to a new one while
its queued occurrences kept firing AtracOutput. In Outrun 2006 that was
the vblank, and the game waited for it forever.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
MP3 emulation already went through FFmpeg, leaving MiniMp3Audio dead.
The one live user was loading MP3 UI sound effects (custom achievement
sounds), which now splits the file into frames and decodes them with
the FFmpeg MP3 decoder.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
gpu/rendertarget/copy no longer prints a million pixels one at a time,
and runs in 0.17s under the interpreter rather than ~4.5s.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
gpu/clut/offset and gpu/commands/material take over 1s each in a Debug
build, and every so often pushed past the 5s limit mid-run, which reads
as a failure with truncated output.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Neither is serialized, and both went stale on load. The ME busy time was
measured against the pre-load clock, so loading an earlier state made the
next SAS/codec job wait until the old time came around, freezing the game.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The write pass stores memory with the JIT's emuhacks cleared, but the
verify pass compared against memory that still had them, so it would
report a mismatch under a JIT. Only EMULATOR_DEVCTL__VERIFY_STATE runs it,
and nothing currently does. It also counted as a save in the generation.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Some DoState code meant for after a load ran on every save:
- scePower reset the bus frequency a game set (and with a locked CPU
speed, applied the current setting to the clock).
- sceDisplay reset the lag sync baseline, and could schedule lag sync in
the measuring pass only, which failed the save.
- GPUState dirtied the texture, sceUmd notified the UI, and sceMpeg
dropped a pending ringbuffer fix-up for an old state.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
CoreTiming::DoState replaced every event's callback with the anti-crash
one in every mode, relying on each module's restore to put it back. A save
that failed partway never got to those, and left the running game with
events that break into the debugger. The missing-section fallbacks then
also ran on the save: cheats and the mic re-registered events into the
wrong slots, and achievements reset the runtime.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
SetError overwrote the first bad section with whichever section a later
error came from, and an error before any section read an uninitialized
curTitle_.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
DoMap and DoSet cleared them, but only once the count had been read. A
state truncated right there left the deleted pointers in place, to be
freed again when the failed load reset the game.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The list belongs to the sceGe call still in progress, whose end would have
run on the loaded CPU state. Also stop the camera and GPS when a state has
them off, don't restart capture when saving, and fix a double free of the
pmp frame queue (it only holds the media engine's own frame).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Reject sizes past the end of the state before allocating (FPL, PGF,
achievements, SAS grain, savedata list, the memory fast path), fail
instead of desyncing on a SAS voice count mismatch, and free what old
states' paths and shrinking pointer containers dropped.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Containers of pointers are filled with nullptr and then DoClass'd, and
once an error switches the load to MODE_NOOP, every remaining element
called DoState on null.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
DirectoryFileSystem reused one entry across files, so a failed reopen
could seek another file's handle. VirtualDiscFileSystem leaked every open
handle on each load. MemoryStick ignored the saved free space basis.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
They were recreated without block size or extradata, which Atrac3 needs,
so it stayed silent after a load. Also drop the old decoders when the
state has none, and don't overflow on v1 states.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Delete the old HLE mips call actions instead of leaking them or keeping
stale ones, fail the load on an unknown action type instead of crashing,
and derive the exit-callback-pending flag from the loaded state.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The function list is from before the load, where other code (an overlay
module) may have been. Hashing it again from the loaded memory keeps the
hooks to code that actually matches.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With WLAN on and no server reachable, sceNetAdhocctlInit keeps its thread
waiting for the login. The load dropped the request, and the wait ended in
BUSY, which Init can't retry. Splinter Cell then ran its failure path with
a deleted event flag. Also stop freeing matching event buffers into the
restored allocator, and take the event lock when clearing.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The old icons stayed in PPGe's decimation list with kernel addresses from
before the load, and got freed out of whatever the loaded state had there.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Both drained only in their own DoState, after memory and Atrac contexts
had already been replaced under a mix or read still in flight. Also fix
sceUmd loading umdActivated into the wrong variable, and count each save
once in saveStateGeneration (it also bumped in the measuring pass).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
KernelHeap was missing from CreateByIDType, and an unknown type returned
without an error, desyncing the rest of the load. States from before exit
callbacks kept the boot-time action slot, which sceMpeg's restore took over.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
DoState replaced native_ while holding a lock_guard on its mutex, so the
old NativeInput could be freed before the guard unlocked it. Crashed every
state load (e.g. Splinter Cell Essentials, anywhere in the game).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The blit rates were measured with nothing else running. In a game, threads
waking up and SAS mixing on the Media Engine compete with the GE for main RAM:
Star Wars: Lethal Alliance's movie blit takes 8.65ms alone and 10.3ms in the
game. We don't model that load, so RAM texture fetches get a fixed 1.17x for a
typical one. With it, that game's long movie plays at 30 fps as on hardware.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The SAS mix estimate was a guess capped at 1200us. Measured on a PSP
(pspautotests audio/timing/sastiming), a mix costs 110us plus 0.49us per grain
sample, plus per voice and sample 0.445us + 0.0675us per unit of pitch ratio
(VAG; PCM and noise slightly less), plus 0.64us per sample with a reverb type
set. Linear to 1% across 64-2048 samples and 0-32 voices: 32 VAG voices at 512
samples take 8.7ms, not 1.2.
The mix runs on the Media Engine, as do video and audio decoding, so they now
queue behind each other there (MEScheduleJob). A game that keeps SAS running
during a movie - Star Wars: Lethal Alliance - decodes more slowly for it, like
on hardware.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Full-screen clears measured on a PSP (pspautotests gpu/timing/blittiming):
0.49ms on a 16-bit framebuffer whatever is cleared, 0.69ms on 8888, 1.02ms on
8888 with depth. Charging them may help games that spin hard on an empty
screen, but it's off (chargeClearTime) until tried on some. The video blit
cost moves into the same function, now EstimateFillCycles.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The blit cost remembered only the last buffer a decoder wrote into, forever.
Move the texture cache's video list (with its ageing out a few flips after
the last write) into GPUCommon, so the texture cache, the blit cost and
SoftGPU all share one. That also counts both of a double-buffered player's
frames, which exposed that a clear drawn with texturing still enabled was
being charged as a blit - skip clears and draws without texture coordinates.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Measured on a PSP (pspautotests gpu/timing/blittiming), a full-screen blit
from an unswizzled texture costs what the texture fetch costs: 16-bit formats
half of 32-bit, VRAM a fifth of RAM, and rectangles wider than ~128 texels
~7.5x as much as narrow strips, from texture cache thrashing. Framebuffer
format, filtering and blending don't matter. Ys I & II draws its movie as one
full-width sprite from a 565 texture in RAM, which takes 33ms - that, not the
decode, is what holds it to 30 fps.
All of the ME and GE costs speed up with the clock (2/3 as long at 333/166),
since the whole system runs from the one PLL.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
sceMp4AacDecode of an AAC-LC stereo frame takes about 1.7ms on a PSP, measured
with pspautotests video/mp4/mp4timing.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Movie players like the one in Star Wars: Lethal Alliance present every decoded
frame after a single vblank wait, with no clock or timestamp check, so the
frame rate depends on decode, CSC, ATRAC decode and the GE blit adding up to
more than a vblank. We charged nearly nothing for any of them, so such movies
ran at 60 fps until the ringbuffer's slack ran out.
Costs measured on a PSP with a copy of that player (pspautotests
video/mpeg/playertiming), for a 480x272 frame:
- sceVideocodecDecode: 3.4ms (sceMpegAvcDecode 5.8ms less sceMpegAvcCsc 2.4ms)
- sceMpegBaseCscAvc: 2.4ms, was a flat 4ms
- sceAudiocodecDecode, ATRAC3+ only: 2.5ms per frame
- GE: 9.7ms for a through-mode rectangle blit from a decoded video frame,
charged by area, only for textures in the buffer a decoder last wrote.
GE time also now carries across stall address updates. Before, a list sent
in stalled chunks only had its last chunk's time counted, so sceGeDrawSync
returned 39us after a blit that takes 9.7ms. This affects every game that
builds its lists incrementally, so GE-timing-sensitive games need checking.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
While apctl was in JOINING, every Update restarted the fade-in, so the dialog
sat at its first, barely visible step until the state moved on to getting an
IP. It now keeps the fade-in it started with.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- The result was only ever written on cancel, so a connection that worked
gave the game back whatever was in the field, which some fill with -1.
- Infrastructure mode drew a Cancel button that did nothing, so if the
access point never gave an IP there was no way out. Cancelling now also
disconnects the connect it started, so the game isn't left connected
after being told the dialog was aborted.
- An unknown netAction went down neither path and never finished.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Loading a utility module without room for our (rough) memory block stored
(u32)-1 as its address, which av_atrac3plus then memset. It now loads
without the block, which HLE doesn't need.
- LoadNetModule's module id was unsigned, so a load error was passed on to
sceKernelStartModule as an id.
- The dialog helper threads put the game's priorities straight into ORI
immediates; ones that aren't a priority at all now fall back to 0x20.
- A fade never finished with an animSpeed of 0 or less.
- MsgDialog V3 button captions filling all 64 bytes ran on into the next
field.
- GamedataInstall: a file shorter than it claimed was retried forever, Abort
worked in any state and wrote through an unchecked pointer, and the game
and data names were read as C strings from fixed-size fields.
- NpSignin: a cancel was overwritten with SUCCESS in the same frame, so the
game saw a sign-in. The status is reset on start, so a reused struct no
longer hangs.
- Unloading av_atrac3plus never reached __AtracNotifyUnloadModule, leaving
the atrac state pointing at freed memory.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- InitStart checks the address and size the way sceUtility_Driver does
(0x40 and 0x44 are both fine, utility/dialog/sizes) and that the field
struct is in memory.
- The input text was read up to a terminator with no limit and no memory
check, and the output was written through an unchecked pointer every
Update. Both are bounded and checked now.
- Converting a string to UTF-8 checked for room before each character, then
wrote up to three bytes and a terminator, overflowing a 2048-byte stack
buffer on long non-ASCII text (both conversions did).
- An output buffer of length 0 made FieldMaxLength wrap around, and the
keyboard preview then indexed the text at -1.
- The native input box's callbacks captured the dialog and could write into
it after it was deleted, and its status was read and written without the
lock. They now share a small state object instead, a new one per start
and per state load (unless a box is still open, which answers into the
loaded state).
- Savestates keep the current keyboard and the Korean combining state. With
older ones, it picks a keyboard the field allows, as Init does.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- The Yugioh savedata workaround force-stopped the other dialog without
releasing the volatile memory it held, so the savedata helper then waited
for it forever.
- Screenshot's ShutdownStart and Update succeeded when no screenshot was
running, putting it back into SHUTDOWN.
- Savestates: Netconf didn't keep its request address, and a DNS json
download going on was gone after a load, so it could wait for it forever;
it now fetches the json again (it's cached). NpSignin didn't keep its
request address either. Both restart their timeouts instead of timing out
at once. With states from before, they keep the current address as they
used to.
- Loading a state from before NpSignin, GameSharing or HtmlViewer were saved
resets them rather than keeping this session's state (including the
HtmlViewer's memory block), and without Shutdown's side effects, which
would write to the loaded memory and release its volatile lock.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
InitStart sizes: Netconf and NpSignin accepted any size, then wrote
common.size bytes back from a 64-68 byte host struct, copying host memory
into PSP RAM. GamedataInstall looked for install files before checking the
size, and the HtmlViewer read options before checking the whole request was
in memory. All the dialogs now check the address, then the sizes
sceUtility_Driver accepts (utility/dialog/sizes), before anything else, as
the firmware does (a bad address is INVALID_ADDRESS), and write back no more
than the struct.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Replaces the temporary fix that did the hidden modes' IO inside Update.
The IO thread read and wrote the dialog's request, display state and save
list, all shared with the emulator thread, which kept using them to draw the
dialog and reload the request from the game.
Now the IO thread works on its own copy of the request, its own SavedataParam
and directory names resolved up front, and shares nothing else with the
emulator thread but the (locked) file system, MemoryStick_FreeSpace's cached
use and sceChnnlsv's scratch buffer and kirk state, the last two now under
locks too. It still reads and writes the game's buffers directly, like a PSP's
utility threads and sceIoReadAsync do, so a savestate waits for it before it
saves or loads memory. Save
bookkeeping, the save indicator and display changes happen on the emulator
thread when the results are taken, and only the request fields the IO changed
are copied back, so a game's own edits in the meantime survive.
Hidden modes take the results at the next Update (or, with Host IO timing,
the first Update that finds them done). The visible dialogs keep drawing and
take them once the IO is done; save and load used to stall the emulator
thread for the whole operation. Savestates keep results that haven't been
taken yet.
When the results land in PSP memory doesn't matter to games, so
utility/savedata/filelist now only prints them once the utility has finished.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Until then the dialog runs normally (and writes result = 0, which an
immediate abort skipped). Measured with one Update per vblank; at one every
other vblank a PSP took 6, so it isn't purely a count.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>