Commit Graph
212 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 23fc0cd422 Interrupts: Refuse sub-interrupt handlers where the firmware does
Only the GE and vblank interrupts take user sub-interrupt handlers, and
vblank only in slots 0-15, with some of the rest already held by the
kernel. The errors follow interruptman.prx's checks, and which interrupts
have handlers at all is read back from pspautotests intr/registersub and
intr/releasesub, which now pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik RydgårdandClaude Opus 5.5 5674c789ef sceUmd: Match hardware's parameter checks and wait timeouts
- A timeout of 0 to sceUmdWaitDriveStatWithTimer/CB means no timeout,
  not a tiny one (or 8ms for the CB version).
- Timeouts round like the event flag wait does.
- A wait with no timeout no longer times out right after a callback.
- sceUmdRegisterUMDCallBack only accepts callbacks.
- sceUmdActivate requires the name to be exactly "disc0:", and it and
  sceUmdDeactivate/sceUmdGetDiscInfo reject kernel pointers.
- sceUmdDeactivate needs a name in mode 2.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 09:19:32 -06:00
Henrik Rydgård eab0b53a7a Merge pull request #22371 from hrydgard/audiocodec-fixes
sceAudiocodec and sceVideocodec timing and Atrac3+ fixes
2026-09-28 17:26:04 -06:00
Henrik RydgårdandClaude Opus 5.5 c162eb3d74 sceAudiocodec: Match hardware setup, framing, errors and timing; fix Atrac3 polarity
Checked against pspautotests audio/audiocodec, recorded on a PSP.

- Atrac3+: at3Related selects headered (mpeg.prx) or raw (libatrac3plus)
  frames, instead of sniffing for the sync word. The header's size field
  is 10 bits, as the context's. Header errors 0x211/0x213, bad frames
  0x20a, all returning SCE_AVCODEC_ERROR_INVALID_DATA with nothing read.
- The first successfully decoded Atrac3+ frame, and the first two AAC
  frames, produce no output. Checked sample-for-sample against hardware.
- Atrac3: the parameter at 0x28 selects the frame layout, as
  libatrac3plus.prx's table maps it. We used to read its low bit as a
  joint-stereo flag, which decoded mono (0x0F) streams as stereo garbage.
  AtracCtx2 had the table's fields swapped the same way.
- at3_standalone's Atrac3 output was inverted relative to the PSP's
  (sceAtrac too). Negate the IMDCT scale.
- CheckNeedMem sizes (AAC is 0x658c), codec 0x1004/0x1005, Init
  validation (AAC sample rate, Atrac3 parameter, Atrac3+ channels), and
  ReleaseEDRAM clearing edramAddr.
- Every call that reaches the ME now blocks for its measured time, and
  decode time is modelled per codec and frame size.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 16:57:41 -06:00
Henrik RydgårdandClaude Opus 5.5 c11e46aeea LwMutex: Take the lock after a callback if it was released during it
The end callback checked the kernel object's lockThread, which for an
lwmutex is only refreshed by sceKernelReferLwMutexStatus. The lock state
lives in the workarea, so an unlock during the callback left the waiter
waiting forever. Verified against pspautotests threads/lwmutex/callbacks.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 15:59:25 -06:00
Henrik RydgårdandClaude Opus 5.5 eda92c4ef2 Callbacks: Track callback nesting per thread, one level deep
Replaces the global in-callback counter with each thread's own mipscall
chain, so several threads can be inside callbacks at once, and other
threads' callbacks (better priority ones right away) run while one is.
Verified against pspautotests threads/callbacks/otherthread, recursion
and intrnotify:

- A callback nests only one level: a CB wait that would go deeper never
  returns on hardware, so the callback is left pending instead.
- A non-CB wait inside a callback no longer runs callbacks because of
  the CB wait the callback interrupted.
- Callbacks for a waiting thread are only taken when it beats both the
  running thread and every ready one. After an interrupt (which runs on
  the idle thread) that's the thread about to resume, not idle.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 15:59:25 -06:00
Henrik RydgårdandClaude Opus 5.5 acd2738b7a Callbacks: Deliver to other threads by priority, fix sceKernelCancelCallback
Verified against pspautotests threads/callbacks/delivery:

- Notifying the callback of a better priority thread in a CB wait runs
  it right away. Callbacks of other waiting threads stay pending until
  those threads would get to run, rather than being taken at any
  reschedule, so they can still be counted or canceled.
- sceKernelCancelCallback clears the notify count, not just the arg.

threads/callbacks/cancel, count and umd/wait/wait now pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 15:59:25 -06:00
Henrik RydgårdandClaude Opus 5.5 3bf8b5b5c9 Callbacks: Run nested callbacks from CB waits, match hardware ordering
Verified against new pspautotests threads/callbacks/afterwait and nested:

- A thread inside a callback runs its own pending callbacks (even the
  same one again) nested, when it enters a CB wait. Waits paused by a
  nested callback are keyed by the outer callback's id.
- sceKernelCheckCallback inside a callback returns ILLEGAL_CONTEXT
  without running anything.
- sceKernelSleepThreadCB with a queued wakeup runs pending callbacks
  before consuming it.
- A thread whose wait ended during a callback keeps the CPU, instead of
  queueing behind threads of the same priority.

threads/callbacks/notify now passes too.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 15:59:25 -06:00
Henrik RydgårdandClaude Opus 5.5 44d0c401d5 test.py: Give Debug builds three times the wall-clock timeout
gpu/clut/offset and gpu/commands/material take over 1s each in a Debug
build, and every so often pushed past the 5s limit mid-run, which reads
as a failure with truncated output.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:03:53 -06:00
Henrik RydgårdandClaude Opus 5.5 9f182a2a1b Update pspautotests: compact rendertarget test output
gpu/rendertarget/copy no longer prints a million pixels one at a time,
and runs in 0.17s under the interpreter rather than ~4.5s.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 10:32:42 -06:00
Henrik RydgårdandClaude Opus 5.5 6dffcc91d2 Utility: Check request sizes like the firmware
InitStart sizes: Netconf and NpSignin accepted any size, then wrote
common.size bytes back from a 64-68 byte host struct, copying host memory
into PSP RAM. GamedataInstall looked for install files before checking the
size, and the HtmlViewer read options before checking the whole request was
in memory. All the dialogs now check the address, then the sizes
sceUtility_Driver accepts (utility/dialog/sizes), before anything else, as
the firmware does (a bad address is INVALID_ADDRESS), and write back no more
than the struct.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 12:50:41 -06:00
Henrik RydgårdandClaude Opus 5.5 1f1c478c92 MsgDialog: An Abort takes effect from the 8th Update, like on a PSP
Until then the dialog runs normally (and writes result = 0, which an
immediate abort skipped). Measured with one Update per vblank; at one every
other vblank a PSP took 6, so it isn't purely a count.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 12:50:23 -06:00
Henrik RydgårdandClaude Opus 5.5 5f261bd7ff Utility: Stand in for the HtmlViewer, offering to open the page in a browser
Instead of the PSP's web browser, a dialog shows the URL the game wants
and opens it in the host's browser on X, or backs out on O. Either way the
game sees the browser closed normally. Platforms that can't open a URL
(the new SYSPROP_CAN_LAUNCH_URL) only offer to back out. Only plain
printable-ASCII http(s) addresses are handed over, and on Linux without a
shell.

What the firmware does (sceUtility_Driver, 6.61, plus
utility/dialog/htmlviewer): the HtmlViewer has its own state apart from the
other dialogs, so they don't block each other, and its calls return
WRONG_TYPE until one has started. The request size picks the 2.00 to 3.00
layout, and InitStart allocates 3.5MB of user memory (4.5MB with options
bit 0x400 from 2.70 on), failing with 800200d9 without it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 12:50:23 -06:00
Henrik RydgårdandClaude Opus 5.5 3c15b2b112 Utility: One dialog at a time whatever its type, like the firmware
On a PSP, every InitStart fails with INVALID_STATUS until the last dialog
started is back at NONE, including while it's shutting down, and before
its params are checked. A failed InitStart leaves the current type alone.
We returned WRONG_TYPE instead, and a failed InitStart (e.g. a bad size)
still switched the current type, so every later dialog was refused. (One of
ours that fails after already starting, as savedata can, still becomes the
current type, since the game may poll it.)

The busy check applies status changes that are due, but doesn't use up an
auto status dialog's one-time INITIALIZE/SHUTDOWN reports; one that only
waits to report SHUTDOWN is let finish. Auto status dialogs now release
volatile memory on the way to NONE, including when the game saw RUNNING
before the init thread was done, which used to leave it locked for the next
dialog. GamedataInstall no longer requires currentDialogActive, which its
ShutdownStart cleared even when it then failed, so it could never finish -
and would now have blocked every other dialog.

Also: MsgDialog accepts exactly the three sizes sceUtility_Driver does (we
memcpy'd whatever size was given), and HtmlViewer GetStatus answers
WRONG_TYPE.

Adds utility/dialog/status and utility/dialog/priority, recorded on a PSP.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 12:50:23 -06:00
Henrik RydgårdandClaude Opus 5.5 feb6caa3c4 Implement mfic and mtic
Both were no-ops, so mfic left its destination unchanged. They read and
write the interrupt enable flag that sceKernelCpuSuspendIntr/ResumeIntr
use. Only bit 0 counts for mtic, which also goes for
sceKernelCpuResumeIntr, since on hardware it's just mtic.

Adds the intr/mfic test, recorded on hardware.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 11:08:13 -06:00
Henrik RydgårdandClaude Opus 5.5 f47269864a _sceKernelAllocateTlspl: Check user pointers, support the timeout
The third argument is a timeout pointer, as threadman.prx shows. A kernel
address from user mode is ILLEGAL_ADDR there; we used to write through it.
Also, no lookup by index: the syscall requires the exact uid.

Adds the threads/tls/allocate test, recorded on hardware.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 10:05:07 -06:00
Henrik RydgårdandClaude Opus 5.5 6fc4eb19df VFPU: Fix vrot with the angle in a destination lane
The cosine is then taken of what vrot wrote to that lane: the sine, or zero.
The IR looked at the sine lane instead of the lane holding the angle, and the
legacy JITs ignored the overlap. The assembler refuses such a vrot, so those
now leave it to the interpreter, and don't pair one with the vrot before it.

Covered by the new cpu/vfpu/vrot test.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:22:41 -06:00
Henrik RydgårdandClaude Opus 5.5 ff7371dcc5 Atrac: Write decoded samples when sceAtracDecodeData returns, not when called
On a PSP the Media Engine decodes, and the samples land in the output
buffer as the call returns, a couple of milliseconds in. We wrote them at
once and only then delayed the thread. Since sceAudio plays straight out
of game memory, that matters: Fired Up decodes each chunk to 0x40 bytes into
one of its two buffers, running over the first 16 samples of the other one,
which it has just queued, and relies on the mixer having read those first.
Writing early replaced them about 21 times a second, which is the constant
crackle in its music and intro (it showed up with the sceAudio buffering
rework, which stopped copying buffers at enqueue).

Now the decoder's output is set aside, the old contents put back, and a
CoreTiming event writes the samples just before the thread wakes. Pending
writes are kept in savestates.

Adds audio/blocking/parked, recorded on a PSP: a blocking output that had
to wait returns before any of its buffer has played, so the game really
does depend on the decode's latency.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 14:49:09 -06:00
Henrik RydgårdandClaude Opus 5.5 ba77f2259d Add cpu/vfpu/exact to tests_good
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 12:24:42 -06:00
Henrik RydgårdandClaude Opus 5.5 74c5cbf503 Savedata: Leave GetSize's needed strings alone when nothing is needed
Matches hardware; moves utility/savedata/getsize to tests_good.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 19:25:24 -06:00
Henrik RydgårdandClaude Fable 5.1 43255afe9f test.py: cpu/fpu/roundmode is a known failure on riscv64
The ISA returns the canonical NaN from every operation, so a negative or
signaling NaN operand loses its sign and payload where the PSP keeps them.
Not worth a check per FP op in the JIT.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 15:04:05 -06:00
Henrik RydgårdandClaude Fable 5.1 5d35425acd Add cpu/fpu/roundmode, fpu_branch and cpu/lsu/llsc to tests_good
fpu_branch_hazard (the compare-to-branch hazard), cacheop (the write-back
data cache seen through the uncached mirror) and fpu_nan (which NaN 0/0
makes, host dependent on x86) go to tests_next. cpu/fpu/fpu is re-recorded
from a binary built with the current toolchain, which prints -nan.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:25:42 -06:00
Henrik RydgårdandClaude Fable 5.1 dc983bb1c5 Add cpu/vfpu/specials and overlap_vcrsp to tests_next, update overlap
specials stays in tests_next for vcmp on denormals and the NaN
canonicalization and denormal flush in vbfy/vocp/vavg/vfad/vsocp, which
overlap the USE_VFPU_DOT accuracy switch. overlap_vcrsp is vcrsp with an
overlapping destination, which the assembler refuses and the hardware
doesn't read-before-write for.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 5564ef3ec3 Add cpu/vfpu/overlap to tests_good
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 a6e29f6972 Add cpu/vfpu/vrnd to tests_good
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 55672ebd99 Add cpu/vfpu/vbranch to tests_good and vbranch_hazard to tests_next
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 ba491301ae Add cpu/vfpu/prefix_unpack to tests_next
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:45:25 -06:00
Henrik RydgårdandClaude Fable 5.1 4e0319ffd9 test.py: cpu/vfpu/prefix_ctrl passes now
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:37:49 -06:00
Henrik RydgårdandClaude Fable 5.1 de884aba02 Add cpu/vfpu/prefix_sat to tests_next
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:30:38 -06:00
Henrik RydgårdandClaude Fable 5.1 acb81a049e Add the VFPU prefix tests to tests_next
prefix_branch passes on every core. prefix_ctrl fails on the IR path
(out-of-size swizzle lanes) and the arm64 JIT (that, plus mtvc not
masking). prefix_consume fails everywhere: the interpreter and IR in lane
w of nine ops where the T prefix holds a constant 0, the arm64 JIT on
nineteen ops.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:21:16 -06:00
Henrik RydgårdandClaude Fable 5.1 a5c96e054e Add cpu/vfpu/minmax_tie to tests_next
vmin/vmax return the second operand on a -0/+0 tie. The classic
interpreter does that; the IR interpreter and the JITs return the first.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:09:47 -06:00
Henrik RydgårdandClaude Fable 5.1 6e44931f5b test.py: cpu/cpu_alu/cpu_div passes now
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:02:35 -06:00
Henrik RydgårdandClaude Fable 5.1 5b507e5bb3 Add pspautotests for the div-by-zero, rounding, scaled convert and call-out JIT paths
Covers the cases fixed on the riscv-loongarch-fixes branch, none of which
the suite reached before. Five go in tests_good, and two in tests_next:

- cpu/cpu_alu/cpu_div: hardware leaves HI = 0 for INT_MIN / -1. The x86
  JIT, both interpreters and every IR backend set it to -1 on purpose;
  the classic arm64 JIT passes by not special-casing it at all.
- cpu/vfpu/minmax_zero: signed zero and denormals in vmin/vmax.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 10:59:17 -06:00
Henrik Rydgård 7b95ff2808 Merge pull request #22325 from hrydgard/vertex-decoder-jit-match
Vertex decoder: New test, make the JITs match the C++ decoder closely
2026-09-21 16:28:27 -06:00
Henrik RydgårdandClaude Opus 5 88b80011e4 Run pspautotests under qemu for loongarch64 and riscv64
Only the IR JIT: it's the sole native backend these two have and the one
thing here that isn't shared code, and the x86-64 and arm64 runners
already cover all four backends. Every run costs emulated wall clock, so
the timeout goes up to match.

test.py grows a per-architecture known-failure list, selected with
--known-failures=<arch>, so this can guard against new breakage while the
four outstanding ones stay outstanding. Each entry carries its reason.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 15:36:49 -06:00
Henrik RydgårdandClaude Fable 5.1 a36d09daac sceGe: what callbacks see, and a full GE reset on sceKernelLoadExec
From another pass over ge.prx against our code, each checked on a PSP with
gpu/ge/callbackstate except the last:

- A list's context is restored after its finish callback, which sees the
  state the list left. We restored at the FINISH, before it. Now that nothing
  runs until InterruptEnd(), that's where it happens.
- sceGeSaveContext/RestoreContext only fail while the GE is executing. It's
  stopped during a finish callback and a SUSPEND signal callback, however
  much is queued, so they work there. We said busy whenever a list existed.
- sceGeListDeQueue emptying the queue doesn't turn completed lists into
  nothing, only sceGeDrawSync does. CheckDrawSync() is gone.
- The "break in progress" flag that makes sceGeContinue only requeue the
  list is cleared by an interrupt that follows the break at once, so it's
  only seen from a callback or with interrupts off. Ours lasted until the
  next GE interrupt of any kind.
- sceKernelLoadExec restarts the GE driver, which begins by zeroing every
  register and matrix. Reinitialize() now does too, so a program started that
  way finds the same GE as one booted directly, rather than its launcher's.
  Not testable on hardware: nothing after the restart can report back.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-21 13:21:03 -06:00
Henrik RydgårdandClaude Fable 5.1 91c9c4d14e sceGe: sceGeBreak(1) takes pending interrupts with it
Resetting the GE also gets rid of an interrupt that was raised but not taken
yet, so a list that reached its FINISH just before never gets its finish
callback. We delivered one anyway, for a list that no longer existed. If the
break comes from inside a GE callback, the interrupt being handled is kept,
since its handler still has to return.

Found by gpu/ge/intrsuspend, which also confirms from a thread, with
interrupts suspended, that nothing moves along the queue until the FINISH
interrupt has been taken.

Savestates: bump GPUCommon to 7. We didn't use to mark a PAUSE signal as
delivered, which sceGeContinue now goes by, so a state saved with a list
paused that way would load into a game that could never continue it. Fixed
up on load.

gpu/signals/handlercalls goes in as known failing: with an old SDK version, a
stall address set from inside a SUSPEND callback doesn't reach the GE, which
we can't express with just the one stall address per list. See docs/sceGe.md.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-21 12:48:03 -06:00
Henrik RydgårdandClaude Fable 5.1 8a23e633a1 sceGe: keep a finished list on the queue until its interrupt is done
On hardware the GE stops at every SIGNAL and FINISH, and it's the interrupt
that gets it going again: on the same list after a signal, on the next one
after a FINISH - once the finish callback has run, with the finished list
still at the head of the queue. We ran the next list right away and dropped
the finished one at once, so a finish callback saw an empty queue. A list
enqueued from there was started instead of queued, and then couldn't be
dequeued, which hung the new gpu/ge/queue2 test.

ProcessDLQueue() now runs nothing while the head of the queue has an
interrupt pending, and InterruptEnd() is what takes a finished list off the
queue. This also keeps a stall update from restarting a list that's stopped
at a signal before the handler has run. drawCompleteTicks is still set when
the last list reaches its FINISH, so a sceGeDrawSync in between doesn't wait.

Other things gpu/ge/queue2 and gpu/ge/breakwait showed, all from a real PSP:

- sceGeListEnQueue compares against the address a list was enqueued with
  (or stopped at by sceGeBreak), mirrors included, not against its current pc.
  We had that the wrong way around.
- The stack-in-use check only applies to lists that have started executing.
  This is probably what IgnoreEnqueue was added for (Metal Gear Acid 2,
  #10906). The flag stays until someone has checked the game without it.
- A PAUSE signal makes the list PAUSED at once, before the FINISH delivers it.
  In between, sceGeContinue and sceGeBreak say BUSY, and updating the stall
  address does nothing, so a list that stalls there is stuck.
- A completed list can't be dequeued, with or without a context.
- sceGeDrawSync(1) looked at currentList instead of the list it had found.
- sceGeBreak(1) doesn't wake anyone, and a late interrupt for a list it reset
  no longer marks that list completed. Threads in sceGeDrawSync are woken
  before the ones waiting for the last list.

Also fixes currentList being lost when loading a state where it's list 0,
and makes ge_pending_cb a plain std::list - nothing else touches it, and the
GPU thread it was shared with is long gone. Same savestate format.

See docs/sceGe.md.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-21 12:21:19 -06:00
Henrik RydgårdandClaude Opus 5 bb556bf481 CI: run pspautotests on all four CPU backends
The headless tests only ever exercised the JIT, since that's what headless
defaults to. Run all four on the Linux runners, and add an arm64 Linux lane
so the arm64 JIT is covered too - nothing else in the matrix tested it.

test.py scales the wall clock to the backend instead of raising it for
everyone: the interpreter needs 20s for gpu/rendertarget/copy, which does
over a million guest-side vsprintf calls, while a hang under the JIT is
still caught in five seconds.

The frametest report artifact needs a per-OS name now that two Linux legs
upload one.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-19 12:20:48 -06:00
Henrik RydgårdandClaude Opus 5 cf1d513c9e Don't default to building x64 on a machine that isn't
The MSBuild examples passed /p:Platform=x64 and the run lines pointed at
Windows/x64/..., so following them on an ARM64 machine produced an x64 build -
which then runs anyway under emulation, so nothing looks wrong. It is slower
than the native build, it isn't the code ARM users get, and a benchmark taken
from it measures the emulator: the colour conversion benchmark this was noticed
on reads 200 MPix/s emulated against 300 native.

The examples now say <platform> rather than either value, so there is no default
to follow and the machine has to be looked up. Also note that
$PROCESSOR_ARCHITECTURE describes the shell, not the host, and says AMD64 from
an emulated shell.

test.py searched only Windows\x64 for the headless binary, so on Windows-on-ARM
it would silently test an emulated build; it now looks for the native one first.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 11:38:49 -06:00
Henrik RydgårdandClaude Opus 5 e8fa4e3f56 headless: split --timeout into --timeout-wall and --timeout-emulated
--timeout was wall-clock seconds, which is what CI wants but not what you want
when the question is whether the game has had long enough to get somewhere: a
heavy scene runs many times slower than real time and a near-idle one much
faster, so the same budget means very different amounts of game time. Booting a
firmware VSH is a good example - 10 emulated seconds is about 25 real ones on
6.61 and about 7 on 2.00, and judging those two by the same wall-clock number
makes a working shell look stuck.

Both limits can be set at once and whichever is reached first ends the run,
which also says which one it was. --timeout still works as the old name for
--timeout-wall. The IsDebuggerPresent() exemption stays on the wall-clock check
only; the emulated one doesn't need it, since sitting at a native breakpoint
burns no emulated time.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-17 16:01:47 -06:00
Henrik RydgårdandClaude Opus 5 84bd8459e9 sceAudio: implement sceAudioOneshotOutput, and cover the rest of the channel rules
Another pass looking for gaps, all of it now recorded by audio/blocking/channels
and audio/blocking/oneshot.

sceAudioOneshotOutput was the last unimplemented entry in the module, returning
"library not linked" to anyone who called it. It plays one buffer on a channel it
never reserves, so the channel frees itself when the buffer runs out, and its
argument checks are their own set: any positive sample count, aligned or not, no
upper bound, a negative volume rejected rather than skipped, and no busy check at
all. No game is known to use it; it is implemented because tracing it turned out
to be cheap, not because anything needed it.

The channels test confirms four behaviors: sceAudioChReserve(-1) skips a released
channel that is still playing while a reserve of it succeeds, a second channel
joining a running mixer doesn't lose a block unlike a first, mono counts down in
the same 64-sample steps over the same time as stereo, and the panned blocking
output has no extra delay on the high channels.

Also: the SRC resampler now interpolates into the next buffer at a buffer join
instead of holding the last sample, since the codec reads the two descriptors as
one stream. That only shows up at non-native rates and no test can see it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 17:03:07 -06:00
Henrik RydgårdandClaude Opus 5 412d11534c sceAudio: drain on vaudio release, and charge what an output call really costs
Two things the last pass looked at and left, now measured on hardware by the new
audio/blocking/vaudio and audio/blocking/overhead.

sceVaudioChRelease is not shaped like the Output2 and SRC releases. It hands the
channel a null pointer first, which waits for a buffer to finish, so it blocks for
one buffer and returns 0 where the others would refuse - and the buffer is played
out instead of being dropped, which is what we were doing. The reservation goes
now rather than when the drain finishes: the caller is parked either way and gets
the same answer at the same time, and the buffers keep playing because the mixer
does not look at the reservation.

The same test turned up two more: a *failed* sceVaudioChReserve still marks vaudio
reserved, so a caller that lost the channel to Output2 is told 0x80000021 next
time until a release clears it; and sceVaudioChRelease ignores that flag entirely
and releases whatever holds the SRC channel, which pspautotests already called the
"wrong release".

The flat 10000 cycles charged to every Output2 and SRC output turns out to be
right only for the refused case. Every other outcome ends up querying the codec
and costs over 100us on hardware, including finding the channel unreserved, so
those now cost 25000. Mixer channels are the other way round - cheap to refuse,
expensive on the one output that starts the DMA - so that charge moved to the DMA
start. F1 2009 behaves identically and Burnout Dominator's mixed output is
bit-identical to before.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 17:03:07 -06:00
Henrik RydgårdandClaude Opus 5 ac50cd5759 sceAudio: model the real buffering, so contending threads get told "busy"
I deduced that this was the case, and attempted implementing this path long ago,
but I could never quite get it to work in all games. Set Claude on a quest to
research and implement it, and lo and behold, it works. A bit sobering.

Fixes #12888 and likely more. Additionally, audio latency is likely slightly
improved overall, and memory usage is down by 4.6MB.

Claude says:

The blocking output calls are not a queue that callers line up behind. Each
mixer channel holds exactly one buffer and at most one parked thread; a second
thread arriving while the first is waiting is told the channel is busy and is
expected to skip its turn. The Output2/SRC channel holds two DMA descriptors and
refuses a third caller outright, without waiting at all.

We blocked everyone instead, so a game running a movie thread and a sound-effect
thread over one output made the two alternate - a frame of movie audio, a frame
of effects silence - and the movie played at half rate. That is #12888, seen in
F1 2009 and Colin McRae: DiRT 2. With this, the movie thread keeps the channel
for the whole cutscene and the effects thread is refused, which is what the
hardware trace shows.

The driver also never copies a buffer on the way in: it stores the pointer and
its mixer walks it forward 64 samples at a time out of the game's own memory.
Modelling that fixes #20095 as a side effect, and drops the 4.6MB of per-channel
sample rings we were carrying. The mix event is re-phased to the moment a DMA
starts, since the mixer thread outranks its caller and gets a block in before the
output call returns.

Along the way: the two rest-length calls differ after a null-pointer output,
sceAudioChRelease reports not-reserved rather than not-init,
sceAudioChangeChannelConfig validates the format,
sceAudioChangeChannelVolume validates nothing, and sceAudioOutput2ChangeLength
takes a range of 17..4111. Details in docs/sceAudio.md.

Savestates: AudioChannel goes to version 4. Older ones stored mixed samples that
can't become a pointer and a position again, so they load with the pending audio
dropped and any parked threads released.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-10 17:03:07 -06:00
Henrik Rydgård c71fa5e32b sceIo: stop clobbering st_private, report FAT permissions, implement sceIoChstat
Cashing in the io/stat and io/shortname recordings.

__IoGetStat began with memset(stat, 0xfe, sizeof(SceIoStat)), which destroyed 24 bytes of the
caller's buffer that a real PSP never touches - it writes only as far as the timestamps and
leaves all six st_private words exactly as it found them. It also wrote a made-up sector number
into st_private[0] on the memory stick. That word carries the LBN on a UMD, which games read to
build disc0:/sce_lbn paths, so it stays for non-FAT and is left alone otherwise.

FAT has no permissions of its own and everything reads back as 0777. We were passing the host's
idea of the file through instead. The existing "all files look executable on FAT" hack for Beats
(issue #14812) was right in substance but lived only in sceIoDread, so sceIoGetstat and
sceIoDread disagreed about the same file where hardware has them agree. Both now go through one
path, which also gets the read-only case right: no write bits means mode 0555 and attr 0x21.

sceIoGetstat on the root of a volume is refused, as on hardware.

sceIoChstat was a logging stub. It now applies the read-only flag, which is what st_mode's write
bits and st_attr's 0x01 both mean on FAT - setting either produces both, and it's reversible.
That needs a new IFileSystem::SetFileWritable, defaulting to "can't" so read-only filesystems and
hosts that can't express it (Android content URIs) are unaffected; the call still succeeds there,
since hardware would have.

GenerateFatShortNames now accounts for capitalisation. FAT keeps a lowercase flag for the base and
another for the extension, but the PSP only honours the base one, so "shrt" becomes SHRT while
"readme.txt" becomes README~1.TXT. We were only adding a counter on collision. The unit test
carries the whole recorded set, including the corrected README~1.MD.

io/shortname stays in tests_next: its d_name column can't match while SimulateVFATBug is
uppercasing lowercase 8.3 names, which is deliberate and load-bearing for homebrew.
2026-09-10 09:52:01 -06:00
Henrik Rydgård ad3444ce50 Memory partitions: correct the range, the error codes, and the heap
The new sysmem tests run the same partition sweep from both privilege levels, which settles
several things that were guesses:

The valid range is 1-6, not 1-9-except-7. sceKernelCreateVpl, CreateFpl, CreateMsgPipe and
AllocPartitionMemory all let 8 and 9 through to the permission check, so a caller asking for
partition 8 got ILLEGAL_PERM where hardware says ILLEGAL_ARGUMENT. Privilege changes the
permission check, not the range - 1, 3 and 4 are refused from user mode and work from kernel mode
in every one of these APIs, which is the evidence the earlier BlockAllocatorFromID change was
missing.

sceKernelAllocPartitionMemory reports an out-of-range partition differently depending on which
entry point was used - ILLEGAL_ARGUMENT through SysMemUserForUser, ILLEGAL_PARTITION through
SysMemForKernel. Both NIDs land on the same function here, and hleIsKernelMode() is precisely
"came in through the kernel NID", so it picks the right one.

sceKernelCreateHeap had four "TODO: Validate error code" comments and no test at all - it's
kernel-only, which is why. All four are now recorded: out-of-range partitions are
ILLEGAL_PARTITION, a size of zero or less is HEAPBLOCK_ALLOC_FAILED before anything is allocated,
a NULL name is refused with ERROR, and flags really are ignored. sceKernelAllocHeapMemoryWithOption
had its validation backwards: the option struct's size field isn't checked at all, while the
alignment must be a power of two from 4 to 0x80.

Not fixed, and split into sysmem/kernel/heapgrow in tests_next: a real heap will hand out a block
larger than the heap itself, so the size isn't a cap. Ours is a fixed allocator over the reserved
block. Worth establishing how far the real one grows before implementing that.

Risk: the range change makes partitions 8 and 9 fail earlier and with a different code than
before. Nothing in tests_good depended on the old behaviour except two expectations that had
drifted from hardware, corrected in the submodule.
2026-09-10 09:52:01 -06:00
Henrik Rydgård 41144e3b35 Give the memory partitions the caller's privilege, not the syscall's
PPSSPP decided whether a caller was privileged with hleIsKernelMode(), which reports whether the
syscall being executed is itself a kernel-only export. That's a different question from the one
the hardware answers: on a PSP the privilege belongs to the calling module, and a kernel module
reaches sceKernelCreateTlspl through the ordinary ThreadManForUser NID like anything else. So a
kernel module asking for partition 1, 3 or 4 got ILLEGAL_PERM where a real PSP hands it over,
which the new threads/tls/kernel/partition test shows directly.

BlockAllocatorFromID now also accepts a caller whose thread belongs to a kernel module, via a new
__KernelCurThreadIsKernelMode(). It checks the thread's own attribute first and then the owning
module, because a kernel module's main thread isn't necessarily flagged kernel - the attribute
comes from PSP_MAIN_THREAD_ATTR, which needn't set it. That mirrors how sceKernelCreateThread
already works out allowKernel.

This only ever widens access, and only for threads belonging to kernel modules, so games are
unaffected - they run in user modules and see exactly what they saw before.
2026-09-08 15:18:02 -06:00
Henrik Rydgård 705ea7eb02 Keep the SHA-1 context in game memory too, and scope the Tlspl partition range to user mode
sceKernelUtilsSha1Block* had the same single global context that MD5 did, so it gets the same
treatment: state, counters and block buffer now live at ctxAddr in the layout hash/sha1ctx
records off hardware. Unlike MD5, SHA-1 does not stream whole blocks through buf, which happens
to be what our sha1_update already does - so no fill-in step is needed there.

The Tlspl partition range from the last commit was too broad a cut. Hardware says only 1-6 exist,
but that recording is from user mode, and BlockAllocatorFromID deliberately maps 8 and 10 to the
user partition for a kernel-mode caller - rejecting them outright would have taken that away.
The tightened range now applies to user mode only and kernel mode keeps what it had.
threads/tls/partition also shows the answer doesn't depend on the compiled SDK version, checked
across 1.00 through 6.06, and that partition 5 is accepted - which no test had covered.
2026-09-08 15:18:02 -06:00
Henrik Rydgård 8a02d1ee0f Keep the MD5 context in game memory, and make MT19937 actually be MT19937
Three fixes, all of them things the new hardware tests turned up.

sceMd5Block* and sceKernelUtilsMd5Block* shared one static md5_context and ignored the context
pointer the caller passed in, with a TODO saying it would do "unless games do several MD5
concurrently". hash/md5ctx shows a real PSP keeps everything in the caller's 96 bytes and happily
runs two digests at once, so do that instead: the state, the counters and the block buffer now
live at ctxAddr in the game's own memory, in the layout the test pins down. Two interleaved
digests come out right, and a context that gets copied mid-digest carries on correctly. As a
side effect the state is now covered by savestates, which a file-static never was.

MersenneTwister masked both halves with 0x80000000 where the low half needs 0x7FFFFFFF, so
sceMt19937UInt and sceKernelUtilsMt19937UInt were returning a sequence that isn't MT19937 at
all - every number differed from hardware from the first draw. hash/mt19937ctx computes the
reference sequence itself and confirms the PSP is plain MT19937; with the mask fixed we match it
for both seeds tested. Init also twists the array immediately, as hardware does, so a context
that has been seeded but not drawn from now holds what a real one would.

sceKernelCreateTlspl accepted partitions up to 9 before falling through to the permission check.
Hardware draws the line at 6 - threads/tls/create records 7, 8, 9 and 10 all returning
ILLEGAL_ARGUMENT - so 8 and 9 were coming back ILLEGAL_PERM. Note this is genuinely different
from sceKernelCreateVpl right above it, which does let 8 and 9 through to ILLEGAL_PERM; the two
had been sharing a check that was only ever right for Vpl.

Risk worth naming: the MT19937 change alters the numbers any game gets from these calls. That's
the point - they were wrong - but a savestate taken mid-sequence will resume with a generator
that behaves differently from the one that made it.
2026-09-08 15:18:01 -06:00
Henrik Rydgård 4d795f5130 Merge pull request #22248 from hrydgard/fat-short-names
sceIo: generate and resolve FAT 8.3 short names
2026-09-08 14:42:18 -06:00