Without SSE or NEON, triangle pixels got the secondary color in place of
the primary one plus it, so lit triangles came out black. It showed as
the "unexplained" known failures on riscv64 and loongarch64, and broke
the new gpu/lighting/shademap there. Reproduced on arm64 by building
without NEON.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Recorded on hardware (pspautotests threads/tls/timeout), a Tlspl
allocation follows the same timeout rule as the other waits, including
failing at once for 0 and 1us without writing the timeout back, which the
shared rule it moved to in the last commits didn't give it yet. Before
that it waited the raw timeout, ~30us short.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware a thread waiting for vblank returns ~53us after it, where we
had it back in ~5us, and the first of four waiters runs after ~85us: each
waiter beyond the first adds ~9us (pspautotests threads/scheduling/
vblankwake). The waiters are now released by a separate event 48us + 9us
per extra waiter after the vblank. Which vblank a wait is for is still
decided at the vblank, so a thread that starts waiting in between still
waits a whole frame (sceDisplay section version 8).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Timed on a PSP, each way of handing the CPU to another thread (the call
and the switch together):
hardware before now
rotate to an equal thread 7 14 7
signal, better thread runs 10 17 10
it waits again, back to caller 10 19 12
wakeup, better thread runs 8 13 6
it sleeps again, back to caller 7 12 6
start a better thread, entry 30 28 30
thread ends, back to its waiter 21 13 20
notify, better thread's callback 14 13 14
A switch between two threads now costs 1150 cycles instead of 2700.
Starting a better thread costs 2000 cycles more, ending a thread 3300,
and setting up a callback 1800.
Also splits a wait timeout's ~30us into the deadline being taken 12us
into the call and the timeout going off 18us after it. That only changes
the time left written back, which threads/semaphores/wait and
threads/fpl/cancel pin between them. intr/vblank is re-recorded so it
no longer depends on the phase of the frame.
threads/callbacks/combos now passes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware, notifying a callback of a thread in a CB wait takes it out of
the wait at once, even though the callback only runs when the thread would
get the CPU. A semaphore signalled in between doesn't end the wait: the
callback runs first, then the wait resumes and takes it (pspautotests
threads/callbacks/combos). We left the thread on the wait list until the
callback started, so the signal ended the wait and the callback didn't
run.
The notify now pauses the wait, as starting a callback used to. If the
callbacks are canceled before the thread's turn comes, the wait just
resumes (Thread savestate section version 7).
threads/callbacks/combos goes in the to-do list: a callback returning to
the thread that notified it still takes ~13us where hardware takes ~9,
part of the context switch cost.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Every wait with a timeout behaves the same on hardware (pspautotests
threads/scheduling/waittimeouts). The deadline is taken, and the alarm set
up a moment later. If the deadline has passed by then, the wait fails with
WAIT_TIMEOUT at once, without yielding or writing the timeout back. That's
usual for 0us, half the time for 1us, and rare after; AllocateVpl does more
first. Otherwise it ends max(t, 205us) + ~35us after the call. Each object
had its own guess (24/245, 25/250, 20/250 and so on), and only MsgPipe had
the immediate case.
__KernelWaitTimesOutAtOnce() and __KernelWaitTimeoutUs() now do it for
semaphores, event flags, mutexes, lwmutexes, mbx, msgpipes, fpl, vpl and
WaitThreadEnd. The latency past the deadline isn't counted in the time
left written back.
Outcomes that hardware decides by the clock's phase (these, and
sceKernelDelayThread returning at once) go with the likelier one. Ones
between 50% and certain are instead spread evenly over calls, so a polling
loop can't lock into never yielding (sceKernelThread section version 7).
This replaces the pseudo-random choice for delays.
Also adds threads/scheduling/readyqueue, which already passes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware sceVaudioChReserve takes ~260us once it gets past the busy
check, succeed or fail, and worse threads can run meanwhile. Releasing
takes ~25us and doesn't wait. We returned at once, which is what
audio/sceaudio/reserve's [r] markers showed; it now passes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A delay's deadline is now + usec, and the clock is read again when the
alarm is set. If the deadline has passed by then, the call returns 0 at
once without giving up the CPU. On hardware that makes
sceKernelDelayThread(0) return at once about 60% of the time. On a thread's
first wait after it starts, a delay of 1 does so about two times in three
as well (pspautotests threads/scheduling/delayzero). We always waited at
least 210us.
The choice is pseudo-random off the tick count, not the tick phase, since
our cycle counts are regular enough for a polling loop to lock into never
yielding. Threads remember whether they've waited since starting (Thread
savestate section version 6).
Also moves threads/vpl/create into the passing tests: re-recorded on 6.61,
it agrees with what we do for partitions 8 and 9.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With dispatch suspended, IO fails in the driver when it tries to wait. The
memory stick driver returns SCE_KERNEL_ERROR_CAN_NOT_WAIT, while usbhostfs,
which serves host0: under PSPLink, returns -1. host0: is mostly what
homebrew developers run from, so it now does the same.
Also from threads/scheduling/dispatch, which now passes:
- sceIoRead reports an async operation still in progress before failing
on suspended dispatch.
- A write to stdout or stderr doesn't give up the CPU.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Measured with pspautotests display/vblanklen: 730-770us from
sceDisplayWaitVblankStart returning to the end of vblank, with an hcount
of up to 14 inside it. The old value dated from the first source drop
and left the highest hcount at 13. display/hcount now passes (with the
test fixed not to depend on where a line boundary falls).
Booting 75 games against master shows no difference.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The time queries rescheduled since 2013, so that a game spinning on the
clock would let a thread that a timing event had woken run (it fixed
audio in Crimson Gem Saga and Where Is My Heart?). In 2014 audio and
delay wakeups started rescheduling themselves, but IO completion never
did, and a movie reader thread in Driver 76 was only getting in through
the time queries. Now IO completion dispatches like any other wakeup.
The PSP doesn't dispatch in a time query, and doing so let a thread
that a terminate woke run too early. threads/threads/terminate now
passes.
Checked by booting 75 games against master: the same in all of them,
with Asphalt Urban GT2 getting further in the same time.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Only the GE and vblank interrupts take user sub-interrupt handlers, and
vblank only in slots 0-15, with some of the rest already held by the
kernel. The errors follow interruptman.prx's checks, and which interrupts
have handlers at all is read back from pspautotests intr/registersub and
intr/releasesub, which now pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- A timeout of 0 to sceUmdWaitDriveStatWithTimer/CB means no timeout,
not a tiny one (or 8ms for the CB version).
- Timeouts round like the event flag wait does.
- A wait with no timeout no longer times out right after a callback.
- sceUmdRegisterUMDCallBack only accepts callbacks.
- sceUmdActivate requires the name to be exactly "disc0:", and it and
sceUmdDeactivate/sceUmdGetDiscInfo reject kernel pointers.
- sceUmdDeactivate needs a name in mode 2.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Checked against pspautotests audio/audiocodec, recorded on a PSP.
- Atrac3+: at3Related selects headered (mpeg.prx) or raw (libatrac3plus)
frames, instead of sniffing for the sync word. The header's size field
is 10 bits, as the context's. Header errors 0x211/0x213, bad frames
0x20a, all returning SCE_AVCODEC_ERROR_INVALID_DATA with nothing read.
- The first successfully decoded Atrac3+ frame, and the first two AAC
frames, produce no output. Checked sample-for-sample against hardware.
- Atrac3: the parameter at 0x28 selects the frame layout, as
libatrac3plus.prx's table maps it. We used to read its low bit as a
joint-stereo flag, which decoded mono (0x0F) streams as stereo garbage.
AtracCtx2 had the table's fields swapped the same way.
- at3_standalone's Atrac3 output was inverted relative to the PSP's
(sceAtrac too). Negate the IMDCT scale.
- CheckNeedMem sizes (AAC is 0x658c), codec 0x1004/0x1005, Init
validation (AAC sample rate, Atrac3 parameter, Atrac3+ channels), and
ReleaseEDRAM clearing edramAddr.
- Every call that reaches the ME now blocks for its measured time, and
decode time is modelled per codec and frame size.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The end callback checked the kernel object's lockThread, which for an
lwmutex is only refreshed by sceKernelReferLwMutexStatus. The lock state
lives in the workarea, so an unlock during the callback left the waiter
waiting forever. Verified against pspautotests threads/lwmutex/callbacks.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Replaces the global in-callback counter with each thread's own mipscall
chain, so several threads can be inside callbacks at once, and other
threads' callbacks (better priority ones right away) run while one is.
Verified against pspautotests threads/callbacks/otherthread, recursion
and intrnotify:
- A callback nests only one level: a CB wait that would go deeper never
returns on hardware, so the callback is left pending instead.
- A non-CB wait inside a callback no longer runs callbacks because of
the CB wait the callback interrupted.
- Callbacks for a waiting thread are only taken when it beats both the
running thread and every ready one. After an interrupt (which runs on
the idle thread) that's the thread about to resume, not idle.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Verified against pspautotests threads/callbacks/delivery:
- Notifying the callback of a better priority thread in a CB wait runs
it right away. Callbacks of other waiting threads stay pending until
those threads would get to run, rather than being taken at any
reschedule, so they can still be counted or canceled.
- sceKernelCancelCallback clears the notify count, not just the arg.
threads/callbacks/cancel, count and umd/wait/wait now pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Verified against new pspautotests threads/callbacks/afterwait and nested:
- A thread inside a callback runs its own pending callbacks (even the
same one again) nested, when it enters a CB wait. Waits paused by a
nested callback are keyed by the outer callback's id.
- sceKernelCheckCallback inside a callback returns ILLEGAL_CONTEXT
without running anything.
- sceKernelSleepThreadCB with a queued wakeup runs pending callbacks
before consuming it.
- A thread whose wait ended during a callback keeps the CPU, instead of
queueing behind threads of the same priority.
threads/callbacks/notify now passes too.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
gpu/clut/offset and gpu/commands/material take over 1s each in a Debug
build, and every so often pushed past the 5s limit mid-run, which reads
as a failure with truncated output.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
gpu/rendertarget/copy no longer prints a million pixels one at a time,
and runs in 0.17s under the interpreter rather than ~4.5s.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
InitStart sizes: Netconf and NpSignin accepted any size, then wrote
common.size bytes back from a 64-68 byte host struct, copying host memory
into PSP RAM. GamedataInstall looked for install files before checking the
size, and the HtmlViewer read options before checking the whole request was
in memory. All the dialogs now check the address, then the sizes
sceUtility_Driver accepts (utility/dialog/sizes), before anything else, as
the firmware does (a bad address is INVALID_ADDRESS), and write back no more
than the struct.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Until then the dialog runs normally (and writes result = 0, which an
immediate abort skipped). Measured with one Update per vblank; at one every
other vblank a PSP took 6, so it isn't purely a count.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Instead of the PSP's web browser, a dialog shows the URL the game wants
and opens it in the host's browser on X, or backs out on O. Either way the
game sees the browser closed normally. Platforms that can't open a URL
(the new SYSPROP_CAN_LAUNCH_URL) only offer to back out. Only plain
printable-ASCII http(s) addresses are handed over, and on Linux without a
shell.
What the firmware does (sceUtility_Driver, 6.61, plus
utility/dialog/htmlviewer): the HtmlViewer has its own state apart from the
other dialogs, so they don't block each other, and its calls return
WRONG_TYPE until one has started. The request size picks the 2.00 to 3.00
layout, and InitStart allocates 3.5MB of user memory (4.5MB with options
bit 0x400 from 2.70 on), failing with 800200d9 without it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On a PSP, every InitStart fails with INVALID_STATUS until the last dialog
started is back at NONE, including while it's shutting down, and before
its params are checked. A failed InitStart leaves the current type alone.
We returned WRONG_TYPE instead, and a failed InitStart (e.g. a bad size)
still switched the current type, so every later dialog was refused. (One of
ours that fails after already starting, as savedata can, still becomes the
current type, since the game may poll it.)
The busy check applies status changes that are due, but doesn't use up an
auto status dialog's one-time INITIALIZE/SHUTDOWN reports; one that only
waits to report SHUTDOWN is let finish. Auto status dialogs now release
volatile memory on the way to NONE, including when the game saw RUNNING
before the init thread was done, which used to leave it locked for the next
dialog. GamedataInstall no longer requires currentDialogActive, which its
ShutdownStart cleared even when it then failed, so it could never finish -
and would now have blocked every other dialog.
Also: MsgDialog accepts exactly the three sizes sceUtility_Driver does (we
memcpy'd whatever size was given), and HtmlViewer GetStatus answers
WRONG_TYPE.
Adds utility/dialog/status and utility/dialog/priority, recorded on a PSP.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Both were no-ops, so mfic left its destination unchanged. They read and
write the interrupt enable flag that sceKernelCpuSuspendIntr/ResumeIntr
use. Only bit 0 counts for mtic, which also goes for
sceKernelCpuResumeIntr, since on hardware it's just mtic.
Adds the intr/mfic test, recorded on hardware.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The third argument is a timeout pointer, as threadman.prx shows. A kernel
address from user mode is ILLEGAL_ADDR there; we used to write through it.
Also, no lookup by index: the syscall requires the exact uid.
Adds the threads/tls/allocate test, recorded on hardware.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The cosine is then taken of what vrot wrote to that lane: the sine, or zero.
The IR looked at the sine lane instead of the lane holding the angle, and the
legacy JITs ignored the overlap. The assembler refuses such a vrot, so those
now leave it to the interpreter, and don't pair one with the vrot before it.
Covered by the new cpu/vfpu/vrot test.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On a PSP the Media Engine decodes, and the samples land in the output
buffer as the call returns, a couple of milliseconds in. We wrote them at
once and only then delayed the thread. Since sceAudio plays straight out
of game memory, that matters: Fired Up decodes each chunk to 0x40 bytes into
one of its two buffers, running over the first 16 samples of the other one,
which it has just queued, and relies on the mixer having read those first.
Writing early replaced them about 21 times a second, which is the constant
crackle in its music and intro (it showed up with the sceAudio buffering
rework, which stopped copying buffers at enqueue).
Now the decoder's output is set aside, the old contents put back, and a
CoreTiming event writes the samples just before the thread wakes. Pending
writes are kept in savestates.
Adds audio/blocking/parked, recorded on a PSP: a blocking output that had
to wait returns before any of its buffer has played, so the game really
does depend on the decode's latency.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The ISA returns the canonical NaN from every operation, so a negative or
signaling NaN operand loses its sign and payload where the PSP keeps them.
Not worth a check per FP op in the JIT.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
fpu_branch_hazard (the compare-to-branch hazard), cacheop (the write-back
data cache seen through the uncached mirror) and fpu_nan (which NaN 0/0
makes, host dependent on x86) go to tests_next. cpu/fpu/fpu is re-recorded
from a binary built with the current toolchain, which prints -nan.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
specials stays in tests_next for vcmp on denormals and the NaN
canonicalization and denormal flush in vbfy/vocp/vavg/vfad/vsocp, which
overlap the USE_VFPU_DOT accuracy switch. overlap_vcrsp is vcrsp with an
overlapping destination, which the assembler refuses and the hardware
doesn't read-before-write for.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
prefix_branch passes on every core. prefix_ctrl fails on the IR path
(out-of-size swizzle lanes) and the arm64 JIT (that, plus mtvc not
masking). prefix_consume fails everywhere: the interpreter and IR in lane
w of nine ops where the T prefix holds a constant 0, the arm64 JIT on
nineteen ops.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
vmin/vmax return the second operand on a -0/+0 tie. The classic
interpreter does that; the IR interpreter and the JITs return the first.
Co-Authored-By: Claude Fable 5.1 <[email protected]>