Commit Graph
48130 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 c04d425b21 Update frametests: regenerated GachiTora reference
The spline pole fix changes the lighting in this frame.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 15:16:48 -06:00
Henrik RydgårdandClaude Opus 5.5 d3ea417241 Unit tests: Allow rsqrt precision in spline normal lengths
x86 normalizes with a bare _mm_rsqrt_ps, so normal lengths are only good
to about 3.7e-4 there and the 1e-4 check failed on every x86 CI job.
Compare direction after normalizing instead.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 13:29:12 -06:00
Henrik RydgårdandClaude Opus 5.5 0715f43ab8 Splines: Detect poles relative to the other derivative
With animated control points the pole is only nearly degenerate, so the
vanishing derivative is rounding noise rather than exactly zero, and the
absolute threshold missed it. The resulting random normals still showed
as dark patches on Pac-Man Arrangement's ghosts (#12354).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 13:20:01 -06:00
Henrik RydgårdandClaude Opus 5.5 d32636e91b Unit tests: Add a spline/Bezier tessellation test
Compares the software tessellator's positions, normals, UVs and colors
against a longhand double-precision reference (Bernstein polynomials,
Cox-de Boor with clamped knots, finite-difference normals), across
Bezier and spline surfaces, edge types, poles and patch facing, so it
can be optimized with something independent to be wrong against. Also
reports vertices per second.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 11:11:23 -06:00
Henrik RydgårdandClaude Opus 5.5 22c2cea923 Splines: Fix a single patch with both edges open
With only one patch, the open-last-edge adjustment assumed the first edge
was closed, so one knot interval came out as 2 instead of 1, and the
patch wasn't the Bezier patch a fully clamped cubic is. Positions were off
by up to about 7% of the patch.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 11:11:23 -06:00
Henrik RydgårdandClaude Opus 5.5 e8c39ed1a7 Unit tests: Share the benchmark timing loop
Four tests had their own copy of "call this until N seconds have passed,
then divide". CallsPerSecond in UnitTest.h does it; each keeps its old
duration and batch size.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 11:11:22 -06:00
Henrik RydgårdandClaude Opus 5.5 7a06e25aa0 GPU: Remove leftovers from hardware tessellation
The GLES sampler uniforms and texture slots for the control points and
weights, the Vulkan storage buffer bindings, and the u_spline_counts
uniform, which becomes padding (the C++ side already was).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 11:00:02 -06:00
Henrik RydgårdandClaude Opus 5.5 f87a07b1a3 Splines: Use the limit normal at a pole instead of NaN
Where all the control points along a patch edge meet at one point, like
the top of a dome, one derivative is zero and so is the cross product,
and normalizing it gave NaN. Use the limit instead, built from the mixed
second derivative. Fixes the dark spots on the ghosts' heads in Pac-Man
Arrangement (#12354).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 10:57:05 -06:00
Henrik Rydgård 1b4eca0184 Merge pull request #22395 from hrydgard/thread-timing
Claude hardware testing: Thread timing improvements
2026-09-30 10:40:22 -06:00
Henrik RydgårdandClaude Opus 5.5 3ac43a70a6 Add the utility/savedata/shutdownstatus test
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:58:37 -06:00
Henrik RydgårdandClaude Opus 5.5 cc23f4cf13 AGENTS.md: Hardware claims need a test, one PSP operation at a time
Also that unresolved scePsmfPlayer imports early in a game run are
expected.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:58:37 -06:00
Henrik RydgårdandClaude Opus 5.5 a974440c79 Debugger: Time input.buttons.press in emulated vblanks
The release was counted down on the WebSocket thread, one step per poll
of host time however many vblanks had passed, so how long a scripted
press lasted depended on how fast the emulator ran, and scripted runs
went different ways. sceCtrl now releases it after that many vblank
samples, on the emulator thread; the debugger only reports when it's done.

Also: wsdbg's :screenshot works in headless with Vulkan.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:58:37 -06:00
Henrik RydgårdandClaude Opus 5.5 00c9a2d389 Utility: A savedata shutdown ends at priority 0x20
On hardware the last part of a savedata shutdown runs at priority 0x20,
whatever the dialog's own thread priorities, so a caller at 0x20 gets the
CPU back first and sees SHUTDOWN, and one at 0x21 or worse only sees NONE
(pspautotests utility/savedata/shutdownstatus). We ended it at the access
thread's priority, so Freak Out, which calls ShutdownStart from 0x20 and
waits for SHUTDOWN, sat at 'Please press START' forever. NFL Street 3,
which calls it from 111 and then InitStart straight away, still gets NONE.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:58:37 -06:00
Henrik RydgårdandClaude Opus 5.5 51bac9349e docs: Never run two PSP hardware tests at once
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 54e4be47aa CoreTiming: Grow the event table for events a savestate lacks
Taking the highest unused slot instead could steal one that a state event
restores later, as VBlankWake did to MicBlockingResume, which then had
nowhere to go. Also name the event in the assert.

AGENTS.md: When a savestate fails to load, suspect the branch first.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 d86896bed9 Tlspl: Time out at once like other waits
Recorded on hardware (pspautotests threads/tls/timeout), a Tlspl
allocation follows the same timeout rule as the other waits, including
failing at once for 0 and 1us without writing the timeout back, which the
shared rule it moved to in the last commits didn't give it yet. Before
that it waited the raw timeout, ~30us short.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 80938529a9 Kernel waits: Share waiter ordering and clearing between objects
Priority-ordered waiting lists were sorted with a comparator wrapper per
object (msgpipe, fpl, vpl), or searched with a copy of the same function
(mutex, mbx). HLEKernel::SortWaitingThreadsByPriority() and
FindBestPriorityWaiter() now do both for any waiting list, of thread ids
or of structs with a threadID.

HLEKernel::ClearWaitingThreads() replaces the identical cancel/delete
loops in semaphores, event flags, fpl and vpl.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 cc5ee42dda Kernel waits: One timeout event for every kind of object
Semaphores, event flags, mutexes, lwmutexes, mbx, msgpipes, fpl, vpl,
tlspl and WaitThreadEnd each had their own CoreTiming event, handler
registration and savestate entry for wait timeouts, and their own function
to schedule one. Now one event (WaitThreadEnd's, renamed) times out all of
them, keyed by thread, and dispatches on the thread's wait type to a
timeoutFunc registered alongside the begin/end callback functions.
__KernelWaitCurThreadWithTimeout() starts such a wait, and the HLEKernel
helpers have overloads that use the shared event.

Old savestates still load: each object's section reads its old event id
and points it at the shared handler, so a timeout pending in the state
goes off as before. Checked with a state saved mid-wait by the previous
build, and with four games.

The one behaviour change: tlspl timeouts now follow the same hardware
rule as the others, where they used the raw timeout.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 0c9438e60c CoreTiming: Load states that have event types we no longer register
A state with more event types than are registered now was refused, so no
event could ever be removed or merged. Loading now keeps the extra slots:
modules that still know an old event restore it to a handler, and the rest
stay placeholders that do nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 804dd59df0 Interrupts: Charge for alarm handlers, and stop parking threads on idle
An interrupt with no handler to run, a vblank with none registered for
example, switched the running thread off to idle and left it there until
some later event rescheduled: ~775us of every frame in a game that spins
without a vblank handler. It now reschedules at once. Taking an interrupt
also clears the ll bit directly, which that switch had been doing.

Interrupt handlers can now carry a cost before they run and after the
last queued one returns. Alarms use it: on hardware a thread that keeps
running loses ~70us to an alarm handler, and a thread the handler wakes
runs ~50us after it (pspautotests threads/scheduling/alarmcosts), so
17us in and 40us out. sceKernelSetAlarm's 40us is split evenly around the
deadline, keeping the handler ~1040us after a 1000us alarm.

A handler's return value re-arms its alarm counting from the previous
deadline, so a repeating alarm doesn't drift by those costs, unless
that's already past, as after interrupts were suspended for a while.

The vblank's own cost (~62us of CPU on hardware) isn't charged yet: with
it, a waiter ~90us after the vblank still reads hcount 1 on hardware, but
line 2 here. Hardware evidently raises the interrupt ~40us before the
line count wraps. That's noted where the waiters are released.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 02b2b7b699 sceDisplay: Release vblank waiters ~50us after the vblank
On hardware a thread waiting for vblank returns ~53us after it, where we
had it back in ~5us, and the first of four waiters runs after ~85us: each
waiter beyond the first adds ~9us (pspautotests threads/scheduling/
vblankwake). The waiters are now released by a separate event 48us + 9us
per extra waiter after the vblank. Which vblank a wait is for is still
decided at the vblank, so a thread that starts waiting in between still
waits a whole frame (sceDisplay section version 8).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 f26ca37dad Threads: Context switches cost what they do on hardware
Timed on a PSP, each way of handing the CPU to another thread (the call
and the switch together):

                                     hardware   before   now
  rotate to an equal thread              7        14       7
  signal, better thread runs            10        17      10
  it waits again, back to caller        10        19      12
  wakeup, better thread runs             8        13       6
  it sleeps again, back to caller        7        12       6
  start a better thread, entry          30        28      30
  thread ends, back to its waiter       21        13      20
  notify, better thread's callback      14        13      14

A switch between two threads now costs 1150 cycles instead of 2700.
Starting a better thread costs 2000 cycles more, ending a thread 3300,
and setting up a callback 1800.

Also splits a wait timeout's ~30us into the deadline being taken 12us
into the call and the timeout going off 18us after it. That only changes
the time left written back, which threads/semaphores/wait and
threads/fpl/cancel pin between them. intr/vblank is re-recorded so it
no longer depends on the phase of the frame.

threads/callbacks/combos now passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 f5b4fd739d test.py: Add threads/scheduling dispatchwake and mutexhandoff
Both already pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 a17a9eb147 Callbacks: A notify takes the thread out of its CB wait right away
On hardware, notifying a callback of a thread in a CB wait takes it out of
the wait at once, even though the callback only runs when the thread would
get the CPU. A semaphore signalled in between doesn't end the wait: the
callback runs first, then the wait resumes and takes it (pspautotests
threads/callbacks/combos). We left the thread on the wait list until the
callback started, so the signal ended the wait and the callback didn't
run.

The notify now pauses the wait, as starting a callback used to. If the
callbacks are canceled before the thread's turn comes, the wait just
resumes (Thread savestate section version 7).

threads/callbacks/combos goes in the to-do list: a callback returning to
the thread that notified it still takes ~13us where hardware takes ~9,
part of the context switch cost.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 fb9ac8397e Threads: Wait timeouts work like hardware's, one rule for all of them
Every wait with a timeout behaves the same on hardware (pspautotests
threads/scheduling/waittimeouts). The deadline is taken, and the alarm set
up a moment later. If the deadline has passed by then, the wait fails with
WAIT_TIMEOUT at once, without yielding or writing the timeout back. That's
usual for 0us, half the time for 1us, and rare after; AllocateVpl does more
first. Otherwise it ends max(t, 205us) + ~35us after the call. Each object
had its own guess (24/245, 25/250, 20/250 and so on), and only MsgPipe had
the immediate case.

__KernelWaitTimesOutAtOnce() and __KernelWaitTimeoutUs() now do it for
semaphores, event flags, mutexes, lwmutexes, mbx, msgpipes, fpl, vpl and
WaitThreadEnd. The latency past the deadline isn't counted in the time
left written back.

Outcomes that hardware decides by the clock's phase (these, and
sceKernelDelayThread returning at once) go with the likelier one. Ones
between 50% and certain are instead spread evenly over calls, so a polling
loop can't lock into never yielding (sceKernelThread section version 7).
This replaces the pseudo-random choice for delays.

Also adds threads/scheduling/readyqueue, which already passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 f731e45f40 Alarms: Setting one takes ~40us, and it can't go off within ~215us
On hardware sceKernelSetAlarm takes about 40us, and the handler never runs
sooner than about 215us after the alarm is set, however short it asked
for (pspautotests threads/scheduling/alarmcosts). Also clamps huge
sysclock alarms before converting to cycles; LONG_LONG_MAX used to
overflow and go off at once, which the late-firing events had hidden.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik RydgårdandClaude Opus 5.5 93a0d50384 CoreTiming: An event due before the slice ends shortens it
Each slice is sized to end at the next queued event, but scheduling a
sooner one didn't touch it, so the new event waited for the old slice to
run out. An alarm set by a thread that kept running went off 175-440us
late. GE enqueues worked around this with hleCoreTimingForceCheck(); now
every caller gets it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 09:06:30 -06:00
Henrik Rydgård 3118490ecf Merge pull request #22391 from hrydgard/atrac3-joint-stereo
Atrac3: Correct the decoder setup for mono streams (fixes LocoRoco 2 MuiMui house music)
2026-09-30 08:59:00 -06:00
Henrik RydgårdandClaude Opus 5.5 80b7d27c12 Add the audio/atrac/c0mono and audio/audiocodec/at3param tests
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 17:14:03 -06:00
Henrik RydgårdandClaude Opus 5.5 9ba40416da sceAtrac: Charge ME time when SetData's first frame doesn't decode
That failure comes after the codec is set up and the frame has been tried
on the ME, so the thread waits for both, unlike the other SetData errors.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 17:14:03 -06:00
Henrik RydgårdandClaude Opus 5.5 b9ea98d2c5 Atrac3: Set up the decoder the way libatrac3plus does
libatrac3plus picks the codec parameter from the frame size and the
header's joint stereo flag, and the channel count plays no part. We
guessed joint stereo from the frame size and channel count instead, so
LocoRoco 2's MuiMui house music never started: the game writes a 2-channel
normal-stereo header for every track it streams, and that 0xC0 track holds
one mono sound unit per frame. Taken as joint stereo, its first frame failed
during setup, and the game retried forever.

Now the joint stereo flag comes from the track header, and the decoder's
channel count from the parameter it maps to, so that track decodes as mono
into both output channels, as on hardware. Low-level decoding, which has
no header, still goes by the frame size. Atrac2 saves the flag; older
states fall back to the guess.

Fixes #8647.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 17:14:00 -06:00
Henrik Rydgård 3bcd93d594 Merge pull request #22394 from hrydgard/gameinfo-priorities
GameInfoCache: Prioritize loads, so launching doesn't wait behind search
2026-09-29 17:05:27 -06:00
Henrik RydgårdandClaude Opus 5.5 3aaea4c0f9 GameInfoCache: Prioritize loads, so launching doesn't wait behind search
Type-to-search asks for the title of every game in the list, queueing a
load for each, and launching a game whose info wasn't loaded yet then
blocked until the whole queue had drained. Search's loads are now LOW,
and launching asks at HIGH, which queues its own load rather than wait
for a pending one.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 16:34:38 -06:00
Henrik Rydgård ae460c9e1a Merge pull request #22390 from hrydgard/gpu-review-leftovers
Claude code review of GPU: Fix leftover findings
2026-09-29 15:43:31 -06:00
Henrik Rydgård 68755d7f63 Merge pull request #22392 from hrydgard/headless-vulkan-screenshot-assert
Headless: Take the timeout screenshot before ending the draw frame
2026-09-29 15:41:12 -06:00
Henrik RydgårdandClaude Opus 5.5 f6e70b88ad Headless: Take the timeout screenshot before ending the draw frame
On Vulkan, reading back the display framebuffer after EndDrawFrame hit
the insideFrame_ assert in CopyFramebufferToMemory, so any run that
timed out with --screenshot-save crashed in debug builds.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 14:05:00 -06:00
Henrik Rydgård c9c26ce0fd Merge pull request #22389 from hrydgard/savedata-icon-black-background
Savedata dialog: Draw icons over black, with alpha blending
2026-09-29 13:52:34 -06:00
Henrik Rydgård 949423d8e9 Merge pull request #22388 from hrydgard/update-xxhash
Update xxhash to v0.8.4, use XXH3 for texture hashing
2026-09-29 13:52:00 -06:00
Henrik Rydgård b9c5b28b8f Merge pull request #22386 from hrydgard/adhoc-shutdown-race
Adhoc: Fix shutdown hanging when it comes right after adhoc init
2026-09-29 13:29:20 -06:00
Henrik RydgårdandClaude Opus 5.5 b9501df529 TextureReplacer: Fix stale and shared lookup results
- Reloading the ini clears the per-key lookup caches, which could keep
  saying "no replacement" for textures the new ini replaces.
- With ignoreAddress, hash ranges were skipped when sizing the
  replacement, though ComputeHash applies them. cache_ is now keyed by
  the full key, so its lookups hit too.
- Textures sharing files but differing in size, hash range or filtering
  no longer share one ReplacedTexture (the first one's settings won).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:28:21 -06:00
Henrik RydgårdandClaude Opus 5.5 1288c294fd Savedata dialog: Draw icons over black, with alpha blending
The PSP blends save icons over a black background, so their transparent
parts come out black. 0a5fa27957 turned blending off instead, which is
wrong for icons that rely on it. Revert that, and draw a black rectangle
under each icon. Fixes #22280.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:21:24 -06:00
Henrik RydgårdandClaude Opus 5.5 bd27ceb669 Headless: Reject --debugger together with --debugger-run
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:17:01 -06:00
Henrik RydgårdandClaude Opus 5.5 b8789af997 GPU: Draw frames displayed from RAM in non-buffered mode
They were uploaded but never drawn, and the block transfer hack drew
whatever source was left over. Draw them straight into the backbuffer
pass, without post shaders, which would need to bind their own targets.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:17:01 -06:00
Henrik RydgårdandClaude Opus 5.5 4bfa4e14ec Vulkan: Make WaitForPipelines wait for queued compiles too
Pipelines still in the compile queue weren't in flight yet, so the
shader cache load could stop waiting (and the spinner) early.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:17:01 -06:00
Henrik RydgårdandClaude Opus 5.5 39b6cc0c0c Vulkan: Fix the debugger screenshot in headless
There's no swapchain, so reading the backbuffer asserted. Fail the
readback, and have gpu.buffer.screenshot fall back to the displayed
framebuffer.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:17:00 -06:00
Henrik RydgårdandClaude Opus 5.5 7e553b0732 GE debugger: Lock when adding command breakpoints
Debuggers add them from their own threads, racing ClearTempBreakpoints
on the emu thread.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:17:00 -06:00
Henrik RydgårdandClaude Opus 5.5 035c98d14f D3D11: Re-read the device and context on DeviceRestore
The restored draw context can be a new device, so the cached pointers
could go stale.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:16:59 -06:00
Henrik RydgårdandClaude Opus 5.5 ba37f327b0 D3D11: Check Map() results
Map fails after device removal, leaving pData garbage. Skip the upload or
draw instead of writing through it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:16:59 -06:00
Henrik RydgårdandClaude Opus 5.5 bc581349fd GPU: Install the draw engines' invalidation callback from BeginFrame
The draw engine is created on the loader thread while the UI thread may
already be rendering and invoking the callback.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:16:42 -06:00
Henrik RydgårdandClaude Opus 5.5 759b494b6a GPU: Check post shaders once per host frame, after resizes
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:16:41 -06:00