With depth clip off, the GE doesn't clip triangles and lines at the near
plane: the part beyond it is drawn with extrapolated depth, and a vertex behind
the camera (w <= 0) culls the whole primitive instead. With it on, they're
clipped, and a vertex that gets clipped away can't cull the primitive by being
outside the screen range (w = 0 included). Rects, which aren't clipped, are
culled by a vertex behind the camera either way. Measured with gpu/probe.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Document pspautotests' ppdmp-playback and its new run.py, which replays a
frame dump over PSPLink and compares it with PPSSPPHeadless.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The GE counts a vertex as outside when |z| > w, comparing the float24 clip
coordinates without dividing, and culls a primitive only when all its vertices
are outside, with depth clamp on or off. Points were never culled, and lines,
triangles and rectangles were culled on a single outside vertex with clamp off.
Measured with gpu/probe for every primitive type, both clamp modes and a
range of w.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The GE converts the fog factor to 8 bits per vertex, min(floor(256 * f), 255),
and interpolates that linearly in screen space. Clamping after interpolation
instead was up to 127 levels off on triangles that cross the fog range.
Measured with gpu/probe; points and wacky values still match exactly.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Truncating both terms to the precision of the larger one is clearing the
low 8 + (exponent difference) bits of each, after which a float sum is
exact. Only the products still go through doubles, since a product of two
float24s needs up to 32 bits before it's truncated. Same results.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Measured on hardware (gpu/depth/transformprecision, gpu/depth/precision):
clip Z and W come out of the combined matrix as float24s, z/w uses the
GE's reciprocal, which linearly interpolates 128 table segments and isn't
quite 1/w, and the viewport's z * scale + center adds without guard bits:
the smaller term is truncated to the precision of the larger one.
Fixes the Test Drive Unlimited map (#12786) and the missing background in
Fuuun Shinsengumi Bakumatsuden (from #16131), where vertices land right at
the minimum depth.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Includes the GachiTora reference from the spline lighting branch, which
is to be merged first.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Read as an integer, a float's bits are its log2 with the mantissa a
straight line between powers of two, scaled by 2^23, and writing an
integer back is the matching exp2. That's exactly the GE's
approximation, without log2/exp2/floor. pspPow now also returns 1 for
e <= 0 itself, so the callers drop their checks. Shader languages
without integers fall back to a true pow.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Without SSE or NEON, triangle pixels got the secondary color in place of
the primary one plus it, so lit triangles came out black. It showed as
the "unexplained" known failures on riscv64 and loongarch64, and broke
the new gpu/lighting/shademap there. Reproduced on arm64 by building
without NEON.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Environment map S and T are (N.L + 1) / 2 with L the light's vector as
lighting uses it: from the vertex to the light for point and spot lights,
a zero vector staying zero, and the half vector for a light that does
specular. Whether lighting or the light is enabled still doesn't matter
(gpu/lighting/shademap).
The vertex shader ID now carries the type and computation of the shade
mapping lights (the ubershader reads them from u_lightControl), so both
shader caches get a new version.
Fixes the hair shine in iDOLM@STER SP (#12376).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The viewer is at infinity along view space +z, so in world space, where
lighting happens, it's the view matrix's third column rather than
(0,0,1): turning the camera moves the highlights (gpu/lighting/specular).
The shaders read it from u_view.
In a Need for Speed Carbon frame replayed on a PSP, this and the pow
bring the error of the cars from MSE 957 to 120 (Vulkan).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The shaders and the software renderer only add specular when N.L >= 0,
as the GE does (gpu/lighting/specular); the CPU lighter, used for points,
lines and rectangles, didn't check.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The GE computes these powers as exp2(e * log2(x)), with log2 and exp2
each a straight line between powers of two (Mitchell's approximation),
and only uses the top 4 bits of the specular coefficient's mantissa.
Through a highlight's falloff a true pow is 10-30 steps of 255 brighter
at the exponents games use. Measured in gpu/lighting/specular.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
x86 normalizes with a bare _mm_rsqrt_ps, so normal lengths are only good
to about 3.7e-4 there and the 1e-4 check failed on every x86 CI job.
Compare direction after normalizing instead.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With animated control points the pole is only nearly degenerate, so the
vanishing derivative is rounding noise rather than exactly zero, and the
absolute threshold missed it. The resulting random normals still showed
as dark patches on Pac-Man Arrangement's ghosts (#12354).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Software transform expands each point, line and rectangle to four
vertices, and gave up on the whole draw when that didn't fit
VERTEX_BUFFER_MAX. Batching only counted input vertices, so a batch over
16384 points (or 32768 line or rectangle vertices) vanished silently,
whether it came from one PRIM or several merged ones. Count the expanded
vertices when batching, and submit a PRIM too big on its own in parts.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Compares the software tessellator's positions, normals, UVs and colors
against a longhand double-precision reference (Bernstein polynomials,
Cox-de Boor with clamped knots, finite-difference normals), across
Bezier and spline surfaces, edge types, poles and patch facing, so it
can be optimized with something independent to be wrong against. Also
reports vertices per second.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With only one patch, the open-last-edge adjustment assumed the first edge
was closed, so one knot interval came out as 2 instead of 1, and the
patch wasn't the Bezier patch a fully clamped cubic is. Positions were off
by up to about 7% of the patch.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Four tests had their own copy of "call this until N seconds have passed,
then divide". CallsPerSecond in UnitTest.h does it; each keeps its old
duration and batch size.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The GLES sampler uniforms and texture slots for the control points and
weights, the Vulkan storage buffer bindings, and the u_spline_counts
uniform, which becomes padding (the C++ side already was).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Where all the control points along a patch edge meet at one point, like
the top of a dome, one derivative is zero and so is the cross product,
and normalizing it gave NaN. Use the limit instead, built from the mixed
second derivative. Fixes the dark spots on the ghosts' heads in Pac-Man
Arrangement (#12354).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The release was counted down on the WebSocket thread, one step per poll
of host time however many vblanks had passed, so how long a scripted
press lasted depended on how fast the emulator ran, and scripted runs
went different ways. sceCtrl now releases it after that many vblank
samples, on the emulator thread; the debugger only reports when it's done.
Also: wsdbg's :screenshot works in headless with Vulkan.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware the last part of a savedata shutdown runs at priority 0x20,
whatever the dialog's own thread priorities, so a caller at 0x20 gets the
CPU back first and sees SHUTDOWN, and one at 0x21 or worse only sees NONE
(pspautotests utility/savedata/shutdownstatus). We ended it at the access
thread's priority, so Freak Out, which calls ShutdownStart from 0x20 and
waits for SHUTDOWN, sat at 'Please press START' forever. NFL Street 3,
which calls it from 111 and then InitStart straight away, still gets NONE.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Taking the highest unused slot instead could steal one that a state event
restores later, as VBlankWake did to MicBlockingResume, which then had
nowhere to go. Also name the event in the assert.
AGENTS.md: When a savestate fails to load, suspect the branch first.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Recorded on hardware (pspautotests threads/tls/timeout), a Tlspl
allocation follows the same timeout rule as the other waits, including
failing at once for 0 and 1us without writing the timeout back, which the
shared rule it moved to in the last commits didn't give it yet. Before
that it waited the raw timeout, ~30us short.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Priority-ordered waiting lists were sorted with a comparator wrapper per
object (msgpipe, fpl, vpl), or searched with a copy of the same function
(mutex, mbx). HLEKernel::SortWaitingThreadsByPriority() and
FindBestPriorityWaiter() now do both for any waiting list, of thread ids
or of structs with a threadID.
HLEKernel::ClearWaitingThreads() replaces the identical cancel/delete
loops in semaphores, event flags, fpl and vpl.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Semaphores, event flags, mutexes, lwmutexes, mbx, msgpipes, fpl, vpl,
tlspl and WaitThreadEnd each had their own CoreTiming event, handler
registration and savestate entry for wait timeouts, and their own function
to schedule one. Now one event (WaitThreadEnd's, renamed) times out all of
them, keyed by thread, and dispatches on the thread's wait type to a
timeoutFunc registered alongside the begin/end callback functions.
__KernelWaitCurThreadWithTimeout() starts such a wait, and the HLEKernel
helpers have overloads that use the shared event.
Old savestates still load: each object's section reads its old event id
and points it at the shared handler, so a timeout pending in the state
goes off as before. Checked with a state saved mid-wait by the previous
build, and with four games.
The one behaviour change: tlspl timeouts now follow the same hardware
rule as the others, where they used the raw timeout.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A state with more event types than are registered now was refused, so no
event could ever be removed or merged. Loading now keeps the extra slots:
modules that still know an old event restore it to a handler, and the rest
stay placeholders that do nothing.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
An interrupt with no handler to run, a vblank with none registered for
example, switched the running thread off to idle and left it there until
some later event rescheduled: ~775us of every frame in a game that spins
without a vblank handler. It now reschedules at once. Taking an interrupt
also clears the ll bit directly, which that switch had been doing.
Interrupt handlers can now carry a cost before they run and after the
last queued one returns. Alarms use it: on hardware a thread that keeps
running loses ~70us to an alarm handler, and a thread the handler wakes
runs ~50us after it (pspautotests threads/scheduling/alarmcosts), so
17us in and 40us out. sceKernelSetAlarm's 40us is split evenly around the
deadline, keeping the handler ~1040us after a 1000us alarm.
A handler's return value re-arms its alarm counting from the previous
deadline, so a repeating alarm doesn't drift by those costs, unless
that's already past, as after interrupts were suspended for a while.
The vblank's own cost (~62us of CPU on hardware) isn't charged yet: with
it, a waiter ~90us after the vblank still reads hcount 1 on hardware, but
line 2 here. Hardware evidently raises the interrupt ~40us before the
line count wraps. That's noted where the waiters are released.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware a thread waiting for vblank returns ~53us after it, where we
had it back in ~5us, and the first of four waiters runs after ~85us: each
waiter beyond the first adds ~9us (pspautotests threads/scheduling/
vblankwake). The waiters are now released by a separate event 48us + 9us
per extra waiter after the vblank. Which vblank a wait is for is still
decided at the vblank, so a thread that starts waiting in between still
waits a whole frame (sceDisplay section version 8).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Timed on a PSP, each way of handing the CPU to another thread (the call
and the switch together):
hardware before now
rotate to an equal thread 7 14 7
signal, better thread runs 10 17 10
it waits again, back to caller 10 19 12
wakeup, better thread runs 8 13 6
it sleeps again, back to caller 7 12 6
start a better thread, entry 30 28 30
thread ends, back to its waiter 21 13 20
notify, better thread's callback 14 13 14
A switch between two threads now costs 1150 cycles instead of 2700.
Starting a better thread costs 2000 cycles more, ending a thread 3300,
and setting up a callback 1800.
Also splits a wait timeout's ~30us into the deadline being taken 12us
into the call and the timeout going off 18us after it. That only changes
the time left written back, which threads/semaphores/wait and
threads/fpl/cancel pin between them. intr/vblank is re-recorded so it
no longer depends on the phase of the frame.
threads/callbacks/combos now passes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware, notifying a callback of a thread in a CB wait takes it out of
the wait at once, even though the callback only runs when the thread would
get the CPU. A semaphore signalled in between doesn't end the wait: the
callback runs first, then the wait resumes and takes it (pspautotests
threads/callbacks/combos). We left the thread on the wait list until the
callback started, so the signal ended the wait and the callback didn't
run.
The notify now pauses the wait, as starting a callback used to. If the
callbacks are canceled before the thread's turn comes, the wait just
resumes (Thread savestate section version 7).
threads/callbacks/combos goes in the to-do list: a callback returning to
the thread that notified it still takes ~13us where hardware takes ~9,
part of the context switch cost.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Every wait with a timeout behaves the same on hardware (pspautotests
threads/scheduling/waittimeouts). The deadline is taken, and the alarm set
up a moment later. If the deadline has passed by then, the wait fails with
WAIT_TIMEOUT at once, without yielding or writing the timeout back. That's
usual for 0us, half the time for 1us, and rare after; AllocateVpl does more
first. Otherwise it ends max(t, 205us) + ~35us after the call. Each object
had its own guess (24/245, 25/250, 20/250 and so on), and only MsgPipe had
the immediate case.
__KernelWaitTimesOutAtOnce() and __KernelWaitTimeoutUs() now do it for
semaphores, event flags, mutexes, lwmutexes, mbx, msgpipes, fpl, vpl and
WaitThreadEnd. The latency past the deadline isn't counted in the time
left written back.
Outcomes that hardware decides by the clock's phase (these, and
sceKernelDelayThread returning at once) go with the likelier one. Ones
between 50% and certain are instead spread evenly over calls, so a polling
loop can't lock into never yielding (sceKernelThread section version 7).
This replaces the pseudo-random choice for delays.
Also adds threads/scheduling/readyqueue, which already passes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware sceKernelSetAlarm takes about 40us, and the handler never runs
sooner than about 215us after the alarm is set, however short it asked
for (pspautotests threads/scheduling/alarmcosts). Also clamps huge
sysclock alarms before converting to cycles; LONG_LONG_MAX used to
overflow and go off at once, which the late-firing events had hidden.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Each slice is sized to end at the next queued event, but scheduling a
sooner one didn't touch it, so the new event waited for the old slice to
run out. An alarm set by a thread that kept running went off 175-440us
late. GE enqueues worked around this with hleCoreTimingForceCheck(); now
every caller gets it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>