Includes the GachiTora reference from the spline lighting branch, which
is to be merged first.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Read as an integer, a float's bits are its log2 with the mantissa a
straight line between powers of two, scaled by 2^23, and writing an
integer back is the matching exp2. That's exactly the GE's
approximation, without log2/exp2/floor. pspPow now also returns 1 for
e <= 0 itself, so the callers drop their checks. Shader languages
without integers fall back to a true pow.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Without SSE or NEON, triangle pixels got the secondary color in place of
the primary one plus it, so lit triangles came out black. It showed as
the "unexplained" known failures on riscv64 and loongarch64, and broke
the new gpu/lighting/shademap there. Reproduced on arm64 by building
without NEON.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Environment map S and T are (N.L + 1) / 2 with L the light's vector as
lighting uses it: from the vertex to the light for point and spot lights,
a zero vector staying zero, and the half vector for a light that does
specular. Whether lighting or the light is enabled still doesn't matter
(gpu/lighting/shademap).
The vertex shader ID now carries the type and computation of the shade
mapping lights (the ubershader reads them from u_lightControl), so both
shader caches get a new version.
Fixes the hair shine in iDOLM@STER SP (#12376).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The viewer is at infinity along view space +z, so in world space, where
lighting happens, it's the view matrix's third column rather than
(0,0,1): turning the camera moves the highlights (gpu/lighting/specular).
The shaders read it from u_view.
In a Need for Speed Carbon frame replayed on a PSP, this and the pow
bring the error of the cars from MSE 957 to 120 (Vulkan).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The shaders and the software renderer only add specular when N.L >= 0,
as the GE does (gpu/lighting/specular); the CPU lighter, used for points,
lines and rectangles, didn't check.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The GE computes these powers as exp2(e * log2(x)), with log2 and exp2
each a straight line between powers of two (Mitchell's approximation),
and only uses the top 4 bits of the specular coefficient's mantissa.
Through a highlight's falloff a true pow is 10-30 steps of 255 brighter
at the exponents games use. Measured in gpu/lighting/specular.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
x86 normalizes with a bare _mm_rsqrt_ps, so normal lengths are only good
to about 3.7e-4 there and the 1e-4 check failed on every x86 CI job.
Compare direction after normalizing instead.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With animated control points the pole is only nearly degenerate, so the
vanishing derivative is rounding noise rather than exactly zero, and the
absolute threshold missed it. The resulting random normals still showed
as dark patches on Pac-Man Arrangement's ghosts (#12354).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Software transform expands each point, line and rectangle to four
vertices, and gave up on the whole draw when that didn't fit
VERTEX_BUFFER_MAX. Batching only counted input vertices, so a batch over
16384 points (or 32768 line or rectangle vertices) vanished silently,
whether it came from one PRIM or several merged ones. Count the expanded
vertices when batching, and submit a PRIM too big on its own in parts.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Compares the software tessellator's positions, normals, UVs and colors
against a longhand double-precision reference (Bernstein polynomials,
Cox-de Boor with clamped knots, finite-difference normals), across
Bezier and spline surfaces, edge types, poles and patch facing, so it
can be optimized with something independent to be wrong against. Also
reports vertices per second.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With only one patch, the open-last-edge adjustment assumed the first edge
was closed, so one knot interval came out as 2 instead of 1, and the
patch wasn't the Bezier patch a fully clamped cubic is. Positions were off
by up to about 7% of the patch.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Four tests had their own copy of "call this until N seconds have passed,
then divide". CallsPerSecond in UnitTest.h does it; each keeps its old
duration and batch size.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The GLES sampler uniforms and texture slots for the control points and
weights, the Vulkan storage buffer bindings, and the u_spline_counts
uniform, which becomes padding (the C++ side already was).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Where all the control points along a patch edge meet at one point, like
the top of a dome, one derivative is zero and so is the cross product,
and normalizing it gave NaN. Use the limit instead, built from the mixed
second derivative. Fixes the dark spots on the ghosts' heads in Pac-Man
Arrangement (#12354).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The release was counted down on the WebSocket thread, one step per poll
of host time however many vblanks had passed, so how long a scripted
press lasted depended on how fast the emulator ran, and scripted runs
went different ways. sceCtrl now releases it after that many vblank
samples, on the emulator thread; the debugger only reports when it's done.
Also: wsdbg's :screenshot works in headless with Vulkan.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware the last part of a savedata shutdown runs at priority 0x20,
whatever the dialog's own thread priorities, so a caller at 0x20 gets the
CPU back first and sees SHUTDOWN, and one at 0x21 or worse only sees NONE
(pspautotests utility/savedata/shutdownstatus). We ended it at the access
thread's priority, so Freak Out, which calls ShutdownStart from 0x20 and
waits for SHUTDOWN, sat at 'Please press START' forever. NFL Street 3,
which calls it from 111 and then InitStart straight away, still gets NONE.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Taking the highest unused slot instead could steal one that a state event
restores later, as VBlankWake did to MicBlockingResume, which then had
nowhere to go. Also name the event in the assert.
AGENTS.md: When a savestate fails to load, suspect the branch first.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Recorded on hardware (pspautotests threads/tls/timeout), a Tlspl
allocation follows the same timeout rule as the other waits, including
failing at once for 0 and 1us without writing the timeout back, which the
shared rule it moved to in the last commits didn't give it yet. Before
that it waited the raw timeout, ~30us short.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Priority-ordered waiting lists were sorted with a comparator wrapper per
object (msgpipe, fpl, vpl), or searched with a copy of the same function
(mutex, mbx). HLEKernel::SortWaitingThreadsByPriority() and
FindBestPriorityWaiter() now do both for any waiting list, of thread ids
or of structs with a threadID.
HLEKernel::ClearWaitingThreads() replaces the identical cancel/delete
loops in semaphores, event flags, fpl and vpl.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Semaphores, event flags, mutexes, lwmutexes, mbx, msgpipes, fpl, vpl,
tlspl and WaitThreadEnd each had their own CoreTiming event, handler
registration and savestate entry for wait timeouts, and their own function
to schedule one. Now one event (WaitThreadEnd's, renamed) times out all of
them, keyed by thread, and dispatches on the thread's wait type to a
timeoutFunc registered alongside the begin/end callback functions.
__KernelWaitCurThreadWithTimeout() starts such a wait, and the HLEKernel
helpers have overloads that use the shared event.
Old savestates still load: each object's section reads its old event id
and points it at the shared handler, so a timeout pending in the state
goes off as before. Checked with a state saved mid-wait by the previous
build, and with four games.
The one behaviour change: tlspl timeouts now follow the same hardware
rule as the others, where they used the raw timeout.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A state with more event types than are registered now was refused, so no
event could ever be removed or merged. Loading now keeps the extra slots:
modules that still know an old event restore it to a handler, and the rest
stay placeholders that do nothing.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
An interrupt with no handler to run, a vblank with none registered for
example, switched the running thread off to idle and left it there until
some later event rescheduled: ~775us of every frame in a game that spins
without a vblank handler. It now reschedules at once. Taking an interrupt
also clears the ll bit directly, which that switch had been doing.
Interrupt handlers can now carry a cost before they run and after the
last queued one returns. Alarms use it: on hardware a thread that keeps
running loses ~70us to an alarm handler, and a thread the handler wakes
runs ~50us after it (pspautotests threads/scheduling/alarmcosts), so
17us in and 40us out. sceKernelSetAlarm's 40us is split evenly around the
deadline, keeping the handler ~1040us after a 1000us alarm.
A handler's return value re-arms its alarm counting from the previous
deadline, so a repeating alarm doesn't drift by those costs, unless
that's already past, as after interrupts were suspended for a while.
The vblank's own cost (~62us of CPU on hardware) isn't charged yet: with
it, a waiter ~90us after the vblank still reads hcount 1 on hardware, but
line 2 here. Hardware evidently raises the interrupt ~40us before the
line count wraps. That's noted where the waiters are released.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware a thread waiting for vblank returns ~53us after it, where we
had it back in ~5us, and the first of four waiters runs after ~85us: each
waiter beyond the first adds ~9us (pspautotests threads/scheduling/
vblankwake). The waiters are now released by a separate event 48us + 9us
per extra waiter after the vblank. Which vblank a wait is for is still
decided at the vblank, so a thread that starts waiting in between still
waits a whole frame (sceDisplay section version 8).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Timed on a PSP, each way of handing the CPU to another thread (the call
and the switch together):
hardware before now
rotate to an equal thread 7 14 7
signal, better thread runs 10 17 10
it waits again, back to caller 10 19 12
wakeup, better thread runs 8 13 6
it sleeps again, back to caller 7 12 6
start a better thread, entry 30 28 30
thread ends, back to its waiter 21 13 20
notify, better thread's callback 14 13 14
A switch between two threads now costs 1150 cycles instead of 2700.
Starting a better thread costs 2000 cycles more, ending a thread 3300,
and setting up a callback 1800.
Also splits a wait timeout's ~30us into the deadline being taken 12us
into the call and the timeout going off 18us after it. That only changes
the time left written back, which threads/semaphores/wait and
threads/fpl/cancel pin between them. intr/vblank is re-recorded so it
no longer depends on the phase of the frame.
threads/callbacks/combos now passes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware, notifying a callback of a thread in a CB wait takes it out of
the wait at once, even though the callback only runs when the thread would
get the CPU. A semaphore signalled in between doesn't end the wait: the
callback runs first, then the wait resumes and takes it (pspautotests
threads/callbacks/combos). We left the thread on the wait list until the
callback started, so the signal ended the wait and the callback didn't
run.
The notify now pauses the wait, as starting a callback used to. If the
callbacks are canceled before the thread's turn comes, the wait just
resumes (Thread savestate section version 7).
threads/callbacks/combos goes in the to-do list: a callback returning to
the thread that notified it still takes ~13us where hardware takes ~9,
part of the context switch cost.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Every wait with a timeout behaves the same on hardware (pspautotests
threads/scheduling/waittimeouts). The deadline is taken, and the alarm set
up a moment later. If the deadline has passed by then, the wait fails with
WAIT_TIMEOUT at once, without yielding or writing the timeout back. That's
usual for 0us, half the time for 1us, and rare after; AllocateVpl does more
first. Otherwise it ends max(t, 205us) + ~35us after the call. Each object
had its own guess (24/245, 25/250, 20/250 and so on), and only MsgPipe had
the immediate case.
__KernelWaitTimesOutAtOnce() and __KernelWaitTimeoutUs() now do it for
semaphores, event flags, mutexes, lwmutexes, mbx, msgpipes, fpl, vpl and
WaitThreadEnd. The latency past the deadline isn't counted in the time
left written back.
Outcomes that hardware decides by the clock's phase (these, and
sceKernelDelayThread returning at once) go with the likelier one. Ones
between 50% and certain are instead spread evenly over calls, so a polling
loop can't lock into never yielding (sceKernelThread section version 7).
This replaces the pseudo-random choice for delays.
Also adds threads/scheduling/readyqueue, which already passes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware sceKernelSetAlarm takes about 40us, and the handler never runs
sooner than about 215us after the alarm is set, however short it asked
for (pspautotests threads/scheduling/alarmcosts). Also clamps huge
sysclock alarms before converting to cycles; LONG_LONG_MAX used to
overflow and go off at once, which the late-firing events had hidden.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Each slice is sized to end at the next queued event, but scheduling a
sooner one didn't touch it, so the new event waited for the old slice to
run out. An alarm set by a thread that kept running went off 175-440us
late. GE enqueues worked around this with hleCoreTimingForceCheck(); now
every caller gets it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
That failure comes after the codec is set up and the frame has been tried
on the ME, so the thread waits for both, unlike the other SetData errors.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
libatrac3plus picks the codec parameter from the frame size and the
header's joint stereo flag, and the channel count plays no part. We
guessed joint stereo from the frame size and channel count instead, so
LocoRoco 2's MuiMui house music never started: the game writes a 2-channel
normal-stereo header for every track it streams, and that 0xC0 track holds
one mono sound unit per frame. Taken as joint stereo, its first frame failed
during setup, and the game retried forever.
Now the joint stereo flag comes from the track header, and the decoder's
channel count from the parameter it maps to, so that track decodes as mono
into both output channels, as on hardware. Low-level decoding, which has
no header, still goes by the frame size. Atrac2 saves the flag; older
states fall back to the guess.
Fixes#8647.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Type-to-search asks for the title of every game in the list, queueing a
load for each, and launching a game whose info wasn't loaded yet then
blocked until the whole queue had drained. Search's loads are now LOW,
and launching asks at HIGH, which queues its own load rather than wait
for a pending one.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On Vulkan, reading back the display framebuffer after EndDrawFrame hit
the insideFrame_ assert in CopyFramebufferToMemory, so any run that
timed out with --screenshot-save crashed in debug builds.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Reloading the ini clears the per-key lookup caches, which could keep
saying "no replacement" for textures the new ini replaces.
- With ignoreAddress, hash ranges were skipped when sizing the
replacement, though ComputeHash applies them. cache_ is now keyed by
the full key, so its lookups hit too.
- Textures sharing files but differing in size, hash range or filtering
no longer share one ReplacedTexture (the first one's settings won).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The PSP blends save icons over a black background, so their transparent
parts come out black. 0a5fa27957 turned blending off instead, which is
wrong for icons that rely on it. Revert that, and draw a black rectangle
under each icon. Fixes#22280.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
They were uploaded but never drawn, and the block transfer hack drew
whatever source was left over. Draw them straight into the backbuffer
pass, without post shaders, which would need to bind their own targets.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Pipelines still in the compile queue weren't in flight yet, so the
shader cache load could stop waiting (and the spinner) early.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
There's no swapchain, so reading the backbuffer asserted. Fail the
readback, and have gpu.buffer.screenshot fall back to the displayed
framebuffer.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Debuggers add them from their own threads, racing ClearTempBreakpoints
on the emu thread.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Map fails after device removal, leaving pData garbage. Skip the upload or
draw instead of writing through it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The draw engine is created on the loader thread while the UI thread may
already be rendering and invoking the callback.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Flush() on an empty queue now trims the state and CLUT rings, since its
callers flush because one is full and push right after.
- BinQueue::Full() uses >=, so an overshoot can't go unnoticed.
- IsExactSelfRender compares against the target the queued draws were
binned for, not gstate, which already has the next one during a flush.
- A depth test without depth writes marks the depth buffer as read.
- The DarkStalkers untextured sprite recomputes the binner state around it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With xxhash 0.8.4, XXH3 is faster than StableQuickTexHash on ARM64
(about 37 vs 24 GB/s on a Snapdragon X), and its scalar path, which
RISC-V builds get, is about as fast as the quick hash's. It doesn't
collide the way the quick hash does (#8249).
Texture replacement still uses its own hash setting, so texture packs
are unaffected.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Keeps our two local changes (the ppsspp_config.h include and the ARM32
prefetch in the XXH32 loop). Hash values are unchanged.
XXH3 is much faster on ARM64 now: about 37 GB/s against 13.5 with v0.8.1
on a Snapdragon X, where our quick texture hash does 24.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Homebrew written for a PSP-2000+ under custom firmware can use the top
32MB of RAM directly without setting MEMSIZE, which leaves the user
partition at its normal size. NJEMU's slim builds do this (#8925).
Previously we didn't map that memory at all for PBPs; MEMSIZE=1 isn't a
workaround either, since it grows the partition and the heap and stacks
land where the program writes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
NJEMU's SystemButtons.prx kernel plugin imports it when the firmware
reports 3.71 or later, and polls it every frame to read HOME/volume.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The net tests need WLAN on, but that also made every game run log into a
real adhoc server on the internet, so results depended on the network.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
sceNetAdhocctlInit waits until the friend finder has logged into the adhoc
server, polling in emulated time but giving up only after 5s of wall time.
The friend finder makes one attempt per login request, and it too waited
the full 5s on a connection that had already been refused, because it
only looked for success. Now it checks the socket error and gives up at
once, and records the failure so the wait ends with it.
In headless, which runs far ahead of real time, Gods Eater Burst sat in
this wait for its whole run.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Start from the quality that fit the previous frame instead of the top, so
most frames encode once. Step down while a frame is too big, and step back
up when one comes out under half the limit. Windows and the recompression
fallback share the logic.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Android, iOS, macOS and Linux encode camera frames at a fixed quality, so a
detailed frame can exceed the game's framesize just as on Windows before.
pushCameraImage now decodes such a frame and re-encodes it at lower quality
until it fits.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The friend finder thread set friendFinderRunning itself, after a DNS
lookup of the adhoc server. A shutdown in that window cleared the flag
first; the thread then set it again, looped forever, and the join in
NetAdhocctl_Term() never returned. Gods Eater Burst hit this in about one
headless run in six. The flag is now set before the thread is created,
and a finished thread is joined before a new one replaces it, which
would otherwise call std::terminate.
The built-in adhoc server thread had the same race, behind its check for
an existing server, and gets the same fix.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The PSP camera keeps JPEG frames within the framesize from the video setup.
Go!Edit stores frames in 15KB slots and only takes frames that fit, so our
uncompressed-quality webcam frames were dropped and one frame got repeated
for the whole clip. Lower the JPEG quality until a frame fits.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With a host microphone present, a blocking read waited until the host had
delivered all the data. If it never did, the thread waited forever: Go!Edit's
sound thread stalled that way and its video recording never advanced. The
PSP mic streams in real time, so wake at the scheduled time and fill what
the host didn't deliver with silence.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
:screenshot saves gpu.buffer.screenshot as a PNG without dumping the data
URI into the output. The docs claimed nested parameters need a raw JSON
line, which gets no ticket; a single-quoted JSON value in key=value form
works and keeps it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
It returned at once. Go!Edit reads frames in a loop on a high-priority
thread (bhCameraGetJpeg) that only yields to its own priority level, so
after the "Loading complete" dialog it spun and starved the rest of the
game. Return at the camera's next frame tick, at the rate from the setup
params.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Only the rounding mode, flags, enables, cause, FCC and FS bits can be
written (0x0181FFFF, pspautotests cpu/fpu/fcr), as the interpreter, IR
and x86 already have it. Both ARM JITs stored the whole value.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware sceVaudioChReserve takes ~260us once it gets past the busy
check, succeed or fail, and worse threads can run meanwhile. Releasing
takes ~25us and doesn't wait. We returned at once, which is what
audio/sceaudio/reserve's [r] markers showed; it now passes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A delay's deadline is now + usec, and the clock is read again when the
alarm is set. If the deadline has passed by then, the call returns 0 at
once without giving up the CPU. On hardware that makes
sceKernelDelayThread(0) return at once about 60% of the time. On a thread's
first wait after it starts, a delay of 1 does so about two times in three
as well (pspautotests threads/scheduling/delayzero). We always waited at
least 210us.
The choice is pseudo-random off the tick count, not the tick phase, since
our cycle counts are regular enough for a polling loop to lock into never
yielding. Threads remember whether they've waited since starting (Thread
savestate section version 6).
Also moves threads/vpl/create into the passing tests: re-recorded on 6.61,
it agrees with what we do for partitions 8 and 9.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
With dispatch suspended, IO fails in the driver when it tries to wait. The
memory stick driver returns SCE_KERNEL_ERROR_CAN_NOT_WAIT, while usbhostfs,
which serves host0: under PSPLink, returns -1. host0: is mostly what
homebrew developers run from, so it now does the same.
Also from threads/scheduling/dispatch, which now passes:
- sceIoRead reports an async operation still in progress before failing
on suspended dispatch.
- A write to stdout or stderr doesn't give up the CPU.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
On hardware a load costs an open and a read of the file, plus about 1ms
and 30us per KB of loader work (pspautotests threads/scheduling/callcosts),
with the caller waiting throughout. It now charges sceIoOpen's and
sceIoRead's estimates for the file plus that, instead of a flat 500us.
Also notes why sceKernelLoadModuleByID fails from a game's own fd on ms0:
or host0: on hardware, which we don't emulate.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
These own GPU objects, memory or refcounts in their destructors (or assert
there that they were torn down), so a copy would double-free. Nothing copies
them today; this keeps it that way. The manager base classes cover every
backend's subclass.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- The texture and fragment test caches dropped their GLRTexture objects on
DeviceLost without queueing them for deletion, leaking them on every
Android background/resume. The deleter already skips the GL calls when
the context is gone.
- GLRenderManager::ThreadEnd cleared unsubmitted init and render steps
without freeing the data they own. Run them through the dry run instead,
which now also frees stereo matrices and shader code.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Release the file reference when loading a level fails or finds nothing,
since only a loaded level takes ownership of it.
- Free the PNG image when the size changed since the header was read.
- Delete the VFS when a pack without an ini has no hash-named textures.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Put the anisotropy level in the sampler key, so changing it applies on
Vulkan and D3D11.
- Release CLUT textures at shutdown on GLES and D3D11.
- Fix the depth readback viewport, which squeezed the image whenever the
read rectangle was smaller than the fbo.
- Test the computed depth, not the unset result, in the equal-depth clear
check.
- Don't read back a CLUT from a framebuffer without an fbo.
- Tolerate null entries when releasing post-shader objects and CLUT
textures after a failed creation.
- ImGe: Don't crash on a framebuffer without an fbo, or on GetVFB under the
software renderer.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
DeviceLost saves the cache and clears everything, and the next save wrote
back only what was drawn since, so each Android background/resume cycle
shrank the cache.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
RecordNextFrame now refuses while a recording is active and during frame
dump playback (which asserted on the next replay). The callback handoff to
the CPU thread is locked, and a failed file open no longer crashes.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- The block transfer overlap check passed the stride in pixels where bytes
are expected, so it only covered part of the rectangle.
- A selfrender/selfdepth flush in UpdateState dropped the current draw's
pending writes and reads, so later transfers didn't wait for it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- A list dropped for a bad pc or a GE error stayed RUNNING: its ID was never
freed and sceGeListSync on it never returned. Complete it like a finished
one.
- FlushImm switches to through mode and another vertex decoder, so flush the
queued draws first even when the immediate flags match.
- Clear leftover temporary GE breakpoints when setting or clearing the next
break, so a step that never got there doesn't trip later.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Flush before the queued draws would decode more than VERTEX_BUFFER_MAX
vertices. The batch was limited by index count, which doesn't bound a
sparse index range, and DecodeVerts silently stopped while DecodeInds
still emitted indices for the undecoded draws.
- Give TestBoundingBox its own scratch buffer. It used offsets in decoded_,
which can hold decoded vertices that aren't flushed yet.
- Read 32-bit indices the way the PSP does, ignoring the upper 16 bits.
IndexConverter and the fast bounding box test used all 32, so a game
setting them indexed far past the decoded vertices.
- D3D11: Flush in FinishDeferred like the other backends, since indices
are still read from PSP memory at flush time (#10095).
- Don't JIT new vertex decoders once the code space is full.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Setting data decodes and throws away the frames before the first sample,
so on hardware it costs a decoder setup plus that decode, with the caller
waiting: ~900us for mono Atrac3 and ~3.5ms for stereo Atrac3+, whatever the
buffer size (pspautotests threads/scheduling/callcosts). It was charged
100us.
Decoding a frame now costs what the same frame costs through
sceAudiocodecDecode, through the shared ME queue, instead of a flat
2300us. That's about the same for stereo Atrac3+ and less for Atrac3
(685us mono, ~1100us stereo). The first Atrac3+ frames after setup still
come out ~500us short.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>