Without SSE or NEON, triangle pixels got the secondary color in place of
the primary one plus it, so lit triangles came out black. It showed as
the "unexplained" known failures on riscv64 and loongarch64, and broke
the new gpu/lighting/shademap there. Reproduced on arm64 by building
without NEON.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Environment map S and T are (N.L + 1) / 2 with L the light's vector as
lighting uses it: from the vertex to the light for point and spot lights,
a zero vector staying zero, and the half vector for a light that does
specular. Whether lighting or the light is enabled still doesn't matter
(gpu/lighting/shademap).
The vertex shader ID now carries the type and computation of the shade
mapping lights (the ubershader reads them from u_lightControl), so both
shader caches get a new version.
Fixes the hair shine in iDOLM@STER SP (#12376).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The viewer is at infinity along view space +z, so in world space, where
lighting happens, it's the view matrix's third column rather than
(0,0,1): turning the camera moves the highlights (gpu/lighting/specular).
The shaders read it from u_view.
In a Need for Speed Carbon frame replayed on a PSP, this and the pow
bring the error of the cars from MSE 957 to 120 (Vulkan).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The GE computes these powers as exp2(e * log2(x)), with log2 and exp2
each a straight line between powers of two (Mitchell's approximation),
and only uses the top 4 bits of the specular coefficient's mantissa.
Through a highlight's falloff a true pow is 10-30 steps of 255 brighter
at the exponents games use. Measured in gpu/lighting/specular.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- Flush() on an empty queue now trims the state and CLUT rings, since its
callers flush because one is full and push right after.
- BinQueue::Full() uses >=, so an overshoot can't go unnoticed.
- IsExactSelfRender compares against the target the queued draws were
binned for, not gstate, which already has the next one during a flush.
- A depth test without depth writes marks the depth buffer as read.
- The DarkStalkers untextured sprite recomputes the binner state around it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
These own GPU objects, memory or refcounts in their destructors (or assert
there that they were torn down), so a copy would double-free. Nothing copies
them today; this keeps it that way. The manager base classes cover every
backend's subclass.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- The block transfer overlap check passed the stride in pixels where bytes
are expected, so it only covered part of the rectangle.
- A selfrender/selfdepth flush in UpdateState dropped the current draw's
pending writes and reads, so later transfers didn't wait for it.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Skipping them is only a speed-up, but it changes how states get batched and
optimized, which changes the rendered output (NBA 2K13 and Virtua Tennis
frame dumps). Keep the old behaviour until that's understood.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A code space clear frees the functions the current state points to, but
the state was kept as long as the GE registers didn't change. Track the
clear generations and recompute, also when a compile during the state
computation clears the caches.
Also skip the binner flush for compiles on builds without the software
JIT, where Compile() does nothing.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Full-screen clears measured on a PSP (pspautotests gpu/timing/blittiming):
0.49ms on a 16-bit framebuffer whatever is cleared, 0.69ms on 8888, 1.02ms on
8888 with depth. Charging them may help games that spin hard on an empty
screen, but it's off (chargeClearTime) until tried on some. The video blit
cost moves into the same function, now EstimateFillCycles.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The blit cost remembered only the last buffer a decoder wrote into, forever.
Move the texture cache's video list (with its ageing out a few flips after
the last write) into GPUCommon, so the texture cache, the blit cost and
SoftGPU all share one. That also counts both of a double-buffered player's
frames, which exposed that a clear drawn with texturing still enabled was
being charged as a blit - skip clears and draws without texture coordinates.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Movie players like the one in Star Wars: Lethal Alliance present every decoded
frame after a single vblank wait, with no clock or timestamp check, so the
frame rate depends on decode, CSC, ATRAC decode and the GE blit adding up to
more than a vblank. We charged nearly nothing for any of them, so such movies
ran at 60 fps until the ringbuffer's slack ran out.
Costs measured on a PSP with a copy of that player (pspautotests
video/mpeg/playertiming), for a 480x272 frame:
- sceVideocodecDecode: 3.4ms (sceMpegAvcDecode 5.8ms less sceMpegAvcCsc 2.4ms)
- sceMpegBaseCscAvc: 2.4ms, was a flat 4ms
- sceAudiocodecDecode, ATRAC3+ only: 2.5ms per frame
- GE: 9.7ms for a through-mode rectangle blit from a decoded video frame,
charged by area, only for textures in the buffer a decoder last wrote.
GE time also now carries across stall address updates. Before, a list sent
in stalled chunks only had its last chunk's time counted, so sceGeDrawSync
returned 39us after a blit that takes 9.7ms. This affects every game that
builds its lists incrementally, so GE-timing-sensitive games need checking.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Fix#21363https://chat.deepseek.com/share/980xe4c2ue35te8cy2
Performance impact is minimal because sprites that require depth/fog are relatively rare, and the fallback still uses the per-pixel rectangle loop (not the full triangle rasterizer)