Commit Graph
4086 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Opus 5.5 a6849661ba Shade mapping: Use the light vector as lighting sees it
Environment map S and T are (N.L + 1) / 2 with L the light's vector as
lighting uses it: from the vertex to the light for point and spot lights,
a zero vector staying zero, and the half vector for a light that does
specular. Whether lighting or the light is enabled still doesn't matter
(gpu/lighting/shademap).

The vertex shader ID now carries the type and computation of the shade
mapping lights (the ubershader reads them from u_lightControl), so both
shader caches get a new version.

Fixes the hair shine in iDOLM@STER SP (#12376).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 15:47:16 -06:00
Henrik RydgårdandClaude Opus 5.5 d7a96875fb Lighting: Use the GE's approximate pow for specular, diffuse and spot
The GE computes these powers as exp2(e * log2(x)), with log2 and exp2
each a straight line between powers of two (Mitchell's approximation),
and only uses the top 4 bits of the specular coefficient's mantissa.
Through a highlight's falloff a true pow is 10-30 steps of 255 brighter
at the exponents games use. Measured in gpu/lighting/specular.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 15:47:16 -06:00
Henrik Rydgård a7ec706a4c Merge pull request #22400 from hrydgard/pacman-spline-lighting
Spline lighting fixes
2026-09-30 15:46:20 -06:00
Henrik RydgårdandClaude Opus 5.5 aa38fafd77 DrawEngine: Don't drop large batches of points, lines and rectangles
Software transform expands each point, line and rectangle to four
vertices, and gave up on the whole draw when that didn't fit
VERTEX_BUFFER_MAX. Batching only counted input vertices, so a batch over
16384 points (or 32768 line or rectangle vertices) vanished silently,
whether it came from one PRIM or several merged ones. Count the expanded
vertices when batching, and submit a PRIM too big on its own in parts.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 12:04:53 -06:00
Henrik RydgårdandClaude Opus 5.5 7a06e25aa0 GPU: Remove leftovers from hardware tessellation
The GLES sampler uniforms and texture slots for the control points and
weights, the Vulkan storage buffer bindings, and the u_spline_counts
uniform, which becomes padding (the C++ side already was).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-30 11:00:02 -06:00
Henrik RydgårdandClaude Opus 5.5 bc581349fd GPU: Install the draw engines' invalidation callback from BeginFrame
The draw engine is created on the loader thread while the UI thread may
already be rendering and invoking the callback.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 13:16:42 -06:00
Henrik RydgårdandClaude Opus 5.5 93f57a4f38 GPU: Delete copy operations on classes that own resources
These own GPU objects, memory or refcounts in their destructors (or assert
there that they were torn down), so a copy would double-free. Nothing copies
them today; this keeps it that way. The manager base classes cover every
backend's subclass.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 39049a67fd GLES: Free textures on device lost, and unsubmitted step data at exit
- The texture and fragment test caches dropped their GLRTexture objects on
  DeviceLost without queueing them for deletion, leaking them on every
  Android background/resume. The deleter already skips the GL calls when
  the context is gone.
- GLRenderManager::ThreadEnd cleared unsubmitted init and render steps
  without freeing the data they own. Run them through the dry run instead,
  which now also frees stereo matrices and shader code.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 16ea2cbf88 GPU: Fix draw engine buffer overruns and stale vertex data
- Flush before the queued draws would decode more than VERTEX_BUFFER_MAX
  vertices. The batch was limited by index count, which doesn't bound a
  sparse index range, and DecodeVerts silently stopped while DecodeInds
  still emitted indices for the undecoded draws.
- Give TestBoundingBox its own scratch buffer. It used offsets in decoded_,
  which can hold decoded vertices that aren't flushed yet.
- Read 32-bit indices the way the PSP does, ignoring the upper 16 bits.
  IndexConverter and the fast bounding box test used all 32, so a game
  setting them indexed far past the decoded vertices.
- D3D11: Flush in FinishDeferred like the other backends, since indices
  are still read from PSP memory at flush time (#10095).
- Don't JIT new vertex decoders once the code space is full.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 10:33:32 -06:00
Henrik RydgårdandClaude Opus 5.5 3f44da709b Clamp sampling of video textures and direct-displayed video to the 480x272 frame
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:22:01 -06:00
KailashandClaude Opus 5 eb2f5dfc61 GLES: Pass known index range to glDrawRangeElements
DrawEngineGLES already knows that every index it generates is below the
decoded/transformed vertex count, but only passed the count to
glDrawElements. Drivers that need the vertex range before vertex
shading, like Mesa's Panfrost, then scan the index data on the CPU for
every draw, and their min/max caches can't help since the index push
buffer is rewritten every frame.

Only used for single-instance draws on desktop GL or GLES3, and only
when the caller provides the range, so other DrawIndexed callers are
unchanged.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-16 21:01:54 +05:30
Henrik Rydgård 256984d561 Fix some Claude-isms 2026-09-04 10:41:31 -06:00
Henrik Rydgård ac016201dc GLES: Small cleanups
Scissor the stencil readback to the region actually being read back - latent,
every caller passes a zero origin today.

Remove a DecodeVerts call that can never do anything: both branches above it
have already advanced decodeVertsCounter_ to numDrawVerts_. Worse than useless,
since in the non-skinning branch the vertices went to the push buffer, so
decoded_ doesn't hold them.
2026-09-04 10:41:31 -06:00
Henrik Rydgård 78692deca2 GLES: Set IS_3D when creating the 3D texture, not after uploading it
The out-of-memory bail-out added in the previous commit returned before the
status flag was set, leaving a GL_TEXTURE_3D object bound while ApplyTexture
told the shader generator it was a 2D texture. The entry stays cached, so it
would repeat every frame, not just the one that failed to allocate.
2026-09-04 10:41:31 -06:00
Henrik Rydgård 89ebb8acc6 GLES: Actually apply anisotropic filtering
TextureCacheGLES passed a hardcoded 0.0f instead of key.aniso, so the
Anisotropic Filtering setting did nothing at all on the OpenGL backend, even
though GPU_USE_ANISOTROPY was advertised and D3D11/Vulkan both honor it. Looks
like it was left behind by the 2017 render manager refactor.

The queue runner now clamps to the device maximum it already queried into
maxAnisotropyLevel_ (until now unused), and only touches the parameter when the
extension is actually supported - the anisotropy branch there has been dead
since every caller passed 0.0f, so this is the first time it runs.

0.0f keeps its meaning of "don't care" for the CLUT/fragment-test/thin3d
callers; the texture cache now passes 1.0f when the setting is off, so turning
it off takes effect on already-uploaded textures instead of only new ones.

TexCache: Never use anisotropic filtering for CLUT8-indexed textures

What gets sampled for those is palette indices, depalettized by the shader
afterwards - averaging indices across an anisotropic footprint produces garbage
colors. Affects all backends, not just the GL one that just started honoring
key.aniso.

TexCache: Clear key.aniso wherever filtering is forced to nearest

It was only cleared in the two places inside the AUTO_MAX_QUALITY branch, so the
TEX_FILTER_AUTO path (pixel-mapped textures, the ugly color test heuristic), the
FORCE_NEAREST setting and the replacement-texture override could all end up
requesting nearest filtering with anisotropy still on.

Doing it in the switch that applies forceFiltering covers every path, so it
can't drift apart again.

GLES: Only record the applied anisotropy, and log skipped draws

The queue runner updated tex->anisotropy even when it skipped the call because
the value was 0.0f ("don't care") - harmless while nothing ever set anisotropy,
but now it would make the tracked state disagree with GL, so a later request for
the value it thinks is set would be wrongly skipped.

Also log when a draw is skipped for a missing vertex shader. The failure is
cached per shader ID, so without it geometry silently disappears for the rest of
the session after the one-shot OSD message.
2026-09-04 10:41:15 -06:00
Henrik Rydgård 388eef9d88 GLES: Harden the shader disk cache loader and the 3D texture upload path
The cache loader indexed &vec[0] on vectors that can legitimately be empty (a
header-sized file with zero counts passes both sanity checks), and the counts
are signed ints where only the upper bound was checked - a negative count would
reach resize() as a huge size_t.

The 3D texture branch had the out-of-memory assert but not the bail-out the 2D
branch has, so an ignored assert fell straight into memset(nullptr).
2026-09-03 20:43:06 -06:00
Henrik Rydgård ad131e522f GLES: Handle shader compilation failure in the hardware transform path
ApplyVertexShader can return null - if the requested shader fails to compile it
retries with a software transform ID, and if that fails too it returns (and
caches) null. We then called UseHWTransform() on it.

ApplyFragmentShader can likewise return null, and the hardware path ignored it,
unlike the software path. Without a linked shader nothing binds a program for
this render pass, so the draw would have gone through with whatever program a
previous pass left bound.
2026-09-03 20:43:06 -06:00
Henrik Rydgård 245ed61c0f GLES: Actually apply the stencil write mask in ApplyDrawStateLate
53aa2cc596 changed the first argument from "true" to stencilState_.writeMask,
but that slot is "bool enabled" - the writeMask argument stayed hardcoded to
0xFF, so the mask still never reached GL. The clear-mode call just above gets
the slots right.

Reachable because SoftwareTransformCommon refuses the fast clear path when the
stencil write mask is partial, so exactly those clears end up here.
2026-09-03 20:43:06 -06:00
Henrik Rydgård 3eb056ad86 Move Common/GraphicsContext.h to Common/GPU/GraphicsContext.h 2026-07-26 13:58:17 +02:00
Henrik Rydgård cff09041c9 Completely rework the flow in TextureCacheCommon::ApplyTexture 2026-07-18 11:57:48 +02:00
Henrik Rydgård 1ffdc52111 Remove another "next" variable, cleanup 2026-07-18 11:57:48 +02:00
Henrik Rydgård 40bce17f53 Some cleanup 2026-07-16 20:27:44 +02:00
Henrik Rydgård 0528646ea3 Invert the relationship between the two functions 2026-07-16 19:08:18 +02:00
Henrik Rydgård e81055db16 Remove the need to call SetTexture() externally. 2026-07-16 18:32:47 +02:00
Henrik Rydgård b4d1b3d469 Reshuffle so that SetTexture and ApplyTexture are always called together (will merge them later) 2026-07-16 18:23:58 +02:00
Henrik Rydgård 7c73ed71bf Remove the ability to "set safe size" from SoftwareTransformCommon, now redundant 2026-07-16 18:05:59 +02:00
Henrik Rydgård fef50f5348 Remove obsolete software transform vertex decoder hack that forced software skinning 2026-07-16 17:52:14 +02:00
Henrik Rydgård 92ec6a992b Remove "pixelMapped" from gstate_c. 2026-07-16 17:45:22 +02:00
Henrik Rydgård 6f4f0c41b4 Split ApplyTexture into ApplyTexture and ApplySampler 2026-07-16 17:45:22 +02:00
Henrik Rydgård 2e012b8966 Unify the "BindSampler" function between the backends 2026-07-16 17:45:22 +02:00
Henrik Rydgård 66014870d2 Remove unnecessary "DirtyLastShader" mechanism. 2026-07-16 15:47:59 +02:00
Henrik Rydgård 57e6bafb62 Remove largely ineffective and nearly not functional "low memory mode". Unlikely to save us. 2026-07-16 15:28:40 +02:00
Henrik Rydgård 672b549109 Make GetFramebufferSamplingParams a loose function. 2026-07-16 15:28:33 +02:00
Henrik Rydgård d2ae8af2bf Start splitting apart Texture vs Sampler binding code 2026-07-16 15:03:23 +02:00
Henrik Rydgård ea95eb420d Clean up GetFramebufferSamplingParams 2026-07-16 15:03:20 +02:00
Henrik Rydgård c3eafeb005 Remove obsolete vertex decoder code and options 2026-07-14 17:26:00 +02:00
Henrik Rydgård 61e1ef8f7a Remove the "Software skinning" option. Now always on. 2026-07-14 17:02:57 +02:00
Henrik Rydgård 38b0b0b126 Eliminate two divisions and multiplication from all vertex shaders 2026-07-13 12:11:31 +02:00
Henrik Rydgård db712016bb Comment improvements, rename a dirty-flag 2026-07-13 11:10:31 +02:00
Henrik Rydgård a207a46fba VertexShaderGenerator: The hasColor bit is only relevant in hw transform. 2026-07-10 12:49:57 +02:00
Henrik Rydgård 9e687c3282 Make Description a member function of VShaderID / FShaderID. Assorted cleanup 2026-07-09 17:46:07 +02:00
Henrik Rydgård 2e9ee49679 Reorganize the FS shader bits too, for a better sort order in the shader viewer 2026-07-07 20:16:37 +02:00
Henrik Rydgård 93b570f8c7 Reverse the bits so we actually get the desired shader sorting order in the debug shader viewer 2026-07-07 20:16:37 +02:00
Henrik Rydgård 6508ae0aae Reorganize the vertex ShaderId bits, better sorting order 2026-07-07 20:16:32 +02:00
Henrik Rydgård ebc564a465 Centralize handling of another GPU flag 2026-06-27 13:50:03 +02:00
Henrik Rydgård d2f1193afc Remove another unused GPU_USE_* flag 2026-06-27 13:19:13 +02:00
Henrik Rydgård 623545bd24 Delete the "Hardware tessellation" feature.
Very hard to maintain and debug, not worth it.
2026-06-16 16:17:14 +02:00
Henrik Rydgård e0634f3df9 Assorted cleanup and tweaks 2026-06-13 13:34:41 +02:00
Henrik Rydgård c5931ea690 Unify more shader uniform update code, fix bug in fallback for Uint8x3ToFloat4. 2026-06-13 10:14:39 +02:00
Henrik Rydgård 40a345bff8 Fix a filtering issue, enable this for Fushigi no Dungeon 4. 2026-06-12 23:52:50 +02:00