Commit Graph
47820 Commits
Author SHA1 Message Date
Henrik RydgårdandClaude Fable 5.1 fc9dbf5ff5 FPU: saturate float-to-int in the interpreters
The C cast is undefined past the int32 range, and x86 makes it INT_MIN, so
round/trunc/ceil/floor/cvt.w.s of anything from 2^31 up gave 0x80000000 on
x86 hosts while the PSP saturates to 0x7fffffff (cpu/fpu/roundmode). Route
all of them through SaturatedFloatToInt, which also covers NaN and inf, and
drop the special cases that did.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:25:41 -06:00
Henrik RydgårdandClaude Fable 5.1 dc983bb1c5 Add cpu/vfpu/specials and overlap_vcrsp to tests_next, update overlap
specials stays in tests_next for vcmp on denormals and the NaN
canonicalization and denormal flush in vbfy/vocp/vavg/vfad/vsocp, which
overlap the USE_VFPU_DOT accuracy switch. overlap_vcrsp is vcrsp with an
overlapping destination, which the assembler refuses and the hardware
doesn't read-before-write for.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 613422c15f IR: vi2s/vi2us with the destination inside the source
The second pack read a lane the first one had just written
(vi2s.q C002, C000). Pack into temps when the outputs overlap the inputs.
Found by the corrected cpu/vfpu/overlap.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 7d69ba76f2 x64 IR JIT: fix vcmp EN/NN, NI and TR
EN/NN compared the source against an uninitialized temp, usually the
previous lane's all-ones mask (a NaN), so every lane came out as NaN. NI
used a less-than compare, which is false for a NaN. TR set the bit and
then replaced it with a garbage bit from the temp. Found by
cpu/vfpu/specials.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 b16cd9a734 VFPU: vsgn of a denormal is zero
In the interpreter, the IR interpreter, every backend's FSign and the x86
JIT's own vsgn. Recorded in cpu/vfpu/specials.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 2c801b8a31 VFPU: vsrt1-4 order NaN, denormals and signed zero like the hardware
The same ordering as vmin/vmax (sign and magnitude, denormals compare as
zero), and on a tie vsrt1/2 keep the lower lane of each pair, vsrt3/4 the
upper one. Recorded in cpu/vfpu/specials.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 ed6d13e15f VFPU: bit-exact vh2f and vf2h
vf2h truncates the mantissa (no rounding), gives a signed zero below 2^-14
and a signed inf from 65536 up, and keeps the low ten mantissa bits of a
NaN, so one with those clear becomes inf. vh2f flushes a subnormal half to
a signed zero and keeps inf/NaN mantissa bits unshifted. The x86 JIT's own
vh2f gets the subnormal flush; the other backends go through the
interpreter. Recorded in cpu/vfpu/specials.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 5564ef3ec3 Add cpu/vfpu/overlap to tests_good
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 a6e29f6972 Add cpu/vfpu/vrnd to tests_good
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 55672ebd99 Add cpu/vfpu/vbranch to tests_good and vbranch_hazard to tests_next
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik RydgårdandClaude Fable 5.1 ae81a20c74 arm64 JIT: mtvc to the VFPU condition register keeps six bits
The other control registers were masked already; the CC path wasn't, so
a value with bits 6 or 7 set made bvt 6 and bvt 7 branch. Hardware never
sets those (cpu/vfpu/vbranch), and the IR masks them.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 14:10:22 -06:00
Henrik Rydgård ffc32d02fe Merge pull request #22336 from hrydgard/misc-work
Code cleanup, add a dismissable reminder about the 32-bit version on a 64-bit OS.
2026-09-22 13:15:12 -06:00
Henrik Rydgård c14455ef11 Merge pull request #22334 from hrydgard/div-overflow-hi
Fix issues with div overflow and VFPU prefixes
2026-09-22 13:14:43 -06:00
Henrik RydgårdandClaude Fable 5.1 5e5ecc5389 x64 IR JIT: A NaN in the second operand of vmin/vmax took the fast path
The NaN check compared src1 with itself, so only a NaN in src1 reached
the slow path with the hardware's integer-compare rule. A NaN in src2 went
to MINSS/MAXSS, which just return it. Compare src1 with src2, as the arm64
backend does. Caught by cpu/vfpu/minmax.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 12:49:06 -06:00
Henrik RydgårdandClaude Fable 5.1 9f003f7773 x86 JIT: Gate the VFPU ops with odd prefix rules the way the IR does, and mask mtvc
The same set of gates the arm64 JIT got: vabs/vneg, vsat1, the vrcp
group, vzero/vone/vidt, vdiv, vscl, vfad/vavg, vrot, vuc2i, and now also
vcrs/vcrsp/vqmul, vbfy and vtfm, which the IR frontend hands to the
interpreter as well. Swizzles naming a lane past the op's size go there
too. mtvc now keeps only the bits the hardware does (six for CC, the low
20 of a prefix, 0x3F8 on top of the RNG state). vcst's SIMD path skipped
the D prefix.

Fixes cpu/vfpu/prefix_ctrl on Windows CI. Verified on an x86_64 build
under Rosetta: prefix_ctrl and the full tests_good pass, prefix_consume
is down to the same three ops that differ on every backend.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 12:49:06 -06:00
Henrik RydgårdandClaude Fable 5.1 8d94b245ad Headless: Key the no-SDL path off HEADLESS_CROSS, not the architecture
The riscv64/loongarch64 cross builds have no SDL, and Headless.cpp had a
matching "no window" path gated on those two architectures. Gate it on
the option instead (a HEADLESS_NO_SDL define), so any headless-only build
without a usable SDL gets it - an x86_64 build on an arm64 Mac, say, where
Homebrew's SDL3 is arm64 only. That's how the x86 JIT fixes here were
tested, under Rosetta.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 12:49:06 -06:00
Henrik RydgårdandClaude Fable 5.1 783892ed25 VFPU: Writing an RNG state register keeps 0x3F8 in the top bits
mtvc to RCX0-7 keeps the low 20 bits of the value and forces the top
twelve to 0x3F8 (cpu/vfpu/vrnd), which is also the form vrnds and vrndi
leave them in. We masked with 0x3FFFFFFF. The generator itself matches
hardware for every seed and state in that test.

Interpreter, IR and the arm64 JIT. The x86 and ARM32 JITs don't mask
control register writes at all and are left as they were.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 12:49:06 -06:00
Henrik Rydgård 85ed46d97e Add new translation string 2026-09-22 12:42:05 -06:00
Henrik RydgårdandClaude Fable 5.1 ba491301ae Add cpu/vfpu/prefix_unpack to tests_next
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:45:25 -06:00
Henrik RydgårdandClaude Fable 5.1 4e0319ffd9 test.py: cpu/vfpu/prefix_ctrl passes now
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:37:49 -06:00
Henrik RydgårdandClaude Fable 5.1 960514d63a arm64 JIT: Gate the VFPU ops with odd prefix rules the way the IR does
The arm64 JIT applied the plain prefix rules to every op. Several ops
don't follow them, and the interpreter and IR frontend know which:

- vabs and vneg are vmov with the abs/negate bit forced, so a negate in
  the S prefix doesn't negate twice.
- vsat1 doesn't apply the D saturation.
- vrcp, vrsq, vsqrt, vnrcp (and the rest of that group) apply S and D to
  the last lane only, from prefix position 0.
- vzero, vone and vidt are changed by a pending S prefix.
- vdiv, vscl, vfad/vavg and vrot each have their own rule.
- vuc2i and friends with an S prefix.

Hand these to the interpreter, exactly where the IR frontend does. Also
mask what mtvc writes to a control register (the low 20 bits of a prefix,
12 of the D prefix), as the IR already did.

Goes from 19 ops differing from hardware in cpu/vfpu/prefix_consume to
the 3 that differ on every core, and clears cpu/vfpu/prefix_ctrl.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:37:48 -06:00
Henrik RydgårdandClaude Fable 5.1 591731cb8e IR: A swizzle naming a lane past the op's size isn't "within size"
IsPrefixWithinSize only looked at prefix positions past the size, not at
in-size positions naming a lane past it, so vadd.p with [z, w, ...] was
compiled with whatever the prefix code substitutes. Hardware writes those
result lanes as zero (cpu/vfpu/prefix_ctrl), which the interpreter does,
so send them there.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:37:48 -06:00
Henrik RydgårdandClaude Fable 5.1 0af831b5d8 VFPU: An invalid swizzle on the one-lane ops zeroes the result lane
vrcp, vrsq, vsin, vcos, vexp2, vlog2, vsqrt, vasin, vnrcp, vnsin, vrexp2
and vdiv apply the S prefix to their last lane only, taking it from prefix
position 0. When that names a lane other than x, we fed the op an
infinity, which happens to give the right answer for vrcp and vexp2 and
the wrong one for the rest. Hardware writes the lane as zero regardless
of the op (cpu/vfpu/prefix_sat, where vlog2 and vsqrt give zero too),
which is also what RetainInvalidSwizzleST already does for the other ops.

Covers the IR path as well, which hands these cases to the interpreter.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:32:42 -06:00
Henrik RydgårdandClaude Fable 5.1 de884aba02 Add cpu/vfpu/prefix_sat to tests_next
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:30:38 -06:00
Henrik RydgårdandClaude Fable 5.1 acb81a049e Add the VFPU prefix tests to tests_next
prefix_branch passes on every core. prefix_ctrl fails on the IR path
(out-of-size swizzle lanes) and the arm64 JIT (that, plus mtvc not
masking). prefix_consume fails everywhere: the interpreter and IR in lane
w of nine ops where the T prefix holds a constant 0, the arm64 JIT on
nineteen ops.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:21:16 -06:00
Henrik RydgårdandClaude Fable 5.1 a5c96e054e Add cpu/vfpu/minmax_tie to tests_next
vmin/vmax return the second operand on a -0/+0 tie. The classic
interpreter does that; the IR interpreter and the JITs return the first.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:09:47 -06:00
Henrik RydgårdandClaude Fable 5.1 6e44931f5b test.py: cpu/cpu_alu/cpu_div passes now
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:02:35 -06:00
Henrik RydgårdandClaude Fable 5.1 9f1f1f23c9 arm64 IR JIT: divu by zero compared the numerator signed
The rule is LO = numerator <= 0xFFFF ? 0xFFFF : -1, unsigned. The IR
backend used CC_LE, so any numerator with the top bit set counted as
small and got 0xFFFF. The other backends compare unsigned. Also caught by
cpu/cpu_alu/cpu_div.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:02:20 -06:00
Henrik RydgårdandClaude Fable 5.1 e0f09c55d5 MIPS: INT_MIN / -1 leaves HI at zero, like the hardware
Every path special-cased the one signed division overflow and set the
remainder to -1. cpu/cpu_alu/cpu_div, recorded on a PSP, says it's 0.
The classic arm64 JIT was the only one that got it right, by not
special-casing it at all.

arm64, RISC-V and LoongArch all produce INT_MIN with remainder 0 natively,
so their fixup blocks go away. x86 has to keep the check (IDIV traps) but
sets HI to 0 now, and so do both interpreters.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 11:02:20 -06:00
Henrik RydgårdandClaude Fable 5.1 5b507e5bb3 Add pspautotests for the div-by-zero, rounding, scaled convert and call-out JIT paths
Covers the cases fixed on the riscv-loongarch-fixes branch, none of which
the suite reached before. Five go in tests_good, and two in tests_next:

- cpu/cpu_alu/cpu_div: hardware leaves HI = 0 for INT_MIN / -1. The x86
  JIT, both interpreters and every IR backend set it to -1 on purpose;
  the classic arm64 JIT passes by not special-casing it at all.
- cpu/vfpu/minmax_zero: signed zero and denormals in vmin/vmax.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 10:59:17 -06:00
Henrik Rydgård 2fe24d1182 Add warning if you run the 32-bit version on 64-bit window. 2026-09-22 10:04:31 -06:00
Henrik Rydgård 8c6ec65b2d Refactor the version reminder bar to make it reusable 2026-09-22 09:48:18 -06:00
Henrik Rydgård b366e8a6d0 Move persistent notifications (the only one right now is upgrade reminder) to the top 2026-09-22 09:34:50 -06:00
Henrik RydgårdandClaude Opus 5 28b5d0ee5c RISC-V, LoongArch64: fix caller-saved FPRs, div by zero, and more
RiscVRegCache::FlushBeforeCall claimed the caller-saved X and F registers
"match between X0 and F0". They don't: X0-X4 are zero/ra/sp/gp/tp and are
never allocated, but F0-F4 (ft0-ft4) are caller-saved and are in the FPR
allocation order. A block that spilled into them and then called out (for
vsin or vcos, say) had the callee destroy registers the cache still
thought were live and dirty. LoongArch and arm64 get this right.

The LoongArch signed div-by-zero sequence applied its ADDI_D(-1) after
the jump target, so it ran on both paths: a negative numerator ended up
with a quotient of 0 instead of 1, and a non-negative one borrowed into
the remainder in the high half, giving hi = num - 1. Insert the -1 into
the low word rather than subtracting, and branch around the other case.

FCvtScaledWS loaded 0x7FFFFFFF into SCRATCHF1 and then immediately
overwrote it with the multiplier, while the FSELs picked SCRATCHF2 - a
register the function never writes. A NAN input produced whatever was
left there.

Also:
- RISC-V had no SyscallUnresolved case, falling through to INVALIDOP,
  which is a live assert in release builds. Added it, matching arm64.
- applyRoundingMode_ fed fcr31 straight to FSRM, which only takes three
  bits - the flags at 2-6 could turn a mode of 0 into 4. Mask with 3
  first, as MIPS.cpp and LoongArch64Asm.cpp do.
- OverwriteExit lacked arm64's assert that the exit fits the 8-byte hole,
  which matters more here since QuickJ is 4, 8 or 12 bytes.
- Both backends allocated the FMin/FMax temp GPR inside the conditional
  NAN path, where a spill store would only run on one side while the
  regcache assumed it always ran. Hoisted above the branch.

Both backends build now, and pspautotests and the unit tests pass under
qemu on each - but nothing in the suite reaches the paths above, so the
two worth checking were checked by running the emitted sequences
directly. The div-by-zero one is wrong for all eight numerators tried and
right for all eight after. The rounding one turns a guest mode of 0 into
RMM whenever fcr31 has its inexact flag set, so 0.5 converts to 1 rather
than 0. The rest are still by inspection: they need register pressure, an
unresolved import, or a debug build to reach.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-22 08:39:25 -06:00
Henrik Rydgård 2a1700a7ee Merge pull request #22332 from sum2012/remove-hack
Remove hack
2026-09-22 07:13:27 -06:00
sum2012 f5b6b45b5b Remove IgnoreEnqueue hack 2026-09-22 20:19:57 +08:00
sum2012 14fd9d1752 Remove Fontltn12Hack 2026-09-22 20:10:33 +08:00
Henrik Rydgård ab6ccbd08c Merge pull request #22329 from hrydgard/vertex-decoder-arm64-morph
Vertex decoder: JIT morph on arm64
2026-09-21 17:56:15 -06:00
Henrik Rydgård 7a9a354e50 Add a MIPSState context pointer to the JitAt call 2026-09-21 17:34:10 -06:00
Henrik RydgårdandClaude Opus 5 f28896e1ed Vertex decoder: JIT morph on arm64
The arm64 vertex JIT had no morph support (the table entries were commented
out), so every morphed format fell back to the step functions. Add the same
morph steps x86 has: texcoords (plain and prescaled), normals, positions and
the four color formats.

Each step sums its frames the way its step function does as compiled: fused
where that's plain C++, which the compiler turns into fused multiply-adds,
and separate multiply and add where it's CrossSIMD. The unit test confirms
they match bit for bit. Morph formats decode about 2.5-3x faster than with
the steps.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 16:51:00 -06:00
Henrik Rydgård d21467beb0 De-claude some commments. 2026-09-21 16:50:22 -06:00
Henrik Rydgård 7b95ff2808 Merge pull request #22325 from hrydgard/vertex-decoder-jit-match
Vertex decoder: New test, make the JITs match the C++ decoder closely
2026-09-21 16:28:27 -06:00
Henrik Rydgård be62c09fe5 Merge pull request #22327 from hrydgard/sce-ge-queue
Claude code review: Rework sceGe
2026-09-21 16:28:15 -06:00
Henrik Rydgård f40ab22fd6 Merge pull request #22328 from hrydgard/lle-firmware-module-unload
LLE firmware module unload
2026-09-21 16:23:30 -06:00
Henrik RydgårdandClaude Opus 5 23278cb2fd Headless: turn the draw frame over each emulated frame, not once per run
A long headless run on Vulkan dies in VulkanPushPool::CreateBlock. Watching the
allocator, it makes a fresh 8MB block roughly twice a second and garbage
collects none of them - about 13MB a second of device memory, which runs out
after a minute or two.

Nothing is leaking as such. The push buffers are recycled by BeginFrame, which
walks the blocks belonging to the current frame index and marks them unused.
Headless called draw->BeginFrame() once before the run loop and draw->EndFrame()
once after, so that recycling pass ran exactly once for the whole run and every
allocation after the first had to take a new block.

This is the same mistake one level up from the host frame, which already turns
over per emulated frame for the same reason - the comment there says a single
host frame spanning the run meant the texture cache and framebuffer manager
never decayed anything. The draw context needs the same treatment, nested the
way the app nests them: draw frame outside, host frame inside.

Verified on a two-minute Tekken 6 run: new blocks created goes from around 200
to zero, and the run ends on its timeout instead of asserting. Framedump
rendering tests are unchanged - the same 23 of 30 fail before and after, which
is a separate pre-existing matter on this platform.

Also taught frametests.py to look for an ARM64 build, gated on the machine's own
architecture the way test.py already does. It was picking a stale x64 Debug
binary, which is exactly the trap that makes a rendering comparison meaningless.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 15:47:40 -06:00
Henrik RydgårdandClaude Opus 5 88b80011e4 Run pspautotests under qemu for loongarch64 and riscv64
Only the IR JIT: it's the sole native backend these two have and the one
thing here that isn't shared code, and the x86-64 and arm64 runners
already cover all four backends. Every run costs emulated wall clock, so
the timeout goes up to match.

test.py grows a per-architecture known-failure list, selected with
--known-failures=<arch>, so this can guard against new breakage while the
four outstanding ones stay outstanding. Each entry carries its reason.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 15:36:49 -06:00
Henrik RydgårdandClaude Opus 5 46632812cd LoongArch64: keep LSX detected, and map lanes the way the compilers expect
Two bugs stacked on each other, and between them every vector path in
this backend was either dead or miscompiled.

The register cache declared mapFPUSIMD unconditionally, but every
compiler here picks its path from cpu_info.LOONGARCH_LSX at runtime.
Without LSX the scalar fallbacks ran against a SIMD mapping and reached
lanes with F(reg + n), which addresses nothing there - ApplyMapping only
allocates per-lane registers when mapFPUSIMD is false. Vec4Unpack8To32
and Vec4DuplicateUpperBitsAndShift1 came out with only their first lane,
Vec2Unpack16To32 put both halves in one register, and the pack ops read
back whatever was next door. riscv64 sets the flag false and has never
had the problem. Tie the two together and the fallbacks are correct as
they stand.

That path was reachable because of the second one: the USE_CPU_FEATURES
block assigned over the hwcaps, and GetLoongArchInfo reads /proc/cpuinfo
and nothing else. Under qemu-user that file belongs to the host, so it
found no features and turned LSX off - leaving the LSX paths dead and the
untested scalar ones running, which is not what any Loongson 3A5000 or
later does. Let it only ever add to what the hwcaps found.

cpu/vfpu/convert, gum and matrix pass now, both with LSX and with it
masked off, putting loongarch64 at 338/342 - the same four as riscv64.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 15:36:49 -06:00
Henrik RydgårdandClaude Opus 5 89c8109823 Honor the guest FPU rounding mode outside x86 and arm64
ApplyHostRoundingMode and RestoreHostRoundingMode only had SSE2 and ARM64
branches, so everywhere else the guest's rounding mode was silently
ignored and ceil and floor rounded to nearest. Fall back to fesetround,
which covers riscv64 and loongarch64 and whatever comes next.

That clears both rounding mismatches in cpu/fpu/fpu on those two. What's
left there is flush-to-zero, which neither ISA has any control for, so
the test still fails - a denormal result survives where the PSP would
have flushed it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 15:36:49 -06:00
Henrik RydgårdandClaude Opus 5 333d84935e Add --force-hle, and stop reporting a normal sceAudiocodec re-init
--disable-hle had no counterpart, which made "is this our fault or the game's?"
awkward to answer: the only ways to put our HLE back were a per-game config or
hiding the firmware, and neither works from a script - the setting is per-game
and the firmware gets found anyway. --force-hle takes the same bitmask and runs
our HLE for those libraries even where the real module is now the default, so
the same repro can be run both ways and the logs diffed.

Used it on the warnings left over in Tekken 6 under the real mpeg.prx. Three of
them appear identically with --force-hle=16, so they are the game's own and
match what the hardware answers: sceKernelChangeThreadPriority(-1) eight times
in a row (pspautotests/threads/threads/change says hardware returns
UNKNOWN_THID for -1 too), sceAtracAddStreamData on a released id with a zero
byte count, and a sceKernelDeleteMutex on a garbage id.

The fourth only happens under the real module and is ours. mpeg.prx sizes its
allocation through a scratch context in its own bss and calls
sceAudiocodecReleaseEDRAM on that one, while decoding through a different
context that never gets released - so the next movie's sceAudiocodecInit finds
a live decoder and replaces it. That is once per video on every game running
the real module, and it was a WARN_LOG_REPORT, so it would have reported from
everyone's machine. It is bounded - removeDecoder deletes the old one and Init
makes exactly one more - so it is an INFO_LOG now.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 15:26:51 -06:00
Henrik RydgårdandClaude Opus 5 626a442c34 Stop shouting about two things the hardware does too
Both of these fire once per video in Tekken 6, and neither is a fault.

sceAudiocodecReleaseEDRAM warned "failed to remove decoder" whenever there was
no decoder to remove. There usually isn't: mpeg.prx calls CheckNeedMem and
GetEDRAM to size the allocation, and only creates a decoder if the stream turns
out to need one, so releasing without ever having made one is the normal path.
Demoted to debug.

While there, its signature was one argument too long. audiocodec_260.prx's own
sceAudiocodecReleaseEDRAM (080007c0) reads only a0 and sets up a1 through a3
itself, so the "id" we took was whatever the caller happened to leave in the
register - which is how the log came to show an sceMpeg error code as the
second parameter of an audio call.

sceUtilityLoadModule logged MODULE_ALREADY_LOADED at error level. It is a
normal answer that games rely on: Tekken 6 asks for av_avcodec three times and
never unloads it, ignoring the result each time. Only that one code is demoted;
everything else from a module load is still an error.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 15:01:03 -06:00