Commit Graph
15548 Commits
Author SHA1 Message Date
Henrik Rydgård 2fe24d1182 Add warning if you run the 32-bit version on 64-bit window. 2026-09-22 10:04:31 -06:00
Henrik Rydgård 7a9a354e50 Add a MIPSState context pointer to the JitAt call 2026-09-21 17:34:10 -06:00
Henrik Rydgård d21467beb0 De-claude some commments. 2026-09-21 16:50:22 -06:00
Henrik Rydgård 7b95ff2808 Merge pull request #22325 from hrydgard/vertex-decoder-jit-match
Vertex decoder: New test, make the JITs match the C++ decoder closely
2026-09-21 16:28:27 -06:00
Henrik Rydgård be62c09fe5 Merge pull request #22327 from hrydgard/sce-ge-queue
Claude code review: Rework sceGe
2026-09-21 16:28:15 -06:00
Henrik RydgårdandClaude Opus 5 46632812cd LoongArch64: keep LSX detected, and map lanes the way the compilers expect
Two bugs stacked on each other, and between them every vector path in
this backend was either dead or miscompiled.

The register cache declared mapFPUSIMD unconditionally, but every
compiler here picks its path from cpu_info.LOONGARCH_LSX at runtime.
Without LSX the scalar fallbacks ran against a SIMD mapping and reached
lanes with F(reg + n), which addresses nothing there - ApplyMapping only
allocates per-lane registers when mapFPUSIMD is false. Vec4Unpack8To32
and Vec4DuplicateUpperBitsAndShift1 came out with only their first lane,
Vec2Unpack16To32 put both halves in one register, and the pack ops read
back whatever was next door. riscv64 sets the flag false and has never
had the problem. Tie the two together and the fallbacks are correct as
they stand.

That path was reachable because of the second one: the USE_CPU_FEATURES
block assigned over the hwcaps, and GetLoongArchInfo reads /proc/cpuinfo
and nothing else. Under qemu-user that file belongs to the host, so it
found no features and turned LSX off - leaving the LSX paths dead and the
untested scalar ones running, which is not what any Loongson 3A5000 or
later does. Let it only ever add to what the hwcaps found.

cpu/vfpu/convert, gum and matrix pass now, both with LSX and with it
masked off, putting loongarch64 at 338/342 - the same four as riscv64.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 15:36:49 -06:00
Henrik RydgårdandClaude Opus 5 89c8109823 Honor the guest FPU rounding mode outside x86 and arm64
ApplyHostRoundingMode and RestoreHostRoundingMode only had SSE2 and ARM64
branches, so everywhere else the guest's rounding mode was silently
ignored and ceil and floor rounded to nearest. Fall back to fesetround,
which covers riscv64 and loongarch64 and whatever comes next.

That clears both rounding mismatches in cpu/fpu/fpu on those two. What's
left there is flush-to-zero, which neither ISA has any control for, so
the test still fails - a denormal result survives where the PSP would
have flushed it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 15:36:49 -06:00
Henrik RydgårdandClaude Opus 5 333d84935e Add --force-hle, and stop reporting a normal sceAudiocodec re-init
--disable-hle had no counterpart, which made "is this our fault or the game's?"
awkward to answer: the only ways to put our HLE back were a per-game config or
hiding the firmware, and neither works from a script - the setting is per-game
and the firmware gets found anyway. --force-hle takes the same bitmask and runs
our HLE for those libraries even where the real module is now the default, so
the same repro can be run both ways and the logs diffed.

Used it on the warnings left over in Tekken 6 under the real mpeg.prx. Three of
them appear identically with --force-hle=16, so they are the game's own and
match what the hardware answers: sceKernelChangeThreadPriority(-1) eight times
in a row (pspautotests/threads/threads/change says hardware returns
UNKNOWN_THID for -1 too), sceAtracAddStreamData on a released id with a zero
byte count, and a sceKernelDeleteMutex on a garbage id.

The fourth only happens under the real module and is ours. mpeg.prx sizes its
allocation through a scratch context in its own bss and calls
sceAudiocodecReleaseEDRAM on that one, while decoding through a different
context that never gets released - so the next movie's sceAudiocodecInit finds
a live decoder and replaces it. That is once per video on every game running
the real module, and it was a WARN_LOG_REPORT, so it would have reported from
everyone's machine. It is bounded - removeDecoder deletes the old one and Init
makes exactly one more - so it is an INFO_LOG now.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 15:26:51 -06:00
Henrik RydgårdandClaude Opus 5 626a442c34 Stop shouting about two things the hardware does too
Both of these fire once per video in Tekken 6, and neither is a fault.

sceAudiocodecReleaseEDRAM warned "failed to remove decoder" whenever there was
no decoder to remove. There usually isn't: mpeg.prx calls CheckNeedMem and
GetEDRAM to size the allocation, and only creates a decoder if the stream turns
out to need one, so releasing without ever having made one is the normal path.
Demoted to debug.

While there, its signature was one argument too long. audiocodec_260.prx's own
sceAudiocodecReleaseEDRAM (080007c0) reads only a0 and sets up a1 through a3
itself, so the "id" we took was whatever the caller happened to leave in the
register - which is how the log came to show an sceMpeg error code as the
second parameter of an audio call.

sceUtilityLoadModule logged MODULE_ALREADY_LOADED at error level. It is a
normal answer that games rely on: Tekken 6 asks for av_avcodec three times and
never unloads it, ignoring the result each time. Only that one code is demoted;
everything else from a module load is still an error.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 15:01:03 -06:00
Henrik RydgårdandClaude Opus 5 1ad17e10d1 Unload the firmware modules we swap in, when the game unloads the library
Tekken 6 never reaches gameplay with the real mpeg.prx: it plays its intro
movie, returns to the title screen, starts loading a demo match and loads
forever. With the sceMpeg HLE it plays fine.

The game unloads its video libraries before the level load and expects the
memory back. It gets most of it - scePsmf and scePsmfPlayer do go away - but
mpeg.prx stays resident, 33KB of it, sitting in the middle of the region the
loader then asks for:

  08c64000 - 09e24000  18.6MB  taken  UserSbrk
  09e24000 - 09ed4000   720KB  free
  09ed4000 - 09edc300  33.5KB  taken  ELF/sceMpeg_library
  09edc300 - 09f44000   425KB  free
  09f44000 - 09f4c000    32KB  taken  UtilityModule/302_av_atrac3plus

0x09f44000 - 0x09e24000 is 0x120000, which is exactly the allocation that
fails. Without mpeg.prx in the way that span is one free block and the level
loads.

sceUtility notifies the per-library hooks with state 1 when a utility module is
loaded and -1 when it is unloaded. The hooks that swap in a firmware module
only ever handled the load, so nothing ever took them back out. That affects
sceMpeg, sceMp3, sceMp4 and sceAtrac alike; Tekken is just the game whose
memory budget is tight enough to notice.

The unload has to take out what we put in and nothing else, so the loaded ids
are remembered rather than looked up by name: a game like Death Jr ships its
own mpeg.prx and loads it itself, and freeing that would be freeing the game's
memory. Verified that Death Jr still decodes its 1033 frames with our loader
never touching its module.

Savestates from before this have no record of what was swapped in, so they keep
the old behaviour of leaving the modules loaded rather than risk freeing
something the game owns.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 14:37:10 -06:00
Henrik RydgårdandClaude Fable 5.1 cc41a23256 GPU: empty the display list queue on sceKernelLoadExec, fixes Crazy Taxi
Reinitialize() wiped the 64 display lists but kept the queue of their ids.
Whatever the old executable still had queued came back as lists with no state
and a pc of 0, behind the first list of the new executable, where they
blocked everything. Crazy Taxi: Fare Wars is a launcher for its two games,
and stopped at a black screen that way. Fixes #19894.

This removes the workaround for it, which dropped such a list but returned
before currentList was cleared, and only worked as long as something else
happened to clear it later. A list with a bad pc is now dropped like one that
ran into an error, instead of sitting at the head of the queue for good.

Also narrows what sceGeBreak(1) throws away to interrupts that have actually
been raised, which is what gpu/ge/intrsuspend shows. The ones we haven't
raised yet are only late because we execute lists ahead of time: a game that
breaks right after its last list, and then waits for what the finish callback
signals, got that callback long ago on hardware.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-21 13:10:02 -06:00
Henrik RydgårdandClaude Fable 5.1 91c9c4d14e sceGe: sceGeBreak(1) takes pending interrupts with it
Resetting the GE also gets rid of an interrupt that was raised but not taken
yet, so a list that reached its FINISH just before never gets its finish
callback. We delivered one anyway, for a list that no longer existed. If the
break comes from inside a GE callback, the interrupt being handled is kept,
since its handler still has to return.

Found by gpu/ge/intrsuspend, which also confirms from a thread, with
interrupts suspended, that nothing moves along the queue until the FINISH
interrupt has been taken.

Savestates: bump GPUCommon to 7. We didn't use to mark a PAUSE signal as
delivered, which sceGeContinue now goes by, so a state saved with a list
paused that way would load into a game that could never continue it. Fixed
up on load.

gpu/signals/handlercalls goes in as known failing: with an old SDK version, a
stall address set from inside a SUSPEND callback doesn't reach the GE, which
we can't express with just the one stall address per list. See docs/sceGe.md.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-21 12:48:03 -06:00
Henrik RydgårdandClaude Fable 5.1 8a23e633a1 sceGe: keep a finished list on the queue until its interrupt is done
On hardware the GE stops at every SIGNAL and FINISH, and it's the interrupt
that gets it going again: on the same list after a signal, on the next one
after a FINISH - once the finish callback has run, with the finished list
still at the head of the queue. We ran the next list right away and dropped
the finished one at once, so a finish callback saw an empty queue. A list
enqueued from there was started instead of queued, and then couldn't be
dequeued, which hung the new gpu/ge/queue2 test.

ProcessDLQueue() now runs nothing while the head of the queue has an
interrupt pending, and InterruptEnd() is what takes a finished list off the
queue. This also keeps a stall update from restarting a list that's stopped
at a signal before the handler has run. drawCompleteTicks is still set when
the last list reaches its FINISH, so a sceGeDrawSync in between doesn't wait.

Other things gpu/ge/queue2 and gpu/ge/breakwait showed, all from a real PSP:

- sceGeListEnQueue compares against the address a list was enqueued with
  (or stopped at by sceGeBreak), mirrors included, not against its current pc.
  We had that the wrong way around.
- The stack-in-use check only applies to lists that have started executing.
  This is probably what IgnoreEnqueue was added for (Metal Gear Acid 2,
  #10906). The flag stays until someone has checked the game without it.
- A PAUSE signal makes the list PAUSED at once, before the FINISH delivers it.
  In between, sceGeContinue and sceGeBreak say BUSY, and updating the stall
  address does nothing, so a list that stalls there is stuck.
- A completed list can't be dequeued, with or without a context.
- sceGeDrawSync(1) looked at currentList instead of the list it had found.
- sceGeBreak(1) doesn't wake anyone, and a late interrupt for a list it reset
  no longer marks that list completed. Threads in sceGeDrawSync are woken
  before the ones waiting for the last list.

Also fixes currentList being lost when loading a state where it's list 0,
and makes ge_pending_cb a plain std::list - nothing else touches it, and the
GPU thread it was shared with is long gone. Same savestate format.

See docs/sceGe.md.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-21 12:21:19 -06:00
Henrik RydgårdandClaude Opus 5 315d128141 VFPU: make vmfvc see a pending prefix, and actually write its result
Comp_Vmfvc read vfpuCtrl[] straight from the context in all four
backends, while the mfvc path in Comp_Mftv flushes first, with the
comment "In case we have a saved prefix" - so "vpfxs X" followed by
"vmfvc sN, $128" returned the stale value in memory. The IR frontend
flushes only for the three prefix registers, which is the tighter form,
so vmfvc does the same there.

On ARM and ARM64 the fix alone wouldn't have been observable: those two
map the destination with no flags, leaving isDirty false, so the flush
dropped the loaded value without storing it. x86 already passes
MAP_DIRTY | MAP_NOINIT. That part is a fix of its own, but the two are
inseparable in this function - a vmfvc that reads the right value and
then throws it away is no better.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 11:32:24 -06:00
Henrik RydgårdandClaude Opus 5 6c7fc8df37 JitBlockCache: fix the invalidation margin, and unlink destroyed blocks
InvalidateICache widened its search by MAX_BLOCK_INSTRUCTIONS, but
block_map_ is keyed by byte addresses - a block can span four times that
many bytes. A block longer than 0x4000 bytes starting just below the end
of the invalidated range could therefore be missed and left behind after
the code under it changed. Multiply by 4.

DestroyBlock's "invalid original address" safety check returned after
setting b->invalid but before UnlinkBlock() and the checkedEntry
poisoning, so other blocks kept jumping directly into a block the cache
had already written off. The check is there to guard the memory access
that restores the original opcode, so only guard that.

Also: codeSize is a byte count assigned from a full pointer difference,
so a block over 64KB truncated it (affecting GetAddressFromBlockPtr and
the disassembly view); widen it to u32. The RAMTOP range check missed a
block sitting exactly on the user-memory midpoint, since the RAMBOTTOM
check next to it is strict. And GetBlockDebugInfo disassembled one
instruction past the end of the block.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 11:32:24 -06:00
Henrik Rydgård 12e7d222a7 Merge pull request #22324 from hrydgard/minor-changes-rollup
Decide vertex decoder based on JIT ability, not CPU core. Plus misc small changes.
2026-09-21 11:09:03 -06:00
Henrik Rydgård 0a5fa27957 Disable alpha for save/load icons in sceUtility UI 2026-09-21 10:41:35 -06:00
Henrik RydgårdandClaude Opus 5 0b5a8f537e Decide on the vertex decoder JIT with CoreParameter, not the CPU core
The vertex decoder JIT was enabled only when g_Config.iCpuCore was one of the
JIT cores, which tied an unrelated GPU path to the CPU setting. Headless leaned
on that by forcing iCpuCore to the interpreter after ApplyToConfig(), to keep
the decoder JIT off in the tests.

Add CoreParameter::bUseVertexDecoderJit instead. Standalone and libretro set it
wherever the host can JIT, so IR interpreter users on such hosts now get the
decoder JIT too. Headless sets it false, which keeps the test results as they
were and lets --cpu go through ApplyToConfig() like every other option.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 10:31:58 -06:00
Henrik Rydgård 6964cb0489 Merge pull request #22322 from hrydgard/pace-video-decode
Add PaceVideocodecDecode compat.ini option
2026-09-21 09:41:18 -06:00
Henrik Rydgård 2a52072132 Merge pull request #22320 from 4RH1T3CT0R7/fix/savedata-date-from-sfo
Savedata: use PARAM.SFO's time as the save date in the dialog
2026-09-21 09:07:07 -06:00
Henrik RydgårdandClaude Opus 5 282647803c Add PaceVideocodecDecode compat.ini option
There are a few games that seem to rely on performance problems or
timing quirks to hold the correct pace when playing sceMpeg - have not
been able to nail down any other timing mechanism.

This adds a timing enforcement function.

Coded by Claude.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-21 09:00:48 -06:00
Henrik Rydgård 4d259a6f95 De-claude the comments. 2026-09-20 16:39:35 -06:00
Henrik RydgårdandClaude Opus 5 0b2c6812e1 Headless: refuse a run that can't honour an explicit --disable-hle
When the firmware module a --disable-hle bit asks for is neither installed nor
on the disc, that library silently runs our HLE instead. For the emulator that
is the right thing; for a test tool it means the run measures something other
than what was asked for and says so only as one INFO line, which is easy to
grep past and easy to never see. It cost a round of wrong results here, where
the memory stick headless defaults to (beside the exe, not the app's) had no
firmware, so a comparison against the real mpeg.prx was quietly a comparison
against the HLE it was supposed to be measured against.

g_unavailableDisableFlags already records exactly which flags fell back, so it
just needed an accessor. Headless now names each one, prints the flash0:/kd and
memory stick it looked in, and fails the run.

Only an explicit --disable-hle binds. sceMpeg and sceMp4 are LLE by default, and
falling back is the correct and expected behaviour wherever no firmware is
installed - making that fatal would fail every run on such a machine.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-20 16:39:35 -06:00
Henrik RydgårdandClaude Opus 5 1a08f3c88c Give the Media Engine an address space instead of a table of packets
sceMpegBasePESpacketCopy used to gather each DMA into a buffer filed under the
address of its first block, and sceVideocodecDecode looked one up by the address
it was handed. That works only while an access unit arrives in a single call. It
doesn't: mpeg.prx routinely DMAs one in several pieces at consecutive addresses,
and the table then kept whichever piece happened to start where the decode later
asked, and dropped the rest.

So the DMA now writes each block at the address it names. The Media Engine's own
memory moves to its real addresses - low ones, which mpeg.prx bounds-checks
against 0x3FFFFF - and pieces written at consecutive addresses end up
consecutive, which is all this ever needed. Our own allocations move to the top
half, away from the addresses mpeg.prx picks for itself.

That memory is the ME's, not the Allegrex's, and the two are separate address
spaces: on hardware only the ME reaches it, so here it is a buffer of our own
that nothing emulated can address. It gets its own MEIsValidRange and
MEGetPointerRange rather than borrowing Memory::, which answers for the
Allegrex's memory and has nothing to say about this one.

An address only means something together with the space it came from, and
mpeg.prx hands over bare integers. The three places that have to work it out
from the value now ask the ME first, where they used to ask Memory:: first and
so read every address as the CPU's. It separates the addresses these callers
actually pass (ME memory well down in the first megabyte, PSP RAM at 0x08000000
and up), but that's those callers rather than a rule.

Savestates from before this carry the old packet table; it's read and dropped,
since the Media Engine's memory is saved and that's where the payloads live now.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-20 16:39:35 -06:00
Henrik RydgårdandClaude Opus 5 8a343bec57 AvcDecoder: let ffmpeg say when a frame came out damaged
Nothing on this path ever called InitFFmpeg, so ffmpeg's log callback was never
installed and everything it had to say about the bitstream went nowhere. Our own
return codes only report that a frame arrived, not what it was built from, so a
video could fall apart on screen while the log showed two warnings in two
thousand frames.

Also report decode_error_flags and AV_FRAME_FLAG_CORRUPT per frame, which covers
the other half: a frame that decodes "successfully" out of concealment. Coalesced
into one line per run of damage, since a missing reference makes every frame
damaged until the next keyframe.

Not setting AV_EF_EXPLODE, which would make the decoder fail where it currently
conceals - that is a behaviour change, and this is meant to be a reporting one.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-20 16:39:35 -06:00
Henrik Rydgård daed3ec408 sceReg logging is not that interesting anymore, demote to Debug from Info level. 2026-09-19 13:02:44 -06:00
Henrik Rydgård 11c6111027 Only show the relay mode notification if networking is enabled 2026-09-19 13:02:44 -06:00
Henrik Rydgård af929f8e45 Merge pull request #22312 from hrydgard/ir-fpcond-and-vec-fixes
Claude code review: IR: fix FpCondFromReg operand slot, and some smaller IR bugs
2026-09-19 11:50:45 -06:00
Henrik RydgårdandClaude Opus 5 029d17edfc IR: fix FpCondFromReg operand slot, and some smaller IR bugs
FpCondFromReg's meta is "_G", so the register is in src1 - which is where
every native backend, PropagateConstants and ReorderLoadStore read it.
But IRInterpret read it from dest, and Comp_VecDo3 wrote it to dest to
match. The other emitter passes (0, MIPS_REG_ZERO), so both fields are
zero there and the disagreement stayed hidden.

The result was that vsge/vslt, which save fpcond to IRTEMP_0 and restore
it afterwards, restored r0 (always zero) on the native IR JITs instead of
the saved value, losing any c.cond.s result live across them. Fixed both
the emitter and the interpreter to use src1.

Vec4Pack31To8's SSE2 path computed (v >> 24) << 1, which is (v >> 23)
with bit 23 forced to zero - the scalar and NEON paths both do
(v >> 23) & 0xFF. Shift left first instead. The pspautotests inputs all
happen to have bit 23 clear after the clamp, so this wasn't caught.

Also:
- Evaluate() didn't mask constant-folded shift amounts to 5 bits, though
  the neighbouring one-immediate path does.
- ApplyMemoryValidation only invalidated its address-check cache for ops
  with a 'G' destination, while Interpret and CallReplacement can write
  any GPR - both are barriers, so drop the whole cache when we see one.
- ReduceVec4Flush indexed isVec4 with (src2 & 3) instead of (src2 & ~3)
  for Vec4Scale; a missed optimization rather than a miscompile.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-19 10:31:33 -06:00
Henrik RydgårdandClaude Opus 5 b03bbdbca3 MIPSTracer: don't index trace_info with stale block indices
flush_to_file() ends with clear(), but tracing stays on and the blocks
already compiled keep the index baked into their LogIRBlock instruction.
A second flush then indexed an emptied trace_info out of bounds.

Skip indices that no longer refer to anything and say so in the log, and
give the LogIRBlock placeholder an explicit invalid index so a block that
prepare_block failed to record doesn't dump block 0 instead.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-19 10:02:54 -06:00
Henrik RydgårdandClaude Opus 5 675ad0acd6 Interpreter: assorted fixes found while reviewing Core/MIPS
WriteMMIO_U32 was missing the return after the kernel-mode check, so a
user-mode write raised the exception and then went through anyway. The
other five MMIO accessors already return here; this one lost it when the
GPIO/syscon branches were added.

Int_Vrot scanned all four entries of dregs, but GetVectorRegs only fills
the first n - so a vrot with vs == 0 (S000) matched a lane that isn't
there and took the cosine from it. The IR backend gets this right via
IsOverlapSafe, so the two disagreed.

The breakpoint checks in MIPSInterpret and RunUntilDowncountZeroWithChecks
read instr->flags without checking for null, which MIPSGetInstruction
returns for the eight primary opcodes (and many subops) that don't decode.
With a memcheck or register breakpoint active, landing on one of those
crashed instead of raising ILLEGAL.

Also: GetVectorOverlap decoded the second vector with size1 (no callers
today), and RegisterFunction left most of its AnalyzedFunction
uninitialized, including the size that ends up in knownfuncs.ini.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-19 10:02:46 -06:00
Artem Lytkin 4c749a31cf Savedata: use PARAM.SFO's time as the save date in the dialog
GetSaveInfo() took the date from whichever file the directory listing
returned first, so a save whose ICON0.PNG is older than the rest showed
the icon's date in the save/load dialog. The savedata manager already
uses PARAM.SFO's time. Use that here too, keeping the first file as the
fallback when there's no PARAM.SFO.
2026-09-19 17:17:55 +03:00
Henrik Rydgård 99fecc1874 Merge pull request #22310 from hrydgard/kprintf
KDebugForKernel: implement Kprintf
2026-09-18 16:59:23 -06:00
Henrik Rydgård 8f3da1a84a Merge pull request #22233 from hrydgard/combo-suppress-singles
Don't fire single-button mappings while a combo using them is held
2026-09-18 16:58:46 -06:00
Henrik Rydgård 5c9eda7cd8 KDebugForKernel: implement Kprintf
Outrun 2006's USB kernel module calls it on a failed sceUsbbdRegister,
so booting the game printed "Unimplemented function Kprintf". It's the
kernel's debug printf - on hardware it goes to whatever
sceKernelRegisterKprintfHandler() installed, which on a retail PSP is
the serial port, so this is just debug output the game shipped with.
We format it and log it, which is the useful thing to do with it.

The vararg walker that sysclib's sprintf/snprintf already had is now
HLEFormatPrintf() in HLE.cpp, so there's one of these rather than a
second copy. It takes the index of the first vararg (counting a0 as 0)
instead of an offset from a2, which is the same mapping written in
absolute terms - sprintf passes 2 and snprintf 3, where they passed
0 and 1 before.

Kprintf replaces the existing nullptr entry in the KDebugForKernel
table, so no indices move and savestates are unaffected.
2026-09-18 16:26:41 -06:00
Henrik RydgårdandClaude Opus 5 19bac4894c Let sceMpeg LLE override the ForceHLEPsmf compat flag
Our psmf and psmfPlayer HLE plays video by calling our sceMpeg HLE, so it has
nothing to talk to when the real mpeg.prx is running. The flag now gives way
when sceMpeg is LLE.

Moved below the force-enable and unavailable masks so it tests what sceMpeg
actually ended up as rather than what was asked for.

Also remove a bad assert.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 15:15:45 -06:00
Henrik RydgårdandClaude Opus 5 9b78155ebd Settle the wording of the two video firmware messages
The sceMpeg one goes away: our HLE handles almost everything, so ending up on it
isn't worth interrupting the player over. It stays in the log, where it explains
why a video might not look the way it does with the real module.

The sceMp4 one is the one that matters, since there is no working HLE behind it,
and it now says "installed firmware" rather than "firmware dump" - PPSSPP
installs firmware much as a PSP does these days, so that is the wrong word.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 13:50:25 -06:00
Henrik RydgårdandClaude Opus 5 32017202bf Use the disc's own mpeg.prx, and say when sceMp4 won't work
A game that ships MPEG.PRX doesn't need a firmware installed to run the real
module - Death Jr. loads PSP_GAME/USRDIR/MODULES/MPEG.PRX itself - but the flag
came off anyway, because the decision has to be made before the game's imports
resolve and nothing had looked on the disc yet. So look: a bounded walk for a
file of that name, only when flash0 comes up empty, so the usual case pays
nothing.

The name is the easy part - about a quarter of discs ship one and it is called
mpeg.prx in every case seen - but the directory is not. MODULE and MODULES are
the common ones, with KMODULE, PRX, DATA/MODULE, and more at five levels deep,
hence the generous depth limit. Matching on the name rather than reading each
PRX to see what it exports means guessing wrong only costs us the real module.

sceMp4 is the other half of this. Our HLE of it is nearly all stubs, so dropping
the flag for want of firmware doesn't rescue anything - those libraries only
exist in firmware 6.00 and later, and without them MP4 playback simply isn't
available. Since almost nothing uses sceMp4, warning about that every boot would
be noise, so NotifyLoadStatusMp4 says it instead, which only something actually
asking for MP4 reaches.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 13:49:10 -06:00
Henrik RydgårdandClaude Opus 5 c39db565da Run sceMpeg and sceMp4 for real by default
Both now graduate to AlwaysDisableHLEFlags, so a firmware dump gets the real
mpeg.prx, libmp4.prx and mp4msv.prx without anyone having to find the setting.
The checkbox moves from "Disable HLE" to "Force-enable HLE" on its own, so a
game that regresses still has a way back.

Neither can be counted on being there, so HLECheckModuleAvailability drops the
flag when the module is missing and the HLE serves as before - the same thing
sceFont does when flash0:/font is empty. Its sceMp4 check asked the setting,
which is no longer where the answer is now that the default is on; it asks
AlwaysDisableHLEFlags instead, and sceMpeg gets a check of its own.

Both are quiet about it now. Missing firmware used to mean a request we couldn't
honour, which was worth a warning on screen; now it just means the user has no
dump, which is the ordinary way to run PPSSPP.

The cost is that a disc carrying its own mpeg.prx also falls back when there's
no firmware, though it would have run fine. The choice has to be made before the
game's imports are resolved and there's no telling then whether a module will
appear later, and getting it wrong leaves the game importing from a module that
never loads.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 12:55:08 -06:00
Henrik RydgårdandClaude Opus 5 5c5c7dec86 Include what these files use
UnitTest.h's EXPECT_ macros all call printf and EXPECT_EQ_MEM calls memcmp, but
it included neither <cstdio> nor <cstring> - it has been relying on whatever the
including file happened to pull in first, and TestMpegCsc was the first not to.
The same shape in sceMpegbase.cpp and sceVideocodec.cpp, which use std::min,
std::move and memcpy without saying where they come from.

Also drop an abs() from TestMpegCsc rather than include <cstdlib> for one
subtraction.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 11:59:48 -06:00
Henrik RydgårdandClaude Opus 5 8fa0826bdb sceMpegbase: init and shut down from the kernel, like everything else
__MpegBaseInit was called by __MpegInit rather than by __KernelInit, and
__MpegBaseShutdown had just been added the same way. Every other module is
started and stopped directly by the kernel - __MpegBaseDoState already was -
so do the same here and let sceMpeg.cpp mind only its own state.

The order is unchanged: __MpegBaseInit ran first inside __MpegInit and now sits
just before it, __MpegBaseShutdown ran last and now sits just after.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 11:38:49 -06:00
Henrik RydgårdandClaude Opus 5 dc743d4fce sceMpegbase: free the conversion's memory on shutdown
__MpegBaseShutdown, hanging off __MpegShutdown the way __MpegBaseInit hangs off
__MpegInit, so the de-tiling scratch and the swscale context go back when the
game stops rather than only when the next one starts. Between them they are a
few hundred kilobytes that a game which played one video early on has no further
use for.

Also name the swscale flags rather than passing SWS_POINT inline, and say next
to it what the choice actually decides - with equal sizes in and out it is only
how chroma gets to full resolution, and SWS_BILINEAR (what the HLE uses) is a
one-line swap. Worth a real option one day; not adding one now.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 11:38:49 -06:00
Henrik RydgårdandClaude Opus 5 4cb01fea67 sceMpegbase: convert with swscale, keeping the scalar path as the fallback
The planes the de-tiling produces are already the YUV420P swscale wants, and our
sceMpeg HLE converts the same frames the same way, so the pixel formats and the
studio-range setup come straight from MediaEngine::getSwsFormat. It is 3-4x
quicker than going a pixel at a time: 0.37-0.44ms a frame becomes 0.09-0.12ms,
which is the whole reason sceMpegBaseCscAvc was at the top of a profile.

Chroma is upsampled with SWS_POINT rather than the HLE's SWS_BILINEAR, since
replicating is what the scalar path does and, being a fixed-function block,
almost certainly what the hardware does.

It is not bit-identical - swscale rounds its own way. TestMpegCsc measures the
gap per channel rather than per byte, so the number means something for a packed
16-bit pixel: worst 1 step of 31 for 5650 and 5551, 2 of 15 for 4444, 3 of 255
for 8888, with means around a fifth of a step. The scalar path stays as what the
longhand reference is checked against, and takes anything swscale won't - an odd
range origin, or a build without ffmpeg.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 11:38:49 -06:00
Henrik RydgårdandClaude Opus 5 02906f0510 sceMpegbase: write alpha as zero, and stop rebuilding the planes every frame
The colour conversion was writing alpha fully set - 0xFF000000, or the top bit
for 5551 - where the hardware writes zero. Our sceMpeg HLE already masks it off
and names Sword Art Online as a game that depends on it: it doesn't clear the
alpha in the buffer it hands over, and expects the video not to set it. The two
paths now agree.

The de-tiling ahead of it becomes UntileYCbCr, taking the eight buffers already
resolved, so it can be measured and compared against the original longhand
version in TestMpegCsc. Its planes move to scratch that persists between calls -
a movie converts one frame per displayed frame, and this was allocating and
clearing about 200KB every time - and the per-pixel bounds checks in the chroma
loop, which only depend on the group of eight, are hoisted out of it.

That last part is worth 2529 -> 3201 MPix/s, but the point of measuring was to
find out whether it mattered, and it doesn't much: de-tiling is 0.04ms of a
frame against the conversion's 0.4ms. The conversion is where the time is.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 11:38:49 -06:00
Henrik RydgårdandClaude Opus 5 a6996b2c3f sceMpegbase: pull the colour conversion out, and measure it
sceMpegBaseCscAvc is the top of a profile during video playback, so the loop
that does the work becomes MpegCscRange - a pure function with the HLE plumbing
left behind - and TestMpegCsc measures and checks it.

The measuring half reports megapixels per second for a 480x272 frame in each of
the four pixel formats. The checking half compares against the conversion
written out longhand, over whole frames and over partial ranges with odd offsets
and sizes, plus one-pixel, one-row and one-column ranges and one that reaches
the far edge of the frame. Those are the cases an optimized version gets wrong:
chroma is half resolution, so an odd left edge starts mid-sample, and anything
handling two pixels at a time has to deal with the leftover. The destination is
padded and prefilled, so writing outside the range fails too.

This is only the move - the loop is the same one, so the numbers it gives are
the baseline to improve on. On a Snapdragon X Elite it runs at about 300 MPix/s,
0.44ms for a frame.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 11:38:49 -06:00
Henrik Rydgård 21c7312486 Merge pull request #22306 from hrydgard/assorted-fixes
Various fixes to sceMpeg LLE and sceMp3 LLE
2026-09-18 10:11:22 -06:00
Henrik RydgårdandClaude Opus 5 01be454b7a sceMp3: keep accepting sample rates a PSP would refuse, on purpose
libmp3.prx only accepts a rate other than 44.1kHz from a game built with SDK
3.09.05 or later, and audio/mp3/init has the hardware's answers for the rest.
PPSSPP has nonetheless accepted them from every game that declares an SDK
version, because the threshold was written as decimal 3090500 rather than
0x03090500 - an accident, but one people have come to rely on. Beats and games
like it build levels out of MP3s the user supplies, and refusing an ordinary
48kHz file looks like a bug to whoever supplied it.

So the hardware answer now goes only to something that declares no SDK version
at all, which in practice means the test, and the comment says that is a choice
rather than an oversight. Being strict again is a one-line change, with the two
lines it would cost named next to it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 09:45:11 -06:00
Henrik RydgårdandClaude Opus 5 7ed0598433 sceAudiocodec: report the decoded size in bytes, not samples
The field at 0x24 is a byte count, like srcBytesRead next to it, and
sceAudiocodecGetOutputBytes describes the same quantity the same way (0x1200
for MPEG1 MP3). We were putting the sample count there, a quarter of the value,
and libmp3.prx takes it as the length of the PCM to pass on. Renamed to
dstBytesWritten so it reads like what it is.

Nothing on our side consumed the field, so this only changes what the firmware
modules see. mpeg.prx ignores it, which is why Atrac3+ playback was unaffected
either way, but libatrac3plus.prx does read it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 09:44:39 -06:00
Henrik RydgårdandClaude Opus 5 08573668c5 sceAudiocodec: fill in the MP3 version in GetInfo, log what GetInfo read out
sceAudiocodecInit puts 9999 in the version field to mean "not known yet", and
filling it in is what this call is for - libmp3.prx reads it straight back out.
We wrote every other MP3 field and left that one alone, so the real libmp3.prx
got as far as GetInfo and then stopped without ever asking for a decode.

While here, read the fields off the frame rather than claiming 128kbps 44.1kHz
stereo unconditionally, which is what the hardware does with them. They are the
raw MPEG header fields apart from the version index, which has its own
numbering. The old fixed values stay as the fallback for when there is no
readable frame to look at.

Also carries a CheckNeedMem log tweak that was already in the tree.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 09:43:39 -06:00
Henrik RydgårdandClaude Opus 5 b77e7c4e66 sceVideocodec: implement CopyYCbCr
This is what sceMpegAvcCopyYCbCr is built on, and a game that wants the raw
YCbCr rather than letting sceMpegbase convert to RGB uses it and nothing else.

mpeg.prx builds the descriptor on its own stack and avcodec.prx reads it back at
0x800015c4. Dimensions in pixels at 0x00/0x04, the eight frame buffers from 0x0c
but ordered 0,2,4,6 then 1,3,5,7, and from 0x2c the destination Y with Cb and Cr
following it contiguously - ordinary planar YUV420. Checked against what the
game passes: the eight addresses are exactly the buffers we handed out, and the
three destinations are spaced width*height and width*height/4 apart.

Un-tiling is the same operation the colour conversion already does, so that
moves out of sceMpegbase.cpp as ReadTiledYCbCr rather than being written twice.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-09-18 09:43:15 -06:00