On a cold start, onCreate passed the intent's shortcut parameter to
NativeApp.init and then also posted it as a "shortcutParam" message, which
is only meant for when the native side is already up. For a game shortcut
the second boot request was rejected with an error. For the Args extra,
the raw command line got booted as if it was a game path, so something
like --loglevel=4 applied fine and then popped up a load failure.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
EmuThread_Join drains the render queue on the calling thread. On Android
that is either the UI thread with no GL context (NativeApp.shutdown), or
the GL thread with a brand new context after the old one was lost
(displayInit the second time). The skip flag was only set once the emu
thread reached NativeShutdownGraphics, so everything queued before that
ran for real, against nothing or against the wrong context.
Add GraphicsContext::NotifyContextLost() and call it before the join in
both places. skipGLCalls_ is now atomic, since it's set from one thread
and read from another.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
The surface may not be used after surfaceDestroyed returns. We relied on
onPause having joined the thread first, which is the usual order but not a
guaranteed one. When it didn't hold, the thread kept spinning on a lost
surface, and the replacement surface was refused as "Already running".
Also forget a surface deferred while paused if it gets destroyed before
the resume that would have used it.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
If InitSurface failed, the thread returned but stayed joinable, so the next
surface was refused as "Already running" until the next pause. Tell a live
thread from a finished one with renderLoopRunning, and reap the latter.
Also make VulkanGraphicsContext::InitSurface clean up when it fails after
creating the draw context. Callers don't call ShutdownSurface on failure,
so draw_ (with its render manager) and the surface leaked, and the next
attempt overwrote the pointer.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
requestExitVulkanRenderLoop went by renderLoopRunning, which the thread
only set once it got going. A pause landing between thread creation and
that point returned without signalling or joining, leaving a render thread
running while paused, and every later start refused as "Already running".
Go by whether there is a thread to join instead, set the running flag
before the thread is created, and let the requester reset exitRenderLoop
after the join rather than the thread itself.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
When the firmware module a --disable-hle bit asks for is neither installed nor
on the disc, that library silently runs our HLE instead. For the emulator that
is the right thing; for a test tool it means the run measures something other
than what was asked for and says so only as one INFO line, which is easy to
grep past and easy to never see. It cost a round of wrong results here, where
the memory stick headless defaults to (beside the exe, not the app's) had no
firmware, so a comparison against the real mpeg.prx was quietly a comparison
against the HLE it was supposed to be measured against.
g_unavailableDisableFlags already records exactly which flags fell back, so it
just needed an accessor. Headless now names each one, prints the flash0:/kd and
memory stick it looked in, and fails the run.
Only an explicit --disable-hle binds. sceMpeg and sceMp4 are LLE by default, and
falling back is the correct and expected behaviour wherever no firmware is
installed - making that fatal would fail every run on such a machine.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Running a commercial game through headless to measure something has four ways to
report a clean pass for a run that proved nothing, and they compound: --log is
needed before anything is printed (and it goes to stderr, so grepping forces the
streams together), the memory stick defaults to one beside the exe rather than
the app's, which means no installed firmware and a silent LLE-to-HLE fallback,
an unrecognised parameter then hides in the merged output, and a count of zero
errors reads identically whether the code ran or never got there.
All four of these cost a round of wrong results while investigating ME memory.
Also records that the line-ending check has to be done on the bytes, since Git
Bash's grep normalises them: the obvious `grep -c $'\r$'` for a stray LF reports
every file clean, which is how a handful of LF lines got into two CRLF files
here without the usual whole-file-rewrite tell in git diff --stat.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
sceMpegBasePESpacketCopy used to gather each DMA into a buffer filed under the
address of its first block, and sceVideocodecDecode looked one up by the address
it was handed. That works only while an access unit arrives in a single call. It
doesn't: mpeg.prx routinely DMAs one in several pieces at consecutive addresses,
and the table then kept whichever piece happened to start where the decode later
asked, and dropped the rest.
So the DMA now writes each block at the address it names. The Media Engine's own
memory moves to its real addresses - low ones, which mpeg.prx bounds-checks
against 0x3FFFFF - and pieces written at consecutive addresses end up
consecutive, which is all this ever needed. Our own allocations move to the top
half, away from the addresses mpeg.prx picks for itself.
That memory is the ME's, not the Allegrex's, and the two are separate address
spaces: on hardware only the ME reaches it, so here it is a buffer of our own
that nothing emulated can address. It gets its own MEIsValidRange and
MEGetPointerRange rather than borrowing Memory::, which answers for the
Allegrex's memory and has nothing to say about this one.
An address only means something together with the space it came from, and
mpeg.prx hands over bare integers. The three places that have to work it out
from the value now ask the ME first, where they used to ask Memory:: first and
so read every address as the CPU's. It separates the addresses these callers
actually pass (ME memory well down in the first megabyte, PSP RAM at 0x08000000
and up), but that's those callers rather than a rule.
Savestates from before this carry the old packet table; it's read and dropped,
since the Media Engine's memory is saved and that's where the payloads live now.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Nothing on this path ever called InitFFmpeg, so ffmpeg's log callback was never
installed and everything it had to say about the bitstream went nowhere. Our own
return codes only report that a frame arrived, not what it was built from, so a
video could fall apart on screen while the log showed two warnings in two
thousand frames.
Also report decode_error_flags and AV_FRAME_FLAG_CORRUPT per frame, which covers
the other half: a frame that decodes "successfully" out of concealment. Coalesced
into one line per run of damage, since a missing reference makes every frame
damaged until the next keyframe.
Not setting AV_EF_EXPLODE, which would make the decoder fail where it currently
conceals - that is a behaviour change, and this is meant to be a reporting one.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The contract is only that a bad lane becomes something that yields zero
when multiplied by zero, by whatever route is cheapest. SSE2 clamps to
+-FLT_MAX, NEON, LSX and the scalar fallback zero the lane - both fine.
Check that property, and that good lanes are untouched.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
EncodeCR, EncodeCI and EncodeCSS shifted the RiscVReg enum value straight
into the instruction without DecodeReg(), which every 32-bit encoder in
the file uses. FPRs are 0x20..0x3F in that enum, so bit 5 is set for all
of them and spilled into a neighbouring field - for the CI format, into
bit 12, which holds imm[5].
The visible effect: the dispatcher's epilogue restores the callee-saved
FP registers with c.fldsp, and every offset whose bit 5 was clear came
out 32 bytes too high. fs3-fs6 were restored from the wrong slots and
fs11 from 224(sp), past the 208-byte frame. So the native JIT corrupted
whatever the C++ caller had in those registers - which, in the headless
test runner, was the wall-clock deadline, making every test report an
instant TIMEOUT.
pspautotests on riscv64 under qemu goes from 0/342 to 338/342 with this,
matching the IR interpreter exactly.
Found with the disassembly the enableDisasm flag in RiscVAsm.cpp emits -
the save offsets and the restore offsets simply didn't match.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
g_Config.iCpuCore does double duty: it picks the MIPS core, and it gates the
vertex decoder JIT in the DrawEngineCommon constructor. Headless forces it to
INTERPRETER to keep that decoder off, but it did so before ApplyToConfig(), so
--cpu=jit and --cpu=jit-ir overwrote it and enabled a GPU path the tests aren't
recorded against - 16 of them failed on x86-64, gpu/vertices/morph among them.
Force it after ApplyToConfig() instead. The core the tests actually run on
comes from CoreParameter, straight off the command line, so the backends are
still tested; only the GPU side is pinned.
Also limit the frametests to x86-64. The reference images don't match the arm64
software renderer - 23 of 30 dumps differ, reproducible on any arm64 host. That
predates the new runner, which just gave it somewhere to show up.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Same treatment the loongarch64 cross build already gets - its sysroot has
no SDL3, so there's no windowed graphics context to create. Both run fine
under qemu with --graphics=software, which needs no context at all, so
correct the comment that called this compile-tested only.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
CrossSIMD has four independent implementations and only one is compiled
per machine, so each architecture has to check its own copy against
hand-worked results. Running this under qemu gives us coverage of the
targets we have no hardware for - it's already caught two LSX bugs.
Covers Vec4S32 and Vec4F32 arithmetic, comparisons and the mask helpers,
the lane accessors and shuffles, transpose, the various loads (including
the 24-bit and normalizing ones), NaN/Inf handling, and the matrix
routines the old test already had.
Verified on three of the four implementations: NEON natively, the scalar
fallback via TEST_FALLBACK, and LSX under qemu-loongarch64. SSE2 is left
to CI.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
LoadConvertU8, StoreConvertToU8 and LoadTranspose existed in the SSE2,
NEON and LSX implementations but not in the scalar one, so anything using
them wouldn't build on a target without SIMD.
The two LSX bugs were found by the new CrossSIMD unit test, run under
qemu-loongarch64:
- Vec4F32::operator[] had a switch with no breaks, so every index fell
through to the default and returned lane 3.
- StoreConvertToU8 narrowed with the logical (unsigned) saturating shifts,
which turn a negative value into a huge unsigned one and saturate it to
255. It should clamp to 0, as the packs/packus pair in the SSE version
does. Narrow signed->signed and then signed->unsigned instead.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The headless tests only ever exercised the JIT, since that's what headless
defaults to. Run all four on the Linux runners, and add an arm64 Linux lane
so the arm64 JIT is covered too - nothing else in the matrix tested it.
test.py scales the wall clock to the backend instead of raising it for
everyone: the interpreter needs 20s for gpu/rendertarget/copy, which does
over a million guest-side vsprintf calls, while a hang under the JIT is
still caught in five seconds.
The frametest report artifact needs a per-OS name now that two Linux legs
upload one.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
MOVE(LoongArch64Reg::X4, arg) targets X4, which is an LASX 256-bit vector
register (0x64), not a0. The LP64D ABI puts the first integer argument in
a0, which is R4.
This is not a corner case: GenerateFixedCode uses QuickCallFunctionR for
CoreTiming::Advance, so the emitter asserted ("DJK instruction rd must be
GPR") while building the dispatcher, before a single block was compiled.
The LoongArch native JIT could not start at all.
Found by running the loongarch64 cross build under qemu-user.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The riscv64 cross build is the first target to take the non-SIMD path,
and it didn't compile:
- LoadF24x4 called LoadR24x3_One, which doesn't exist. It should shift
all four lanes, like the SSE/NEON/LSX versions do.
- isnan/isinf were unqualified, and <cmath> wasn't included.
- WithLane3From and AnyCompareBitsSet were missing entirely.
LoadF24x3_One also left lane 3 as zero, where all three SIMD versions set
it to 1.0f - a behavioural bug that would only have shown up once someone
ran this path.
Verified by flipping TEST_FALLBACK in the header: the whole tree builds,
and with the scalar path in use the unit tests and pspautotests both pass
in full (342/342, software renderer, which leans on this code heavily).
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The pieces were nearly all in place already - CMakeLists has detected
riscv64 and set RISCV64 since the backend landed - so this is mostly a
matter of not hardcoding loongarch64 in the cross-build plumbing:
- setup-loongarch64-cross.sh becomes setup-cross.sh <target>, taking the
triple and dynamic linker name from the argument.
- b.sh gains --riscv64, and the GL stub block is keyed off a compiler
variable rather than a loongarch64-only flag.
- New cmake/Toolchains/riscv64-linux-gnu.cmake, mirroring the existing one.
The SDL carve-out moves from LOONGARCH64_DEVICE to a new HEADLESS_CROSS
option set by b.sh. That flag was keyed off the target architecture, which
is also true when building natively on such a machine - harmless so far,
but riscv64 hardware that people actually build PPSSPP on exists, and it
should still get the normal SDL frontend.
Both toolchain files now find qemu rather than assuming the static build:
Debian ships the static binaries in qemu-user-static and the dynamic ones
in qemu-user, and only the latter is available on some releases. Neither
being present is fine too - it just means no compile-check programs run.
Verified by a clean loongarch64 build; riscv64 is untested so far, the
toolchain isn't installed here yet.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The GL stub generation in b.sh and setup-loongarch64-cross.sh harvested
symbols from hardcoded /usr/lib/x86_64-linux-gnu paths. On any other host
those simply don't exist, and neither script treats that as an error - so
the stub silently fell back to the handful of GLX symbols listed inline,
which isn't enough for GLEW's static archive to link against.
Ask the host compiler for its multiarch triplet instead.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
FpCondFromReg's meta is "_G", so the register is in src1 - which is where
every native backend, PropagateConstants and ReorderLoadStore read it.
But IRInterpret read it from dest, and Comp_VecDo3 wrote it to dest to
match. The other emitter passes (0, MIPS_REG_ZERO), so both fields are
zero there and the disagreement stayed hidden.
The result was that vsge/vslt, which save fpcond to IRTEMP_0 and restore
it afterwards, restored r0 (always zero) on the native IR JITs instead of
the saved value, losing any c.cond.s result live across them. Fixed both
the emitter and the interpreter to use src1.
Vec4Pack31To8's SSE2 path computed (v >> 24) << 1, which is (v >> 23)
with bit 23 forced to zero - the scalar and NEON paths both do
(v >> 23) & 0xFF. Shift left first instead. The pspautotests inputs all
happen to have bit 23 clear after the clamp, so this wasn't caught.
Also:
- Evaluate() didn't mask constant-folded shift amounts to 5 bits, though
the neighbouring one-immediate path does.
- ApplyMemoryValidation only invalidated its address-check cache for ops
with a 'G' destination, while Interpret and CallReplacement can write
any GPR - both are barriers, so drop the whole cache when we see one.
- ReduceVec4Flush indexed isVec4 with (src2 & 3) instead of (src2 & ~3)
for Vec4Scale; a missed optimization rather than a miscompile.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
flush_to_file() ends with clear(), but tracing stays on and the blocks
already compiled keep the index baked into their LogIRBlock instruction.
A second flush then indexed an emptied trace_info out of bounds.
Skip indices that no longer refer to anything and say so in the log, and
give the LogIRBlock placeholder an explicit invalid index so a block that
prepare_block failed to record doesn't dump block 0 instead.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
WriteMMIO_U32 was missing the return after the kernel-mode check, so a
user-mode write raised the exception and then went through anyway. The
other five MMIO accessors already return here; this one lost it when the
GPIO/syscon branches were added.
Int_Vrot scanned all four entries of dregs, but GetVectorRegs only fills
the first n - so a vrot with vs == 0 (S000) matched a lane that isn't
there and took the cosine from it. The IR backend gets this right via
IsOverlapSafe, so the two disagreed.
The breakpoint checks in MIPSInterpret and RunUntilDowncountZeroWithChecks
read instr->flags without checking for null, which MIPSGetInstruction
returns for the eight primary opcodes (and many subops) that don't decode.
With a memcheck or register breakpoint active, landing on one of those
crashed instead of raising ILLEGAL.
Also: GetVectorOverlap decoded the second vector with size1 (no callers
today), and RegisterFunction left most of its AnalyzedFunction
uninitialized, including the size that ends up in knownfuncs.ini.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
GetSaveInfo() took the date from whichever file the directory listing
returned first, so a save whose ICON0.PNG is older than the rest showed
the icon's date in the save/load dialog. The savedata manager already
uses PARAM.SFO's time. Use that here too, keeping the first file as the
fallback when there's no PARAM.SFO.
Outrun 2006's USB kernel module calls it on a failed sceUsbbdRegister,
so booting the game printed "Unimplemented function Kprintf". It's the
kernel's debug printf - on hardware it goes to whatever
sceKernelRegisterKprintfHandler() installed, which on a retail PSP is
the serial port, so this is just debug output the game shipped with.
We format it and log it, which is the useful thing to do with it.
The vararg walker that sysclib's sprintf/snprintf already had is now
HLEFormatPrintf() in HLE.cpp, so there's one of these rather than a
second copy. It takes the index of the first vararg (counting a0 as 0)
instead of an offset from a2, which is the same mapping written in
absolute terms - sprintf passes 2 and snprintf 3, where they passed
0 and 1 before.
Kprintf replaces the existing nullptr entry in the KDebugForKernel
table, so no indices move and savestates are unaffected.
Our psmf and psmfPlayer HLE plays video by calling our sceMpeg HLE, so it has
nothing to talk to when the real mpeg.prx is running. The flag now gives way
when sceMpeg is LLE.
Moved below the force-enable and unavailable masks so it tests what sceMpeg
actually ended up as rather than what was asked for.
Also remove a bad assert.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Follows each language file's own word for firmware rather than imposing one, so
this reads the way the neighbouring strings in [System] already do: fastvare in
Norwegian, systemprogramvara in Swedish, laiteohjelmisto in Finnish, proshivka
in Russian and Ukrainian, and the English word where that is what the language
has settled on.
Thirteen files keep the English string as a placeholder. Those are the ones with
no translation for the firmware strings already in that section either, so there
was no house style to follow and guessing would leave nobody able to check it.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The sceMpeg one goes away: our HLE handles almost everything, so ending up on it
isn't worth interrupting the player over. It stays in the log, where it explains
why a video might not look the way it does with the real module.
The sceMp4 one is the one that matters, since there is no working HLE behind it,
and it now says "installed firmware" rather than "firmware dump" - PPSSPP
installs firmware much as a PSP does these days, so that is the wrong word.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The command took $1/$2/$3 off the front of whatever was typed, so it only got
the right section and key when called in exactly the documented shape. Called
with a sentence - which is the natural way to ask for this - it silently
produced three arbitrary words, and a paragraph warning about that is a poor
substitute for not doing it. It now gets $ARGUMENTS whole and works the three
values out, which is the part that needed a model anyway.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
A game that ships MPEG.PRX doesn't need a firmware installed to run the real
module - Death Jr. loads PSP_GAME/USRDIR/MODULES/MPEG.PRX itself - but the flag
came off anyway, because the decision has to be made before the game's imports
resolve and nothing had looked on the disc yet. So look: a bounded walk for a
file of that name, only when flash0 comes up empty, so the usual case pays
nothing.
The name is the easy part - about a quarter of discs ship one and it is called
mpeg.prx in every case seen - but the directory is not. MODULE and MODULES are
the common ones, with KMODULE, PRX, DATA/MODULE, and more at five levels deep,
hence the generous depth limit. Matching on the name rather than reading each
PRX to see what it exports means guessing wrong only costs us the real module.
sceMp4 is the other half of this. Our HLE of it is nearly all stubs, so dropping
the flag for want of firmware doesn't rescue anything - those libraries only
exist in firmware 6.00 and later, and without them MP4 playback simply isn't
available. Since almost nothing uses sceMp4, warning about that every boot would
be noise, so NotifyLoadStatusMp4 says it instead, which only something actually
asking for MP4 reaches.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Both now graduate to AlwaysDisableHLEFlags, so a firmware dump gets the real
mpeg.prx, libmp4.prx and mp4msv.prx without anyone having to find the setting.
The checkbox moves from "Disable HLE" to "Force-enable HLE" on its own, so a
game that regresses still has a way back.
Neither can be counted on being there, so HLECheckModuleAvailability drops the
flag when the module is missing and the HLE serves as before - the same thing
sceFont does when flash0:/font is empty. Its sceMp4 check asked the setting,
which is no longer where the answer is now that the default is on; it asks
AlwaysDisableHLEFlags instead, and sceMpeg gets a check of its own.
Both are quiet about it now. Missing firmware used to mean a request we couldn't
honour, which was worth a warning on screen; now it just means the user has no
dump, which is the ordinary way to run PPSSPP.
The cost is that a disc carrying its own mpeg.prx also falls back when there's
no firmware, though it would have run fine. The choice has to be made before the
game's imports are resolved and there's no telling then whether a module will
appear later, and getting it wrong leaves the game importing from a module that
never loads.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The legacy Android build has a unit test executable of its own, so a new file in
unittest/ goes in three build files rather than the two the docs named. Missing
the Android one builds fine everywhere it is convenient to try and fails only on
Android CI, which is what happened here - so both AGENTS.md and building.md now
say three, and which one is easy to forget.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>