Two bugs stacked on each other, and between them every vector path in
this backend was either dead or miscompiled.
The register cache declared mapFPUSIMD unconditionally, but every
compiler here picks its path from cpu_info.LOONGARCH_LSX at runtime.
Without LSX the scalar fallbacks ran against a SIMD mapping and reached
lanes with F(reg + n), which addresses nothing there - ApplyMapping only
allocates per-lane registers when mapFPUSIMD is false. Vec4Unpack8To32
and Vec4DuplicateUpperBitsAndShift1 came out with only their first lane,
Vec2Unpack16To32 put both halves in one register, and the pack ops read
back whatever was next door. riscv64 sets the flag false and has never
had the problem. Tie the two together and the fallbacks are correct as
they stand.
That path was reachable because of the second one: the USE_CPU_FEATURES
block assigned over the hwcaps, and GetLoongArchInfo reads /proc/cpuinfo
and nothing else. Under qemu-user that file belongs to the host, so it
found no features and turned LSX off - leaving the LSX paths dead and the
untested scalar ones running, which is not what any Loongson 3A5000 or
later does. Let it only ever add to what the hwcaps found.
cpu/vfpu/convert, gum and matrix pass now, both with LSX and with it
masked off, putting loongarch64 at 338/342 - the same four as riscv64.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
ApplyHostRoundingMode and RestoreHostRoundingMode only had SSE2 and ARM64
branches, so everywhere else the guest's rounding mode was silently
ignored and ceil and floor rounded to nearest. Fall back to fesetround,
which covers riscv64 and loongarch64 and whatever comes next.
That clears both rounding mismatches in cpu/fpu/fpu on those two. What's
left there is flush-to-zero, which neither ISA has any control for, so
the test still fails - a denormal result survives where the PSP would
have flushed it.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Those two have no runner of their own, so the jobs only ever compiled -
which is thin coverage for a code generator. Both build PPSSPPUnitTest
now and run it under qemu-user, which is where the vertex decoder and
CrossSIMD bugs on those backends turned up in the first place.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Enables TestVertexJitMatchesSteps on both. riscv64 now passes all 216000
jitted formats and loongarch64 all 204000, against 28906 on arm64.
riscv64:
- Jit_PosFloat didn't clean NaN or infinity at all, it just copied the
three words. Clamp to +-FLT_MAX like the x86 JIT does.
- Jit_PosFloatThrough was missing the truncation of Z to an integer.
- The morph helpers started the sum from the first product rather than
from +0.0, which rounds differently and lets a -0 term through, and
rounded that first product towards zero where the steps round to
nearest. The rest of the sum stays fused, since the compiler contracts
the steps into fused multiply-adds.
- The texcoord prescale and 5551 color morph paths read morph weights
from tables that GetMorphValueUsage never asked to be filled in, so
they used whatever an earlier vertex type had left there.
- The non-Zbb bounds update compared the wrong way around, so through
mode texcoord bounds came out inverted.
loongarch64:
- Jit_PosFloat had a TODO to clean NaN and infinity, and didn't.
- Jit_PosFloatThrough was missing the same Z truncation.
- Jit_WriteMorphColor narrowed with the logical saturating shifts, so a
negative channel became a huge unsigned value and saturated to 255
instead of clamping to 0, and it rounded where the steps truncate. It
also read the packed color back sign-extended, so any alpha above 0x7F
compared as larger than 0xFF000000 and claimed full alpha.
- The three packed color morph formats are rewritten. The LSX versions
built each channel with a chain of inserts, shifts and shuffles that
didn't survive being run; 4444 also broadcast its scale from the mask
register. They now follow the steps channel by channel. Note the
accumulator has to be an LSX scratch register - F4-F7 alias V4-V7,
which hold the skin matrix for the whole vertex.
- The vertex bounds were loaded with a signed halfword load, so the
0xFFFF they start at became -1 and no texcoord was ever below it.
PrescaleUV now fuses on these two as well - both JITs fuse it, and the
compiler would have contracted the plain expression there anyway.
The test tolerates a small relative difference on decoded floats, scaled
by the morph count since each term rounds once. Every bug above was
orders of magnitude larger than that.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The test didn't compile: VERTS is captured by reference into testFormat, and
MSVC won't use a captured constexpr as an array bound. Make it static.
On arm64 it then failed on the UV prescale steps, by one ULP. The arm64 JIT and
the NEON handwritten decoders fuse the multiply-add, and the steps only match
that when the compiler contracts a * b + c - which clang does and MSVC doesn't,
in Debug or Release. Spell out which one happens instead of relying on it.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The two handwritten SIMD decoders are used with or without the JIT, and had
drifted from the step functions:
- The God of War one ignored g_DoubleTextureCoordinates, giving HD Remaster
games the wrong UVs, and passed NaN and infinite positions through. Those
now come out finite like everywhere else.
- The GTA one expanded 5551 colors wrongly on NEON (a left shift where the
SSE version shifts right).
- On ARM64, both now fuse the UV scale and offset, like the JIT and the steps
as the compiler builds them.
Both are back in the unit test, which now also feeds NaN and infinity to plain
float positions and checks they come out finite, and runs only on x86-64 and
arm64 for now.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Add a unit test that decodes every vertex format through both the step
functions and the JIT and requires identical output, side effects included.
Only skinning may differ by rounding, since arm64 accumulates the bone
matrices with fused multiply-adds.
What it found, and fixed:
- x86 morph colors rounded to nearest where the steps truncate, and applied
the scale before the weight, which rounds differently.
- x86 through-mode u16 UV bounds compared signed, so texcoords above 32767
scrambled the bounds for everyone running the x86 JIT.
- x86 morph sums could produce -0 where the steps produce +0.
- Step_NormalS16Morph scaled by 1/32768 twice, giving near-zero normals.
- Step_PosFloatThrough lost the truncation of Z to an integer that the JITs
and the vertex reader did before it moved into the decoder.
- SetVertexType never reset skinInDecode.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The vertex decoder JIT was enabled only when g_Config.iCpuCore was one of the
JIT cores, which tied an unrelated GPU path to the CPU setting. Headless leaned
on that by forcing iCpuCore to the interpreter after ApplyToConfig(), to keep
the decoder JIT off in the tests.
Add CoreParameter::bUseVertexDecoderJit instead. Standalone and libretro set it
wherever the host can JIT, so IR interpreter users on such hosts now get the
decoder JIT too. Headless sets it false, which keeps the test results as they
were and lets --cpu go through ApplyToConfig() like every other option.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
There are a few games that seem to rely on performance problems or
timing quirks to hold the correct pace when playing sceMpeg - have not
been able to nail down any other timing mechanism.
This adds a timing enforcement function.
Coded by Claude.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
When the firmware module a --disable-hle bit asks for is neither installed nor
on the disc, that library silently runs our HLE instead. For the emulator that
is the right thing; for a test tool it means the run measures something other
than what was asked for and says so only as one INFO line, which is easy to
grep past and easy to never see. It cost a round of wrong results here, where
the memory stick headless defaults to (beside the exe, not the app's) had no
firmware, so a comparison against the real mpeg.prx was quietly a comparison
against the HLE it was supposed to be measured against.
g_unavailableDisableFlags already records exactly which flags fell back, so it
just needed an accessor. Headless now names each one, prints the flash0:/kd and
memory stick it looked in, and fails the run.
Only an explicit --disable-hle binds. sceMpeg and sceMp4 are LLE by default, and
falling back is the correct and expected behaviour wherever no firmware is
installed - making that fatal would fail every run on such a machine.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Running a commercial game through headless to measure something has four ways to
report a clean pass for a run that proved nothing, and they compound: --log is
needed before anything is printed (and it goes to stderr, so grepping forces the
streams together), the memory stick defaults to one beside the exe rather than
the app's, which means no installed firmware and a silent LLE-to-HLE fallback,
an unrecognised parameter then hides in the merged output, and a count of zero
errors reads identically whether the code ran or never got there.
All four of these cost a round of wrong results while investigating ME memory.
Also records that the line-ending check has to be done on the bytes, since Git
Bash's grep normalises them: the obvious `grep -c $'\r$'` for a stray LF reports
every file clean, which is how a handful of LF lines got into two CRLF files
here without the usual whole-file-rewrite tell in git diff --stat.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
sceMpegBasePESpacketCopy used to gather each DMA into a buffer filed under the
address of its first block, and sceVideocodecDecode looked one up by the address
it was handed. That works only while an access unit arrives in a single call. It
doesn't: mpeg.prx routinely DMAs one in several pieces at consecutive addresses,
and the table then kept whichever piece happened to start where the decode later
asked, and dropped the rest.
So the DMA now writes each block at the address it names. The Media Engine's own
memory moves to its real addresses - low ones, which mpeg.prx bounds-checks
against 0x3FFFFF - and pieces written at consecutive addresses end up
consecutive, which is all this ever needed. Our own allocations move to the top
half, away from the addresses mpeg.prx picks for itself.
That memory is the ME's, not the Allegrex's, and the two are separate address
spaces: on hardware only the ME reaches it, so here it is a buffer of our own
that nothing emulated can address. It gets its own MEIsValidRange and
MEGetPointerRange rather than borrowing Memory::, which answers for the
Allegrex's memory and has nothing to say about this one.
An address only means something together with the space it came from, and
mpeg.prx hands over bare integers. The three places that have to work it out
from the value now ask the ME first, where they used to ask Memory:: first and
so read every address as the CPU's. It separates the addresses these callers
actually pass (ME memory well down in the first megabyte, PSP RAM at 0x08000000
and up), but that's those callers rather than a rule.
Savestates from before this carry the old packet table; it's read and dropped,
since the Media Engine's memory is saved and that's where the payloads live now.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Nothing on this path ever called InitFFmpeg, so ffmpeg's log callback was never
installed and everything it had to say about the bitstream went nowhere. Our own
return codes only report that a frame arrived, not what it was built from, so a
video could fall apart on screen while the log showed two warnings in two
thousand frames.
Also report decode_error_flags and AV_FRAME_FLAG_CORRUPT per frame, which covers
the other half: a frame that decodes "successfully" out of concealment. Coalesced
into one line per run of damage, since a missing reference makes every frame
damaged until the next keyframe.
Not setting AV_EF_EXPLODE, which would make the decoder fail where it currently
conceals - that is a behaviour change, and this is meant to be a reporting one.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The contract is only that a bad lane becomes something that yields zero
when multiplied by zero, by whatever route is cheapest. SSE2 clamps to
+-FLT_MAX, NEON, LSX and the scalar fallback zero the lane - both fine.
Check that property, and that good lanes are untouched.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
EncodeCR, EncodeCI and EncodeCSS shifted the RiscVReg enum value straight
into the instruction without DecodeReg(), which every 32-bit encoder in
the file uses. FPRs are 0x20..0x3F in that enum, so bit 5 is set for all
of them and spilled into a neighbouring field - for the CI format, into
bit 12, which holds imm[5].
The visible effect: the dispatcher's epilogue restores the callee-saved
FP registers with c.fldsp, and every offset whose bit 5 was clear came
out 32 bytes too high. fs3-fs6 were restored from the wrong slots and
fs11 from 224(sp), past the 208-byte frame. So the native JIT corrupted
whatever the C++ caller had in those registers - which, in the headless
test runner, was the wall-clock deadline, making every test report an
instant TIMEOUT.
pspautotests on riscv64 under qemu goes from 0/342 to 338/342 with this,
matching the IR interpreter exactly.
Found with the disassembly the enableDisasm flag in RiscVAsm.cpp emits -
the save offsets and the restore offsets simply didn't match.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
g_Config.iCpuCore does double duty: it picks the MIPS core, and it gates the
vertex decoder JIT in the DrawEngineCommon constructor. Headless forces it to
INTERPRETER to keep that decoder off, but it did so before ApplyToConfig(), so
--cpu=jit and --cpu=jit-ir overwrote it and enabled a GPU path the tests aren't
recorded against - 16 of them failed on x86-64, gpu/vertices/morph among them.
Force it after ApplyToConfig() instead. The core the tests actually run on
comes from CoreParameter, straight off the command line, so the backends are
still tested; only the GPU side is pinned.
Also limit the frametests to x86-64. The reference images don't match the arm64
software renderer - 23 of 30 dumps differ, reproducible on any arm64 host. That
predates the new runner, which just gave it somewhere to show up.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Same treatment the loongarch64 cross build already gets - its sysroot has
no SDL3, so there's no windowed graphics context to create. Both run fine
under qemu with --graphics=software, which needs no context at all, so
correct the comment that called this compile-tested only.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
CrossSIMD has four independent implementations and only one is compiled
per machine, so each architecture has to check its own copy against
hand-worked results. Running this under qemu gives us coverage of the
targets we have no hardware for - it's already caught two LSX bugs.
Covers Vec4S32 and Vec4F32 arithmetic, comparisons and the mask helpers,
the lane accessors and shuffles, transpose, the various loads (including
the 24-bit and normalizing ones), NaN/Inf handling, and the matrix
routines the old test already had.
Verified on three of the four implementations: NEON natively, the scalar
fallback via TEST_FALLBACK, and LSX under qemu-loongarch64. SSE2 is left
to CI.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
LoadConvertU8, StoreConvertToU8 and LoadTranspose existed in the SSE2,
NEON and LSX implementations but not in the scalar one, so anything using
them wouldn't build on a target without SIMD.
The two LSX bugs were found by the new CrossSIMD unit test, run under
qemu-loongarch64:
- Vec4F32::operator[] had a switch with no breaks, so every index fell
through to the default and returned lane 3.
- StoreConvertToU8 narrowed with the logical (unsigned) saturating shifts,
which turn a negative value into a huge unsigned one and saturate it to
255. It should clamp to 0, as the packs/packus pair in the SSE version
does. Narrow signed->signed and then signed->unsigned instead.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The headless tests only ever exercised the JIT, since that's what headless
defaults to. Run all four on the Linux runners, and add an arm64 Linux lane
so the arm64 JIT is covered too - nothing else in the matrix tested it.
test.py scales the wall clock to the backend instead of raising it for
everyone: the interpreter needs 20s for gpu/rendertarget/copy, which does
over a million guest-side vsprintf calls, while a hang under the JIT is
still caught in five seconds.
The frametest report artifact needs a per-OS name now that two Linux legs
upload one.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
MOVE(LoongArch64Reg::X4, arg) targets X4, which is an LASX 256-bit vector
register (0x64), not a0. The LP64D ABI puts the first integer argument in
a0, which is R4.
This is not a corner case: GenerateFixedCode uses QuickCallFunctionR for
CoreTiming::Advance, so the emitter asserted ("DJK instruction rd must be
GPR") while building the dispatcher, before a single block was compiled.
The LoongArch native JIT could not start at all.
Found by running the loongarch64 cross build under qemu-user.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The riscv64 cross build is the first target to take the non-SIMD path,
and it didn't compile:
- LoadF24x4 called LoadR24x3_One, which doesn't exist. It should shift
all four lanes, like the SSE/NEON/LSX versions do.
- isnan/isinf were unqualified, and <cmath> wasn't included.
- WithLane3From and AnyCompareBitsSet were missing entirely.
LoadF24x3_One also left lane 3 as zero, where all three SIMD versions set
it to 1.0f - a behavioural bug that would only have shown up once someone
ran this path.
Verified by flipping TEST_FALLBACK in the header: the whole tree builds,
and with the scalar path in use the unit tests and pspautotests both pass
in full (342/342, software renderer, which leans on this code heavily).
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The pieces were nearly all in place already - CMakeLists has detected
riscv64 and set RISCV64 since the backend landed - so this is mostly a
matter of not hardcoding loongarch64 in the cross-build plumbing:
- setup-loongarch64-cross.sh becomes setup-cross.sh <target>, taking the
triple and dynamic linker name from the argument.
- b.sh gains --riscv64, and the GL stub block is keyed off a compiler
variable rather than a loongarch64-only flag.
- New cmake/Toolchains/riscv64-linux-gnu.cmake, mirroring the existing one.
The SDL carve-out moves from LOONGARCH64_DEVICE to a new HEADLESS_CROSS
option set by b.sh. That flag was keyed off the target architecture, which
is also true when building natively on such a machine - harmless so far,
but riscv64 hardware that people actually build PPSSPP on exists, and it
should still get the normal SDL frontend.
Both toolchain files now find qemu rather than assuming the static build:
Debian ships the static binaries in qemu-user-static and the dynamic ones
in qemu-user, and only the latter is available on some releases. Neither
being present is fine too - it just means no compile-check programs run.
Verified by a clean loongarch64 build; riscv64 is untested so far, the
toolchain isn't installed here yet.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The GL stub generation in b.sh and setup-loongarch64-cross.sh harvested
symbols from hardcoded /usr/lib/x86_64-linux-gnu paths. On any other host
those simply don't exist, and neither script treats that as an error - so
the stub silently fell back to the handful of GLX symbols listed inline,
which isn't enough for GLEW's static archive to link against.
Ask the host compiler for its multiarch triplet instead.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
FpCondFromReg's meta is "_G", so the register is in src1 - which is where
every native backend, PropagateConstants and ReorderLoadStore read it.
But IRInterpret read it from dest, and Comp_VecDo3 wrote it to dest to
match. The other emitter passes (0, MIPS_REG_ZERO), so both fields are
zero there and the disagreement stayed hidden.
The result was that vsge/vslt, which save fpcond to IRTEMP_0 and restore
it afterwards, restored r0 (always zero) on the native IR JITs instead of
the saved value, losing any c.cond.s result live across them. Fixed both
the emitter and the interpreter to use src1.
Vec4Pack31To8's SSE2 path computed (v >> 24) << 1, which is (v >> 23)
with bit 23 forced to zero - the scalar and NEON paths both do
(v >> 23) & 0xFF. Shift left first instead. The pspautotests inputs all
happen to have bit 23 clear after the clamp, so this wasn't caught.
Also:
- Evaluate() didn't mask constant-folded shift amounts to 5 bits, though
the neighbouring one-immediate path does.
- ApplyMemoryValidation only invalidated its address-check cache for ops
with a 'G' destination, while Interpret and CallReplacement can write
any GPR - both are barriers, so drop the whole cache when we see one.
- ReduceVec4Flush indexed isVec4 with (src2 & 3) instead of (src2 & ~3)
for Vec4Scale; a missed optimization rather than a miscompile.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
flush_to_file() ends with clear(), but tracing stays on and the blocks
already compiled keep the index baked into their LogIRBlock instruction.
A second flush then indexed an emptied trace_info out of bounds.
Skip indices that no longer refer to anything and say so in the log, and
give the LogIRBlock placeholder an explicit invalid index so a block that
prepare_block failed to record doesn't dump block 0 instead.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
WriteMMIO_U32 was missing the return after the kernel-mode check, so a
user-mode write raised the exception and then went through anyway. The
other five MMIO accessors already return here; this one lost it when the
GPIO/syscon branches were added.
Int_Vrot scanned all four entries of dregs, but GetVectorRegs only fills
the first n - so a vrot with vs == 0 (S000) matched a lane that isn't
there and took the cosine from it. The IR backend gets this right via
IsOverlapSafe, so the two disagreed.
The breakpoint checks in MIPSInterpret and RunUntilDowncountZeroWithChecks
read instr->flags without checking for null, which MIPSGetInstruction
returns for the eight primary opcodes (and many subops) that don't decode.
With a memcheck or register breakpoint active, landing on one of those
crashed instead of raising ILLEGAL.
Also: GetVectorOverlap decoded the second vector with size1 (no callers
today), and RegisterFunction left most of its AnalyzedFunction
uninitialized, including the size that ends up in knownfuncs.ini.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
GetSaveInfo() took the date from whichever file the directory listing
returned first, so a save whose ICON0.PNG is older than the rest showed
the icon's date in the save/load dialog. The savedata manager already
uses PARAM.SFO's time. Use that here too, keeping the first file as the
fallback when there's no PARAM.SFO.
Outrun 2006's USB kernel module calls it on a failed sceUsbbdRegister,
so booting the game printed "Unimplemented function Kprintf". It's the
kernel's debug printf - on hardware it goes to whatever
sceKernelRegisterKprintfHandler() installed, which on a retail PSP is
the serial port, so this is just debug output the game shipped with.
We format it and log it, which is the useful thing to do with it.
The vararg walker that sysclib's sprintf/snprintf already had is now
HLEFormatPrintf() in HLE.cpp, so there's one of these rather than a
second copy. It takes the index of the first vararg (counting a0 as 0)
instead of an offset from a2, which is the same mapping written in
absolute terms - sprintf passes 2 and snprintf 3, where they passed
0 and 1 before.
Kprintf replaces the existing nullptr entry in the KDebugForKernel
table, so no indices move and savestates are unaffected.