LoongArch64: keep LSX detected, and map lanes the way the compilers expect

Two bugs stacked on each other, and between them every vector path in
this backend was either dead or miscompiled.

The register cache declared mapFPUSIMD unconditionally, but every
compiler here picks its path from cpu_info.LOONGARCH_LSX at runtime.
Without LSX the scalar fallbacks ran against a SIMD mapping and reached
lanes with F(reg + n), which addresses nothing there - ApplyMapping only
allocates per-lane registers when mapFPUSIMD is false. Vec4Unpack8To32
and Vec4DuplicateUpperBitsAndShift1 came out with only their first lane,
Vec2Unpack16To32 put both halves in one register, and the pack ops read
back whatever was next door. riscv64 sets the flag false and has never
had the problem. Tie the two together and the fallbacks are correct as
they stand.

That path was reachable because of the second one: the USE_CPU_FEATURES
block assigned over the hwcaps, and GetLoongArchInfo reads /proc/cpuinfo
and nothing else. Under qemu-user that file belongs to the host, so it
found no features and turned LSX off - leaving the LSX paths dead and the
untested scalar ones running, which is not what any Loongson 3A5000 or
later does. Let it only ever add to what the hwcaps found.

cpu/vfpu/convert, gum and matrix pass now, both with LSX and with it
masked off, putting loongarch64 at 338/342 - the same four as riscv64.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
This commit is contained in:
Henrik RydgårdandClaude Opus 5 committed 2026-09-21 15:36:49 -06:00
1 parent 89c8109823
commit 46632812cd
2 files changed
+21 -14

No files matched your search

+17 -13
View File
@@ -170,21 +170,25 @@ void CPUInfo::Detect()
LOONGARCH_PTW = ExtensionSupported(hwcap, 13);
#ifdef USE_CPU_FEATURES
// Only ever add to what the hwcaps said. GetLoongArchInfo reads /proc/cpuinfo and nothing
// else, so anywhere that isn't the real thing - under qemu-user, where /proc/cpuinfo belongs
// to the host - it reports no features at all, and letting it assign would turn off LSX and
// with it every vector path in the JIT.
cpu_features::LoongArchInfo info = cpu_features::GetLoongArchInfo();
LOONGARCH_CPUCFG = true;
LOONGARCH_LAM = info.features.LAM;
LOONGARCH_UAL = info.features.UAL;
LOONGARCH_FPU = info.features.FPU;
LOONGARCH_LSX = info.features.LSX;
LOONGARCH_LASX = info.features.LASX;
LOONGARCH_CRC32 = info.features.CRC32;
LOONGARCH_COMPLEX = info.features.COMPLEX;
LOONGARCH_CRYPTO = info.features.CRYPTO;
LOONGARCH_LVZ = info.features.LVZ;
LOONGARCH_LBT_X86 = info.features.LBT_X86;
LOONGARCH_LBT_ARM = info.features.LBT_ARM;
LOONGARCH_LBT_MIPS = info.features.LBT_MIPS;
LOONGARCH_PTW = info.features.PTW;
LOONGARCH_LAM = LOONGARCH_LAM || info.features.LAM;
LOONGARCH_UAL = LOONGARCH_UAL || info.features.UAL;
LOONGARCH_FPU = LOONGARCH_FPU || info.features.FPU;
LOONGARCH_LSX = LOONGARCH_LSX || info.features.LSX;
LOONGARCH_LASX = LOONGARCH_LASX || info.features.LASX;
LOONGARCH_CRC32 = LOONGARCH_CRC32 || info.features.CRC32;
LOONGARCH_COMPLEX = LOONGARCH_COMPLEX || info.features.COMPLEX;
LOONGARCH_CRYPTO = LOONGARCH_CRYPTO || info.features.CRYPTO;
LOONGARCH_LVZ = LOONGARCH_LVZ || info.features.LVZ;
LOONGARCH_LBT_X86 = LOONGARCH_LBT_X86 || info.features.LBT_X86;
LOONGARCH_LBT_ARM = LOONGARCH_LBT_ARM || info.features.LBT_ARM;
LOONGARCH_LBT_MIPS = LOONGARCH_LBT_MIPS || info.features.LBT_MIPS;
LOONGARCH_PTW = LOONGARCH_PTW || info.features.PTW;
#endif
}
@@ -35,7 +35,10 @@ LoongArch64RegCache::LoongArch64RegCache(MIPSComp::JitOptions *jo)
config_.totalNativeRegs = NUM_LAGPR + NUM_LAFPR;
// F regs are used for both FPU and Vec, so we don't need VREGs.
config_.mapUseVRegs = false;
config_.mapFPUSIMD = true;
// Every compiler in this backend picks its path from cpu_info.LOONGARCH_LSX, so the mapping
// has to agree with it. Claiming SIMD here while the scalar paths run leaves them addressing
// lanes with F(reg + n) against a mapping that has no per-lane registers.
config_.mapFPUSIMD = cpu_info.LOONGARCH_LSX;
}
void LoongArch64RegCache::Init(LoongArch64Emitter *emitter) {