Headless: turn the draw frame over each emulated frame, not once per run

A long headless run on Vulkan dies in VulkanPushPool::CreateBlock. Watching the
allocator, it makes a fresh 8MB block roughly twice a second and garbage
collects none of them - about 13MB a second of device memory, which runs out
after a minute or two.

Nothing is leaking as such. The push buffers are recycled by BeginFrame, which
walks the blocks belonging to the current frame index and marks them unused.
Headless called draw->BeginFrame() once before the run loop and draw->EndFrame()
once after, so that recycling pass ran exactly once for the whole run and every
allocation after the first had to take a new block.

This is the same mistake one level up from the host frame, which already turns
over per emulated frame for the same reason - the comment there says a single
host frame spanning the run meant the texture cache and framebuffer manager
never decayed anything. The draw context needs the same treatment, nested the
way the app nests them: draw frame outside, host frame inside.

Verified on a two-minute Tekken 6 run: new blocks created goes from around 200
to zero, and the run ends on its timeout instead of asserting. Framedump
rendering tests are unchanged - the same 23 of 30 fail before and after, which
is a separate pre-existing matter on this platform.

Also taught frametests.py to look for an ARM64 build, gated on the machine's own
architecture the way test.py already does. It was picking a stale x64 Debug
binary, which is exactly the trap that makes a rendering comparison meaningless.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
This commit is contained in:
Henrik RydgårdandClaude Opus 5 committed 2026-09-21 15:47:40 -06:00
1 parent 333d84935e
commit 23278cb2fd
2 files changed
+21 -3

No files matched your search

+7 -1
View File
@@ -28,13 +28,19 @@ import re
import shlex
import shutil
import subprocess
import platform
import sys
import time
from pathlib import Path
# test.py-style candidate paths for the headless binary, relative to the
# current working directory, in preference order.
HEADLESS_CANDIDATES = [
# The machine's own architecture comes first, the same way test.py picks: an x64 build runs on
# Windows-on-ARM too, under emulation, so looking for it first quietly tests the emulated build.
HEADLESS_CANDIDATES = ([
"Windows/ARM64/Debug/PPSSPPHeadless.exe",
"Windows/ARM64/Release/PPSSPPHeadless.exe",
] if platform.machine().lower() in ("arm64", "aarch64") else []) + [
"Windows/x64/Debug/PPSSPPHeadless.exe",
"Windows/Debug/PPSSPPHeadless.exe",
"Windows/x64/Release/PPSSPPHeadless.exe",
+14 -2
View File
@@ -454,12 +454,24 @@ static bool RunAutoTest(GraphicsContext *graphicsContext, CoreParameter &corePar
if (coreState == CORE_NEXTFRAME) {
// INFO_LOG(Log::System, "(frame)");
coreState = CORE_RUNNING_CPU;
// Close and reopen the host frame, which is what the app does once per displayed
// frame. All the GPU's per-frame work hangs off BeginHostFrame - the texture cache's
// Close and reopen the frame, which is what the app does once per displayed frame.
// All the GPU's per-frame work hangs off BeginHostFrame - the texture cache's
// StartFrame and the framebuffer manager's DecimateFBOs - so with a single host frame
// spanning the whole run, none of it ever ran here, and a long test decayed nothing.
//
// The draw context's frame has to turn over too, and for the same reason one level up:
// Vulkan's push buffers are recycled by BeginFrame, so one frame spanning the run means
// nothing is ever reused and every allocation takes a fresh 8MB block - about 13MB a
// second, which runs a long test out of device memory. Draw frame outside, host frame
// inside, the way the app nests them.
if (gpu) {
gpu->EndHostFrame();
}
if (draw) {
draw->EndFrame();
draw->BeginFrame(Draw::DebugFlags::NONE);
}
if (gpu) {
gpu->BeginHostFrame(g_Config.GetDisplayLayoutConfig(DeviceOrientation::Landscape));
}
}