Four tests had their own copy of "call this until N seconds have passed,
then divide". CallsPerSecond in UnitTest.h does it; each keeps its old
duration and batch size.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
UnitTest.h's EXPECT_ macros all call printf and EXPECT_EQ_MEM calls memcmp, but
it included neither <cstdio> nor <cstring> - it has been relying on whatever the
including file happened to pull in first, and TestMpegCsc was the first not to.
The same shape in sceMpegbase.cpp and sceVideocodec.cpp, which use std::min,
std::move and memcpy without saying where they come from.
Also drop an abs() from TestMpegCsc rather than include <cstdlib> for one
subtraction.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The planes the de-tiling produces are already the YUV420P swscale wants, and our
sceMpeg HLE converts the same frames the same way, so the pixel formats and the
studio-range setup come straight from MediaEngine::getSwsFormat. It is 3-4x
quicker than going a pixel at a time: 0.37-0.44ms a frame becomes 0.09-0.12ms,
which is the whole reason sceMpegBaseCscAvc was at the top of a profile.
Chroma is upsampled with SWS_POINT rather than the HLE's SWS_BILINEAR, since
replicating is what the scalar path does and, being a fixed-function block,
almost certainly what the hardware does.
It is not bit-identical - swscale rounds its own way. TestMpegCsc measures the
gap per channel rather than per byte, so the number means something for a packed
16-bit pixel: worst 1 step of 31 for 5650 and 5551, 2 of 15 for 4444, 3 of 255
for 8888, with means around a fifth of a step. The scalar path stays as what the
longhand reference is checked against, and takes anything swscale won't - an odd
range origin, or a build without ffmpeg.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The colour conversion was writing alpha fully set - 0xFF000000, or the top bit
for 5551 - where the hardware writes zero. Our sceMpeg HLE already masks it off
and names Sword Art Online as a game that depends on it: it doesn't clear the
alpha in the buffer it hands over, and expects the video not to set it. The two
paths now agree.
The de-tiling ahead of it becomes UntileYCbCr, taking the eight buffers already
resolved, so it can be measured and compared against the original longhand
version in TestMpegCsc. Its planes move to scratch that persists between calls -
a movie converts one frame per displayed frame, and this was allocating and
clearing about 200KB every time - and the per-pixel bounds checks in the chroma
loop, which only depend on the group of eight, are hoisted out of it.
That last part is worth 2529 -> 3201 MPix/s, but the point of measuring was to
find out whether it mattered, and it doesn't much: de-tiling is 0.04ms of a
frame against the conversion's 0.4ms. The conversion is where the time is.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
sceMpegBaseCscAvc is the top of a profile during video playback, so the loop
that does the work becomes MpegCscRange - a pure function with the HLE plumbing
left behind - and TestMpegCsc measures and checks it.
The measuring half reports megapixels per second for a 480x272 frame in each of
the four pixel formats. The checking half compares against the conversion
written out longhand, over whole frames and over partial ranges with odd offsets
and sizes, plus one-pixel, one-row and one-column ranges and one that reaches
the far edge of the frame. Those are the cases an optimized version gets wrong:
chroma is half resolution, so an odd left edge starts mid-sample, and anything
handling two pixels at a time has to deal with the leftover. The destination is
padded and prefilled, so writing outside the range fails too.
This is only the move - the loop is the same one, so the numbers it gives are
the baseline to improve on. On a Snapdragon X Elite it runs at about 300 MPix/s,
0.44ms for a frame.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>