softgpu: Only optimize rasterizer states while the workers are idle

OptimizePendingStates sat four lines below the `!tasksSplit_ || waitable_->Empty()`
check that makes touching shared state safe. It memcpys a 71-byte PixelFuncID over
a RasterizerState and swaps drawPixel/samplerID, while worker threads copy those
same entries by value to rasterize from - so a primitive could be drawn with the
new drawPixel against the old pixelID bytes.

Moving it inside the guard costs nothing correctness-wise: skipping a round just
means those draws use the unoptimized function, and the next Drain with an empty
waitable picks up the whole accumulated range.

This does not close the whole race - Add* still ORs into states_[stateIndex_].flags
after pushing, which items dispatched by an earlier Drain can be reading. That one
needs the state entries to become copy-on-write once dispatched, which is a bigger
change.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01Vd8ntC2brCUtCrDJMqLbs8
This commit is contained in:
Henrik RydgårdandClaude Opus 5 committed 2026-08-29 12:07:27 +02:00
1 parent f41fe77643
commit 0e62a81fb6
1 file changed
+9 -4
+9 -4
View File
@@ -509,11 +509,16 @@ void BinManager::Drain(bool flushing) {
}
tasksSplit_ = true;
}
// Let's try to optimize states, if we can.
OptimizePendingStates(pendingStateIndex_, stateIndex_);
pendingStateIndex_ = stateIndex_;
// Let's try to optimize states, if we can.
// This has to stay inside the drained check above: it memcpys a whole PixelFuncID over the
// state and swaps drawPixel/samplerID, and worker threads copy those same entries by value
// while rasterizing. Doing it with tasks in flight is a torn read waiting to happen. Skipping
// a round only means those draws run the unoptimized function, which is correct, just slower -
// the next Drain with an empty waitable picks up the whole accumulated range.
OptimizePendingStates(pendingStateIndex_, stateIndex_);
pendingStateIndex_ = stateIndex_;
}
if (taskRanges_.size() <= 1) {
PROFILE_THIS_SCOPE("bin_drain_single");