IR interpreter: Use threaded dispatch with GCC and Clang

Every op ended with a break back to one shared indirect jump, which the
CPU has to predict for every op in the program. With labels as values,
each op jumps through a table from its own site instead, which predicts
much better: an integer-heavy benchmark runs about 13% faster on an M1.
Other compilers keep the switch, and ops missing from the table fall back
to it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
Henrik RydgårdandClaude Opus 5.5 committed 2026-09-25 11:47:17 -06:00
1 parent 7eb371b231
commit d1c04d70a6
1 file changed
+570 -365
File diff suppressed because it is too large. Load diff