Files
Henrik RydgårdandClaude Opus 5.5 d1c04d70a6 IR interpreter: Use threaded dispatch with GCC and Clang
Every op ended with a break back to one shared indirect jump, which the
CPU has to predict for every op in the program. With labels as values,
each op jumps through a table from its own site instead, which predicts
much better: an integer-heavy benchmark runs about 13% faster on an M1.
Other compilers keep the switch, and ops missing from the table fall back
to it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-25 11:47:17 -06:00
..
2026-09-24 16:35:11 -06:00
…