Typically, the inner loop of an interpreter is all about keeping the branch predictor happy and trying to fit in L1 cache as much as possible. So, falling through makes sense. I’m also curious whether a simple check of a flag and a conditional branch (which is highly predictable for a mode flag like whether you’re tracing or not) wouldn’t be faster than one of the indirect branches. That would keep the tracing code out of L1 when not in use (you can locate it in a remote function).
A `bool profile` branch stays expensive even when it predicts perfectly, because it bloats every opcode handler, and on a computed goto interpreter that extra code messes with the branch predictor history each dispatch site builds up, which is the whole reason dispatch is fast. Swapping the entire table dodges that. The part I like is fanning every opcode into one recording instruction and then back out through the real table, which is what keeps the second table from turning into a whole second interpreter like the earlier approach did.
```
```becomes: ```
```This way you could directly swap `dispatch_var` with an array populated with the `RECORD_INST_*` labels, and remove one step at runtime.
Or maybe this is what you are trying to avoid to reduce the binary size ?