mirror of
https://github.com/esphome/esphome.git
synced 2026-09-11 15:27:33 +00:00
CodSpeed showed the single virtual costing 7 to 18 percent on the decode benchmarks. The x86-64 disassembly pointed at the call, not the switch: passing the field number and wire type alongside the tag plus a 16 byte union payload kept five values live across the call, so the compiler spilled this, the end pointer and half of the payload to the stack and reloaded them for every field. decode_field() now takes only the tag, the payload pointer (already the loop cursor) and one scalar that holds the varint or fixed32 value or the payload length. The generated override wraps them in a ProtoFieldValue that never exists in memory. On the host the switch key is the field number derived with one shift and the guard compares the whole tag against the constant the case declares, which is the same two instructions the old per wire type dispatch cost. The loop also handles single byte varints inline instead of going through the parse result struct, which drops the materialized consumed count and its add on every tag and small value.