decode_case() reads wire_type instead of taking it at every call, the
expression and the store statement are two small hooks that repeated
fields override, and message fields build one body. No generated case
declares a local any more, so the braced form and its test go; the
compiler rejects a jump over a local if one ever appears. The wire
type test now proves a dropped frame with an ordering marker instead
of assuming it, and shares StateWaiter.
With one decode_field() switch per message, the three per wire type
content properties only differed in the attribute they read; a single
decode_content built from decode_expr() replaces them, and repeated
fields reuse the element type's expression. Case bodies with several
statements get their block from the body itself instead of a caller
flag, the fixed byte array body copies straight from the payload
instead of through a heap std::string, and the decode comments no
longer restate the switch keying explained next to the macros.
A case body with a declaration or several statements now gets its own
block, as the per wire type overrides had, so no jump to a later case
label crosses an initialization.
CodSpeed showed the single virtual costing 7 to 18 percent on the
decode benchmarks. The x86-64 disassembly pointed at the call, not the
switch: passing the field number and wire type alongside the tag plus
a 16 byte union payload kept five values live across the call, so the
compiler spilled this, the end pointer and half of the payload to the
stack and reloaded them for every field.
decode_field() now takes only the tag, the payload pointer (already
the loop cursor) and one scalar that holds the varint or fixed32 value
or the payload length. The generated override wraps them in a
ProtoFieldValue that never exists in memory. On the host the switch
key is the field number derived with one shift and the guard compares
the whole tag against the constant the case declares, which is the
same two instructions the old per wire type dispatch cost.
The loop also handles single byte varints inline instead of going
through the parse result struct, which drops the materialized consumed
count and its add on every tag and small value.
Every decodable message overrode up to three virtuals, one per wire
type, so each carried a five slot vtable and up to three functions
with their own prologue and return tails. The shared decode loop now
parses the payload for the wire type into a ProtoFieldValue and calls
a single decode_field() virtual with the tag, the field number and the
wire type; the generated override is one switch.
The switch key is chosen per target through PROTO_DECODE_KEY. Embedded
builds compile switches to compare chains (ESP-IDF passes
-fno-jump-tables), so they key on the full wire tag, one compare per
field with no separate wire type check. The host compiler builds a
jump table for the dense field number switch, so there the key is the
field number and PROTO_DECODE_GUARD rejects a mismatched wire type.
Both forms drop a field that arrives with a wire type it does not
declare, exactly as the per wire type virtuals did.
Per decodable message the vtable shrinks from 20 to 12 bytes on
xtensa and the extra decode functions fold into one; the shared loop
shrinks as well. Host instruction counts per decoded field are
unchanged apart from the guard compare, which replaces the prologue of
the separate function it used to call.
Every message sent through send_message or the entity paths needed a
proto_encode_msg<T> thunk (17 bytes on xtensa) and, for entity state
and info messages, a calc_size<T> thunk, because the generated encode
and calculate_size were member functions and the connection code wants
plain function pointers over const void *.
The generator now emits the bodies as static encode_msg(const void *)
and calc_size_msg(const void *) functions, so &T::encode_msg is
already a MessageEncodeFn and the thunks disappear. The member
encode() and calculate_size() remain as inline forwarders for direct
callers. ProtoMessage carries the same static defaults for messages
without fields, which also removes the separate no-op encode thunk.
A call that drops the returned cursor would silently truncate the message, so the
compiler now warns on it and a unit test scans the generated file for the same mistake.
Also corrects the outlining comment for ESP8266, where the inline write is a few byte
stores rather than one, and the RAW_ENCODE_MAP annotation.
_encode_call() owns the cursor assignment and the _force suffix, so
the convention lives in one place instead of at every emission site;
the fixed32 fast path is an arm of the generic encode_content keyed by
a per type value template. write_fixed32_le uses convert_little_endian
instead of its own byte order switch. The integration test shares a
StateWaiter from state_utils and leaves the disconnect to the fixture.
One helper next to the other precomputed tag paths decides how a
single byte tag fixed32 field is written; the float and fixed32 types
only differ in the value expression. Drop the non forced std::string
encode_string overload, which the generator never emits, and build the
generator tests from one block of field type constants.
The ProtoEncode helpers took the write cursor by reference and a
bool force flag. At -Os the compiler outlines most of them, so every
call site had to keep pos in a stack slot and pass its address, plus
a constant for the flag. The helpers now take the cursor by value and
return the advanced cursor, so consecutive calls chain through the
return register; forced fields call a _force overload instead of
passing a flag.
The fixed32 writers use __builtin_memcpy, which stays a builtin under
ESP-IDF's -fno-builtin-memcpy, and are outlined on embedded targets so
each fixed32 or float field is a short call instead of an inline
memcpy call. Non-forced float and fixed32 fields with a single-byte
tag share the same writer behind a zero check.
Generated encode bodies shrink by 18 percent on an ESP32 IDF proxy
build (2360 to 1932 bytes for 27 messages); entity messages gain the
most, for example ListEntitiesSensorResponse::encode 190 to 134 bytes
and SensorStateResponse::encode 78 to 49 bytes.