The dead size parameter goes, the string type reads the inherited flag, the forced short
string path asserts against it, and the generated declarations say why the pointer may be
null.
A StringRef field that is only ever encoded is skipped when empty before its pointer is
read, and the dump helper checks empty() first, so pointing it at "" buys nothing while
costing one store per field in every message constructor. Fields that are decoded or force
encoded keep the empty string default.
The barrier was a command on the aioesphomeapi connection, which nothing orders against the
raw frames under test; a well formed frame on the raw socket itself now closes each block and
the assertion checks the exact states seen since the marker.
Three malformed frames (a tag with a dangling continuation bit, a length prefix past the
payload, a two byte fixed32) must stop the decode loop without taking the connection down;
a generator test records that a double field is rejected before it could reach the loop.
Repeated fields encode their elements through one encode_element() hook instead of two
isinstance ladders, the fixed32 precomputed tag path owns its own guard, the generated
switches drop the dead default case, StringRef takes the byte pointer directly, the
three hand written tag expressions in proto.h go through proto_tag(), and stale comments
about the previous decode design go. The compiled functions are unchanged.
The out parameter form regressed the host, where the 16 byte result already travels in
registers, by about 20 percent on the direct varint parse benchmarks. The void
decode_field and low word bool changes stay.
decode_field() no longer returns a bool that only fed a verbose log; unknown fields are
skipped silently like every other protobuf decoder does, and each message loses the
return value materialisation. Bools read the low 32 bits of the varint, which drops the
second compare on 64 bit varint builds. The multi byte varint path writes its value
through an out parameter instead of returning a 16 byte struct, which takes the spills
out of the decode loop and count_repeated_field.
A type now sets a single decode_expr; the wire type it already declares picks the case
label. Drops the unused force_str helper, two dead decode_length overrides, a duplicate
field builder in the generator tests and a needless list copy in the state waiter. The
generated files are unchanged.
ESP-IDF passes -fno-builtin-memcpy, so the four byte memcpy in the decode loop was an
out of line call on every fixed32 field; host compilers fold the byte loads back into one
load. Checking the varint wire type first keeps the common path to one taken branch on
xtensa.
decode_case() reads wire_type instead of taking it at every call, the
expression and the store statement are two small hooks that repeated
fields override, and message fields build one body. No generated case
declares a local any more, so the braced form and its test go; the
compiler rejects a jump over a local if one ever appears. The wire
type test now proves a dropped frame with an ordering marker instead
of assuming it, and shares StateWaiter.
Hand built frames check that decode_field() takes a field with its
declared wire type, drops the same field sent length delimited or as
fixed32, ignores a varint key, and skips an unknown field with a two
byte tag before decoding the rest. Client commands cover two byte tags
and varints, a two byte length prefix and a negative fixed32 float.
With one decode_field() switch per message, the three per wire type
content properties only differed in the attribute they read; a single
decode_content built from decode_expr() replaces them, and repeated
fields reuse the element type's expression. Case bodies with several
statements get their block from the body itself instead of a caller
flag, the fixed byte array body copies straight from the payload
instead of through a heap std::string, and the decode comments no
longer restate the switch keying explained next to the macros.
A case body with a declaration or several statements now gets its own
block, as the per wire type overrides had, so no jump to a later case
label crosses an initialization.
CodSpeed showed the single virtual costing 7 to 18 percent on the
decode benchmarks. The x86-64 disassembly pointed at the call, not the
switch: passing the field number and wire type alongside the tag plus
a 16 byte union payload kept five values live across the call, so the
compiler spilled this, the end pointer and half of the payload to the
stack and reloaded them for every field.
decode_field() now takes only the tag, the payload pointer (already
the loop cursor) and one scalar that holds the varint or fixed32 value
or the payload length. The generated override wraps them in a
ProtoFieldValue that never exists in memory. On the host the switch
key is the field number derived with one shift and the guard compares
the whole tag against the constant the case declares, which is the
same two instructions the old per wire type dispatch cost.
The loop also handles single byte varints inline instead of going
through the parse result struct, which drops the materialized consumed
count and its add on every tag and small value.
Every decodable message overrode up to three virtuals, one per wire
type, so each carried a five slot vtable and up to three functions
with their own prologue and return tails. The shared decode loop now
parses the payload for the wire type into a ProtoFieldValue and calls
a single decode_field() virtual with the tag, the field number and the
wire type; the generated override is one switch.
The switch key is chosen per target through PROTO_DECODE_KEY. Embedded
builds compile switches to compare chains (ESP-IDF passes
-fno-jump-tables), so they key on the full wire tag, one compare per
field with no separate wire type check. The host compiler builds a
jump table for the dense field number switch, so there the key is the
field number and PROTO_DECODE_GUARD rejects a mismatched wire type.
Both forms drop a field that arrives with a wire type it does not
declare, exactly as the per wire type virtuals did.
Per decodable message the vtable shrinks from 20 to 12 bytes on
xtensa and the extra decode functions fold into one; the shared loop
shrinks as well. Host instruction counts per decoded field are
unchanged apart from the guard compare, which replaces the prologue of
the separate function it used to call.
Every message sent through send_message or the entity paths needed a
proto_encode_msg<T> thunk (17 bytes on xtensa) and, for entity state
and info messages, a calc_size<T> thunk, because the generated encode
and calculate_size were member functions and the connection code wants
plain function pointers over const void *.
The generator now emits the bodies as static encode_msg(const void *)
and calc_size_msg(const void *) functions, so &T::encode_msg is
already a MessageEncodeFn and the thunks disappear. The member
encode() and calculate_size() remain as inline forwarders for direct
callers. ProtoMessage carries the same static defaults for messages
without fields, which also removes the separate no-op encode thunk.
The fixed32 store helper moves to a private section since it neither bounds checks nor
advances the cursor, its comment describes the path each target takes, the generated
file scan flags any ProtoEncode call that does not assign the cursor, and StateWaiter
timeouts can carry a label so gathered waits are told apart.
Cortex-M0+ and ARM9 turn the four byte unaligned store into a memcpy call with a stack
temporary at every fixed32 field, and the outlined helper itself became a memcpy call
there, so the helper now spells out the byte stores. Xtensa and host objects are byte for
byte unchanged; on the RP2040 bench config the api object loses 28 bytes and the fixed32
memcpy calls.
A call that drops the returned cursor would silently truncate the message, so the
compiler now warns on it and a unit test scans the generated file for the same mistake.
Also corrects the outlining comment for ESP8266, where the inline write is a few byte
stores rather than one, and the RAW_ENCODE_MAP annotation.
On the ESP8266 the inline write was already a single store, so the
outlined helper cost a call per fixed32 field: sensor state encode went
from 615 to 864 ns on a d1 mini. ESP32 builds pass -fno-builtin-memcpy,
where the shared copy is both smaller and faster (562 to 328 ns on an
atom), so the gate is now USE_ESP32.
_encode_call() owns the cursor assignment and the _force suffix, so
the convention lives in one place instead of at every emission site;
the fixed32 fast path is an arm of the generic encode_content keyed by
a per type value template. write_fixed32_le uses convert_little_endian
instead of its own byte order switch. The integration test shares a
StateWaiter from state_utils and leaves the disconnect to the fixture.
Covers a zero float that is skipped on the wire, a fixed32 state, a
negative int32, list entity strings and text states whose length
prefix needs two varint bytes, a two byte field tag through the
device info area, and the field free disconnect exchange.
One helper next to the other precomputed tag paths decides how a
single byte tag fixed32 field is written; the float and fixed32 types
only differ in the value expression. Drop the non forced std::string
encode_string overload, which the generator never emits, and build the
generator tests from one block of field type constants.
The macro only exists for the two fixed32 writers in ProtoEncode, so
drop it once the class is complete instead of leaking it into every
translation unit that includes proto.h.
The ProtoEncode helpers took the write cursor by reference and a
bool force flag. At -Os the compiler outlines most of them, so every
call site had to keep pos in a stack slot and pass its address, plus
a constant for the flag. The helpers now take the cursor by value and
return the advanced cursor, so consecutive calls chain through the
return register; forced fields call a _force overload instead of
passing a flag.
The fixed32 writers use __builtin_memcpy, which stays a builtin under
ESP-IDF's -fno-builtin-memcpy, and are outlined on embedded targets so
each fixed32 or float field is a short call instead of an inline
memcpy call. Non-forced float and fixed32 fields with a single-byte
tag share the same writer behind a zero check.
Generated encode bodies shrink by 18 percent on an ESP32 IDF proxy
build (2360 to 1932 bytes for 27 messages); entity messages gain the
most, for example ListEntitiesSensorResponse::encode 190 to 134 bytes
and SensorStateResponse::encode 78 to 49 bytes.