Add a new message-level option (inline_encode) that causes the code
generator to inline sub-message encoding directly into the parent's
encode/calculate_size methods instead of going through the
encode_sub_message function pointer indirection.
When set on a sub-message type, the generator:
- Inlines field encoding directly (no function pointer, no backpatch overhead)
- Inlines size calculation (no separate method call)
- Skips generating standalone encode/calculate_size methods
- Validates at generation time that max encoded size < 128 bytes
Applied to BluetoothLERawAdvertisement which is encoded 12 times per
BLE advertisement batch in a hot loop.
Remove the per-byte `i < length - 1` check in the separator path by
writing the separator unconditionally after each byte and overwriting
the last one with the null terminator.
Single-loop approach keeps code smaller than the original while
eliminating the hot-path branch. Flash: -16 bytes vs baseline.
The loop split in format_hex_internal doubled the call sites for
format_hex_char, causing the compiler to outline it (+30 bytes new
symbol). Mark it always_inline to keep it inlined.
Split the single loop into two paths (with/without separator) to
eliminate the per-byte branch on separator. In the separator path,
write the separator unconditionally and overwrite the last one with
the null terminator. This also lets the compiler use constant stride
values (2 or 3) instead of a runtime variable.
Benchmarks show ~16-20% improvement on the separator path
(format_hex_pretty_to, format_mac_addr_upper).
Add benchmarks for format_hex_to, format_hex_pretty_to,
format_mac_addr_upper, fnv1_hash, fnv1a_hash, fnv1_hash_object_id,
parse_hex, crc8, crc16, value_accuracy_to_buf, int8_to_str, and
base64_decode to establish baseline performance data before optimization.
When the client sends 6+ messages during handshake (HelloReq, AuthReq,
GetTimeResp, SubscribeLogsReq, DeviceInfoReq, ListEntitiesReq),
MAX_MESSAGES_PER_LOOP=5 caused ListEntitiesReq to remain unread in
the socket buffer.
LWIP's rcvevent counter (used by is_socket_ready) tracks pbuf dequeues,
not remaining bytes. When multiple ESPHome messages share a single TCP
segment/pbuf, reading all but the last message exhausts rcvevent to 0
while the last message's data remains in LWIP's internal lastdata cache.
is_socket_ready() then returns false, and the data is never read — the
entity listing silently stalls until something else (like a keepalive
ping) triggers new socket activity.
Fix:
- Track when the read loop hits MAX_MESSAGES_PER_LOOP and retry on the
next iteration without the is_socket_ready() gate
- Increase MAX_MESSAGES_PER_LOOP from 5 to 10 (decode cost is now 3x
cheaper, and we batch 24-34 messages per write)
- Move set_nodelay_for_message(false) before try_to_clear_buffer in
process_batch_() so NODELAY is set before draining overflow
The multi-message batch path in process_batch_multi_ bypassed
set_nodelay_for_message(), so batch data written to the socket would
sit in LWIP's Nagle buffer when log messages had previously enabled
Nagle. The remote wouldn't ACK until it sent its own data (e.g. a
ping), which could take 20+ seconds. This caused log-only API clients
(like esphome logs) to time out waiting for ListEntitiesDoneResponse.
Additionally, when LIST_ENTITIES completed without a state subscription,
the batched entity listing responses were never explicitly flushed —
they waited for the 100ms batch timer instead of being sent immediately
like the INITIAL_STATE completion path already did.
Fixes both issues:
- Call set_nodelay_for_message(false) before write_protobuf_messages
in process_batch_multi_ to ensure TCP_NODELAY is on
- Extract finalize_iterator_sync_() helper and call it when
LIST_ENTITIES completes without state_subscription