Introduce entity_types.inc X-macro and entity_includes.h shared header
to eliminate ~1000 lines of repetitive #ifdef/entity-type blocks across
core files.
entity_types.inc defines each entity type once with two macros:
- ENTITY_TYPE_ for non-controller entities (button, infrared)
- ENTITY_CONTROLLER_TYPE_ for entities with controller callbacks
Each consumer includes the file with appropriate macro definitions.
Sites wanting all entities delegate ENTITY_CONTROLLER_TYPE_ to
ENTITY_TYPE_. Sites wanting only controller entities define
ENTITY_TYPE_ as empty.
entity_includes.h consolidates the conditional entity header includes
shared by application.h, controller.h, and controller_registry.h.
Also makes ComponentIterator::on_media_player pure virtual for
consistency (trivial override added to web_server::ListEntitiesIterator).
Add a new message-level option (inline_encode) that causes the code
generator to inline sub-message encoding directly into the parent's
encode/calculate_size methods instead of going through the
encode_sub_message function pointer indirection.
When set on a sub-message type, the generator:
- Inlines field encoding directly (no function pointer, no backpatch overhead)
- Inlines size calculation (no separate method call)
- Skips generating standalone encode/calculate_size methods
- Validates at generation time that max encoded size < 128 bytes
Applied to BluetoothLERawAdvertisement which is encoded 12 times per
BLE advertisement batch in a hot loop.
- uint32_to_str_(): raw pointer, internal use
- uint32_to_str(): template with compile-time buffer size check
- frac_to_str_(): raw pointer, internal use
- small_pow10(): simplify to ternary chain
Replace snprintf("%.*f") with integer-based formatting for finite
float values with accuracy_decimals 0-3 (covers virtually all sensor
usage). Falls back to snprintf for higher accuracy or NaN/Inf.
Uses lrint() with double cast for the multiply to match snprintf's
rounding behavior exactly. The fast path avoids snprintf's heavy
float formatting machinery entirely.
Also optimizes value_accuracy_with_uom_to_buf to append the UOM
string directly instead of going through snprintf.
Adds C++ unit tests that verify output matches snprintf for a range
of values including edge cases.
Benchmark: 92,961ns -> 6,484ns (14.3x faster, 2000 iterations).
Remove the per-byte `i < length - 1` check in the separator path by
writing the separator unconditionally after each byte and overwriting
the last one with the null terminator.
Single-loop approach keeps code smaller than the original while
eliminating the hot-path branch. Flash: -16 bytes vs baseline.
The loop split in format_hex_internal doubled the call sites for
format_hex_char, causing the compiler to outline it (+30 bytes new
symbol). Mark it always_inline to keep it inlined.
Split the single loop into two paths (with/without separator) to
eliminate the per-byte branch on separator. In the separator path,
write the separator unconditionally and overwrite the last one with
the null terminator. This also lets the compiler use constant stride
values (2 or 3) instead of a runtime variable.
Benchmarks show ~16-20% improvement on the separator path
(format_hex_pretty_to, format_mac_addr_upper).
Add benchmarks for format_hex_to, format_hex_pretty_to,
format_mac_addr_upper, fnv1_hash, fnv1a_hash, fnv1_hash_object_id,
parse_hex, crc8, crc16, value_accuracy_to_buf, int8_to_str, and
base64_decode to establish baseline performance data before optimization.
When the client sends 6+ messages during handshake (HelloReq, AuthReq,
GetTimeResp, SubscribeLogsReq, DeviceInfoReq, ListEntitiesReq),
MAX_MESSAGES_PER_LOOP=5 caused ListEntitiesReq to remain unread in
the socket buffer.
LWIP's rcvevent counter (used by is_socket_ready) tracks pbuf dequeues,
not remaining bytes. When multiple ESPHome messages share a single TCP
segment/pbuf, reading all but the last message exhausts rcvevent to 0
while the last message's data remains in LWIP's internal lastdata cache.
is_socket_ready() then returns false, and the data is never read — the
entity listing silently stalls until something else (like a keepalive
ping) triggers new socket activity.
Fix:
- Track when the read loop hits MAX_MESSAGES_PER_LOOP and retry on the
next iteration without the is_socket_ready() gate
- Increase MAX_MESSAGES_PER_LOOP from 5 to 10 (decode cost is now 3x
cheaper, and we batch 24-34 messages per write)
- Move set_nodelay_for_message(false) before try_to_clear_buffer in
process_batch_() so NODELAY is set before draining overflow