The wrappers called write_raw_inline_ which expanded the full fast
path code, defeating the purpose. Cold callers now call write_raw_slow_
directly with sent=-1, which is the same behavior without duplicating
the fast path at each cold call site.
Hot paths use write_raw_inline_ (ALWAYS_INLINE fast path).
Cold paths (handshake, error handling) use write_raw_ which is
an out-of-line wrapper that calls write_raw_inline_ — same logic,
no code duplication, but the compiler won't expand the fast path
at each cold call site.
Cold paths (write_frame_, bad indicator) no longer need to construct
an iovec just to call the slow path. The single-buffer overload wraps
in iovec internally, keeping call sites clean and avoiding iovec
setup in the caller's stack frame.
write_frame_ (handshake) and bad indicator response are cold paths
that don't benefit from the inlined write_raw_ fast path. Calling
write_raw_slow_ directly avoids expanding the inline fast path code
at these call sites, saving ~70 bytes per site.
Move the happy path (overflow empty + full write succeeds) inline into
the header so it gets inlined at each call site. Both overloads share
a single out-of-line write_raw_slow_ that handles partial writes,
errors, and overflow buffering.
Both write_raw_ overloads now have a clean fast path (direct socket
write) and delegate to a shared noinline enqueue_overflow_ for the
rare overflow case. The multi-iovec path now always calls writev
since single-message callers use the dedicated overload.
The existing write_raw_ takes iovec array + count, but single-message
writes (87-100% of traffic) always pass iovcnt=1. Adding a dedicated
overload that takes (data, len) eliminates iovec construction, the
iovcnt==1 branch, and pointer indirection on the hot path.
Instead of routing single messages through write_protobuf_messages and
branching internally, make write_protobuf_packet a separate virtual
override in each frame helper. The single-message path gets its own
minimal stack frame with no StaticVector allocation, and the batch
path in write_protobuf_messages has no size==1 branch — each caller
picks the right method upfront.
The noinline attribute on the batch path prevented the compiler from
inlining callees like write_plaintext_header and encrypt_noise_message_
into the loop body, causing a 5.6% regression on batch writes. Adding
flatten forces all callees to be inlined within the batch function while
noinline still keeps its large stack frame separate from the single-message
fast path.
Each of these methods called encode_field_raw() then a value encoder,
causing a store-load pair on pos_ between the tag and value writes.
Inline the tag write into the same __restrict__ local scope as the
value write so the compiler can emit tag + value with a single pos_
load at start and store at end. Verified on Xtensa: encode_uint32
now does one load + two writes + one store (was load-store-load-store).
Add sync_debug_check_bounds_() that syncs pos_ from a local pointer
before checking bounds. Use it in encode_varint_raw_64 and
encode_varint_raw_slow_ to restore per-byte bounds checking in
debug mode without breaking the __restrict__ optimization in
production (where it's a no-op).
Previously encode_string called encode_varint_raw(len) then
encode_raw(data, len) as separate methods, each with their own
__restrict__ pos scope. This caused a redundant store-load pair
of pos_ between the two operations.
Inline the length varint write and memcpy under a single local
pos variable so the compiler can keep pos_ in a register across
both operations. Eliminates one load-store pair per string encode.
Apply the same __restrict__ local pointer pattern proven in
encode_varint_raw_64 to all remaining methods that write through
pos_: encode_varint_raw, encode_varint_raw_short, write_raw_byte,
encode_raw, write_tag_and_fixed32, encode_string, encode_bool,
and encode_fixed32.
Each method now hoists pos_ into a __restrict__ local before
writing and stores back once at the end. When the compiler inlines
these into a generated encode() method, it can keep pos_ in a
register across consecutive calls instead of reloading from memory
after every write.
Same optimization as encode_varint_raw_64: hoist pos_ into a
__restrict__ local so the compiler keeps it in a register across
the loop instead of reloading from memory each iteration.
Use a __restrict__ local pointer for the varint write loop so the
compiler can keep it in a register instead of reloading pos_ from
memory on each iteration. Eliminates the store→load dependency
chain that was causing 7 load/store pairs for a typical 48-bit
BLE address varint.
Use a __restrict__ local pointer for the varint write loop so the
compiler can keep it in a register instead of reloading pos_ from
memory on each iteration. Eliminates the store→load dependency
chain that was causing 7 load/store pairs for a typical 48-bit
BLE address varint.
Add encode_varint_raw_short() and ProtoSize::varint_short() that
inline both the 1-byte and 2-byte varint paths, falling back to
the noinline slow path for 3+ bytes.
Use these for sint32 fields (zigzag encoding), where values like
RSSI (-100 to 0) produce zigzag values that are 1-2 bytes. This
avoids a function call for the common case without bloating the
generic encode_varint_raw fast path.
Add CodSpeed benchmarks for BluetoothLERawAdvertisementsResponse
(12 advertisements) covering calculate_size, encode, calc+encode,
and fresh-buffer paths.
Includes a lightweight bluetooth_proxy stub header in
tests/benchmarks/stubs/ so the api component can compile with
USE_BLUETOOTH_PROXY on the host platform without pulling in
ESP32 BLE dependencies.
Remove direct include of build_info_data.h from version_text_sensor.cpp
and use the existing App.get_build_time_string() API instead. This
eliminates duplicate ESPHOME_BUILD_TIME_STR and ESPHOME_COMMENT_STR
symbols that were emitted into every translation unit including the
header (static const in a header = one copy per TU).
Now only application.cpp includes build_info_data.h, which also
improves incremental rebuild times: when build info changes, only
application.cpp needs recompiling instead of both application.cpp
and version_text_sensor.cpp.