Replace two out-of-line encode_varint_to_buffer calls with direct
byte writes using the already-computed varint lengths. Eliminates
function call overhead (register save/restore) from the batch loop
and removes the 34-byte out-of-line encode_varint_to_buffer function
which had no other callers.
Pre-compute the first message's header length before the loop to
initialize write_start/write_end. The first loop iteration naturally
skips memmove since src == write_end.
Extract varint_encoded_length_16/8 and plaintext_header_length as
reusable inline helpers from write_plaintext_header.
Use a single loop for all messages instead of separate first-message
and loop paths. The first iteration skips memmove via the null check.
Eliminates duplicated write_plaintext_header inlining, reducing flash
from 326 to 229 bytes (-30%) while keeping the 64-byte stack frame.
Replace StaticVector<iovec> + writev() scatter-gather in the batch write
path with contiguous single-buffer write() calls.
Plaintext: compact messages via memmove to close 0-3 byte varint header
gaps, then write_raw_fast_buf_. Noise: messages are already contiguous
(fixed 7-byte header + 16-byte MAC fills all reserved space), switch
directly to write_raw_fast_buf_.
Fix LOG_PACKET_SENDING to log after write/enqueue to prevent re-entrant
log sends from corrupting the shared buffer before data is sent.
Consolidate the macro to api_frame_helper.cpp and expose via out-of-line
log_packet_sending_() helper.
Remove write_raw_fast_iov_ (no remaining callers) and change
encrypt_noise_message_ to return uint16_t length instead of iovec.
Move get_component_log_str() from component.cpp into the header
as an always_inline method. This is a trivial one-liner that just
calls component_source_lookup(), but the compiler was keeping it
out-of-line across 8 call sites (13 bytes).
Saves 8 bytes on ESP32.
Inline record_runtime_stats_() to eliminate the non-inlined function
call overhead per component per loop iteration. Extract check_blocking_()
as an inline helper for readability.
The millis() return value stays in the caller's clock domain — only
the micros()-based stats recording is inlined alongside it.
Use a single micros() call in WarnIfComponentBlockingGuard::finish()
for both runtime stats recording and blocking detection. Previously
finish() called millis() for blocking detection and a separate
non-inlined record_runtime_stats_() that called micros() again.
Changes:
- Inline record_runtime_stats_() to avoid function call overhead
- Derive curr_time from micros()/1000 instead of separate millis() call
- Override started_ from micros()/1000 in constructor so both timestamps
use the same clock source (fixes unsigned underflow on platforms where
millis() and micros()/1000 can diverge)
- Extract check_blocking_() as inline helper for readability