- Scheduler_Defer: use nullptr name matching Component::defer(func)
production pattern (skips cancel_item_locked_ entirely)
- Scheduler_Defer_SameID: fixed ID 0 measuring cancel-and-replace
pattern for coalescing rapid updates
- Scheduler_Defer_UniqueID: unique IDs measuring cancel scan overhead
when no match is found
Change kInnerIterations from 2000 to 2100 (divisible by batch sizes 3
and 10) to prevent pool imbalance at iteration boundaries that caused
spurious malloc. Add static_assert to each benchmark to catch this at
compile time.
Replace duplicated pool warmup blocks with a shared warm_pool() helper
that registers and replaces items twice to populate the recycling pool
before the benchmark loop begins.
Add Scheduler_SetTimeout_ExceedPool with batch size 10 (exceeding
MAX_POOL_SIZE=5) to measure the performance impact when the recycling
pool is exhausted and items must be malloc'd/freed each cycle.
Instead of draining after every single registration or batching 5,
use a batch size of 3 which represents a realistic worst case where
multiple components schedule in the same loop iteration while staying
within the recycling pool (MAX_POOL_SIZE=5).
Scheduler registration benchmarks (SetTimeout, SetInterval, Defer) were
not calling scheduler.call() periodically to drain and clean up cancelled
items. In production, call() runs every loop iteration, keeping the
scheduler containers small. Without draining, cancelled items accumulated
causing O(n²) scan cost in cancel_item_locked_ that doesn't reflect
real-world behavior.
- SetTimeout: was only calling process_to_add() (no cleanup), now calls
call() every kKeyCount iterations
- SetInterval: was calling process_to_add() (no cleanup of items_), now
calls call() for proper cleanup
- Defer: was never draining the defer queue, now calls call() to process
deferred items as production does
- All three now advance time (++now) to match production loop behavior
Co-authored-by: pre-commit-ci-lite[bot] <117423508+pre-commit-ci-lite[bot]@users.noreply.github.com>
Co-authored-by: J. Nick Koston <nick+github@koston.org>
Co-authored-by: J. Nick Koston <nick@koston.org>
Replace two out-of-line encode_varint_to_buffer calls with direct
byte writes using the already-computed varint lengths. Eliminates
function call overhead (register save/restore) from the batch loop
and removes the 34-byte out-of-line encode_varint_to_buffer function
which had no other callers.
Pre-compute the first message's header length before the loop to
initialize write_start/write_end. The first loop iteration naturally
skips memmove since src == write_end.
Extract varint_encoded_length_16/8 and plaintext_header_length as
reusable inline helpers from write_plaintext_header.
Use a single loop for all messages instead of separate first-message
and loop paths. The first iteration skips memmove via the null check.
Eliminates duplicated write_plaintext_header inlining, reducing flash
from 326 to 229 bytes (-30%) while keeping the 64-byte stack frame.
Replace StaticVector<iovec> + writev() scatter-gather in the batch write
path with contiguous single-buffer write() calls.
Plaintext: compact messages via memmove to close 0-3 byte varint header
gaps, then write_raw_fast_buf_. Noise: messages are already contiguous
(fixed 7-byte header + 16-byte MAC fills all reserved space), switch
directly to write_raw_fast_buf_.
Fix LOG_PACKET_SENDING to log after write/enqueue to prevent re-entrant
log sends from corrupting the shared buffer before data is sent.
Consolidate the macro to api_frame_helper.cpp and expose via out-of-line
log_packet_sending_() helper.
Remove write_raw_fast_iov_ (no remaining callers) and change
encrypt_noise_message_ to return uint16_t length instead of iovec.