- script/build_helpers.py: when injecting a non-MULTI_CONF component
into the post-validation config, run its CONFIG_SCHEMA with {} so
defaults are populated. Without this, socket got config = {} and
socket.FILTER_SOURCE_FILES crashed with KeyError on
'implementation' (the schema's defaulted key was never filled in).
Falls back to {} if the schema can't validate empty input.
- tests/components/json/__init__.py: enable codegen for json so its
to_code runs during cpp unit test builds, registering the
ArduinoJson library. Required for any api dep test, since
json_util.cpp #includes <ArduinoJson.h>.
Locally verified 'script/cpp_unit_test.py api' now compiles and runs;
ProtoMacVarint test suite (9 cases) passes.
Three fixes so 'script/cpp_unit_test.py api' actually compiles instead
of crashing in build setup:
1. script/build_helpers.py: when adding transitive component
dependencies to the post-validation config, use {} (dict) instead
of [] (list) for non-MULTI_CONF components. socket's
FILTER_SOURCE_FILES (and any other code that subscripts
CORE.config[component] with a string key) was crashing because
socket got config = [] from setdefault.
2. esphome/components/api/api_pb2_service.cpp + the codegen in
script/api_protobuf/api_protobuf.py: wrap the generated
APIConnection::read_message_ definition in #ifdef USE_API. The
class itself is only declared inside #ifdef USE_API in
api_connection.h, so without the guard the .cpp fails to compile
in any build that pulls in the api source files without setting
USE_API (e.g. cpp unit tests of api dependencies).
- Test file: declare proto_debug_end_ locally instead of misusing
PROTO_ENCODE_DEBUG_INIT (which expands to a comma+expression for
appending to a function call, not a standalone statement). Add
NOLINTNEXTLINE on the deterministic mt19937_64 seed so clang-tidy
cert-msc32-c stops failing the build (the seed is intentional for
reproducible test runs).
- socket FILTER_SOURCE_FILES: tolerate non-dict CORE.config['socket']
(e.g. C++ unit-test builds where socket isn't validated as a
mapping). Returning [] is safe -- all impl files are guarded by
USE_SOCKET_IMPL_* defines so only the selected one contributes
code.
Verifies encode_varint_raw_48bit and calc_uint64_48bit_force for the
8 corner-case MAC addresses requested in review:
00:00:00:00:00:00, 11:00:00:00:00:00, 00:AA:00:00:00:00,
00:00:BB:00:00:00, 00:00:00:CC:00:00, 00:00:00:00:DD:00,
00:00:00:00:00:EE, FF:FF:FF:FF:FF:FF
For each value the test asserts byte-identical output to the reference
encode_varint_raw_64 loop, the expected encoded byte length, agreement
with calc_uint64_48bit_force, and round-trip through a generic varint
decoder. Adds a 100-value deterministic-random sample across the full
48-bit space for additional coverage.
Copilot flagged that encode_varint_raw_48bit/calc_uint64_48bit_force
would silently truncate for uint64 values >= 2^48. In practice the
(mac_address) option is only applied to fields populated by the BLE
stack, which always fits in 48 bits -- so the runtime upper-bound
check added in ba362a7c95 regressed CodSpeed by up to 8.7pp on
CalculateSize_BLERawAdvs12 for a scenario that can't happen.
Move the value-fits-in-48-bits check to a debug assert guarded by
ESPHOME_DEBUG_API, and express 48 via MAC_ADDRESS_SIZE * 8 so the
threshold tracks the existing MAC size constant. Release builds are
back to the original fast path; debug builds catch misuse.
Per Copilot review: encode_varint_raw_48bit and calc_uint64_48bit_force
would silently truncate bits 48..63 if ever called with a uint64 that
doesn't fit in 48 bits. Real MAC addresses always fit, but since the
helpers are exposed and the (mac_address) option is generic, narrow
the fast path to [1<<42, 1<<48) so values outside that range fall back
to the general encoder/size helpers.
Re-verified byte-identical output and decoder round-trip across the
1<<48 boundary and all bit positions up to 63.
Adds a (mac_address) field option that switches uint64 fields holding
48-bit MAC addresses to a specialized varint encoder. The fast path
emits exactly 7 bytes when bits [42..47] are non-zero (the common case
for real MACs, since OUIs occupy the top 24 bits) -- one bounds check
and 7 independent stores instead of the 7-iteration shift+branch loop.
calc_uint64_48bit_force mirrors the same fast path in size calculation.
Applied to BluetoothLERawAdvertisement.address, the per-advertisement
address encode in BluetoothLERawAdvertisementsResponse drops from a
serialized per-byte loop to straight-line code.
Fold the to_remove_empty check, cleanup_slow_path_ call, and
MAX_LOGICALLY_DELETED_ITEMS threshold check into a single hot-path
branch + cold outlined helper.
Before this change, the cleanup block in Scheduler::call compiled to
two independent memw + l32i sequences (one for to_remove_empty_() inside
cleanup_(), one for the separate to_remove_count_() check) because GCC
cannot CSE across the memw barriers that std::atomic<uint32_t>::load
emits on Xtensa.
A previous attempt at collapsing these into a single inline branch
(#15985) had the right assembly for the memw count but grew
Scheduler::call by a handful of bytes and rearranged the control flow
enough to nudge sched up ~0.5 us/iter on gatetrigger in practice.
This version goes further: cleanup_slow_combined_ is annotated
noinline + cold so the entire slow path (both reads, both calls) is
pulled out of Scheduler::call entirely. The hot path becomes:
memw; l32i a8, [to_remove_]; beqz skip
call8 cleanup_slow_combined_ ; unlikely, cold
and Scheduler::call's body shrinks 344 B -> 332 B (-12 B, below the dev
baseline). Adjacent code (feed_wdt_slow_, etc.) stays in the same flash
region, avoiding the cache-layout side-effects that made earlier
attempts a net loss on busier configs.
The [[unlikely]] attribute on the branch plus the cold attribute on the
helper give the compiler permission to keep the skip path straight and
push the call out-of-line.
Read to_remove_ once at the top of the cleanup block instead of loading
it twice (once via cleanup_() -> to_remove_empty_(), once again for the
MAX_LOGICALLY_DELETED_ITEMS check). GCC cannot CSE across the memw
barriers that std::atomic<uint32_t>::load emits on Xtensa, so both the
fast-path zero check and the max check were generating independent
memw + l32i sequences:
memw; l32i a8, [to_remove_]; beqz ... ; to_remove_empty_()
memw; l32i a8, [to_remove_]; bltui 5 ; to_remove_count_()
Reading the counter once into a register and branching on the result
collapses the common zero-case to a single memw + l32i + beqz. The
non-zero path pays one extra read after cleanup_slow_path_ (which may
have decremented the counter), but that path already takes lock_ so the
extra load is negligible.
Scheduler::call stays at 344 B (unchanged); no flash layout shift, so
adjacent code (feed_wdt_slow_ etc.) stays in the same cache lines and
the measurement can't be confounded by MMIO timing drift.
An earlier version of this change also marked Scheduler::millis_64()
ESPHOME_ALWAYS_INLINE to drop its out-of-line $isra$0 clone; that
grew Scheduler::call by 56 B, shifted feed_wdt_slow_ in flash, and
measurably regressed the wdt bucket on a busier winefridge.yaml
config, so the inlining is dropped. This change keeps only the memw
collapse.
Read to_remove_ once at the top of the cleanup block instead of loading
it twice (once via cleanup_() -> to_remove_empty_(), once again for the
MAX_LOGICALLY_DELETED_ITEMS check). GCC cannot CSE across the memw
barriers that std::atomic<uint32_t>::load emits on Xtensa, so both the
fast-path zero check and the max check were generating independent
memw + l32i sequences:
memw; l32i a8, [to_remove_]; beqz ... ; to_remove_empty_()
memw; l32i a8, [to_remove_]; bltui 5 ; to_remove_count_()
Reading the counter once into a register and branching on the result
collapses the common zero-case to a single memw + l32i + beqz. The
non-zero path pays one extra read after cleanup_slow_path_ (which may
have decremented the counter), but that path already takes lock_ so the
extra load is negligible.
Scheduler::call stays at 344 B (unchanged); no flash layout shift, so
adjacent code (feed_wdt_slow_ etc.) stays in the same cache lines and
the measurement can't be confounded by MMIO timing drift.
An earlier version of this change also marked Scheduler::millis_64()
ESPHOME_ALWAYS_INLINE to drop its out-of-line $isra$0 clone; that
grew Scheduler::call by 56 B, shifted feed_wdt_slow_ in flash, and
measurably regressed the wdt bucket on a busier winefridge.yaml
config, so the inlining is dropped. This change keeps only the memw
collapse.
Mirrors the existing bk72xx platform override, but auto-scales to the
user-configurable esp32.watchdog_timeout (CONFIG_ESP_TASK_WDT_TIMEOUT_S)
instead of hard-coding a constant: the feed interval is always 1/5 of
the configured task WDT timeout so the safety margin stays constant
across user configurations.
- default 5 s WDT -> 1000 ms feed interval (was 300 ms, -70% hits)
- 10 s WDT -> 2000 ms feed interval
- 60 s WDT (max) -> 12000 ms feed interval
esp_task_wdt_reset() takes a spinlock and walks the WDT task list, so
every call costs tens of microseconds. At the normal ~62 Hz main loop
the old 300 ms cadence produced ~200 feed_wdt_slow_ hits per 60 s
period; the default 5s -> 1000 ms cuts that to ~60 hits (-70%).
Component-level feeds inside Component::loop() and scheduler items are
unaffected; they continue to call arch_feed_wdt after every operation,
so any op exceeding this rate-limit triggers a real feed naturally.
The rate-limit only applies to the outer guard in Application::loop()
that fires when nothing else fed recently.
esp32/__init__.py already constrains watchdog_timeout to >= 5 s (range
5-60 s), and a static_assert guards against anyone who tweaks sdkconfig
below that floor, ensuring the feed interval never drops below the
prior hardcoded 1000 ms value.
Measured on a live ESP32 IDF build (gatetrigger.yaml, default 5 s WDT)
via runtime_stats: wdt bucket dropped from 2.24 us/iter to 1.56 us/iter
- about 2.5 ms saved per 60 s window.
Mirrors the existing bk72xx platform override, but auto-scales to the
user-configurable esp32.watchdog_timeout (CONFIG_ESP_TASK_WDT_TIMEOUT_S)
instead of hard-coding a constant: the feed interval is always 1/5 of
the configured task WDT timeout so the safety margin stays constant
across user configurations.
- default 5 s WDT -> 1000 ms feed interval (was 300 ms, -70% hits)
- 10 s WDT -> 2000 ms feed interval
- 60 s WDT (max) -> 12000 ms feed interval
esp_task_wdt_reset() takes a spinlock and walks the WDT task list, so
every call costs tens of microseconds. At the normal ~62 Hz main loop
the old 300 ms cadence produced ~200 feed_wdt_slow_ hits per 60 s
period; the default 5s -> 1000 ms cuts that to ~60 hits (-70%).
Component-level feeds inside Component::loop() and scheduler items are
unaffected; they continue to call arch_feed_wdt after every operation,
so any op exceeding this rate-limit triggers a real feed naturally.
The rate-limit only applies to the outer guard in Application::loop()
that fires when nothing else fed recently.
esp32/__init__.py already constrains watchdog_timeout to >= 5 s (range
5-60 s), and a static_assert guards against anyone who tweaks sdkconfig
below that floor, ensuring the feed interval never drops below the
prior hardcoded 1000 ms value.
Measured on a live ESP32 IDF build (gatetrigger.yaml, default 5 s WDT)
via runtime_stats: wdt bucket dropped from 2.24 us/iter to 1.56 us/iter
- about 2.5 ms saved per 60 s window.
Mirrors the bk72xx override: the ESP32 task WDT default timeout is 5s
(CONFIG_ESP_TASK_WDT_TIMEOUT_S), so feeding every 1000ms keeps a ~5x
safety margin while cutting per-iteration feed overhead substantially.
esp_task_wdt_reset() takes a spinlock and walks the WDT task list, so
every call costs tens of microseconds. At the normal ~62 Hz main loop
the old 300ms cadence produced ~200 hits per 60s period; 1000ms cuts
that to ~60 hits per 60s (-70%). Measured on a live ESP32 IDF build
(gatetrigger.yaml) via runtime_stats: wdt bucket dropped from 2.24
us/iter to 1.56 us/iter - about 2.5 ms saved per 60s window, or
~0.7us/iter average.
Component-level feeds inside component loop() and scheduler items are
unaffected; they continue to call arch_feed_wdt after every operation,
so any operation that exceeds this rate-limit triggers a real feed
naturally. The rate-limit only applies to the outer guard in
Application::loop() that fires when nothing else fed recently.
Mirrors the bk72xx override: the ESP32 task WDT default timeout is 5s
(CONFIG_ESP_TASK_WDT_TIMEOUT_S), so feeding every 1000ms keeps a ~5x
safety margin while cutting per-iteration feed overhead substantially.
esp_task_wdt_reset() takes a spinlock and walks the WDT task list, so
every call costs tens of microseconds. At the normal ~62 Hz main loop
the old 300ms cadence produced ~200 hits per 60s period; 1000ms cuts
that to ~60 hits per 60s (-70%). Measured on a live ESP32 IDF build
(gatetrigger.yaml) via runtime_stats: wdt bucket dropped from 2.24
us/iter to 1.56 us/iter - about 2.5 ms saved per 60s window, or
~0.7us/iter average.
Component-level feeds inside component loop() and scheduler items are
unaffected; they continue to call arch_feed_wdt after every operation,
so any operation that exceeds this rate-limit triggers a real feed
naturally. The rate-limit only applies to the outer guard in
Application::loop() that fires when nothing else fed recently.
Replace the CallResult struct return with a uint64_t &now_64_out
reference parameter. The struct return forced GCC on Xtensa to emit 8
redundant s32i stores after every call() (once into the sched_result
local plus once into a compiler-chosen temp slot), and the extra ABI
shuffling around the 16-byte return value was tight enough with the
esp_timer_get_time() MMIOs bracketing the wdt bucket that the wdt
measurement regressed from ~21 us/hit to ~36 us/hit on ESP32.
Out-param keeps the scheduler's now_64 write local to Scheduler::call
(written once to *now_64_out) and leaves the return path a single
uint32 in a10. The caller passes &sched_now_64_raw directly; no struct
copy. loop_task shrinks 988 B -> 960 B.
Disassembly verified: after Scheduler::call returns the only
instructions before the next micros() capture are:
mov.n a6, a10 ; save now
call8 esp_timer_get_time ; loop_after_sched_us
No duplicate struct stores.
Follow-up to the CallResult threading: move the 0-sentinel fallback
(if now_64 == 0, read the clock fresh) out of next_schedule_in and into
the single main-loop caller.
Effect on next_schedule_in: drops the uint32_t now parameter entirely
and removes the cold millis_64_from_(now) inlined clock read from the
function body. Verified in disassembly:
Scheduler::next_schedule_in 159 B -> 86 B (-46%)
The hot path is now just defer_empty + cleanup_ + top-of-heap compare +
return. No clock read anywhere in the function.
Application::loop handles the 0-sentinel with a single inline branch
right before the call; when call() fired items it invokes the public
Scheduler::millis_64_from() wrapper (new, exposes the existing protected
millis_64_from_ so callers can do 64-bit extension without platform
knowledge). Fast path (no items fired) stays branch-free.
Net code size: loop_task +64 B, next_schedule_in -73 B,
Scheduler::call unchanged -> -9 B total.
Follow-up to the CallResult threading: move the 0-sentinel fallback
(if now_64 == 0, read the clock fresh) out of next_schedule_in and into
the single main-loop caller.
Effect on next_schedule_in: drops the uint32_t now parameter entirely
and removes the cold millis_64_from_(now) inlined clock read from the
function body. Verified in disassembly:
Scheduler::next_schedule_in 159 B -> 86 B (-46%)
The hot path is now just defer_empty + cleanup_ + top-of-heap compare +
return. No clock read anywhere in the function.
Application::loop handles the 0-sentinel with a single inline branch
right before the call; when call() fired items it invokes the public
Scheduler::millis_64_from() wrapper (new, exposes the existing protected
millis_64_from_ so callers can do 64-bit extension without platform
knowledge). Fast path (no items fired) stays branch-free.
Net code size: loop_task +64 B, next_schedule_in -73 B,
Scheduler::call unchanged -> -9 B total.
Scheduler::call() and Scheduler::next_schedule_in() both computed now_64 via
millis_64_from_(now) at the top of every main-loop iteration. On ESP32 with
USE_NATIVE_64BIT_TIME this is an esp_timer_get_time() MMIO read (~1us);
doing it twice per iteration is wasted work since the two calls happen
microseconds apart.
Thread the value through:
- Scheduler::call() now returns a CallResult { now, now_64 } struct instead
of just the advanced uint32_t now. When items fired, now_64 is set to 0
(sentinel) because execute_item_() only advances the uint32 and the local
now_64 is stale by the time we return.
- Scheduler::next_schedule_in() takes an optional now_64 parameter
(default 0). Non-zero values are trusted and used directly; the 0
sentinel triggers a fresh millis_64_from_(now) read.
- Application::scheduler_tick_() forwards the CallResult through.
- Application::loop() passes sched_result.now_64 to next_schedule_in().
Verified in the disassembly: on the fast path (no items fired),
next_schedule_in branches over its esp_timer_get_time call entirely via
`or a8, a4, a5; bnez a8, ...` on the two halves of now_64, jumping
straight to the next_exec comparison. When call() returned the 0 sentinel,
next_schedule_in falls through to the existing inlined
micros_to_millis(esp_timer_get_time()) sequence as before.
Bench callers (tests/benchmarks/core/bench_scheduler.cpp) and the external
test component (tests/integration/fixtures/.../scheduler_bulk_cleanup_component)
ignore the return value — no changes needed there.
Two small wins on the per-loop Scheduler::call path:
1. Mark the scheduler's `millis_64()` wrapper `ESPHOME_ALWAYS_INLINE`.
The inner `esphome::millis_64()` is already always_inline, but the
wrapper was not, so GCC emitted an out-of-line `$isra$0` clone with
its own `entry`/`retw` register-window prologue/epilogue at every
call site. Inlining removes the wrapper's call overhead and the
register-window thunk on ESP32.
2. Collapse the two atomic reads of `to_remove_` in `Scheduler::call`
into one. The previous sequence
cleanup_(); // loads to_remove_
if (to_remove_count_() >= MAX_...) ... // reloads to_remove_
produced `memw; l32i; beqz; memw; l32i; bltui` on the fast path
because the compiler cannot CSE across the `memw` barriers that
std::atomic<uint32_t>::load emits on Xtensa. Reading the counter
once and branching on the result leaves a single
`memw; l32i; beqz` on the common zero-case; the slow path
(cleanup_slow_path_ + re-read + optional full_cleanup) pays an
extra read but already holds the scheduler mutex.
The static class members in Millis64Impl's ESPHOME_THREAD_SINGLE branch
kept the trailing-underscore convention used for instance members, but
the project's clang-tidy config (readability-identifier-naming.
ClassMemberCase = lower_case) wants static class members without the
suffix. This violation was latent because defines.h previously
hardcoded MULTI_ATOMICS, so the SINGLE branch was never analyzed;
now that defines.h picks the model per platform, the ESP8266 tidy env
analyzes this branch.
Drop the trailing _ from last_millis / millis_major; their types and
initial values are unchanged.
defines.h previously hardcoded ESPHOME_THREAD_MULTI_ATOMICS regardless
of the active USE_<platform>, so the clang-tidy envs ended up
compiling e.g. wake_esp8266.cpp with USE_ESP8266 + MULTI_ATOMICS both
set — a combination that cannot occur in a real build and that
mismatched the extern in wake.h.
Pick the model that real codegen uses:
USE_ESP8266 / USE_RP2040 / USE_NRF52 → SINGLE
USE_BK72XX (ARMv5TE, no LDREX/STREX) → MULTI_NO_ATOMICS
everything else (ESP32, host, RTL87XX, LN882X) → MULTI_ATOMICS
With that, the single-only wake TUs (esp8266, rp2040, generic) can
drop the defensive conditional and just define g_wake_requested as
plain volatile again.
ESPHome's defines.h unconditionally #defines every feature flag for
static analysis, so clang-tidy sees ESPHOME_THREAD_MULTI_ATOMICS even
when building wake_esp8266.cpp / wake_rp2040.cpp / wake_generic.cpp.
In that mode wake.h picks the std::atomic<uint8_t> extern, and the
plain-volatile definitions in those TUs trip a redefinition error.
Guard each TU's g_wake_requested with the same
ESPHOME_THREAD_MULTI_ATOMICS check wake.h uses. Real builds still pick
volatile on every SINGLE platform.
wake.h's platform #include already ensures only one header contributes
code, but all five wake/*.cpp files were still being copied and compiled
(as empty TUs) on every target. Extend FILTER_SOURCE_FILES to map each
wake_<platform>.cpp to the exact set of PlatformFramework values that
need it — same pattern ring_buffer.cpp already uses.
Cover the four paths introduced by the wake split:
- default manifests (recursive_sources=False) skip subdirs entirely,
- opt-in manifests walk non-subpackage subdirs with /-joined paths,
- subdirs containing __init__.py are skipped to avoid double-counting
existing component subpackages, and __pycache__ is always skipped,
- FILTER_SOURCE_FILES entries accept /-joined subpaths.
Move each platform branch out of wake.h/wake.cpp into its own
esphome/core/wake/wake_<platform>.{h,cpp} pair (wake_freertos for
ESP32/LibreTiny, plus wake_esp8266, wake_rp2040, wake_host, and
wake_generic for the Zephyr/NRF52 fallback). wake.h becomes a thin
dispatcher that defines only the cross-platform wake_request_set/
wake_request_take helpers and then #include's the right platform
header based on USE_*. wake.cpp is removed; each platform's
g_wake_requested (and g_main_loop_woke where applicable) now lives
in the translation unit that actually uses it, guarded by the same
USE_* check.
Because esphome/core/ source discovery was previously flat, add an
opt-in recursive_sources flag to ComponentManifest and enable it only
on the core ("esphome") manifest. Subdirectories that are themselves
Python subpackages are still skipped so components (which already
register platform subpackages independently) are unaffected.
Re-merge the updated PR (now keeps esp8266 millis/delay out-of-line so
integration's accumulator + custom delay overlay cleanly, while still
inlining yield/micros/millis_64 on esp8266).
Resolves esp8266/core.cpp conflict by keeping integration's custom
millis() accumulator, custom delay(), and __wrap_millis machinery
out-of-line — those cannot be inlined safely:
- esphome::millis() is the body of __wrap_millis (-Wl,--wrap=millis),
so it must remain a real symbol; inlining would cause infinite
recursion at the wrapped call site.
- delay() intentionally avoids Arduino's __delay → millis path so the
slow Arduino millis body can be stripped from IRAM.
- millis_64() depends on millis() and stays out-of-line for symmetry.
esp8266 yield() and micros() are removed from core.cpp (now inlined in
hal.h alongside the other Arduino-flavored platforms — both are simple
::yield()/::micros() passthroughs).
hal.h is patched so the USE_ESP8266 branch only inlines yield() and
micros(), while delay()/millis()/millis_64() stay as out-of-line
declarations (defined in esp8266/core.cpp).