Mirrors the existing bk72xx platform override, but auto-scales to the
user-configurable esp32.watchdog_timeout (CONFIG_ESP_TASK_WDT_TIMEOUT_S)
instead of hard-coding a constant: the feed interval is always 1/5 of
the configured task WDT timeout so the safety margin stays constant
across user configurations.
- default 5 s WDT -> 1000 ms feed interval (was 300 ms, -70% hits)
- 10 s WDT -> 2000 ms feed interval
- 60 s WDT (max) -> 12000 ms feed interval
esp_task_wdt_reset() takes a spinlock and walks the WDT task list, so
every call costs tens of microseconds. At the normal ~62 Hz main loop
the old 300 ms cadence produced ~200 feed_wdt_slow_ hits per 60 s
period; the default 5s -> 1000 ms cuts that to ~60 hits (-70%).
Component-level feeds inside Component::loop() and scheduler items are
unaffected; they continue to call arch_feed_wdt after every operation,
so any op exceeding this rate-limit triggers a real feed naturally.
The rate-limit only applies to the outer guard in Application::loop()
that fires when nothing else fed recently.
esp32/__init__.py already constrains watchdog_timeout to >= 5 s (range
5-60 s), and a static_assert guards against anyone who tweaks sdkconfig
below that floor, ensuring the feed interval never drops below the
prior hardcoded 1000 ms value.
Measured on a live ESP32 IDF build (gatetrigger.yaml, default 5 s WDT)
via runtime_stats: wdt bucket dropped from 2.24 us/iter to 1.56 us/iter
- about 2.5 ms saved per 60 s window.
Mirrors the existing bk72xx platform override, but auto-scales to the
user-configurable esp32.watchdog_timeout (CONFIG_ESP_TASK_WDT_TIMEOUT_S)
instead of hard-coding a constant: the feed interval is always 1/5 of
the configured task WDT timeout so the safety margin stays constant
across user configurations.
- default 5 s WDT -> 1000 ms feed interval (was 300 ms, -70% hits)
- 10 s WDT -> 2000 ms feed interval
- 60 s WDT (max) -> 12000 ms feed interval
esp_task_wdt_reset() takes a spinlock and walks the WDT task list, so
every call costs tens of microseconds. At the normal ~62 Hz main loop
the old 300 ms cadence produced ~200 feed_wdt_slow_ hits per 60 s
period; the default 5s -> 1000 ms cuts that to ~60 hits (-70%).
Component-level feeds inside Component::loop() and scheduler items are
unaffected; they continue to call arch_feed_wdt after every operation,
so any op exceeding this rate-limit triggers a real feed naturally.
The rate-limit only applies to the outer guard in Application::loop()
that fires when nothing else fed recently.
esp32/__init__.py already constrains watchdog_timeout to >= 5 s (range
5-60 s), and a static_assert guards against anyone who tweaks sdkconfig
below that floor, ensuring the feed interval never drops below the
prior hardcoded 1000 ms value.
Measured on a live ESP32 IDF build (gatetrigger.yaml, default 5 s WDT)
via runtime_stats: wdt bucket dropped from 2.24 us/iter to 1.56 us/iter
- about 2.5 ms saved per 60 s window.
Mirrors the bk72xx override: the ESP32 task WDT default timeout is 5s
(CONFIG_ESP_TASK_WDT_TIMEOUT_S), so feeding every 1000ms keeps a ~5x
safety margin while cutting per-iteration feed overhead substantially.
esp_task_wdt_reset() takes a spinlock and walks the WDT task list, so
every call costs tens of microseconds. At the normal ~62 Hz main loop
the old 300ms cadence produced ~200 hits per 60s period; 1000ms cuts
that to ~60 hits per 60s (-70%). Measured on a live ESP32 IDF build
(gatetrigger.yaml) via runtime_stats: wdt bucket dropped from 2.24
us/iter to 1.56 us/iter - about 2.5 ms saved per 60s window, or
~0.7us/iter average.
Component-level feeds inside component loop() and scheduler items are
unaffected; they continue to call arch_feed_wdt after every operation,
so any operation that exceeds this rate-limit triggers a real feed
naturally. The rate-limit only applies to the outer guard in
Application::loop() that fires when nothing else fed recently.
Mirrors the bk72xx override: the ESP32 task WDT default timeout is 5s
(CONFIG_ESP_TASK_WDT_TIMEOUT_S), so feeding every 1000ms keeps a ~5x
safety margin while cutting per-iteration feed overhead substantially.
esp_task_wdt_reset() takes a spinlock and walks the WDT task list, so
every call costs tens of microseconds. At the normal ~62 Hz main loop
the old 300ms cadence produced ~200 hits per 60s period; 1000ms cuts
that to ~60 hits per 60s (-70%). Measured on a live ESP32 IDF build
(gatetrigger.yaml) via runtime_stats: wdt bucket dropped from 2.24
us/iter to 1.56 us/iter - about 2.5 ms saved per 60s window, or
~0.7us/iter average.
Component-level feeds inside component loop() and scheduler items are
unaffected; they continue to call arch_feed_wdt after every operation,
so any operation that exceeds this rate-limit triggers a real feed
naturally. The rate-limit only applies to the outer guard in
Application::loop() that fires when nothing else fed recently.
Replace the CallResult struct return with a uint64_t &now_64_out
reference parameter. The struct return forced GCC on Xtensa to emit 8
redundant s32i stores after every call() (once into the sched_result
local plus once into a compiler-chosen temp slot), and the extra ABI
shuffling around the 16-byte return value was tight enough with the
esp_timer_get_time() MMIOs bracketing the wdt bucket that the wdt
measurement regressed from ~21 us/hit to ~36 us/hit on ESP32.
Out-param keeps the scheduler's now_64 write local to Scheduler::call
(written once to *now_64_out) and leaves the return path a single
uint32 in a10. The caller passes &sched_now_64_raw directly; no struct
copy. loop_task shrinks 988 B -> 960 B.
Disassembly verified: after Scheduler::call returns the only
instructions before the next micros() capture are:
mov.n a6, a10 ; save now
call8 esp_timer_get_time ; loop_after_sched_us
No duplicate struct stores.
Follow-up to the CallResult threading: move the 0-sentinel fallback
(if now_64 == 0, read the clock fresh) out of next_schedule_in and into
the single main-loop caller.
Effect on next_schedule_in: drops the uint32_t now parameter entirely
and removes the cold millis_64_from_(now) inlined clock read from the
function body. Verified in disassembly:
Scheduler::next_schedule_in 159 B -> 86 B (-46%)
The hot path is now just defer_empty + cleanup_ + top-of-heap compare +
return. No clock read anywhere in the function.
Application::loop handles the 0-sentinel with a single inline branch
right before the call; when call() fired items it invokes the public
Scheduler::millis_64_from() wrapper (new, exposes the existing protected
millis_64_from_ so callers can do 64-bit extension without platform
knowledge). Fast path (no items fired) stays branch-free.
Net code size: loop_task +64 B, next_schedule_in -73 B,
Scheduler::call unchanged -> -9 B total.
Follow-up to the CallResult threading: move the 0-sentinel fallback
(if now_64 == 0, read the clock fresh) out of next_schedule_in and into
the single main-loop caller.
Effect on next_schedule_in: drops the uint32_t now parameter entirely
and removes the cold millis_64_from_(now) inlined clock read from the
function body. Verified in disassembly:
Scheduler::next_schedule_in 159 B -> 86 B (-46%)
The hot path is now just defer_empty + cleanup_ + top-of-heap compare +
return. No clock read anywhere in the function.
Application::loop handles the 0-sentinel with a single inline branch
right before the call; when call() fired items it invokes the public
Scheduler::millis_64_from() wrapper (new, exposes the existing protected
millis_64_from_ so callers can do 64-bit extension without platform
knowledge). Fast path (no items fired) stays branch-free.
Net code size: loop_task +64 B, next_schedule_in -73 B,
Scheduler::call unchanged -> -9 B total.
Scheduler::call() and Scheduler::next_schedule_in() both computed now_64 via
millis_64_from_(now) at the top of every main-loop iteration. On ESP32 with
USE_NATIVE_64BIT_TIME this is an esp_timer_get_time() MMIO read (~1us);
doing it twice per iteration is wasted work since the two calls happen
microseconds apart.
Thread the value through:
- Scheduler::call() now returns a CallResult { now, now_64 } struct instead
of just the advanced uint32_t now. When items fired, now_64 is set to 0
(sentinel) because execute_item_() only advances the uint32 and the local
now_64 is stale by the time we return.
- Scheduler::next_schedule_in() takes an optional now_64 parameter
(default 0). Non-zero values are trusted and used directly; the 0
sentinel triggers a fresh millis_64_from_(now) read.
- Application::scheduler_tick_() forwards the CallResult through.
- Application::loop() passes sched_result.now_64 to next_schedule_in().
Verified in the disassembly: on the fast path (no items fired),
next_schedule_in branches over its esp_timer_get_time call entirely via
`or a8, a4, a5; bnez a8, ...` on the two halves of now_64, jumping
straight to the next_exec comparison. When call() returned the 0 sentinel,
next_schedule_in falls through to the existing inlined
micros_to_millis(esp_timer_get_time()) sequence as before.
Bench callers (tests/benchmarks/core/bench_scheduler.cpp) and the external
test component (tests/integration/fixtures/.../scheduler_bulk_cleanup_component)
ignore the return value — no changes needed there.
Two small wins on the per-loop Scheduler::call path:
1. Mark the scheduler's `millis_64()` wrapper `ESPHOME_ALWAYS_INLINE`.
The inner `esphome::millis_64()` is already always_inline, but the
wrapper was not, so GCC emitted an out-of-line `$isra$0` clone with
its own `entry`/`retw` register-window prologue/epilogue at every
call site. Inlining removes the wrapper's call overhead and the
register-window thunk on ESP32.
2. Collapse the two atomic reads of `to_remove_` in `Scheduler::call`
into one. The previous sequence
cleanup_(); // loads to_remove_
if (to_remove_count_() >= MAX_...) ... // reloads to_remove_
produced `memw; l32i; beqz; memw; l32i; bltui` on the fast path
because the compiler cannot CSE across the `memw` barriers that
std::atomic<uint32_t>::load emits on Xtensa. Reading the counter
once and branching on the result leaves a single
`memw; l32i; beqz` on the common zero-case; the slow path
(cleanup_slow_path_ + re-read + optional full_cleanup) pays an
extra read but already holds the scheduler mutex.
The static class members in Millis64Impl's ESPHOME_THREAD_SINGLE branch
kept the trailing-underscore convention used for instance members, but
the project's clang-tidy config (readability-identifier-naming.
ClassMemberCase = lower_case) wants static class members without the
suffix. This violation was latent because defines.h previously
hardcoded MULTI_ATOMICS, so the SINGLE branch was never analyzed;
now that defines.h picks the model per platform, the ESP8266 tidy env
analyzes this branch.
Drop the trailing _ from last_millis / millis_major; their types and
initial values are unchanged.
defines.h previously hardcoded ESPHOME_THREAD_MULTI_ATOMICS regardless
of the active USE_<platform>, so the clang-tidy envs ended up
compiling e.g. wake_esp8266.cpp with USE_ESP8266 + MULTI_ATOMICS both
set — a combination that cannot occur in a real build and that
mismatched the extern in wake.h.
Pick the model that real codegen uses:
USE_ESP8266 / USE_RP2040 / USE_NRF52 → SINGLE
USE_BK72XX (ARMv5TE, no LDREX/STREX) → MULTI_NO_ATOMICS
everything else (ESP32, host, RTL87XX, LN882X) → MULTI_ATOMICS
With that, the single-only wake TUs (esp8266, rp2040, generic) can
drop the defensive conditional and just define g_wake_requested as
plain volatile again.
ESPHome's defines.h unconditionally #defines every feature flag for
static analysis, so clang-tidy sees ESPHOME_THREAD_MULTI_ATOMICS even
when building wake_esp8266.cpp / wake_rp2040.cpp / wake_generic.cpp.
In that mode wake.h picks the std::atomic<uint8_t> extern, and the
plain-volatile definitions in those TUs trip a redefinition error.
Guard each TU's g_wake_requested with the same
ESPHOME_THREAD_MULTI_ATOMICS check wake.h uses. Real builds still pick
volatile on every SINGLE platform.
wake.h's platform #include already ensures only one header contributes
code, but all five wake/*.cpp files were still being copied and compiled
(as empty TUs) on every target. Extend FILTER_SOURCE_FILES to map each
wake_<platform>.cpp to the exact set of PlatformFramework values that
need it — same pattern ring_buffer.cpp already uses.
Cover the four paths introduced by the wake split:
- default manifests (recursive_sources=False) skip subdirs entirely,
- opt-in manifests walk non-subpackage subdirs with /-joined paths,
- subdirs containing __init__.py are skipped to avoid double-counting
existing component subpackages, and __pycache__ is always skipped,
- FILTER_SOURCE_FILES entries accept /-joined subpaths.
Move each platform branch out of wake.h/wake.cpp into its own
esphome/core/wake/wake_<platform>.{h,cpp} pair (wake_freertos for
ESP32/LibreTiny, plus wake_esp8266, wake_rp2040, wake_host, and
wake_generic for the Zephyr/NRF52 fallback). wake.h becomes a thin
dispatcher that defines only the cross-platform wake_request_set/
wake_request_take helpers and then #include's the right platform
header based on USE_*. wake.cpp is removed; each platform's
g_wake_requested (and g_main_loop_woke where applicable) now lives
in the translation unit that actually uses it, guarded by the same
USE_* check.
Because esphome/core/ source discovery was previously flat, add an
opt-in recursive_sources flag to ComponentManifest and enable it only
on the core ("esphome") manifest. Subdirectories that are themselves
Python subpackages are still skipped so components (which already
register platform subpackages independently) are unaffected.
Re-merge the updated PR (now keeps esp8266 millis/delay out-of-line so
integration's accumulator + custom delay overlay cleanly, while still
inlining yield/micros/millis_64 on esp8266).
Resolves esp8266/core.cpp conflict by keeping integration's custom
millis() accumulator, custom delay(), and __wrap_millis machinery
out-of-line — those cannot be inlined safely:
- esphome::millis() is the body of __wrap_millis (-Wl,--wrap=millis),
so it must remain a real symbol; inlining would cause infinite
recursion at the wrapped call site.
- delay() intentionally avoids Arduino's __delay → millis path so the
slow Arduino millis body can be stripped from IRAM.
- millis_64() depends on millis() and stays out-of-line for symmetry.
esp8266 yield() and micros() are removed from core.cpp (now inlined in
hal.h alongside the other Arduino-flavored platforms — both are simple
::yield()/::micros() passthroughs).
hal.h is patched so the USE_ESP8266 branch only inlines yield() and
micros(), while delay()/millis()/millis_64() stay as out-of-line
declarations (defined in esp8266/core.cpp).
Extends the prior commit to cover more wrappers and platforms:
- ESP32: also inline yield() and delay()
- ESP8266: also inline yield(), delay(), millis(), millis_64()
- LibreTiny: inline yield(), delay(), micros(), per-variant millis()
fast paths, and millis_64() (via Millis64Impl::compute, now reachable
from hal.h since time_64.h dropped its helpers.h dep)
- RP2040: also inline yield(), delay(), micros()
Consolidates the ESP8266/LibreTiny/RP2040 Arduino-flavored ::yield /
::delay / ::micros wrappers into a single shared block in hal.h.
LibreTiny note: the prior IRAM_ATTR on the wrapper was decorative —
::micros(), ::yield(), ::delay() and xTaskGetTickCount all live in
flash on every libretiny family (realtek-amb, beken-72xx,
lightning-ln882h all checked), so an IRAM ISR call would have crashed
the same way an inlined direct call does.
Also drops the helpers.h include from time_64.h (only used for the
ESPHOME_ALWAYS_INLINE macro, replaced with the raw attribute) so
time_64.h is light enough for hal.h to include.
Replaces the per-function ``__attribute__((optimize("O2")))`` approach in
#15693 with a direct inline definition of micros() / millis_64() in hal.h
for ESP32, plus inline definitions for ESP8266 micros() and RP2040
millis()/millis_64().
The original goal of #15693 was to inline micros() into the main loop so
that ``call esphome::micros() → call esp_timer_get_time()`` collapses to
a single ``call esp_timer_get_time``. That wrapper-collapse benefits
every micros() call site, not just loop_task — so the real fix is to
mark the wrapper inline, not to bump the loop's optimization level.
Doing this at the HAL layer also makes runtime_stats measurements more
accurate: each timing read no longer hides a wrapper call/return between
the component end-time capture and the underlying clock read.
Refactor: move ``micros_to_millis<>()`` from helpers.h to a new
lightweight ``time_conversion.h`` so hal.h can include it without
pulling the rest of helpers.h into every TU that includes hal.h.