On ESP32, millis() uses xTaskGetTickCount (tick clock) but
millis_64_from_() discards the 32-bit value and calls millis_64()
(esp_timer clock). Safe because scheduling only compares millis_64
against millis_64. On ESP8266, both use the same accumulator clock.
millis() uses xTaskGetTickCount (tick clock) while millis_64() uses
esp_timer_get_time (hardware timer). Safe because they are never
cross-compared: Scheduler::millis_64_from_() on ESP32 discards the
32-bit millis parameter and calls millis_64() directly, keeping all
64-bit scheduling on the esp_timer clock.
Revert the inline hal.h approach — millis() must remain IRAM_ATTR
because Wiegand and ZyAura call it from IRAM_ATTR ISR handlers on
all platforms including ESP32.
Use xPortInIsrContext() to dispatch to xTaskGetTickCountFromISR()
from ISR or xTaskGetTickCount() from task context, satisfying the
FreeRTOS API contract. ISR check is [[unlikely]] so the hot path
stays fast.
Benchmarked: 686 ns -> 361 ns (1.9x faster). Correctness verified.
Move millis() to an inline function in hal.h that just returns
xTaskGetTickCount() -- eliminates the function call entirely at every
call site. No IRAM consumed, no flash function call.
Drop IRAM_ATTR from the fallback path. The original IRAM placement was
inherited from ESP8266 where Arduino millis() is called from ISR
handlers (Wiegand, ZyAura). ESPHome does not call millis() from ISR
on ESP32.
- Remove 'No __umulsidi3' claim from top comment — the rare path
does use it via /1000. Reworded to 'pure 32-bit ops on the common
path'.
- Extract MILLIS_RARE_PATH_THRESHOLD_US and US_PER_MS constants to
replace magic numbers 10000 and 1000.
Arduino's delay() uses esp_suspend() with a one-shot os_timer for
efficient single-suspension waiting. Our replacement polls millis()
with optimistic_yield(), which also enters esp_schedule/esp_suspend
via yield() but resumes repeatedly. Functionally correct for ESPHome
(delay is cold path, SDK tasks run via yield), just less power-
efficient for long delays.
Split the μs→ms conversion into two paths to keep the interrupt-
disabled critical section bounded:
- Common path (delta < 10 ms): while loop runs at most 10 iterations
(~100 ns). This covers the normal hot-path case where millis() is
called thousands of times per second.
- Rare path (delta >= 10 ms): constant-time multiply-by-reciprocal
via /1000 (compiled to __umulsidi3, ~2.5 μs). Only fires after a
long block (WiFi scan, boot, component stall), where the extra
latency is negligible relative to the block that caused it.
Addresses the concern that a multi-second WiFi scan block could cause
the unbounded while loop to hold interrupts for tens of microseconds.
Merge the separate millis_accumulator() function directly into millis()
to eliminate the extra function call overhead and prologue/epilogue.
Pack the three statics into a struct so the compiler loads one base
address instead of three literal pool entries.
Saves ~12 bytes of IRAM (87+6 → 81 bytes).
Arduino's micros64() reads the same overflow tracking statics that
init() sets up. Wrapping init() as a no-op would cause micros64()
to silently return wrong values after 71 minutes. The overflow timer
cost is negligible (~3 μs per 60 seconds).
Arduino's delay(0) calls yield() once — used as a yield point by some
code. Our millis-based loop would skip yielding entirely since
0 < 0 is false. Add explicit delay(0) → yield() handling.
Arduino's ::delay() is implemented in core_esp8266_wiring.cpp alongside
millis(). Because __delay calls millis() within the same compilation
unit, --wrap=millis can't intercept that intra-object reference, which
prevents the linker from garbage-collecting the original millis body
(~80 bytes of IRAM).
Replace esphome::delay() with a simple yield loop that uses our fast
millis() accumulator. This has the same behavior as Arduino's delay:
feeds the watchdog and processes SDK tasks via yield(), keeping WiFi
alive. With __delay unreferenced, the linker can now GC both __delay
and the original millis function.
Arduino ESP8266's millis() uses 4x 64-bit multiplies with magic
constants to convert system_get_time() to ms while tracking overflow.
On the LX106 (no hardware multiply-high instruction), each 64-bit
multiply goes through the __umulsidi3 software helper — costing
~3.3 us per call.
Replace with a simple accumulator that tracks a running millis counter
from system_get_time() deltas using pure 32-bit integer ops (subtract,
add, compare, subtract). No 64-bit math, no __umulsidi3.
Use -Wl,--wrap=millis to intercept all ::millis() calls site-wide so
Arduino libraries and ISR handlers (Wiegand, ZyAura) also get the fast
version. Brief interrupt disable (~125 ns) protects the static state.
Overflow safety: unsigned 32-bit delta arithmetic handles the 71-minute
system_get_time() wrap correctly for one wrap. ESPHome calls millis()
thousands of times per second, so missing a full wrap is not a realistic
concern. At boot, both the accumulator and system_get_time() start at 0,
so no special initialization is needed.
Benchmarked on real ESP8266 hardware:
Before: 3348 ns/call (Arduino 4x 64-bit multiply)
After: ~800-900 ns/call (accumulator, estimated)
ESPHome sets CONFIG_FREERTOS_HZ=1000, so one FreeRTOS tick equals one
millisecond and xTaskGetTickCount() returns ms directly. This replaces
esp_timer_get_time() + micros_to_millis() (a hardware timer read plus
a 64-bit multiply-shift conversion) with a single volatile memory read.
Benchmarked on real ESP32 hardware:
Before: 686 ns/call (esp_timer_get_time + micros_to_millis)
After: ~20 ns/call (volatile read of xTickCount)
millis() is called 1+N times per main loop iteration (once at the top
and once per component via WarnIfComponentBlockingGuard::finish()), so
on a 5-component device this saves ~3.4 μs per loop iteration.
millis_64() is left on esp_timer_get_time() for full 64-bit μs
precision — it is only called once per loop by the Scheduler.
micros() is also unchanged (runtime_stats needs μs precision).
Falls back to the original implementation if CONFIG_FREERTOS_HZ != 1000
(non-standard user override).
- get_app_state(): say 'STATUS_LED_* only', not 'STATUS_LED_* and lifecycle'
since lifecycle bits are no longer maintained in app_state_.
- STATUS_LED_SETTLE_S: remove incorrect mention of feed_wdt re-dispatch;
status_led_light is driven by the main loop, not feed_wdt.
- snapshot_led service: say 'status_led_light output' not 'pin state'
since the fixture uses a template output, not GPIO.
Extract the 0.3s magic sleep into a named constant explaining why that
duration is chosen (feed_wdt re-dispatches every ~3 ms; 300 ms gives
~100 opportunities). Fix the idle-write check to snapshot AFTER the
clear instead of before it, so writes still in-flight from the error
phase don't inflate the delta.
Verifies that after clearing all status flags, re-setting a flag makes
status_led_light resume writing to its output. Guards against a future
idle optimization (like #15642) where status_led disables its own
loop() when idle: if the re-enable path were broken, the second set
would not produce writes.
Also checks that writes STOP after all flags are cleared (counter
should not keep growing), proving status_led_light correctly stops
blinking in steady state.
Drop the trailing underscore from any_component_has_status_flag now
that the method is public. Trailing underscores in the codebase are
reserved for protected/private members per clang-tidy naming rules,
which caused a CI failure on the previously-named public helper.
Add integration tests covering:
- Single-component status_set/clear for warning and error
- Multi-component OR semantics (both clear orders)
- Warning and error independence
- End-to-end proof that status_led_light::loop() reads App.app_state_
and writes its output when the bits are set (via a fake template
output whose write_action bumps a counter exposed as a sensor)
Application::loop() used to, on every iteration of the inner component
loop, OR each component's state into a local accumulator and then OR
that into this->app_state_, finally overwriting this->app_state_ with
the accumulator after the loop. That was ~5 instructions per component
per loop iteration on the hot path, paying to rebuild state that's
already kept current elsewhere.
STATUS_LED_WARNING / STATUS_LED_ERROR are the only app_state_ bits
anything reads. Component::set_status_flag_() already writes them to
App.app_state_ directly the moment they're set, so any reader (status_led
components, feed_wdt's LED re-dispatch) already sees them without the
per-iter OR. The per-iter OR only ever re-set bits that were already
there — pure dead code on the hot path.
The clear path is handled by moving the work into
status_clear_warning/error_slow_path_(): on a real set->clear transition
the helpers walk App.components_ once to check if any other component
still has the flag, and clear App.app_state_ if not. Clears are rare
relative to loop iterations (a logged transition event), so O(N) per
clear is far cheaper than O(N) per loop.
Setup subtlety: Application::setup()'s slow-setup busy-wait forces
STATUS_LED_WARNING on to blink the LED while a slow component initializes.
A component clear during setup would prematurely wipe that forced bit,
so status_clear_warning_slow_path_() gates its walk-and-clear behind
APP_STATE_SETUP_COMPLETE (a free bit on app_state_ — zero RAM cost).
Application::setup() reconciles the warning state once at the end and
sets the flag. STATUS_LED_ERROR is never forced so its clear path
always walks.
Also fixes a pre-existing bug: the old per-iter accumulation only iterated
looping_components_, so non-looping components' status bits would get
clobbered by the end-of-loop 'app_state_ = new_app_state' overwrite.
The walk-on-clear approach iterates components_ (all of them), so
non-looping components are respected.
Separate Application::feed_wdt() into two entry points so the hot path
callers stop paying for the time==0 check they never trigger:
- feed_wdt_with_time(time): inline, hot path. Rate-limit check in 3
Xtensa instructions (load + sub + branch). [[unlikely]] tells the
compiler the slow branch is rare so the common path stays
fall-through.
- feed_wdt(): cold, out of line. Fetches millis() and forwards through
the same rate limit. Used by setup loops, upload helpers, yield(),
and any other non-hot caller.
feed_wdt_slow_() is now pure worker code — 11 bytes. It just calls
arch_feed_wdt(), updates last_wdt_feed_, and runs the status LED
re-dispatch. Both entries have already confirmed the rate limit was
exceeded before calling.
Hot call sites updated:
- Application::loop() per-component feed
- Scheduler::execute_item_() after each scheduled item runs
- Application::teardown_components() inner loop (already has 'now')
Cleaner than the early-return form — the action (calling feed_wdt_slow_)
reads as the body of the conditional instead of falling through past a
guard clause. Logically identical and compiles to the same code.
The main loop used to feed the watchdog unconditionally right after
Scheduler::call() returned, regardless of whether the scheduler had any
actual work to do. On an idle device this meant every outer loop
iteration paid the inline rate-limit check (load + sub + branch) for no
benefit.
Move the feed into Scheduler::execute_item_() so it fires only after a
scheduled callback actually runs, and covers both the main heap path
and the defer queue path (both go through execute_item_). This also
bounds the max feed gap during a burst of back-to-back scheduled items
by max(item_runtime) instead of sum(item_runtime).
The top-of-loop feed in Application::before_loop_tasks_() is now
unnecessary — when Scheduler::call does no work, the only elapsed time
is the sleep wake plus a few instructions, and when it does have work,
it fed the wdt as it went.
Split Application::feed_wdt() into an ALWAYS_INLINE wrapper that checks
the 3ms rate limit against last_wdt_feed_ and a feed_wdt_slow_() callee
that performs the actual arch_feed_wdt() + status LED re-dispatch.
Callers on the hot path (loop_task before/after each component) that
already have a millis() timestamp in hand now pay only a load + sub +
branch on the no-op path instead of a full call8 / entry / retw.
Moves the rate-limit state from a function-local static to a class
member (last_wdt_feed_) so the inline can access it.
PR #15639 removed Application::register_socket / monitored_sockets_
on the fast-select path before the ota-disable-loop-when-idle merge.
The merge re-introduced the function definitions without the matching
declarations. Drop them.
Following the zephyr_mcumgr precedent, keep the single-instance pointer
as a file-scope static inside ota_esphome.cpp instead of plumbing an
ota_wake_component_ slot and setter through Application. The extern "C"
wake trampoline also lives in the same TU now, so nothing in core/
touches OTA-specific state.
Guard the C and C++ call sites on USE_OTA_PLATFORM_ESPHOME (added as a
-D flag from components/esphome/ota/__init__.py) instead of the broader
ESPHOME_USE_OTA that base ota/ used to emit — the trampoline symbol only
exists when the esphome OTA platform is actually compiled in.
Also drop the dead host yield_with_select_ wake hook (host has no OTA
platform today).