After the Phase A / Phase B split in this PR, an external producer that
called wake_loop_threadsafe() (MQTT RX, USB RX, BLE event, espnow,
camera, mWW, speakers, USB host/CDC, lwip socket, enable_loop_soon_any_context)
only got Phase A — the component phase stayed gated by loop_interval_,
so the producer's component loop() could be delayed by up to
loop_interval_ ms before draining its queued work. That breaks the
long-standing semantic of wake_loop_threadsafe().
Add a wake_request flag set by every wake_loop_* entry point and
exchange-cleared at the gate in Application::loop(). When the flag is
set, force Phase B regardless of loop_interval_.
Storage is conditional on the threading model:
- ESPHOME_THREAD_MULTI_ATOMICS: std::atomic<uint8_t> (uint8_t, not
bool, because GCC on Xtensa generates an indirect call for
atomic<bool> ops — same workaround as scheduler.h)
- ESPHOME_THREAD_SINGLE / ESPHOME_THREAD_MULTI_NO_ATOMICS: volatile
uint8_t (8-bit aligned loads/stores are atomic on every supported
MCU; the platform signal that follows wake_request_set provides the
cross-thread/cross-core memory barrier)
Helpers (wake_request_set / wake_request_take) are always_inline so
IRAM_ATTR call sites stay in IRAM. Set BEFORE the platform signal so the
consumer is guaranteed to see the flag on its next gate check.
Adds an integration test that raises loop_interval_ to 2s, snapshots a
counting component's loop count, spawns a std::thread that calls
App.wake_loop_threadsafe() after 50ms, and asserts the count increments
inside a 500ms observation window. Without the fix the count would not
move for ~2s.
The Phase A/B split cached `HighFrequencyLoopRequester::is_high_frequency()`
once at the top of Application::loop() and reused that value at the sleep
decision after Phase B. If a component calls
HighFrequencyLoopRequester::start() from inside its own loop(), the cached
value is stale when we pick the sleep length, so the request doesn't take
effect until the next tick.
Before the decoupling, the sleep decision called the function fresh at the
same point in the code, so an in-loop HF request was honored immediately.
This caching silently regressed that behavior. The regression matters a lot
more now than it would have previously: loop_interval_ is the power-saving
knob this PR was written to enable (raised multi-second values are a
documented use case), so under the new freedom the worst-case HF-request
latency can stretch from ~16 ms to several seconds.
Re-read is_high_frequency() fresh at the sleep decision. The gate check
above continues to use the cached value — there's no correctness benefit
to re-reading it there, and the "one read covers the whole tick" design
intent is preserved for the gate. The function is a trivial atomic-read
static; a second call is cheaper than the latency it prevents.
Flagged by Copilot review on application.h:728.
Doc and test updates from a code review of this PR:
- Correct the `tail_us == 0 on Phase A-only ticks` claim in the
Application::loop() comment and the RuntimeStatsCollector::record_loop_active
docstring. `loop_tail_start_us` is set to `loop_before_end_us`, and
`loop_now_us` is sampled later, so `tail_us` on Phase A-only ticks is
the small gate-check + record prefix — tiny but non-zero.
(Also flagged by Copilot on application.h:623 and runtime_stats.h:45.)
- Call out ESP8266 as the floor case in the WDT_FEED_INTERVAL_MS margin
table. Its soft WDT (~1.6 s) is the tightest margin at ~5x, so future
changes to the constant need to preserve comfortable headroom there.
- Tighten the test lower bound at tests/integration/test_loop_interval_decoupling.py
from `2 <= loop_delta <= 6` to `3 <= loop_delta <= 6`. Allowing 2 would
let a >50% slowdown from the 4-in-2s nominal pass as CI jitter, which
undermines the regression signal. 3 keeps the test honest while still
absorbing realistic CI jitter.
- Add a second integration test
(test_loop_interval_default_not_pulled_forward) that covers the inverse
direction: at the default loop_interval_ with a fast scheduler item
(5 ms — well under the old delay_time/2 = 8 ms floor), the component
phase must still run at ~62 Hz, not the pre-fix ~128 Hz. This locks
down the original 128 Hz → 62 Hz regression that motivated the PR.
Raising WDT_FEED_INTERVAL_MS to 300 ms in the previous commit also capped
the status_led update cadence to ~3 Hz, because status_led::loop() was
re-dispatched only from inside feed_wdt_slow_(). That distorts the error
blink pattern (ERROR_PERIOD_MS = 250 ms with 150 ms on-window in
status_led.cpp) — at a 300 ms dispatch interval, the LED can be sampled
entirely inside or entirely outside the on-window on any given period,
turning a readable blink into an aliased one.
Split the two rate limits so they evolve independently:
WDT_FEED_INTERVAL_MS = 300 ms — arch_feed_wdt() rate limit
STATUS_LED_DISPATCH_INTERVAL_MS = 100 ms — status_led loop() dispatch
The feed_wdt_with_time() hot path now has two independent gate checks
(a load + sub + branch each). Fast path on both misses is the common case
and remains cheap. feed_wdt_slow_() no longer touches status_led; the
status_led re-dispatch moves into a new service_status_led_slow_() that's
compiled in only when USE_STATUS_LED is set.
No change to behavior on devices without status_led. Devices with
status_led get the intended LED cadence restored (100 ms dispatch sits
below the 150 ms error on-window and well below the 250 ms warning on-
window).
Flagged by Copilot review on application.h:242.
The 3 ms rate limit was tight enough that the outer feed_wdt_with_time()
at the top of Application::loop() hit the slow path on nearly every
iteration. On-device measurements showed 95-99.9% of iterations
triggering esp_task_wdt_reset(), costing ~9-12 us/hit (BT proxies worse
due to WDT spinlock contention with the BT task on the other core).
Much worse under wake-storm workloads: with loop_interval_ raised for
power saving and an external stack (OpenThread in the reported case)
posting frequent wake notifications, the main loop can spin at thousands
of Hz doing almost nothing but feeding the watchdog — burning ~26% CPU
in the extreme case reported on #15792.
Raising the threshold to 300 ms:
- Normal 62 Hz loop: feeds once every ~19 iterations (~3 Hz) instead
of every iteration.
- Wake-storm case: feeds ~3 Hz regardless of wake rate.
- Any operation exceeding 300 ms still triggers a real feed right
after it finishes (the post-component and post-scheduler-item
feeds naturally clear the gate).
Safety margins vs platform watchdog timeouts remain large: 16x on ESP32
(5 s task WDT), 5x on ESP8266 soft WDT (1.6 s), 20x on ESP8266 HW WDT.
scheduler_tick_ now returns the scheduler's advanced timestamp (free via
PR #15830's Scheduler::call return). Previously we still used the
pre-scheduler millis() for `elapsed = now - last_loop_`, which
underestimated elapsed time by whatever the scheduler dispatch took.
Adopt the returned value as `now` so the gate check, WDT feed, runtime
stats, and sleep computation all see consistent post-scheduler time.
Drops the obsolete "we deliberately reuse pre-scheduler now" comment —
that rationale was predicated on saving a millis() call, which no longer
applies.