Rename the seven counter RMW mutators to carry the `_locked_` suffix
that matches the existing convention (pop_raw_locked_,
is_item_removed_locked_, cancel_item_locked_, etc.):
to_add_count_increment_ -> to_add_count_increment_locked_
to_add_count_clear_ -> to_add_count_clear_locked_
defer_count_increment_ -> defer_count_increment_locked_
defer_count_clear_ -> defer_count_clear_locked_
to_remove_add_ -> to_remove_add_locked_
to_remove_decrement_ -> to_remove_decrement_locked_
to_remove_clear_ -> to_remove_clear_locked_
The caller-must-hold-lock contract became load-bearing when the
underlying counters became volatile on NO_ATOMICS: ++/+=/-- compile to
a three-instruction LDR/OP/STR sequence that is not atomic against a
concurrent RMW from another task, so the lock is what keeps the
counter consistent. The new suffix makes the requirement explicit at
every call site, matching how the rest of the scheduler documents the
same invariant.
No behavioural change; all call sites already hold lock_.
The _empty_() helpers (to_add_empty_, defer_empty_, to_remove_empty_)
forced the lock path on ESPHOME_THREAD_MULTI_NO_ATOMICS by hardcoding
`return false`. That made Scheduler::call() pay a FreeRTOS mutex
round-trip for each of process_defer_queue_ / process_to_add /
cleanup_ on every idle tick just to confirm "nothing to do".
On the only NO_ATOMICS target (BK72xx — ARMv5TE, single-core), an
aligned 32-bit load is atomic at the hardware level. Mark the three
skip-work counters volatile so the compiler cannot cache or elide the
read, and let _empty_() compare against zero directly. Writers still
hold lock_ for any RMW — that invariant is unchanged.
A stale 0 is benign: the counter is checked on every Scheduler::call()
iteration, so a missed update is caught next tick. Same pattern as the
NO_ATOMICS reads in time_64.cpp.
On BK72xx at ~3100 iter/min with ~8us/mutex this reclaims roughly
75ms/min of main-loop overhead. Measured on BK7238/BK7231N while
profiling alongside libretiny-eu/libretiny#360.
ATOMICS and SINGLE paths are unchanged (SINGLE keeps plain uint32_t,
no volatile-read overhead).
yield_with_select_ was a trivial one-line passthrough to
esphome::internal::wakeable_delay(). Remove the wrapper and call
wakeable_delay() directly at the two call sites.
socket.h is unused after #15931 moved the host wake mechanism to
wake.cpp. lwip_fast_select.h is already included via application.h
under the same guard.
BK72xx silicon requires a ~200us busy-wait (sctrl_dpll_delay200us) between
two watchdog register key writes on every reload, making each
arch_feed_wdt() call ~300us on BK7231T/N/BK7238. This is hardware errata
in the BDK's wdt_ctrl (beken378/driver/wdt/wdt.c:WCMD_RELOAD_PERIOD) and
cannot be worked around at the SDK level.
LibreTiny initialises the BK72xx HW watchdog at 10000ms, but ESPHome's
generic WDT_FEED_INTERVAL_MS of 300ms was sized for ESP8266's 1.6s soft
watchdog. That left BK72xx over-servicing the watchdog ~33x per timeout
window, paying ~60ms/min of main-loop overhead to no benefit.
Raise the interval to 2000ms on BK72xx only, which keeps a 5x safety
margin on the 10s HW WDT — matching the ESP8266 ratio that originally
motivated the 300ms value — while cutting feed frequency ~6x.
Measured on a NiceMCU XH-WB3S (BK7238) while testing
libretiny-eu/libretiny#360:
Before (300ms interval):
wdt_slow_path: hits=195 (5.9% of iters, avg=315.97us/hit)
main_loop_before_breakdown: sched=73.2ms, wdt=61.6ms, residual=0.0ms
Other platforms retain the existing 300ms value.
BDK 3.0.78 (required by LibreTiny for BK7238 support, see
libretiny-eu/libretiny#360) declares wifi_event_sta_disconnected_t in
wlan_defs_pub.h, which collides with the identically named typedef in
LibreTiny's Arduino WiFi API (WiFiEvents.h). Rename the BDK version
across the include so both headers can coexist. ESPHome only uses
bk_wlan_get_link_status from this header and doesn't reference the
renamed type.
The rename is a no-op on BDK 3.0.33 (BK7231T/N) since that version
doesn't declare the typedef.
The host platform's wake plumbing was split between wake.cpp (which had
wake_loop_threadsafe()) and application.{h,cpp} (which owned the fd
registry, wake socket lifecycle, select() loop, and readiness query).
The split required application.h to friend wake_loop_threadsafe() and
expose the wake socket fd, and forced a host-only out-of-line override
of yield_with_select_() that was equivalent to wakeable_delay() on every
other platform.
Consolidate the host wake mechanism in wake.h/wake.cpp:
- New host-only API: wake_register_fd, wake_unregister_fd, wake_fd_ready,
wake_setup, wake_drain_notifications.
- Host implementation of internal::wakeable_delay now performs the
select() over registered fds, mirroring the role of xTaskNotifyTake on
ESP32 and esp_delay on ESP8266.
- Application drops the host-specific socket_fds_, max_fd_,
wake_socket_fd_, socket_fds_changed_, read_fds_, base_read_fds_
members, the host include block (sys/select.h, sys/socket.h, etc.),
and the friend declarations for wake_loop_threadsafe() and
socket::socket_ready_fd().
- yield_with_select_() collapses to a single inline definition for all
platforms that delegates to wakeable_delay().
- Socket factories (bsd_sockets_impl.cpp, lwip_sockets_impl.cpp) and
socket::socket_ready_fd() now call the wake.h API directly instead of
going through App.