ESPHome's WiFiIPStateListener only notifies on IP acquisition (GOT_IP events), not
on IP loss — on disconnect, only the WiFiConnectStateListener's disconnect path
fires (see wifi_component_esp8266.cpp:952-962 and wifi_component_pico_w.cpp:340).
The previous commit's `ip_was_up_` transition tracking was broken: after the first
IP-up event, `ip_was_up_` latched to true and never reset, so subsequent
disconnect+reconnect cycles would see has_ip=true && ip_was_up_=true and skip
re-arming the polling window.
Fix: always re-arm on any IP notification. The scheduler's set_interval/set_timeout
with a uint32_t ID already performs atomic cancel-and-add for matching IDs
(Scheduler::set_timer_common_ line 232-234), so start_polling_window_ is idempotent
and needs no explicit cancel. Drop the ip_was_up_ field and cancel_polling_window_
helper entirely.
The !has_ip branch (cancel on disconnect) was dead code: it would never fire because
the listener doesn't receive disconnect events. Removing it; the polling window will
naturally expire on its own (at most 12s of harmless MDNS.update() calls during a
disconnect that isn't followed by reconnect within the window).
The Arduino LEAmDNS library only has meaningful timer-driven work during the
~9 s probe+announce phase following MDNS.begin() or _restart(): 3 probes at
250 ms + 8 announcements at 1000 ms, then all internal timeouts are set to
resetToNeverExpires(). Incoming packets are handled via the lwIP UDP RX
callback independently of update(). ESPHome does not issue service queries,
so the query cache path is always a no-op.
The previous implementation ran set_interval(50) forever — ~20 dispatches/sec,
1200+ scheduler calls per minute of pure overhead once probing completed.
This PR arms a bounded MDNS_POLL_WINDOW_MS (12 s) polling window driven by
WiFiIPStateListener events. A fresh window covers each probe/announce cycle
(boot, wifi reconnect, or internal _restart() triggered by netif changes);
outside the window there are zero scheduler dispatches and the scheduler
heap contains no mDNS items.
ESP8266 is WiFi-only in the Arduino build so the path is unconditional.
RP2040 supports W5500 ethernet without WiFi, so the listener is requested
only when WiFi is in the config; ethernet-only RP2040 builds keep the
legacy polling loop.
Scheduler IDs use uint32_t (MDNS_POLL_ID / MDNS_POLL_STOP_ID) to avoid the
name-hash/strcmp cost of string-named timers on the cancel + re-arm paths.
Address copilot review: wrap_to_code is called once per component, so
the lazy `from esphome import yaml_util` inside it was doing a
(cached) sys.modules lookup per component. Move the import up to
generate_cpp_contents, pass the module down, rename the helper to
_wrap_to_code to mark it private.
CI measured 90.6ms after the lazy-import refactor lands. Set baseline to
91ms with 15% margin (ceiling 104.65ms, ~14ms headroom for GHA variance
over the observed measurement).
The top-level `import esphome.core.config` in `esphome/loader.py` (used
only to register the "esphome" pseudo-component in `_COMPONENT_CACHE`)
was pulling `esphome.automation` + `esphome.config_validation` eagerly
into every `esphome` CLI invocation, even fast paths like
`esphome version`. Move the cache registration into `_lookup_module`,
triggered on first `get_component("esphome")`.
Also defers the two `esphome.core.config` call sites inside
`esphome/config.py` (CoreFinalValidateStep and validate_config), which
were the other eager entry point.
Local `import esphome.__main__` drops from ~49ms to ~46ms; more
importantly, `esphome.core.config` / `esphome.automation` no longer
appear in the top offenders table, and `esphome.config_validation`
drops from 4.7ms self to 2.0ms self on CI.
CI now measures 81.2ms after the zeroconf/writer/yaml_util lazy loads.
Reset baseline to 82ms with the 15% margin (ceiling 94.3ms, ~13ms
headroom for GHA variance).
Every \`esphome\` CLI invocation pays the cost of whatever \`esphome/__main__.py\`
imports at module scope before the requested command even runs. This
moves three heavy imports into the functions that actually use them:
- \`esphome.zeroconf\` (discover_mdns_devices): only needed for the
name_add_mac_suffix OTA discovery path.
- \`esphome.writer\`: only needed by compile/clean paths.
- \`esphome.yaml_util\`: only needed by codegen, config dump, and rename.
Local measurement drops from ~75ms to ~47ms (-37%) for a cold
\`python -c 'import esphome.__main__'\`. The zeroconf chain alone accounts
for most of the gain — the case that motivated the CI budget check.
Also adds a module-level note explaining the intent so future PRs don't
innocently promote these back to the top.
Retargets four test patches from \`esphome.__main__.discover_mdns_devices\`
to \`esphome.zeroconf.discover_mdns_devices\` now that the symbol is only
bound inside the function that calls it.
- determine-jobs: trigger on changes to requirements_test.txt too; that
file is hashed into the venv cache key and installed during
restore-python, so a change there can alter the import environment.
- ci.yml: merge the --check and --har steps so we only run
importtime-waterfall once per job (was measuring twice: ~7-8s of
wasted CI time per run). Uses the new script/check_import_time.py
--check --har <path> combination; the HAR reflects the same
measurement that produced the pass/fail decision.
- script/check_import_time.py: refactor the CLI so --har is a standalone
option rather than a mutually-exclusive mode. --check and --update
each accept an optional --har PATH that writes the HAR from the same
subprocess invocation. Plain --har is still supported for local use.
- tests: add tests/script/test_check_import_time.py covering HAR parsing,
root lookup, offender ranking/dedup, budget round-trip, and the three
--check exit paths (pass, regression, missing budget) plus the new
--check --har combined write. Add requirements_test.txt case to the
should_run_import_time parametrized test.
The import-time job previously did a `pip install` on every run (~1s
and some noisy output). Moving the pin to `requirements_test.txt` bakes
it into the venv that `common` already builds and caches, so the job
restores it from the cache and skips the install step.
Seeing the top contributors every run — regression or not — gives
reviewers and future-PR authors a running reference for which imports
are currently most expensive. That's the data we need to prioritize
lazy-import refactors, so there's no reason to hide it on pass.
CI measured 126.2ms against 123ms baseline + 25% ceiling (153.8ms). That's
27.6ms of headroom — enough slack that a real regression could hide under
it. Drop to 15% (ceiling 141.5ms, ~15ms headroom over observed). Still
well above GHA runner variance, tight enough to catch the class of
regression we care about (zeroconf-style top-level import additions).
The previous parser walked `-X importtime` stderr by hand (string-splitting
on `|`, indent-width math). Replace it with a load of the HAR JSON that
importtime-waterfall already produces: each entry carries the module name,
self-time, and cumulative — no tree reconstruction needed. Drops the
hand-rolled retry loop too, since importtime-waterfall does best-of-6
internally.
The committed budget (75.2ms) was seeded on a fast local machine; GHA
runners measured 122.8ms, so the first CI run tripped the check. Reset
the baseline to 123ms with a 25% margin (ceiling ~154ms) to absorb GHA
variance. Tighten later once we have several data points.
Adds unit tests for should_run_import_time across the trigger matrix and
wires the new mock into the existing test_main_* suite.
Adds a CI gate that runs `python -X importtime -c "import esphome.__main__"`
via importtime-waterfall's best-of-N harness, compares the root cumulative
time against a checked-in budget (script/import_time_budget.json, seeded
at 75.2ms with a 15% margin), and fails the build when top-level imports
regress.
The CLI pays this cost on every invocation before the requested command
even runs, so silently adding a top-level dep chain (the recent zeroconf
move in #13135 being the motivating case) hurts every user. The check
gives us a signal without waiting for user reports.
- script/check_import_time.py: --check / --update / --har PATH modes.
On regression, prints a ranked top-15 offenders table by self-time.
- script/import_time_budget.json: baseline + margin_pct.
- script/determine-jobs.py: should_run_import_time() gates the job on
esphome/**/*.py, requirements.txt, requirements_dev.txt, pyproject.toml,
or changes to the check itself.
- .github/workflows/ci.yml: new import-time job, runs when gated and
uploads a waterfall HAR artifact (14-day retention) for inspection.
In time_64.cpp's true-rollover branch, bump the just-loaded `major`
local first and __atomic_store_n that value to millis_major, instead
of reading the global again for the store expression. Equivalent
under the held lock; clearer and avoids a second read.
scheduler.h indent: the NO_ATOMICS #else branch bodies read at 2
spaces while the sibling #ifdef/#elif branches read at 4. clang-format
refuses to normalise these consistently — every manual re-indent to 4
spaces gets reverted by the hook. Leaving as clang-format produces it.
Two fixes Copilot flagged on #15947:
1. Use __atomic_load_n to re-read last_millis under the lock, not a
plain read. The forward-progression branch (else if) writes
last_millis with __atomic_store_n without holding lock, so the
under-lock plain read would race with it and be UB in the C++
memory model.
2. Reload major from millis_major after acquiring the lock. The
unlocked load at the top of the function can be stale by the time
we get the lock: another thread may have completed a rollover
between the unlocked load and the lock acquisition, leaving our
local major behind by one. Without reload the function could
return a 64-bit timestamp that jumps backwards by ~2^32 ms (~49.7
days). The MULTI_ATOMICS branch already handles this via its retry
loop; NO_ATOMICS just reloads under the lock.
Two fixes Copilot flagged on #15947:
1. Use __atomic_load_n to re-read last_millis under the lock, not a
plain read. The forward-progression branch (else if) writes
last_millis with __atomic_store_n without holding lock, so the
under-lock plain read would race with it and be UB in the C++
memory model.
2. Reload major from millis_major after acquiring the lock. The
unlocked load at the top of the function can be stale by the time
we get the lock: another thread may have completed a rollover
between the unlocked load and the lock acquisition, leaving our
local major behind by one. Without reload the function could
return a 64-bit timestamp that jumps backwards by ~2^32 ms (~49.7
days). The MULTI_ATOMICS branch already handles this via its retry
loop; NO_ATOMICS just reloads under the lock.
Switch writer-side plain stores under lock_ to __atomic_store_n with
__ATOMIC_RELAXED on NO_ATOMICS. The input value for RMW is read
plainly (safe — only writers mutate, serialised by the lock; readers
only atomic-load so two reads don't race). Closes the formal C++
memory-model hole where plain-store vs atomic-load was a data race
in the standard even though aligned 32-bit STR/LDR on ARMv5TE is
atomic in practice.
Applies to scheduler.h counter mutators and the under-lock writes
to last_millis / millis_major in time_64.cpp's near-rollover branch.
Same ARMv5TE codegen (plain STR). ATOMICS / SINGLE paths unchanged.
Use `#if defined(X)` / `#elif defined(Y)` / `#else` for the three-way
ATOMICS / NO_ATOMICS / SINGLE split. Also fix the SINGLE-branch body
indentation to match the other branches.
The preceding commit needlessly rewrote comments that were still
accurate. Revert the prose-only changes; keep only the two line-level
code changes (__atomic_load_n on the unlocked reads, __atomic_store_n
on the unlocked write).
Same treatment as the scheduler counters (#15947):
- Unlocked reads of millis_major / last_millis at the top of
Millis64Impl::compute(): switch from plain reads to
__atomic_load_n(&..., __ATOMIC_RELAXED).
- Unlocked write of last_millis in the "normal forward progression"
branch: switch from plain assignment to __atomic_store_n(...,
__ATOMIC_RELAXED). This is the one write that happens without the
lock, so it needs to be formally atomic to pair cleanly with the
unlocked atomic reader in the C++ memory model.
- Writes under `lock` stay plain (millis_major++, last_millis = now
inside the near-rollover branch). The lock serialises them against
other writers.
On ARMv5TE the builtins compile to plain LDR/STR — same codegen, no
libatomic dependency. Updates the "accepting minor races" comment to
describe the formally-defined version of the race.
Walk back the __atomic_store_n on the writer paths — the mutators
already hold lock_, so plain counter_++/=/+=/-- is sufficient to
serialise against other writers. The reader fast-path still uses
__atomic_load_n(&counter, __ATOMIC_RELAXED) to express concurrent-
read intent in the C++ memory model and keep the compiler from
caching/eliding the read. On ARMv5TE it compiles to a plain LDR —
same codegen as before.
Copilot review on #15947 flagged that `volatile uint32_t` is not a
well-defined concurrent access in the C++ memory model — it prevents
the compiler caching/eliding the read, but does not turn a plain
cross-thread read/write pair into a defined access. Technically still
a formal data race even though aligned 32-bit LDR/STR on ARMv5TE is
atomic at the hardware level.
Switch the NO_ATOMICS counter reads and writes to GCC's atomic
builtins with __ATOMIC_RELAXED:
- Readers: __atomic_load_n(&counter, __ATOMIC_RELAXED)
- Writers (under lock_): __atomic_store_n(&counter, new_value,
__ATOMIC_RELAXED)
- Increment/decrement (under lock_): explicit load + compute +
__atomic_store_n. Lock_ serialises the load-modify-store against
other writers; the atomic ops make the write visible to concurrent
readers in the memory model.
On ARMv5TE these builtins compile to plain LDR/STR — same codegen as
the previous volatile approach, and no libatomic dependency (only RMW
builtins like __atomic_fetch_add would need the lib). ATOMICS and
SINGLE paths are unchanged.
Rename the seven counter RMW mutators to carry the `_locked_` suffix
that matches the existing convention (pop_raw_locked_,
is_item_removed_locked_, cancel_item_locked_, etc.):
to_add_count_increment_ -> to_add_count_increment_locked_
to_add_count_clear_ -> to_add_count_clear_locked_
defer_count_increment_ -> defer_count_increment_locked_
defer_count_clear_ -> defer_count_clear_locked_
to_remove_add_ -> to_remove_add_locked_
to_remove_decrement_ -> to_remove_decrement_locked_
to_remove_clear_ -> to_remove_clear_locked_
The caller-must-hold-lock contract became load-bearing when the
underlying counters became volatile on NO_ATOMICS: ++/+=/-- compile to
a three-instruction LDR/OP/STR sequence that is not atomic against a
concurrent RMW from another task, so the lock is what keeps the
counter consistent. The new suffix makes the requirement explicit at
every call site, matching how the rest of the scheduler documents the
same invariant.
No behavioural change; all call sites already hold lock_.
The _empty_() helpers (to_add_empty_, defer_empty_, to_remove_empty_)
forced the lock path on ESPHOME_THREAD_MULTI_NO_ATOMICS by hardcoding
`return false`. That made Scheduler::call() pay a FreeRTOS mutex
round-trip for each of process_defer_queue_ / process_to_add /
cleanup_ on every idle tick just to confirm "nothing to do".
On the only NO_ATOMICS target (BK72xx — ARMv5TE, single-core), an
aligned 32-bit load is atomic at the hardware level. Mark the three
skip-work counters volatile so the compiler cannot cache or elide the
read, and let _empty_() compare against zero directly. Writers still
hold lock_ for any RMW — that invariant is unchanged.
A stale 0 is benign: the counter is checked on every Scheduler::call()
iteration, so a missed update is caught next tick. Same pattern as the
NO_ATOMICS reads in time_64.cpp.
On BK72xx at ~3100 iter/min with ~8us/mutex this reclaims roughly
75ms/min of main-loop overhead. Measured on BK7238/BK7231N while
profiling alongside libretiny-eu/libretiny#360.
ATOMICS and SINGLE paths are unchanged (SINGLE keeps plain uint32_t,
no volatile-read overhead).
yield_with_select_ was a trivial one-line passthrough to
esphome::internal::wakeable_delay(). Remove the wrapper and call
wakeable_delay() directly at the two call sites.