The marker comment was being emitted as the first line *inside* each
IIFE:
[]() {
// === logger ===
// logger:
// ...
...
}();
That works but buries the component label inside the lambda body, so
scanning generated main.cpp to find "where does component X's setup
live" is harder than it needs to be. Emit the marker before and after
the IIFE instead:
// === logger ===
[]() {
// logger:
// ...
...
}();
// === logger ===
Comment-only components (e.g. sha256, async_tcp, empty platforms like
binary_sensor:) don't grow a useless trailing duplicate marker —
when there's no IIFE to bracket, the marker is emitted once.
Some components (sha256, async_tcp, network, empty text_sensor:, etc.)
emit only a ComponentMarker plus config-dump comments and no actual
C++ statements. Wrapping those in a `[]() { ... }();` IIFE is pure
clutter in the generated main.cpp — the IIFE has no body.
When _wrap_in_iifes sees a chunk whose lines are all // comments,
emit them verbatim instead of wrapping. Peak stack and flash are
unchanged on apollo and neargaragedoor since GCC was already
eliding the empty IIFEs; this just makes the generated code read
cleanly to humans.
Additional measurements showed GCC's -Os inliner re-inlines most IIFE
chunks back into setup() by choice, and the structural scoping alone
captures nearly all of the peak-stack benefit on esp32 without the
flash cost of forcing all chunks to stay as real functions.
Apollo (esp32-s3, -Os) with vs without noinline:
peak setup stack 176 B (noinline) vs 304 B (scope-only)
flash delta +388 B (noinline) vs -504 B (scope-only)
chunks kept 86 vs 20
Issue #15796 is an LVGL-setup class of bug that has only surfaced on
esp32 after years in the field; the extra guarantee that noinline
provides is not worth the flash cost in practice. Also rename the
helper from _wrap_in_noinline_iifes to _wrap_in_iifes to match.
The C++ standard-attribute spelling [[gnu::noinline]] placed between a
lambda's parameter list and body binds to the return type, not the
call operator. GCC 14 silently ignores it and emits -Wattributes
warnings at every chunk site. Switch to GCC's __attribute__((...))
syntax which binds to operator() as intended.
Measured impact on apollo-r-pro-1-eth (esp32-s3, -Os) vs the broken
[[gnu::noinline]] version: setup() frame 160 B -> 32 B, peak stack
304 B -> 176 B (another -42%). Flash grows by 888 B because all 86
chunks now stay as separate functions instead of GCC inlining the
small ones (which it was free to do when the attribute was ignored).
Net vs baseline -Os: peak stack 1264 B -> 176 B (-86%); flash
+388 B (<0.05% of a typical esp32 partition).
Generated setup() is a single monolithic function whose stack frame
scales super-linearly with config size. On a 5,943-line apollo build
the frame reached 1,264 B at -Os; extrapolation onto larger configs
(e.g. the 16k-line LVGL config in #15796) plausibly overflows the
8 KB loop task stack before safe_mode can increment its boot counter.
Emit a ComponentMarker sentinel at the start of each component's
to_code output, then have cpp_main_section wrap each component's
block (and sub-splits of up to 50 statements within each block) in a
noinline IIFE lambda. Each lambda's ENTRY frame is released on
return, bounding peak stack to setup() frame + max chunk frame.
Measured on apollo-r-pro-1-eth (esp32-s3, -Os):
setup() frame 1264 B -> 160 B
max chunk frame n/a -> 144 B
peak setup stack 1264 B -> 304 B (-76%)
total flash 792,471 B -> 791,995 B (-476 B)
The brace-depth guard in _wrap_in_noinline_iifes ensures we never
split between the RawStatement("{") / RawStatement("}") pair emitted
by cg.with_local_variable() (currently only wifi), so scoped locals
stay intact.
The ``pylint`` CI job runs ``pylint esphome`` — esphome/ only. The
local pre-commit hook had no ``files:`` restriction, so it also
linted ``tests/``, flagging pre-existing protected-access usage and
class-size issues that CI never sees. That blocked local commits on
warnings CI doesn't gate on.
Add ``files: ^esphome/.+\\.py$`` to the local pylint hook so its scope
matches the CI job exactly.
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: J. Nick Koston <nick@home-assistant.io>
Co-authored-by: J. Nick Koston <nick@koston.org>
The lambda in DelayAction::play_complex must be `mutable` so captured
copies of non-const reference arguments (e.g. std::string& from the
http_request on_response trigger) can bind to the non-const reference
parameters of play_next_(const Ts&...). Without `mutable`, the captured
members are const-qualified and cannot bind to `T&`.
Regression from #14968 which replaced std::bind with a lambda.
Fixes https://github.com/esphome/esphome/issues/15808
Users (notably PollingComponents with update_interval: 0ms) have
historically relied on set_interval(0) as a pseudo-loop() mechanism. A
literal interval=0 causes Scheduler::call() to spin — the item is
always 'due now' after re-scheduling, so the scheduler loop never
returns, starves the main loop, and triggers a WDT reset in the field.
Coerce interval=0 to 1ms at creation so existing code keeps working at
~1kHz instead of spinning, while still emitting the warning pointing
authors at HighFrequencyLoopRequester (the intended mechanism for
running fast in the main loop). Zero-delay timeouts (defer/set_timeout)
remain legitimate one-shots and are unaffected — defer is a one-shot,
not a spin risk.
Restructures Application::loop() into two independent phases to stop the
scheduler from silently pulling the component loop cadence forward.
Before: Application::loop() bounded its sleep by
min(loop_interval_ - elapsed, next_schedule_in())
with a delay_time/2 floor. Any scheduler item due sooner than
loop_interval_/2 dragged the whole component phase with it. On a typical
ESP32 config with default loop_interval_=16ms, combined scheduler
activity from api / esp32_ble / esp32_ble_tracker / debug was keeping
every component's loop() running at ~128 Hz instead of the documented
~62 Hz.
This has become more visible recently as more components convert to
PollingComponent (which uses set_interval internally) and more in-tree
code uses set_interval / set_timeout directly. Adding or removing any
scheduled item silently changed every other component's loop cadence.
App.set_loop_interval() for power savings was also silently defeated.
After:
- Phase A (every tick): drain wake notifications, run scheduler.call(),
feed WDT
- Phase B (gated by loop_interval_ or HighFrequencyLoopRequester):
iterate registered components and update last_loop_
Sleep = min(time-until-next-component-phase, next_schedule_in()). When
a scheduler event wakes us early, Phase A services it and the component
phase stays gated independently. loop_interval_ is now a true minimum
interval between component phases.
The delay_time/2 floor is removed. Any legitimate need to wake faster
than loop_interval_ has proper mechanisms:
- HighFrequencyLoopRequester for sustained fast-loop needs
- Application::wake_loop_threadsafe() from any context (new in 2026.4.0)
for one-shot wake-on-event
Also guards against set_interval(0) misuse — it asks the main loop to
spin forever, which was never the intended API. Warns at creation time
pointing authors at HighFrequencyLoopRequester. set_timeout(0)/defer()
is unaffected; zero-delay one-shots remain legitimate.
Runtime stats: process_pending_stats is now called on every tick (not
just when the component phase runs) so log_interval_ isn't quantized to
the component-phase cadence. Added an inline fast-path gate in
runtime_stats.h that early-outs unless now >= next_log_time_, keeping
Application::loop() slim; the log_stats_ work stays out-of-line.
Ordering constraints preserved:
- defer() callbacks still FIFO before components same-tick (Phase A
runs before Phase B)
- Scheduled items still execute before components when both due
- Scheduled callbacks still run on main thread only
- loop_component_start_time_ is still set fresh at each component's loop
- WDT is still fed at least once per tick
The idempotency check in _inject_keep returns content unchanged when the
marker is already present, so patched stayed 0 and the fail-hard
RuntimeError fired on every incremental rebuild after the first. Check
for the marker before attempting to patch and count it as success.
IRAM_ATTR is a no-op on BK72xx so there is nothing for the script to
patch or summarize. Guard the extra_scripts registration with a
COMPONENT_BK72XX check and drop the BK72xx variants from
KNOWN_VARIANTS.
Post-link was not registered for RTL8710B (condition was too narrow
after removing it from _PATCHERS_BY_VARIANT). Use a BK72xx exclusion
set instead so all IRAM-active variants get the summary.
Also fall back to TOOLCHAIN_PREFIX-nm when $NM is empty, which fixes
the empty summary on LN882H.
RTL8710B's stock linker already consumes *(.image2.ram.text*) into its
.ram_image2.text output (> BD_RAM), so hal.h can place IRAM_ATTR
functions directly into section(".image2.ram.text") without any linker
patching. RTL8720C already worked this way with section(".sram.text").
The patcher is now only needed for LN882H, whose stock linker has no
glob that catches ".sram.text" — we inject KEEP(*(.sram.text*)) into
.flash_copysection (> RAM0 AT> FLASH).
This removes the _RTL8710B_IMAGE2 regex, the RTL8710B entry from
_PATCHERS_BY_VARIANT, and simplifies the header comments.
Linker-generated interworking veneers (e.g. ___ZN...enable_loop_soon_
any_context_veneer at 0x9b062588) contain the same function name
substrings but live at unrelated addresses, producing a bogus multi-GB
range in the summary. Filter them out.