8 frames was cutting off useful call stack information. 16 frames
costs an additional 32 bytes of .noinit RAM and covers typical
ESPHome call chains which can be 10-12 frames deep.
The esphome logs command via API wasn't running process_stacktrace
on received log lines, so crash handler backtrace addresses were
displayed but not decoded with addr2line.
- Make s_raw_crash_data static with inline .noinit definition (no extern needed)
- Remove valid=false clearing from crash_handler_log() so both serial (boot)
and API (subscribe) paths can emit the crash data
- Clear valid flag after logging to prevent re-logging on API reconnects
- Cache esp_cpu_process_stack_pc result to avoid redundant call
- Remove unused <cstdint> include from header
When an ESP32 crashes, the backtrace is printed to UART and lost.
Users without a serial cable never see this diagnostic information.
This adds a crash handler that:
- Intercepts esp_panic_handler() via --wrap linker flag
- Captures the faulting PC and backtrace into .noinit memory
- Supports both Xtensa (ESP32/S2/S3) and RISC-V (C3/C6/H2/C2)
- Logs crash data at boot via ESP_LOGE (serial output)
- Re-logs when HA subscribes to logs (visible in HA log viewer)
- Adds CLI stacktrace decoding for the new log format
tcp_pcb_listen is a smaller struct than tcp_pcb — calling tcp_recv()
or tcp_err() on it writes past the struct boundary. Revert the listen
PCB close to plain tcp_close(), which is synchronous for listen PCBs
(no async callbacks to worry about).
ESP8266's settimeofday() returns EINVAL (22) directly as the return
value when the timezone parameter is non-NULL, rather than following
POSIX convention of returning -1 and setting errno. The previous
fallback code checked errno == EINVAL which never matched because
errno was never set, so the retry with nullptr never triggered.
Fix by always passing nullptr on ESP8266 since the platform requires it.
Clear LWIP callbacks (tcp_arg, tcp_recv, tcp_err) before calling
tcp_close() or tcp_abort() to prevent use-after-free.
After tcp_close(), the PCB remains alive during the TCP close
handshake (FIN_WAIT, TIME_WAIT states). If LWIP calls recv/err
callbacks during this period and the socket object has already
been destroyed, the callback writes to freed memory, corrupting
the heap.
This was observed as umm_malloc_core crashes on ESP8266 during
rapid API client connect/disconnect cycles — the heap free-list
got corrupted by a dangling callback writing to a freed
LWIPRawImpl object.
Extract pcb_detach_abort() and pcb_detach_close() helpers to
ensure all close/abort sites consistently clear callbacks first.
Co-authored-by: Jonathan Swoboda <swoboda1337@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: pre-commit-ci-lite[bot] <117423508+pre-commit-ci-lite[bot]@users.noreply.github.com>
On ESP32 with CONFIG_LWIP_TCPIP_CORE_LOCKING, bypass lwip_setsockopt()
for TCP_NODELAY by directly modifying tcp_pcb->flags under the TCPIP
core lock. This eliminates ~1091 bytes of overhead per call (socket
lookups, hook, switch cascade, refcounting) for what is just a single
bit flip.
The API frame helper toggles Nagle's algorithm on every message send
via set_nodelay_for_message(), making this a hot path. The fast path
reduces set_nodelay_raw_ from calling the full lwip_setsockopt to just
acquiring the mutex, loading 3 pointers, and flipping the TF_NODELAY
bit.
Only enabled when both USE_LWIP_FAST_SELECT (cached lwip_sock pointer)
and CONFIG_LWIP_TCPIP_CORE_LOCKING (real mutex protection) are
available. Falls back to the standard setsockopt call otherwise.