With the lwip lock infrastructure in place from the TCP race fix,
the lock-free SPSC ring buffer's wasted slot is no longer needed.
Switch to a simple rx_count_ approach that uses all 4 queue slots.
Also add LWIP_LOCK() to all UDP methods that call lwip APIs
(bind, close, sendto, setsockopt, getsockopt, recvfrom, factories).
- Document port_host byte order convention in lwip_ip_to_sockaddr
- Note intentional method hiding in LWIPRawUDPRecvImpl
- Note recvfrom truncation differs from POSIX MSG_TRUNC
- Add (void) flags in sendto to clarify flags are ignored
Co-Authored-By: J. Nick Koston <nick@koston.org>
Add native UDP support to the lwip raw TCP socket layer used by
ESP8266 and RP2040, eliminating the need for Arduino WiFiUDP fallback.
Two new classes:
- LWIPRawUDPImpl: send-only UDP (8 bytes overhead)
- LWIPRawUDPRecvImpl: send+recv with fixed-size ring buffer (no heap
allocation in recv callback)
Factory functions: socket_udp(), socket_udp_recv(), socket_ip_udp(),
socket_ip_udp_recv() with UDPSocket/UDPRecvSocket type aliases.
Additive only — no consumer migration in this PR.
Same pattern as writev — write() calls internal_write_() then
internal_output_(), each acquiring the lock separately. Hold
the lock at the outer scope so inner calls just bump the
recursion counter.
tcp_new() is an lwip core API call that must be bracketed with
the lwip lock on RP2040 per pico-sdk docs. Add LWIP_LOCK() to
socket() and socket_listen() factory functions.
Avoid repeated lock acquire/release cycles per iovec element.
The recursive mutex re-entry in inner calls is nearly free (counter
bump), while the outer lock prevents the expensive IRQ disable/enable
on each iteration.
On RP2040 (Pico W), arduino-pico sets PICO_CYW43_ARCH_THREADSAFE_BACKGROUND=1,
which means lwip callbacks (recv_fn, accept_fn, err_fn) run from a PendSV
interrupt — not the main loop. This allows them to preempt read(), write(),
close(), and accept() at any point, causing race conditions on shared state
like the rx_buf_ pbuf chain.
The most critical race: recv_fn calls pbuf_cat(rx_buf_, pb) while read() is
freeing nodes in the same chain, leading to use-after-free and lwip's
"Creating an infinite loop" assertion panic. This is the root cause of #10681.
Fix: implement RP2040's LwIPLock (previously a no-op) to call
cyw43_arch_lwip_begin/end, which acquires the pico-sdk async_context recursive
mutex. Add LWIP_LOCK() guards to all main-loop lwip API call sites in the
socket layer.
On ESP8266, lwip callbacks run cooperatively from the main loop, so
LwIPLock remains a no-op.
Closes#10681
- Document port_host byte order convention in lwip_ip_to_sockaddr
- Note intentional method hiding in LWIPRawUDPRecvImpl
- Note recvfrom truncation differs from POSIX MSG_TRUNC
- Add (void) flags in sendto to clarify flags are ignored
Co-Authored-By: J. Nick Koston <nick@koston.org>
Add native UDP support to the lwip raw TCP socket layer used by
ESP8266 and RP2040, eliminating the need for Arduino WiFiUDP fallback.
Two new classes:
- LWIPRawUDPImpl: send-only UDP (8 bytes overhead)
- LWIPRawUDPRecvImpl: send+recv with fixed-size ring buffer (no heap
allocation in recv callback)
Factory functions: socket_udp(), socket_udp_recv(), socket_ip_udp(),
socket_ip_udp_recv() with UDPSocket/UDPRecvSocket type aliases.
Additive only — no consumer migration in this PR.
- Add SO_SNDTIMEO to OTA socket to prevent blocking writes from
stalling the WDT when the TCP send buffer is full
- Add SO_SNDTIMEO as no-op in raw TCP (writes never block)
- Use delay(0) instead of delay(1) in readall_() since SO_RCVTIMEO
already handles the wait
- Keep delay(1) in writeall_() since raw TCP writes are non-blocking
and would spin on EWOULDBLOCK without it
Replace the non-blocking poll + delay(1) pattern in OTA data transfer
with SO_RCVTIMEO blocking reads. The socket now wakes immediately when
data arrives instead of sleeping 1ms between polls.
Adds SO_RCVTIMEO support to the raw TCP socket implementation (ESP8266,
RP2040) using the existing socket_delay()/socket_wake() infrastructure.
The timeout is stored as a uint8_t in centiseconds, fitting in existing
struct padding with zero RAM cost.
Tested OTA improvements across platforms:
- ESP32-S3: ~15% faster (6.96-7.76s -> 5.87-6.60s)
- LibreTiny RTL: 24% faster (18.84s -> 14.33s)
- LibreTiny BK72xx: 56% faster (55.52s -> 24.38s)
- ESP8266: ~1% faster (compressed OTA, already efficient)
Split the 256-byte UART read buffer into a separate noinline
read_and_send_() helper so the common "no data" path in loop()
only needs a 32-byte stack frame instead of 288 bytes.
Also reorder checks so api_connection_ == nullptr bails out
first, avoiding the more expensive disconnect detection when
no client is subscribed.
APIBuffer::resize() is already just a capacity check + store,
making the outer size-equality guard redundant. Saves ~8 bytes
of code size across both frame helpers.