Every teardown path routes through release_services(), so the
materialized table cannot outlive its link - said at free_service_table_
so the next reviewer need not re-derive it.
- The rejection test registers the hub and matches the extra-keys error,
so the missing-tracker error can no longer satisfy it vacuously
- register_ble_node logs and drops loudly at capacity: a push_back past
a StaticVector's bound is a silent no-op, and an undersized slot count
must show at boot, not as an unresolvable node
- Every walker failure names the failing call and status; the descriptor
cap gets its own message so a >64-descriptor characteristic is
diagnosable instead of a permanent silent backoff loop
The schema split's accept side was tested; the reject side was not - a
key leaking from the legacy arm into the neutral schema would have
passed silently.
- The empty-table bailout ran after state_ went CONNECTED, so its
teardown's terminal report read as a completed connection and fired
on_disconnect with no preceding on_connect - the exact trigger
contract violation the sibling error branch avoids. The check now
runs first and the teardown resolves through connect_failed; a
node-less client still reaches CONNECTED through the single
assignment after the guard
- The esp32 arm claims its slot through consume_gatt_slot like every
other claimant (behavior-identical: it forwards to the same esp32
validator), so the ledger's one-spelling contract holds
- A zero-service table on a client with nodes is treated as a failed
discovery (warn, release, charge the backoff, disconnect): a real GATT
peer always exposes at least GAP/GATT, so an empty table means the
materialization failed and the connection must retry instead of
sitting inert behind a successful on_connect
- The fan-out abort guard also tests cancel_requested_, covering the
normal async teardown a node starts from on_connected (the state-only
check caught just the synchronous refusal)
- nodes_ becomes a StaticVector sized by ESPHOME_BLE_CLIENT_MAX_NODES:
the client requests a baseline slot, each neutral write action (which
registers itself in its constructor) requests one, and the define
rides defines.h for analysis - no realloc machinery on the rp2 target
- The Bluedroid empty-table return logs its quiet cases; the
service_table=False stub notes the misconfiguration it implies
A gate-refused send to a peer that is not api_connection_ (the
subscribe-time path during a handover) must not arm a retry that
loop() would aim at a different connection.
Overkill for its window: the address latch, drain, reuse guard and
session hygiene bought coverage for a case the client's own timeouts
and the retried connections-free state already heal. The notification
goes out when it can; the connections-free retry (the hardware-verified
fix) stays.
- Both retry latches clear on subscribe/unsubscribe: a new subscriber
must not receive a connected=false for an address from the previous
session (it resyncs through its own subscribe requests)
- The app-register failure path states its BLEClientBase parity so the
no-retry behavior reads as deliberate
The stack-down settle skipped release_services() locally, but its
report routes into the wrapper's reset which calls it anyway - the
warn survived. Gating the clean on an active stack inside
release_services() covers every path, and the stack-down branch goes
back to the plain call.
- The stack-down settle resets the stream latches directly instead of
calling release_services(): the dying stack invalidates its own cache,
and the newly checked cache_clean would warn on every OTA or
ble.disable with a live connection
- static_asserts pin the exactly-sized state bitfields so a future
enumerator truncates loudly at compile time
The connections-free retry left its sibling asymmetric: on the same
full buffer the slot could report free while the device still read as
connected, with nothing resending the connected=false. A one-deep latch
(address + error) drains at the same 100 ms cadence and re-latches on a
repeat failure; a re-reservation of the address clears it, since the
client re-requesting the connect has already acted on the disconnect
and a late resend would shadow the new connection. A lost
connected=true stays client-timeout territory.
An unknown or misspelled platform name yielded an empty set - the file
would be filtered out of every build (link failure at the end of a long
compile) and the subset guard passed vacuously.
The one IDF call in the backend that skipped check_and_log_error_. A
failed clean leaves a stale database the next connection could serve as
authoritative, with nothing in the log to point at it. (The call fails
only at dispatch - stack down - not for an empty cache, so this cannot
warn-spam routine teardowns.)
The collapse of the CONNECTING branch dropped a real race fix: HA
re-requesting a connect while a scheduled teardown was still pending
used to cancel the teardown and let the in-flight open complete;
without it the request was ignored and HA paid a teardown plus a fresh
connect during exactly the reconnect churn this path sees most.
cancel_gatt_disconnect() joins the contract: true only for a scheduled
teardown that has not started closing (Bluedroid clears the latched
want_disconnect_); rp2 and the stubs return false since their teardowns
start inside gatt_disconnect(). The wrapper's cancel_teardown() returns
its state to CONNECTING and the proxy handler carries the dev branch
verbatim.
- Both esp_ble_gap_set_prefer_conn_params calls route through
check_and_log_error_ like the class this PR replaces did; a rejected
preference leaves the link on the controller default interval, which
is exactly the WiFi-coex failure the shared constants exist to avoid,
so it must be visible in logs
- The connections-free retry moves below the 100 ms gate so a full TCP
buffer is retried at the loop cadence instead of every iteration
- The late-OPEN_EVT reclaim was the one unchecked IDF call in the
backend; a failed close there leaks a live link nothing tracks
- The mid-stream services_released_ park now leaves a trace like its
API-lost sibling, so a stuck GetServices is explicable from logs
Round 2 found no code issues - the round-1 mechanisms verify clean on
disk (every IDLE transition through the one door, the teardown-guard
trio minimal-complete per event, the error latch airtight, the reset
lists disjoint by construction). What remained was comment drift:
disconnect_pending() spelling, the OPEN_EVT and cached-MTU comments,
the contract header re-wrap, the two claim-a-slot docstrings
disambiguated, and ragged wraps from earlier text excisions.
- set_idle_() is the single door back to IDLE (bootstrap, open-fail and
the DISCOVERED park went through bare set_state), so the per-attempt
reset list holds only per-attempt latches
- CFG_MTU joins OPEN_EVT and SEARCH_CMPL in suppressing reports while a
teardown owns the link - one spelling of the guard across all three
events, and the wrapper's race arm becomes defense instead of the only
cover
- The connections-free retry drain compiles on every proxy build (the
advertisement-only arm sends the message too; the latch already did)
- latch_pending_error_ makes first-cause-wins the mechanism at all three
latch sites; the dead freed-slot refused branch and its retired
vocabulary go; the initiate_connection/start_connect_ pair collapses
- UNSET_CONN_ID hoisted next to the field it initializes; count-status
shadow renamed; stale busy-error rationale replaced with the real one
(a repeat call would re-arm the teardown timer)
- Fixture states the batch-grouping caveat like its rp2 sibling; the
get_service_table stub carries a greppable direct-consumer warning