- The empty-table bailout ran after state_ went CONNECTED, so its
teardown's terminal report read as a completed connection and fired
on_disconnect with no preceding on_connect - the exact trigger
contract violation the sibling error branch avoids. The check now
runs first and the teardown resolves through connect_failed; a
node-less client still reaches CONNECTED through the single
assignment after the guard
- The esp32 arm claims its slot through consume_gatt_slot like every
other claimant (behavior-identical: it forwards to the same esp32
validator), so the ledger's one-spelling contract holds
- A zero-service table on a client with nodes is treated as a failed
discovery (warn, release, charge the backoff, disconnect): a real GATT
peer always exposes at least GAP/GATT, so an empty table means the
materialization failed and the connection must retry instead of
sitting inert behind a successful on_connect
- The fan-out abort guard also tests cancel_requested_, covering the
normal async teardown a node starts from on_connected (the state-only
check caught just the synchronous refusal)
- nodes_ becomes a StaticVector sized by ESPHOME_BLE_CLIENT_MAX_NODES:
the client requests a baseline slot, each neutral write action (which
registers itself in its constructor) requests one, and the define
rides defines.h for analysis - no realloc machinery on the rp2 target
- The Bluedroid empty-table return logs its quiet cases; the
service_table=False stub notes the misconfiguration it implies
A gate-refused send to a peer that is not api_connection_ (the
subscribe-time path during a handover) must not arm a retry that
loop() would aim at a different connection.
Overkill for its window: the address latch, drain, reuse guard and
session hygiene bought coverage for a case the client's own timeouts
and the retried connections-free state already heal. The notification
goes out when it can; the connections-free retry (the hardware-verified
fix) stays.
- Both retry latches clear on subscribe/unsubscribe: a new subscriber
must not receive a connected=false for an address from the previous
session (it resyncs through its own subscribe requests)
- The app-register failure path states its BLEClientBase parity so the
no-retry behavior reads as deliberate
The stack-down settle skipped release_services() locally, but its
report routes into the wrapper's reset which calls it anyway - the
warn survived. Gating the clean on an active stack inside
release_services() covers every path, and the stack-down branch goes
back to the plain call.
- The stack-down settle resets the stream latches directly instead of
calling release_services(): the dying stack invalidates its own cache,
and the newly checked cache_clean would warn on every OTA or
ble.disable with a live connection
- static_asserts pin the exactly-sized state bitfields so a future
enumerator truncates loudly at compile time
The connections-free retry left its sibling asymmetric: on the same
full buffer the slot could report free while the device still read as
connected, with nothing resending the connected=false. A one-deep latch
(address + error) drains at the same 100 ms cadence and re-latches on a
repeat failure; a re-reservation of the address clears it, since the
client re-requesting the connect has already acted on the disconnect
and a late resend would shadow the new connection. A lost
connected=true stays client-timeout territory.
An unknown or misspelled platform name yielded an empty set - the file
would be filtered out of every build (link failure at the end of a long
compile) and the subset guard passed vacuously.
The one IDF call in the backend that skipped check_and_log_error_. A
failed clean leaves a stale database the next connection could serve as
authoritative, with nothing in the log to point at it. (The call fails
only at dispatch - stack down - not for an empty cache, so this cannot
warn-spam routine teardowns.)
The collapse of the CONNECTING branch dropped a real race fix: HA
re-requesting a connect while a scheduled teardown was still pending
used to cancel the teardown and let the in-flight open complete;
without it the request was ignored and HA paid a teardown plus a fresh
connect during exactly the reconnect churn this path sees most.
cancel_gatt_disconnect() joins the contract: true only for a scheduled
teardown that has not started closing (Bluedroid clears the latched
want_disconnect_); rp2 and the stubs return false since their teardowns
start inside gatt_disconnect(). The wrapper's cancel_teardown() returns
its state to CONNECTING and the proxy handler carries the dev branch
verbatim.
- Both esp_ble_gap_set_prefer_conn_params calls route through
check_and_log_error_ like the class this PR replaces did; a rejected
preference leaves the link on the controller default interval, which
is exactly the WiFi-coex failure the shared constants exist to avoid,
so it must be visible in logs
- The connections-free retry moves below the 100 ms gate so a full TCP
buffer is retried at the loop cadence instead of every iteration
- The late-OPEN_EVT reclaim was the one unchecked IDF call in the
backend; a failed close there leaks a live link nothing tracks
- The mid-stream services_released_ park now leaves a trace like its
API-lost sibling, so a stuck GetServices is explicable from logs
Round 2 found no code issues - the round-1 mechanisms verify clean on
disk (every IDLE transition through the one door, the teardown-guard
trio minimal-complete per event, the error latch airtight, the reset
lists disjoint by construction). What remained was comment drift:
disconnect_pending() spelling, the OPEN_EVT and cached-MTU comments,
the contract header re-wrap, the two claim-a-slot docstrings
disambiguated, and ragged wraps from earlier text excisions.
- set_idle_() is the single door back to IDLE (bootstrap, open-fail and
the DISCOVERED park went through bare set_state), so the per-attempt
reset list holds only per-attempt latches
- CFG_MTU joins OPEN_EVT and SEARCH_CMPL in suppressing reports while a
teardown owns the link - one spelling of the guard across all three
events, and the wrapper's race arm becomes defense instead of the only
cover
- The connections-free retry drain compiles on every proxy build (the
advertisement-only arm sends the message too; the latch already did)
- latch_pending_error_ makes first-cause-wins the mechanism at all three
latch sites; the dead freed-slot refused branch and its retired
vocabulary go; the initiate_connection/start_connect_ pair collapses
- UNSET_CONN_ID hoisted next to the field it initializes; count-status
shadow renamed; stale busy-error rationale replaced with the real one
(a repeat call would re-arm the teardown timer)
- Fixture states the batch-grouping caveat like its rp2 sibling; the
get_service_table stub carries a greppable direct-consumer warning
Log lines ride the same API connection; a debug line at the exact
moment the TCP buffer is full adds traffic when it can least afford it
(the api layer's own buffer-full log is verbose for the same reason).
Field bug on a Pico W: the slot freed device-side but the API client
kept free=0 with the address still allocated. send_message drops the
response when the TCP buffer is full (a boot storm makes that likely
right when HA reconnects), and nothing ever resent it - the client's
slot state is verbatim the last response received, so it stayed stale
until reboot. The cached response is current by construction, so a
pending bit retried from loop() is an idempotent resync; the
subscribe-time send heals through the same bit.
- The esp32 arm's ESPHOME_BLE_GATT_CLIENT_COUNT shadowed the unanchored
regex; the pin now searches the USE_RP2 platform block
- The contract states what the wrapper relies on: nonzero from
gatt_disconnect means nothing to tear down, an accepted teardown
always reaches a terminal report
A client with only connect/disconnect automations paid the build/free
cycle on every reconnect for nothing - post-setup heap churn on the
Bluedroid direct-consumer path. release_services() stays unconditional.
notify_characteristic() had no completion path to the node that asked;
the hook lands now while the surface has no external users, with the
client fanning out after its breadcrumb.
The contract test's own #define only covers its translation unit;
ble_gatt_client.cpp needs the define from the build to link the
find_characteristic/find_cccd definitions the lookup tests call.
The Bluedroid on-demand materializer, the neutral lookup helpers' host
tests, and the USE_BLE_GATT_SERVICE_TABLE define move here, next to
their first consumer (the neutral engine's table resolution and the
service_table codegen flag). One behavior fix rides along: hitting
MAX_DESCRIPTORS_PER_CHARACTERISTIC now fails the walk like every other
inconsistency instead of truncating the table silently.
discover_services() already returned the contract's not-connected value;
the read/write/notify/pair ops fell through to an arbitrary stack error
instead. Dead for the proxy (the wrapper gates on connected()), live for
a direct consumer racing an op against a teardown.
Nothing in this PR emits USE_BLE_GATT_SERVICE_TABLE, so the Bluedroid
materializer, the neutral find_service/find_characteristic/find_cccd
helpers, their host tests, and the defines.h entry were scaffolding with
no caller here. They move to the PR that introduces their first
consumer. The contract op stays: get_service_table() keeps its {} stub
(the proxy streams in place).