The collapse of the CONNECTING branch dropped a real race fix: HA
re-requesting a connect while a scheduled teardown was still pending
used to cancel the teardown and let the in-flight open complete;
without it the request was ignored and HA paid a teardown plus a fresh
connect during exactly the reconnect churn this path sees most.
cancel_gatt_disconnect() joins the contract: true only for a scheduled
teardown that has not started closing (Bluedroid clears the latched
want_disconnect_); rp2 and the stubs return false since their teardowns
start inside gatt_disconnect(). The wrapper's cancel_teardown() returns
its state to CONNECTING and the proxy handler carries the dev branch
verbatim.
- Both esp_ble_gap_set_prefer_conn_params calls route through
check_and_log_error_ like the class this PR replaces did; a rejected
preference leaves the link on the controller default interval, which
is exactly the WiFi-coex failure the shared constants exist to avoid,
so it must be visible in logs
- The connections-free retry moves below the 100 ms gate so a full TCP
buffer is retried at the loop cadence instead of every iteration
- The late-OPEN_EVT reclaim was the one unchecked IDF call in the
backend; a failed close there leaks a live link nothing tracks
- The mid-stream services_released_ park now leaves a trace like its
API-lost sibling, so a stuck GetServices is explicable from logs
Round 2 found no code issues - the round-1 mechanisms verify clean on
disk (every IDLE transition through the one door, the teardown-guard
trio minimal-complete per event, the error latch airtight, the reset
lists disjoint by construction). What remained was comment drift:
disconnect_pending() spelling, the OPEN_EVT and cached-MTU comments,
the contract header re-wrap, the two claim-a-slot docstrings
disambiguated, and ragged wraps from earlier text excisions.
- set_idle_() is the single door back to IDLE (bootstrap, open-fail and
the DISCOVERED park went through bare set_state), so the per-attempt
reset list holds only per-attempt latches
- CFG_MTU joins OPEN_EVT and SEARCH_CMPL in suppressing reports while a
teardown owns the link - one spelling of the guard across all three
events, and the wrapper's race arm becomes defense instead of the only
cover
- The connections-free retry drain compiles on every proxy build (the
advertisement-only arm sends the message too; the latch already did)
- latch_pending_error_ makes first-cause-wins the mechanism at all three
latch sites; the dead freed-slot refused branch and its retired
vocabulary go; the initiate_connection/start_connect_ pair collapses
- UNSET_CONN_ID hoisted next to the field it initializes; count-status
shadow renamed; stale busy-error rationale replaced with the real one
(a repeat call would re-arm the teardown timer)
- Fixture states the batch-grouping caveat like its rp2 sibling; the
get_service_table stub carries a greppable direct-consumer warning
Log lines ride the same API connection; a debug line at the exact
moment the TCP buffer is full adds traffic when it can least afford it
(the api layer's own buffer-full log is verbose for the same reason).
Field bug on a Pico W: the slot freed device-side but the API client
kept free=0 with the address still allocated. send_message drops the
response when the TCP buffer is full (a boot storm makes that likely
right when HA reconnects), and nothing ever resent it - the client's
slot state is verbatim the last response received, so it stayed stale
until reboot. The cached response is current by construction, so a
pending bit retried from loop() is an idempotent resync; the
subscribe-time send heals through the same bit.
- The esp32 arm's ESPHOME_BLE_GATT_CLIENT_COUNT shadowed the unanchored
regex; the pin now searches the USE_RP2 platform block
- The contract states what the wrapper relies on: nonzero from
gatt_disconnect means nothing to tear down, an accepted teardown
always reaches a terminal report
A client with only connect/disconnect automations paid the build/free
cycle on every reconnect for nothing - post-setup heap churn on the
Bluedroid direct-consumer path. release_services() stays unconditional.
notify_characteristic() had no completion path to the node that asked;
the hook lands now while the surface has no external users, with the
client fanning out after its breadcrumb.
The contract test's own #define only covers its translation unit;
ble_gatt_client.cpp needs the define from the build to link the
find_characteristic/find_cccd definitions the lookup tests call.
The Bluedroid on-demand materializer, the neutral lookup helpers' host
tests, and the USE_BLE_GATT_SERVICE_TABLE define move here, next to
their first consumer (the neutral engine's table resolution and the
service_table codegen flag). One behavior fix rides along: hitting
MAX_DESCRIPTORS_PER_CHARACTERISTIC now fails the walk like every other
inconsistency instead of truncating the table silently.
discover_services() already returned the contract's not-connected value;
the read/write/notify/pair ops fell through to an arbitrary stack error
instead. Dead for the proxy (the wrapper gates on connected()), live for
a direct consumer racing an op against a teardown.
Nothing in this PR emits USE_BLE_GATT_SERVICE_TABLE, so the Bluedroid
materializer, the neutral find_service/find_characteristic/find_cccd
helpers, their host tests, and the defines.h entry were scaffolding with
no caller here. They move to the PR that introduces their first
consumer. The contract op stays: get_service_table() keeps its {} stub
(the proxy streams in place).
- on_notify_state and on_pairing_result are handled on the neutral
client: a failed registration or pairing logs instead of vanishing
into the default no-op on a frozen surface
- A node tearing the link down during the connect fan-out leaves a
warning, since on_disconnect then fires with no preceding on_connect
- An INVALID_OFFSET/NOT_FOUND before the count from the same cache is a
contradiction, not an end-of-range: the walker fails the build and the
streamer aborts with the real cause instead of silently truncating
(the walker's descriptor loop keeps its terminator - the 64 cap is its
only bound)
- unconditional_disconnect_'s unset-conn-id path drives the slot to a
terminal reported state instead of leaning on the scheduled-teardown
timer
- find_characteristic/find_cccd log a corrupt range before returning the
not-found sentinel (moved out of line - no ESP_LOG in headers)
- connect()/disconnect() log a rejected user action like the legacy
engine instead of silently ignoring it
- The service_table docstring states the flag is Bluedroid-only forward
scaffolding and that rp2's proxy hub must keep its materializer
regardless (it streams through get_service_table())
- Review-round comments tightened to repo style
- handle_search_cmpl_ takes the event status: a failed discovery no
longer reads as a clean zero-service success (the empty cached
database satisfies the count calls), and the table invalidation plus
param step-down still run on the failure path
- A late successful OPEN_EVT on a slot the teardown net already gave up
is closed instead of leaking a live controller link with no conn id
- The counting pass logs its failure like the two paths below it
- static_assert pins the wrapper's compile-time streamer detection; a
signature drift would silently fall back to a table proxy builds
compile without
- find_characteristic/find_cccd range math widened to 32-bit (correct by
type; the wrapped-end case was already loop-safe)
- The fan-out early return releases the borrowed table itself instead of
relying on the backend's own teardown release
- A registered non-esp32 backend platform without a slot cap now fails
validation loudly instead of failing open (the test pin remains)
- new_gatt_backend's docstring says the service_table flag is forward
scaffolding for the first esp32 direct consumer, not load-bearing on
any current build
Both backends report the no-response completion (Bluedroid at
WRITE_CHAR_EVT, rp2 synthesized or via the can-send path); the neutral
write action already relies on it.
The measured layout showed 56+48 = 104 B/slot vs dev's 96. Two changes
that only pay together (the wrapper is 8-aligned, so nothing under a
full 8 helps):
- The tail packs into 2 bytes of bitfields; remote_addr_type_ becomes a
start_connect_ parameter (written and read on adjacent lines only)
- The wrapper's duplicate 10 s teardown timer is gone: the backend owns
the whole safety window. Its timer now also arms for a teardown
scheduled during CONNECTING (the one case the wrapper covered alone),
a late OPEN_EVT on a given-up slot closes the link instead of
resurrecting it, and the wrapper's transient-refusal branch collapses
because both backends return nonzero only when already idle
48 + 48 = 96 B/slot, the split at dev parity.
The measured layout showed 56+48 = 104 B/slot vs dev's 96. Two changes
that only pay together (the wrapper is 8-aligned, so nothing under a
full 8 helps):
- The tail packs into 2 bytes of bitfields; remote_addr_type_ becomes a
start_connect_ parameter (written and read on adjacent lines only)
- The wrapper's duplicate 10 s teardown timer is gone: the backend owns
the whole safety window. Its timer now also arms for a teardown
scheduled during CONNECTING (the one case the wrapper covered alone),
a late OPEN_EVT on a given-up slot closes the link instead of
resurrecting it, and the wrapper's transient-refusal branch collapses
because both backends return nonzero only when already idle
48 + 48 = 96 B/slot, the split at dev parity.
- SEARCH_CMPL honors the event's own status: a failed discovery leaves
an empty cached database that the count calls read as a clean zero,
which would have become an authoritative empty service list
- The loop stays enabled while a link exists and settles only at IDLE -
the stack-down recovery needs a tick to run in ESTABLISHED, which the
settle-on-established optimization was silently blocking
- The pre-started search uses the OPEN_EVT conn id (replaced-class
parity)
- read_descriptor and notify_characteristic passthroughs join the frozen
surface so the first migrated node can subscribe without extending the
client; the comment now says which ops have callers today
- The choke-point test also validates through the public
BLE_CLIENT_SCHEMA, so removing the cv.All wiring fails a test
- DOMAIN no longer sits between the BTstack comment and the constant it
explains
- A refused MTU request no longer wedges the connection: OPEN_EVT
reports with the default MTU so the consumer proceeds (previously no
CFG_MTU_EVT meant no connected report and a slot that never freed)
- Disabling the BLE stack settles a live link with a connected=false
report before forgetting it, so the consumer frees its slot instead of
holding a phantom connection
- A refused security response answers the pairing request with the
failure instead of hanging it
- The service-table walk mismatch logs its reason before discarding
- The shared streamer's two bounds-check aborts use abort_service_stream
with an ATT Unlikely Error cause (new shared GATT_ERR_UNLIKELY)
- A refused MTU request no longer wedges the connection: OPEN_EVT
reports with the default MTU so the consumer proceeds (previously no
CFG_MTU_EVT meant no connected report and a slot that never freed)
- Disabling the BLE stack settles a live link with a connected=false
report before forgetting it, so the consumer frees its slot instead of
holding a phantom connection
- A refused security response answers the pairing request with the
failure instead of hanging it
- The service-table walk mismatch logs its reason before discarding
- The shared streamer's two bounds-check aborts use abort_service_stream
with an ATT Unlikely Error cause (new shared GATT_ERR_UNLIKELY)
The three coordinated flags become one SearchState (NONE / PRESTARTED /
PRESTART_DONE / CLAIMED / REPORT_PENDING): illegal combinations are
unrepresentable, delivery and the two reset sites collapse to single
assignments, and the requested-meets-done subtlety becomes a named
state. A second discover_services() while one is in flight now returns 0
instead of issuing a duplicate search (one completion is already owed).
Same RAM: the enum sits in a 4-bit bitfield in the same flags byte.
- The connected fan-out stops if a node tears the link down mid-loop
(latent until a backend settles a teardown synchronously)
- A synchronously refused disconnect logs its own warning - a backend/
client state divergence is no longer the most quietly logged outcome
- An action-initiated connect before any sighting warns that the public
address type is assumed (an opaque repeated failure becomes
self-diagnosing); the connect() doc no longer mentions a config key
that does not exist
- The explicit-connections ledger test asserts the exact charged count;
the host-test override comment names the current gate macro