[scheduler] Enable lock-free fast-path on ESPHOME_THREAD_MULTI_NO_ATOMICS

The _empty_() helpers (to_add_empty_, defer_empty_, to_remove_empty_)
forced the lock path on ESPHOME_THREAD_MULTI_NO_ATOMICS by hardcoding
`return false`. That made Scheduler::call() pay a FreeRTOS mutex
round-trip for each of process_defer_queue_ / process_to_add /
cleanup_ on every idle tick just to confirm "nothing to do".

On the only NO_ATOMICS target (BK72xx — ARMv5TE, single-core), an
aligned 32-bit load is atomic at the hardware level. Mark the three
skip-work counters volatile so the compiler cannot cache or elide the
read, and let _empty_() compare against zero directly. Writers still
hold lock_ for any RMW — that invariant is unchanged.

A stale 0 is benign: the counter is checked on every Scheduler::call()
iteration, so a missed update is caught next tick. Same pattern as the
NO_ATOMICS reads in time_64.cpp.

On BK72xx at ~3100 iter/min with ~8us/mutex this reclaims roughly
75ms/min of main-loop overhead. Measured on BK7238/BK7231N while
profiling alongside libretiny-eu/libretiny#360.

ATOMICS and SINGLE paths are unchanged (SINGLE keeps plain uint32_t,
no volatile-read overhead).
This commit is contained in:
J. Nick Koston
2026-04-23 05:47:52 -05:00
parent 43a371caab
commit c414cc393f
+23 -23
View File
@@ -524,30 +524,30 @@ class Scheduler {
std::vector<SchedulerItem *> to_add_;
#ifndef ESPHOME_THREAD_SINGLE
// Fast-path counter for process_to_add() to skip taking the lock when there is
// nothing to add. Uses std::atomic on platforms that support it, plain uint32_t
// otherwise. On non-atomic platforms, callers must hold the scheduler lock when
// mutating this counter. Not needed on single-threaded platforms where we can
// check to_add_.empty() directly.
// Fast-path counter for process_to_add() to skip taking the lock when there
// is nothing to add. std::atomic on ATOMICS; volatile uint32_t on NO_ATOMICS
// (aligned 32-bit reads are atomic on ARMv5TE — BK72xx — and volatile
// prevents the compiler caching/eliding the read). On NO_ATOMICS, callers
// must hold lock_ for any RMW mutation. Not needed on SINGLE.
#ifdef ESPHOME_THREAD_MULTI_ATOMICS
std::atomic<uint32_t> to_add_count_{0};
#else
uint32_t to_add_count_{0};
volatile uint32_t to_add_count_{0};
#endif
#endif /* ESPHOME_THREAD_SINGLE */
// Fast-path helper for process_to_add() to decide if it can try the lock-free path.
// - On ESPHOME_THREAD_SINGLE: direct container check is safe (no concurrent writers).
// - On ESPHOME_THREAD_MULTI_ATOMICS: performs a lock-free check via to_add_count_.
// - On ESPHOME_THREAD_MULTI_NO_ATOMICS: always returns false to force the caller
// down the locked path; this is NOT a lock-free emptiness check on that platform.
// Fast-path helper for process_to_add() to decide if it can skip the lock.
// - SINGLE: direct container check (no concurrent writers).
// - ATOMICS: lock-free load of to_add_count_.
// - NO_ATOMICS: volatile read. A stale 0 is benign — next call() iteration
// observes the update; RMW mutation is still under lock_.
bool to_add_empty_() const {
#ifdef ESPHOME_THREAD_SINGLE
return this->to_add_.empty();
#elif defined(ESPHOME_THREAD_MULTI_ATOMICS)
return this->to_add_count_.load(std::memory_order_relaxed) == 0;
#else
return false;
return this->to_add_count_ == 0;
#endif
}
@@ -580,20 +580,20 @@ class Scheduler {
std::vector<SchedulerItem *> defer_queue_; // FIFO queue for defer() calls
size_t defer_queue_front_{0}; // Index of first valid item in defer_queue_ (tracks consumed items)
// Fast-path counter for process_defer_queue_() to skip lock when nothing to process.
// Fast-path counter for process_defer_queue_() to skip lock when nothing to
// process. See to_add_count_ above for the volatile rationale on NO_ATOMICS.
#ifdef ESPHOME_THREAD_MULTI_ATOMICS
std::atomic<uint32_t> defer_count_{0};
#else
uint32_t defer_count_{0};
volatile uint32_t defer_count_{0};
#endif
bool defer_empty_() const {
// defer_queue_ only exists on multi-threaded platforms, so no ESPHOME_THREAD_SINGLE path
// ESPHOME_THREAD_MULTI_NO_ATOMICS: always take the lock
#ifdef ESPHOME_THREAD_MULTI_ATOMICS
return this->defer_count_.load(std::memory_order_relaxed) == 0;
#else
return false;
return this->defer_count_ == 0;
#endif
}
@@ -615,23 +615,23 @@ class Scheduler {
#endif /* ESPHOME_THREAD_SINGLE */
// Counter for items marked for removal. Incremented cross-thread in cancel_item_locked_().
// On ESPHOME_THREAD_MULTI_ATOMICS this is read without a lock in the cleanup_() fast path;
// on ESPHOME_THREAD_MULTI_NO_ATOMICS the fast path is disabled so cleanup_() always takes the lock.
// Counter for items marked for removal. Incremented cross-thread in
// cancel_item_locked_(). See to_add_count_ above for the volatile rationale
// on NO_ATOMICS.
#ifdef ESPHOME_THREAD_MULTI_ATOMICS
std::atomic<uint32_t> to_remove_{0};
#elif defined(ESPHOME_THREAD_MULTI_NO_ATOMICS)
volatile uint32_t to_remove_{0};
#else
uint32_t to_remove_{0};
uint32_t to_remove_{0};
#endif
// Lock-free check if there are items to remove (for fast-path in cleanup_)
bool to_remove_empty_() const {
#ifdef ESPHOME_THREAD_MULTI_ATOMICS
return this->to_remove_.load(std::memory_order_relaxed) == 0;
#elif defined(ESPHOME_THREAD_SINGLE)
return this->to_remove_ == 0;
#else
return false; // Always take the lock path
return this->to_remove_ == 0;
#endif
}