[api] Revert push_item inline — CodSpeed shows 46% regression

Inlining push_item into add_item bloated the dedup loop, causing
the compiler to make worse inlining decisions. Total time went from
44µs to 64.5µs. Restore out-of-line with __attribute__((flatten))
which keeps add_item's dedup loop tight while still inlining
push_back's callees into push_item.
This commit is contained in:
J. Nick Koston
2026-04-01 09:51:55 -10:00
parent 2b8d4af1e7
commit 18b00fb3fc
2 changed files with 6 additions and 1 deletions
@@ -2072,6 +2072,8 @@ void APIConnection::on_fatal_error() {
this->flags_.remove = true;
}
void __attribute__((flatten)) APIConnection::DeferredBatch::push_item(const BatchItem &item) { items.push_back(item); }
void APIConnection::DeferredBatch::add_item(EntityBase *entity, uint8_t message_type, uint8_t estimated_size,
uint8_t aux_data_index) {
// Check if we already have a message of this type for this entity
+4 -1
View File
@@ -647,7 +647,10 @@ class APIConnection final : public APIServerConnectionBase {
uint8_t aux_data_index = AUX_DATA_UNUSED);
// Add item to the front of the batch (for high priority messages like ping)
void add_item_front(EntityBase *entity, uint8_t message_type, uint8_t estimated_size);
void push_item(const BatchItem &item) { items.push_back(item); }
// Out-of-line with flatten: inlines push_back callees into push_item,
// but keeps push_item itself as a call from add_item. This prevents
// the compiler from bloating add_item's dedup loop and degrading icache.
void push_item(const BatchItem &item);
// Clear all items
void clear() {