Add the missing BRPC_WITH_UBRING Bazel configuration, wire it through
the library and example targets, and document the supported build
commands.
Enable RDMA and UBRING in CI jobs whose all-options configurations did
not exercise those features. Run both feature suites in the existing
Bazel unit-test jobs and enable RDMA in the Make unit-test job.
Use unsigned literals for UBRING atomic counter operations to match the
counter type and avoid template deduction failures on stricter
compilers.
Signed-off-by: Zhengwei Zhu <141622927+ZhengweiZhu@users.noreply.github.com>
* add ubring transport
* fix the bug for ub ring transport
* fix the bug for ub ring transport and other
* add the license for ub transport
* Modifying the variable naming style
* optimize the log message and some field name
* fix some bug for ubring
* add some log ubring endpoint
* add todo
* fix the bug for handshake for ub endpoint
* fix the bug for client ub endpoint
* optimize the iobuf file code
* modify the log level
* fix the declare_shm_ubs define not found bug
* add the timer_mgr support for macos and format code style
* fix the bug for macos epoll
* fix the timespece bug
* adaptor the itimerspec for macos platform
* optimize the cmakelist config
* add the ubring docs for ubring transport
* modify some file name and directory structure
* modify some code style and optimize code logical
* add the cmake ubshm transport ci/testing
* add the dependency header
* bug fix the code
* Fix UB shm allocation cleanup crashes (#11)
Co-authored-by: 郭业昌 <lvpengfei@MacBook-Air.local>
* remove the brpc_ubshm_unittest
* Found and fixed two stability issues (#13)
* 修复ubring server端关闭连接coredump问题
* 修复PollIn/PollOut解引用已释放Socket指针的问题
PollIn/PollOut通过ep->_socket(裸指针)读取data socket,当data socket
被销毁时该指针悬空,导致Socket::Address读到垃圾id触发SIGSEGV。
改为存储_socket_id(SocketId),用Address获取引用计数的Socket,
并在整个回调期间持有该引用,避免解引用悬空指针。
* 修复client非正常退出导致UBRING shm残留的问题
client被强杀(SIGTERM/崩溃/OOM)时teardown没跑完,localShm(_C)的
shm_unlink未执行,导致/dev/shm残留_C文件。server的remoteShm只munmap
不unlink(正确),无法清理client的名字。
在握手ESTABLISHED时(client/server都确认对方已mmap自己的localShm)
立即unlink localShm名字。此时对端已持有mmap引用,unlink只删名字不
影响通信;进程任意时刻退出都不会残留文件名。
* Address chenBright's review: use English comments and BAIDU_CACHELINE_ALIGNMENT
- Convert all Chinese comments in ubshm to English (per chenBright's
'Please use English' on ub_endpoint.cpp:723, ub_ring.cpp:337,
shm_ubs.cpp:316, and similar)
- Replace __attribute__((aligned(64))) with BAIDU_CACHELINE_ALIGNMENT
in ubr_msg.h (per chenBright's comment on ubr_msg.h:41)
- Remove unnecessary TODO comment in ub_ring.cpp:551 (per chenBright's
'Unnecessary comments, please delete')
* Remove unused lock macros in thread_lock.h
Per chenBright's review, the functions and macros defined in
thread_lock.h are largely unused. Verified usage across ubshm:
- LOCK_GUARD / UnlockMutex: 8 call sites in shm_ubs.cpp and
ub_ring_manager.cpp, kept.
- SPIN_LOCK_GUARD, R_LOCK_GUARD, W_LOCK_GUARD, SEMAPHORE_WAIT_GUARD,
SEMAPHORE_WAIT_GUARD_WITH_CLOSE and their helper functions
(UnlockSpinLock, UnlockRWLock, PostSem, PostSemWithClose): 0 call
sites, removed.
* Apply chenBright's review on timer_mgr globals
Per chenBright's review on timer_mgr.cpp:32-37:
- Add explicit default values to uninitialized globals
(g_total_timer_num=0, g_max_system_fd=0, g_epoll_execute_thread=0,
g_timer_module_initialized=0)
- Rename globals to snake_case (g_epollFd -> g_epoll_fd,
g_totalTimerNum -> g_total_timer_num, g_timerFdCtxMap ->
g_timer_fd_ctx_map, maxSystemFd -> g_max_system_fd,
g_epollExecuteThread -> g_epoll_execute_thread,
g_timerModuleInitialized -> g_timer_module_initialized)
- maxSystemFd also gains the g_ prefix to match global naming style
Also fix the missing std:: qualifier on atomic_fetch_sub/add/load
(per chenBright's earlier comment on timer_mgr.cpp:80).
* Change CloseTimerFd fd type from uint32_t to int
Per chenBright's review on timer_mgr.cpp:399 (uint32_t -> int).
fd is a system file descriptor; POSIX APIs use int and -1 denotes an
invalid fd, which uint32_t cannot represent. Changed the CloseTimerFd
signature (header + definition) and removed the now-unnecessary
(uint32_t) casts at the two call sites.
* Use BAIDU_LIKELY/BAIDU_UNLIKELY instead of custom __builtin_expect
Per chenBright's review on common.h:27. Rather than redefine the
macros with __builtin_expect directly, forward LIKELY/UNLIKELY to
brpc's standard BAIDU_LIKELY/BAIDU_UNLIKELY (from butil/compiler_specific.h).
The 122 call sites keep using LIKELY()/UNLIKELY() unchanged; only the
macro bodies change, preserving semantics.
* Add unit tests for UBShmEndpoint
Per chenBright's request to add unit tests for UBShmTransport in this
PR (rather than a follow-up).
Adds test/brpc_ubring_unittest.cpp with tests covering the public
interface of UBShmEndpoint under the g_skip_ub_init=true mode (which
skips real shared-memory/poller setup):
- construct_and_destruct: lifecycle safety
- is_writable_false_when_skip_init: skip-mode behavior
- reset_is_idempotent: Reset() is safe to call repeatedly
The file follows the brpc_*_unittest.cpp naming convention so it is
auto-collected by test/CMakeLists.txt's file(GLOB). Verified: compiles,
links, and all 3 tests pass (g++ 15.2, C++17, gtest, BRPC_WITH_UBRING=ON).
* Rewrite UBShmEndpoint unit tests with real coverage
Per chenBright's feedback that the previous tests were too simple and
did not cover the main methods.
Source changes to enable testing:
- Move HelloMessage struct declaration from ub_endpoint.cpp to
ub_endpoint.h so tests can access it
- Expose private members under #ifdef UNIT_TEST (precedent:
butil/containers/stack_container.h) so tests can call
AllocateClientResources without -Dprivate=public (which breaks
GCC 15 + new libstdc++ <any>/<sstream>)
Tests (9, all passing on Ubuntu 26.04 g++ 15.2 C++17 gtest):
HelloMessageTest (5): serialize/deserialize roundtrip, network byte
order verification, uint64 max boundary, full shm_name, toString
UBShmEndpointTest (4): construct, real IPC shm
AllocateClientResources (g_skip_ub_init=false), reset cleanup, reset
idempotency
* rename the variable to snake_case style
---------
Co-authored-by: YeChang Guo <52730608+YChange01@users.noreply.github.com>
Co-authored-by: 郭业昌 <lvpengfei@MacBook-Air.local>
Co-authored-by: gure <740684863@qq.com>
Previously BringUpQp only did a local ibv_query_ece + ibv_set_ece
roundtrip and never exchanged ECE capabilities with the peer.
This patch wires up the standard requestor/responder ECE negotiation
flow on top of the existing v3 handshake without adding any extra
round trip:
1. Client queries local ECE, advertises it in its v3 hello.
2. Server applies the client's ECE in INIT->RTR (set_ece), then
after RTS queries the reduced/negotiated ECE and sends it back
in the reply hello.
3. Client applies the server's reduced ECE in INIT->RTR.
Background
==========
The legacy v2 handshake ("RDMA" magic + 36B fixed binary HelloMessage)
had correctness bugs that made the wire format effectively
unevolvable.
v2 bugs
=======
A. Client never drained the "unknown tail" bytes when a peer sent
msg_len > HELLO_MSG_LEN_MIN(40). Leftover bytes stayed in the
socket recv buffer and silently corrupted the next ReadFromFd
(the ACK).
B. Server had the symmetric version of A.
C. Server computed the body read length from its LOCAL
g_rdma_hello_msg_len, implicitly assuming the peer's hello is the
same length as its own. A longer peer left bytes behind; a shorter
peer made the read block. Client used the correct compile-time
constant; the two sides were not symmetric.
Combined, A/B/C meant v2 could not safely append a single byte to the
hello -- even an "optional hint" appended by a newer sender would
mis-align an older receiver's next read.
v3 design
=========
Wire format, magic-namespace-isolated from v2:
[ "RDM3" 4B ][ pb_size 4B big-endian ][ RdmaHello protobuf bytes ]
with pb_size in (0, 4096]. RdmaHello carries the same 6 base fields
as v2 plus room to append future capabilities.
Why protobuf (and not "v2 plus length prefix")
----------------------------------------------
- Variable-length fields are coming. Future capability fields will
include strings (rdma_device_name, netdev_name, ...) and other
variable-length data. Supporting them on a fixed-binary protocol
forces us to invent and maintain a TLV layer (per-field type +
length + value framing, plus version-aware deserialization). That
is reimplementing protobuf badly. Using protobuf from day one
costs nothing and is the canonical answer.
- Fixes v2 bug A/B/C generically: pb_size makes the wire
self-describing, so the receiver never needs to guess the length
or know the peer's schema version to read the body cleanly.
- Append-only field evolution out of the box: new optional fields
cost old receivers nothing -- they're skipped as unknown protobuf
fields. v2 with a hand-rolled length prefix would still need
per-field opt-in code on every side.
- Built-in validation: ParseFromArray fails fast on malformed input;
required-field presence is enforced at the parse layer, not by
ad-hoc has_xxx() checks scattered through wire code.
- bRPC already depends on protobuf -- no new build dependency.
Why a NEW MAGIC rather than a version field inside protobuf
-----------------------------------------------------------
- Forces "breaking change" to be a deployment decision visible at
the wire level. You cannot accidentally ship a backwards-
incompatible patch via a field-semantics tweak.
- Server-side dispatch routes by magic to fully independent state
machines that can't entangle (no `if (version == X)` branches
anywhere -- this is the abstraction's red line).
- Any future breaking change bumps the magic ("RDM4", "RDM5", ...).
v3 fields, once shipped, never change semantics.
Rollout
=======
Server-side ALWAYS accepts both v2 and v3 (no gflag, no kill-switch);
magic routes to fully independent code paths. A single rolling upgrade
enables v3 fleet-wide.
Client-side picks the wire protocol via gflag with a safe default:
FLAGS_rdma_client_handshake_version (default 2)
2 = "RDMA" legacy (zero-regression default)
3 = "RDM3" protobuf (opt-in once target servers support v3)
Sub-second rollback is one flag flip away. v3 client to v2-only legacy
server is NOT guaranteed to transparently fall back on the same
connection -- the supported migration is "upgrade servers first, then
opt-in clients".
* ci(linux): bring up redis-server and mysql-server for unittests
The clang-unittest and clang-unittest-asan jobs run the full unit test
suite via test/run_tests.sh, which includes backend integration tests
(e.g. brpc_redis_unittest) that fork a real server when its binary is
present and otherwise silently short-circuit to a passing result. Since
CI never installed those servers, the redis backend tests reported
PASSED while doing nothing (7 of 14 RedisTest cases skip-as-pass).
Install redis-server and mysql-server before running the tests in both
unittest jobs so these backend tests execute against a live server. The
binaries are added only in the unittest jobs, not in the shared
install-essential-dependencies action used by compile-only jobs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci(linux): skip flaky redis client tests under ASan
brpc_redis_unittest forks a real redis-server and waits a fixed 50ms before
connecting; under ASan redis starts too slowly, causing flaky connection-refused.
Skip RedisTest.* in the ASan job only (still covered by clang-unittest).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci(linux): install redis-server and mysql-server in install-essential-dependencies
Move the redis-server/mysql-server install out of the individual
clang-unittest and clang-unittest-asan jobs and into the shared
install-essential-dependencies composite action, so every job that
installs dependencies has the servers available (and the unittest jobs
no longer carry a bespoke install step).
Under ASan the redis integration tests (sanity, keys_with_spaces,
incr_and_decr, by_components, auth) fork a real redis-server and connect
after a fixed 50ms wait; redis starts too slowly there and they flake
with connection refused. Filter just those out under ASan -- the redis
codec/server tests still run, and the full suite runs in clang-unittest.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci(linux): run forked-server integration tests under bazel
The bazel unittest jobs ran //test/... without redis-server/mysql-server
installed, so brpc_redis_unittest forked nothing and its integration cases
(sanity/keys_with_spaces/incr_and_decr/by_components/auth) early-returned as
passes -- green but vacuous.
- Install redis-server/mysql-server in the (non-ASan) bazel test jobs via the
shared install-essential-dependencies action (the same one the make jobs
use), so the servers are on PATH for the forked tests.
- Tag brpc_redis_unittest "external" + "local" so bazel always re-runs it
(never serves a cached skip) and runs it outside the sandbox where the
PATH-located redis-server is visible and loopback works. Threaded through
generate_unittests via a new per_test_tags arg.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: rajvarun77 <287367605+rajvarun77@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: batch create and accept stream
* fix protobuf 22.5 compilation error related to thread_local in MacOS
* refine style
* modify code based on the code review feedback and add more tests
modify code based on the code review feedback and add more tests
* add bzlmod support
* drop buggy clang-11 which dont suport -fno-access-control correctly
* fix typo in ci-linux.yml
* give up trying to test bazel in clang
* Support Protobuf 22
* config_brpc.sh: support to build with protobuf 22
* ci-macos: compatible with Apple silicon
* ci: compile with protobuf 22+ using macos