Table of Contents
For decades, Linux high-performance networking relied on event loop mechanisms like epoll paired with non-blocking sockets. While epoll allows an application to monitor thousands of file descriptors, executing an actual read or write operation (read(), write(), send(), recv()) still requires making a system call (syscall).
At millions of requests per second, CPU overhead is dominated by user-space to kernel-space context switches and Meltdown/Spectre page-table isolation (KPTI) mitigation overhead.
In Linux 5.1, kernel developer Jens Axboe introduced io_uringāa fundamental redesign of Linux asynchronous I/O.
io_uring eliminates system call overhead by creating two lockless circular ring buffers shared directly between user memory and kernel memory: the Submission Queue (SQ) and the Completion Queue (CQ).
This article details the architecture, memory model, and performance mechanics of io_uring.
io_uring Shared Ring Buffer Architecture
How user space and kernel space communicate asynchronously via shared memory ring buffers without syscalls:
Core io_uring Mechanics
- Submission Queue Entry (SQE): A 64-byte descriptor submitted by user space containing the I/O opcode (
IORING_OP_READV,IORING_OP_ACCEPT), file descriptor, memory buffer address, and user metadata tag. - Completion Queue Entry (CQE): A 16-byte descriptor written by the kernel containing the result (bytes read/written or negative error code) and the matching
user_datatag. - SQPOLL (Submission Queue Polling) Mode: In SQPOLL mode, a dedicated kernel thread continuously polls the Submission Queue ring. User space appends SQEs to the ring in memory without invoking a single
sys_entersyscall, enabling Zero-Syscall I/O.
Python Implementation: io_uring Lockless Ring Buffer Engine
Here is a production-grade Python simulation of the io_uring Submission Queue (SQ) and Completion Queue (CQ) lockless ring buffer architecture:
import time
from typing import Dict, List, Optional, Any
from pydantic import BaseModel, Field
class SubmissionQueueEntry(BaseModel):
opcode: str # IORING_OP_READ, IORING_OP_WRITE, IORING_OP_ACCEPT
fd: int # Target file descriptor
buffer_addr: str # Memory buffer identifier
len: int # Bytes to read/write
user_data: int # Unique tag to correlate CQE result
class CompletionQueueEntry(BaseModel):
user_data: int # Matching tag from SQE
res: int # Number of bytes transferred or negative error code
flags: int = 0
class IOUringEngine:
"""
Simulates the Linux io_uring lockless ring buffer subsystem.
Shared memory rings are represented using fixed-size queues.
"""
def __init__(self, queue_depth: int = 16):
self.queue_depth = queue_depth
# Shared Ring Buffers
self.sq_ring: List[Optional[SubmissionQueueEntry]] = [None] * queue_depth
self.cq_ring: List[Optional[CompletionQueueEntry]] = [None] * queue_depth
# Pointers (Head/Tail Indices)
self.sq_head = 0
self.sq_tail = 0
self.cq_head = 0
self.cq_tail = 0
self.sqpoll_active = True
def submit_sqe(self, opcode: str, fd: int, buffer_addr: str, length: int, user_data: int) -> bool:
"""User-space action: Appends an SQE to the Submission Queue ring."""
next_tail = (self.sq_tail + 1) % self.queue_depth
if next_tail == self.sq_head:
print(" ā ļø [io_uring SQ] Submission Queue Full!")
return False # SQ Ring Full
sqe = SubmissionQueueEntry(
opcode=opcode,
fd=fd,
buffer_addr=buffer_addr,
len=length,
user_data=user_data
)
self.sq_ring[self.sq_tail] = sqe
self.sq_tail = next_tail
print(f" š„ [User Space] Prepared SQE (Opcode: {opcode}, FD: {fd}, Tag: {user_data}) -> SQ Tail: {self.sq_tail}")
return True
def kernel_sqpoll_loop_step(self, simulated_disk: Dict[int, bytes]):
"""
Kernel-space SQPOLL thread: Polling SQ ring, executing I/O, and appending CQEs.
Exits zero syscalls!
"""
while self.sq_head != self.sq_tail:
sqe = self.sq_ring[self.sq_head]
self.sq_ring[self.sq_head] = None
self.sq_head = (self.sq_head + 1) % self.queue_depth
if not sqe:
continue
print(f" ā” [Kernel SQPOLL Thread] Processing SQE Tag #{sqe.user_data} ({sqe.opcode} on FD {sqe.fd})...")
# Execute Simulated Async I/O Operation
res_code = 0
if sqe.opcode == "IORING_OP_READ":
data = simulated_disk.get(sqe.fd, b"")
res_code = min(len(data), sqe.len)
elif sqe.opcode == "IORING_OP_WRITE":
res_code = sqe.len
# Append Completion Queue Entry (CQE) to CQ Ring
cqe = CompletionQueueEntry(user_data=sqe.user_data, res=res_code)
self.cq_ring[self.cq_tail] = cqe
self.cq_tail = (self.cq_tail + 1) % self.queue_depth
print(f" ā³ [Kernel] I/O Complete. Pushed CQE (Tag: {sqe.user_data}, Res: {res_code} bytes) -> CQ Tail: {self.cq_tail}")
def reap_cqe(self) -> Optional[CompletionQueueEntry]:
"""User-space action: Reaps a completed I/O entry from the CQ ring."""
if self.cq_head == self.cq_tail:
return None # CQ Ring Empty
cqe = self.cq_ring[self.cq_head]
self.cq_ring[self.cq_head] = None
self.cq_head = (self.cq_head + 1) % self.queue_depth
return cqe
# Demonstration Execution
if __name__ == "__main__":
ring = IOUringEngine(queue_depth=8)
simulated_nvme_disk = {
3: b"Linux Kernel io_uring High Performance Storage Payload\n",
4: b"Network Socket Stream Payload\n"
}
print("š Demonstrating Linux Kernel io_uring Architecture...")
print("=" * 75)
# 1. User space submits 3 async I/O requests without syscalls
ring.submit_sqe("IORING_OP_READ", fd=3, buffer_addr="0x7fff001", length=64, user_data=1001)
ring.submit_sqe("IORING_OP_READ", fd=4, buffer_addr="0x7fff002", length=32, user_data=1002)
ring.submit_sqe("IORING_OP_WRITE", fd=3, buffer_addr="0x7fff003", length=128, user_data=1003)
# 2. Kernel SQPOLL worker processes Submission Queue in memory
print("\nš Executing Kernel SQPOLL Loop (Zero Syscall Context Switches)...")
ring.kernel_sqpoll_loop_step(simulated_nvme_disk)
# 3. User space reaps completed CQEs from Completion Queue ring
print("\nš¦ User Space Reaping Completed I/O Results from CQ Ring...")
while True:
cqe = ring.reap_cqe()
if not cqe:
break
print(f" ā
[User Space] Harvested CQE Tag #{cqe.user_data} -> Transferred {cqe.res} bytes")
io_uring Implementation Gotchas
When designing storage and networking systems around io_uring:
Use Fixed Buffers and Registered Files: Standard io_uring requests still perform page-table lookups for buffers and file descriptors. To achieve maximum throughput, pre-register file descriptors (IORING_REGISTER_FILES) and pre-pin memory buffers (IORING_REGISTER_BUFFERS) to eliminate kernel virtual memory mapping overhead entirely.
Handle Short Reads and Linked Requests: Storage or network reads submitted via io_uring may complete partially (short reads). Use IOSQE_IO_LINK flags to enforce sequential ordering when chaining dependent async I/O requests (e.g. Read Header ā Read Body).
Real-World Enterprise Impact
Databases and web servers adopting io_uring (such as RocksDB, Netty, and ScyllaDB) report:
- Over 2,000,000 IOPS per CPU Core: Achieving more than double the IOPS of traditional
epollorlibaiodrivers. - 50% Reduction in p99 Latency: Eliminating syscall context switches stabilizes tail latencies under extreme network concurrency.

Discussion & Comments