Table of Contents
In ultra-low latency, high-throughput storage engines (ScyllaDB, RocksDB, TigerBeetle, ClickHouse), relying on standard OS POSIX file APIs (read(), write(), pread()) imposes a massive CPU overhead bottleneck.
Standard POSIX disk reads pass data through the Linux kernel Page Cache. For high-scale database engines that manage their own specialized block caches, Page Cache buffering causes:
- Double-Buffering Memory Waste: Storing identical $4\text{ KB}$ data blocks in both the kernel Page Cache and user-space database RAM.
- Unpredictable Latency Spikes: Background OS dirty page writebacks (
pdflush/flush) lock kernel memory pages, triggering severe p99 latency spikes. - High System Call Overhead: Invoking a system call per I/O operation incurs CPU context-switch context costs ($200\text{ns}$ per syscall).
To bypass kernel overhead and harness modern NVMe SSD speeds (capable of $>1,000,000\text{ IOPS}$), modern storage engines utilize Direct I/O (O_DIRECT) and Jens Axboe's revolutionary Linux io_uring subsystem.
This article details O_DIRECT page cache bypass, lockless ring buffer queues, and zero-syscall io_uring submission pipelines.
Linux io_uring Shared Memory Ring Buffer Architecture
How io_uring uses kernel-shared Submission (SQ) and Completion (CQ) ring buffers for zero-syscall I/O:
Core Asynchronous Direct I/O Principles
- Direct I/O (
O_DIRECT) Page Cache Bypass: Opening a file descriptor with theO_DIRECTflag instructs the OS kernel to bypass the Page Cache entirely. Data transfers directly between user-space memory buffers and the NVMe physical storage controller via Direct Memory Access (DMA).- Requirement: User-space memory buffers must be strictly aligned to hardware disk sector boundaries (typically $4,096$-byte alignment).
- The
io_uringRevolution: Introduced in Linux kernel 5.1,io_uringreplaces legacylibaio. It establishes two shared memory ring buffers mapped directly between user-space and kernel memory viammap():- Submission Queue (SQ): User-space writes I/O requests (Submission Queue Entries / SQEs) into the SQ ring buffer.
- Completion Queue (CQ): The Linux kernel writes completed I/O results (Completion Queue Entries / CQEs) into the CQ ring buffer.
- Zero-Syscall Polled Mode (
IORING_SETUP_SQPOLL): InSQPOLLmode, a dedicated kernel thread continuously polls the Submission Queue ring buffer for new SQEs. The application submits thousands of I/O operations per second simply by writing entries to RAM without executing a singleio_uring_enter()system call! - Fixed Buffers & Files (
IORING_REGISTER_BUFFERS): Applications can pre-register memory buffers and file descriptors with the kernel at startup, eliminating per-I/O virtual memory page pinning and file descriptor lookup overheads.
Python Implementation: Asynchronous io_uring Storage Engine Simulator
Here is a production-grade Python implementation of an Asynchronous Storage Engine Simulator featuring O_DIRECT Page-Aligned Memory Buffers and io_uring Lockless Ring Buffers:
import mmap
import time
from typing import List, Dict, Optional
from pydantic import BaseModel
class SubmissionQueueEntry(BaseModel):
sqe_id: int
opcode: str # "READ_DIRECT", "WRITE_DIRECT"
file_offset: int
data_buffer: str
class CompletionQueueEntry(BaseModel):
sqe_id: int
res_code: int # 0 for SUCCESS, negative for ERROR
payload: Optional[str] = None
class IoUringRingBufferSimulator:
"""
Simulates Linux io_uring Shared Lockless Ring Buffers & O_DIRECT DMA Storage Engine.
"""
def __init__(self, queue_depth: int = 4, page_alignment: int = 4096):
self.queue_depth = queue_depth
self.page_alignment = page_alignment
# Shared Ring Buffers
self.submission_queue: List[SubmissionQueueEntry] = []
self.completion_queue: List[CompletionQueueEntry] = []
self.sqe_counter = 0
def prepare_sqe(self, opcode: str, file_offset: int, data: str) -> int:
"""User-space: Prepares and pushes SQE to Submission Queue (No syscall!)."""
self.sqe_counter += 1
sqe = SubmissionQueueEntry(
sqe_id=self.sqe_counter, opcode=opcode, file_offset=file_offset, data_buffer=data
)
self.submission_queue.append(sqe)
print(f" š„ [io_uring SQ] Pushed SQE #{sqe.sqe_id} ('{opcode}' Offset: {file_offset}) to Shared Ring Buffer")
return sqe.sqe_id
def submit_and_poll_kernel(self) -> int:
"""Kernel-space: Simulates SQPOLL thread processing SQEs zero-copy via DMA."""
print(f"\n āļø [Linux Kernel io_uring] Processing {len(self.submission_queue)} SQEs from Ring Buffer...")
processed_count = 0
while self.submission_queue:
sqe = self.submission_queue.pop(0)
# Simulate O_DIRECT Page-Aligned DMA Transfer to NVMe
time.sleep(0.01) # Simulate 10ms hardware DMA latency
cqe = CompletionQueueEntry(
sqe_id=sqe.sqe_id,
res_code=0, # Success
payload=f"DMA_BYTES[{sqe.data_buffer}]@OFFSET_{sqe.file_offset}"
)
self.completion_queue.append(cqe)
processed_count += 1
print(f" ā” [Kernel DMA Complete] Executed SQE #{sqe.sqe_id} -> Pushed CQE to Completion Ring Buffer")
return processed_count
def reap_cqe(self) -> Optional[CompletionQueueEntry]:
"""User-space: Reaps completed I/O from Completion Queue ring buffer."""
if self.completion_queue:
cqe = self.completion_queue.pop(0)
print(f" š¤ [io_uring CQ] Reaped Completed CQE #{cqe.sqe_id} (Res: {cqe.res_code})")
return cqe
return None
# Demonstration Execution
if __name__ == "__main__":
ring = IoUringRingBufferSimulator(queue_depth=4)
print("š Demonstrating Direct I/O (O_DIRECT) & io_uring Storage Engine...")
print("=" * 75)
# 1. User-space submits 3 O_DIRECT Read/Write Requests without system calls!
id1 = ring.prepare_sqe("WRITE_DIRECT", file_offset=0, data="BLOCK_DATA_PART_1")
id2 = ring.prepare_sqe("WRITE_DIRECT", file_offset=4096, data="BLOCK_DATA_PART_2")
id3 = ring.prepare_sqe("READ_DIRECT", file_offset=0, data="READ_BUFFER")
# 2. Kernel SQPoll Thread processes queued SQEs in batch via Direct DMA
ring.submit_and_poll_kernel()
# 3. User-space reaps completed results from Completion Ring Buffer
print("\nš Reaping Completion Queue Entries (CQEs):")
while True:
cqe = ring.reap_cqe()
if not cqe:
break
print(f" ⢠Result Payload: {cqe.payload}")
Direct I/O & io_uring Gotchas & Best Practices
When building storage engines with io_uring:
Use posix_memalign() or mmap() for Buffer Allocation: Passing non-aligned memory buffers to an O_DIRECT file descriptor results in immediate EINVAL kernel errors. Always allocate memory buffers using posix_memalign() or page-aligned mmap().
Set RLIMIT_MEMLOCK Limits appropriately: Using registered fixed buffers (IORING_REGISTER_BUFFERS) pins physical RAM pages into memory, bypassing OS swap space. Ensure system ulimit -l settings allow sufficient locked memory allocations.
Real-World Enterprise Impact
Storage engines leveraging O_DIRECT and io_uring (such as ScyllaDB and TigerBeetle) report:
- Over 2,000,000 IOPS per Server Node: Eliminating system call overhead and Page Cache lock contention unlocks maximum physical NVMe device speed.
- $5\times$ Lower p99 Tail Latency: Bypassing OS dirty page background writebacks eliminates sudden multi-millisecond disk latency spikes.

Discussion & Comments