Timeline
Timeline
2026-07-22
init
2026-08-15
Supplement the VRing structure (the three major regions of split ring), data transfer flow, notification suppression mechanism, Packed Virtqueue, virtqueue implementations in Linux kernel and OpenAMP, and examples of Linux kernel using VRing (virtio-rng, virtio-blk)
This article introduces the basic concepts and core architecture of VirtIO (Virtual Input/Output). VirtIO is an abstraction layer between virtual machines and host devices. It achieves data transfer through minimalist VirtIO devices, offloading most hardware processing work to the host, thereby improving efficiency. The article elaborates on three key parts of the VirtIO architecture: the front-end driver (a kernel module in the Guest kernel), the back-end device (such as the VirtIO device in QEMU), and VirtQueues/VRings. Among them, VRing is the actual data structure, with the classic split ring layout, consisting of three contiguous memory regions: the descriptor table, available ring, and used ring. It uses a lock-free SPSC mechanism to achieve data sharing and notification between the driver and device. The article also mentions the standardization significance of the VirtIO specification and the flexibility in different implementations (such as Linux kernel and OpenAMP).
VirtIO introduction
VirtIO, i.e., Virtual Input/Output, is an abstraction layer for virtual machines on host devices. Essentially, it is an interface,that allows virtual machines to use host devices through minimal virtual devices (called VirtIO devices). These VirtIO devices are very minimalist because they only implement the basic requirements for sending and receiving data.
This is because in VirtIO, we let the host handle most of the setup, maintenance, and processing of its actual hardware devices.The role of the VirtIO device is basically to transfer data to the host’s actual physical hardware。

- VM: I want to go to google.com. Hey virtio-net, can you tell the host to fetch this webpage for me?
- Virtio-net: Okay. Host, can you help us fetch this webpage?
- Host: Okay. I am now fetching the webpage data.
- Host: This is the requested webpage data
- Virtio-net: Thank you. Hey, VM, this is the webpage you requested
Although this is an oversimplified example, it still illustrates the core idea. That is, let the host’s hardware do as much work as possible, and let VirtIO handle sending and receiving data. Offloading most of the work to the host makes execution on the virtual machine faster and more efficient than emulated devices.
Another important aspect of VirtIO is that its core framework has been standardized into the official VirtIO specification. The VirtIO specification defines the standard requirements that VirtIO devices and drivers must meet (e.g., feature bits, status, configuration, common operations, etc.). This is important because it means that regardless of the environment or operating system using VirtIO, the core framework of its implementation must be the same.
Although VirtIO implementations have a certain degree of compliance, there is also flexibility in organization and setup. For example, the
virtqueuestructure in the Linux kernel and QEMU’sVirtQueuestructure is different.
VirtIO architecture
The architecture of VirtIO consists of three key parts:
- Front-end driver
- Back-end device
- VirtQueues and VRings
In the figure below, you can see the position of VirtIO in a typical host and guest configuration (without vHost, SR-IOV, etc.).

The front-end VirtIO driver resides in the Guest kernel, and the back-end VirtIO device resides in the virtual machine (QEMU),communication between them is handled in the data plane through VirtQueue and VRing。
We can also see notifications from the VirtIO driver and device (such as VMExit, vCPU IRQ), which are routed to KVM interrupts.
VirtIO Driver (Front-End)
In a typical VirtIO host and guest configuration, the VirtIO driver resides in the Guest kernel. In the Guest OS, each VirtIO driver is treated as a kernel module. The core responsibilities of the VirtIO driver include:
Accept I/O requests from user processes.
Transfer these I/O requests to the corresponding backend VirtIO device.
Retrieve completed requests from the corresponding VirtIO device.
For example:
A virtio-scsi I/O request might be a user wanting to retrieve a document from storage.
The virtio-scsi driver accepts the request to fetch the document and sends the request to the virtio-scsi device (backend).
After the VirtIO device completes the request, the document is made available to the VirtIO driver. The VirtIO driver retrieves the document and makes it available to the user.
VirtIO Device (Back-End)
Additionally, in a typical Host and Guest configuration, VirtIO devices exist in the virtual machine. In the above diagram, we will use QEMU as our (Type 2) hypervisor. This means our VirtIO device will exist within the QEMU process. The core responsibilities of a VirtIO device include:
- Accept I/O requests from the corresponding front-end VirtIO driver.
- Process requests by offloading I/O operations to the host’s physical hardware.
- Make the data of the request being processed available to the VirtIO driver.
Back to the virtio-scsi example:
- The virtio-scsi driver notifies its device counterpart that it needs to fetch the requested document from storage on the actual physical hardware.
- The virtio-scsi device accepts the request and makes the necessary calls to retrieve the data from the physical hardware.
- Finally, the device makes the data available to the driver by placing the retrieved data into the shared VirtQueue.
VirtQueues
The last key part of the VirtIO architecture is the VirtQueue, which are data structures that help devices and drivers perform various I/O operations.
VirtQueues are shared in Guest physical memory, meaning each VirtIO driver and device pair accesses the same pages in RAM. In other words, the VirtQueues of the driver and device are nottwo different synchronized regions.
There are many inconsistencies online regarding VirtQueue. Some use it synonymously with virtual rings (VRings, or Virtio-rings), while others describe them separately. This is because VRings are the main feature of a VirtQueue; the VRing is the actual data structure that facilitates data transfer between VirtIO devices and drivers. We will describe them separately here, because a VirtQueue is more than just a VRing.
VRing: Split Ring Layout
The classic VRing layout defined by the VirtIO 1.0 specification is called split ring, which consists of three contiguous memory regions:
| Region | Chinese name | Producer | Consumer |
|---|---|---|---|
| Descriptor Table | Descriptor table | Driver (allocate/reclaim) | Device (read) |
| Avail Ring | Available ring | Driver (write) | Device (read) |
| Used Ring | Used ring | Device (write) | Driver (read) |
The layout of the three areas in memory is as follows (vring_init()Invring_alignaligned placement):
123456 | 低地址 ────────────────────────────────────────────────────► 高地址┌─────────────────────┬──────────────────┬─── 对齐填充 ───┬──────────────────┐│ Descriptor Table │ Avail Ring │ │ Used Ring ││ vring_desc[num] │ vring_avail │ │ vring_used ││ 16 * num 字节 │ 6 + 2*num 字节 │ │ 6 + 8*num 字节 │└─────────────────────┴──────────────────┴────────────────┴──────────────────┘ |
Note that the entire VRing is located in Guest physical address space(in the OpenAMP scenario, it is the DDR address space shared by A7 and M4), the buffer address recorded in the descriptor is also an address in this address space, and the peer directly accesses the actual data buffer through it—The data itself is transferred only once in the shared memory; what is passed in the VRing is just “pointer + length”.。
Descriptor Table
The descriptor table is an array, each element describes a data buffer:
1234567 | /* include/uapi/linux/virtio_ring.h */struct vring_desc { __virtio64 addr; /* Address of the buffer */ __virtio32 len; /* buffer length */ __virtio16 flags; /* Flags */ __virtio16 next; /* Index of the next descriptor in the linked list */}; |
flagsThree flags are defined:
123 |
- An I/O request often consists of multiple fragments that are not contiguous in memory (scatter-gather). The driver fills each fragment into a descriptor and uses
NEXTflags to chain them together.Descriptor chain;WRITEThe flags determine whether each segment is device-read (out, e.g., data for transmitted packets) or device-write (in, e.g., buffers for received packets); directions cannot be mixed on the same chain. INDIRECTThis is a performance optimization: the driver can allocate an ‘indirect descriptor table’ containing only one chain, occupying only one descriptor slot in the main table, avoiding long chains filling up the main table.
Avail Ring and Used Ring
The avail ring is produced by the driver and tells the device ‘which descriptor chains are ready’:
12345 | struct vring_avail { __virtio16 flags; /* VRING_AVAIL_F_NO_INTERRUPT, etc. */ __virtio16 idx; /* The position where the driver will next write to ring[] */ __virtio16 ring[]; /* Array of descriptor chain heads (descriptor indices) */}; |
The used ring is produced by the device and tells the driver ‘which descriptor chains have been processed’:
12345678910 | struct vring_used_elem { __virtio32 id; /* Index of the descriptor chain head */ __virtio32 len; /* Number of bytes actually written by the device */};struct vring_used { __virtio16 flags; /* VRING_USED_F_NO_NOTIFY, etc. */ __virtio16 idx; /* The position where the device will next write to ring[] */ struct vring_used_elem ring[];}; |
Both rings areLock-free SPSC (single-producer, single-consumer) ring queues: each side only writes the ring it produces (idxmonotonically increasing, wrapping), and only reads the ring produced by the other side, so no synchronization primitives are needed; it relies only on memory barriers to ensure visibility. This is the key design for VirtQueue’s high throughput.
Data transfer flow
Taking RPMsg sending a message (A7 -> M4, i.e.virtqueue_add_outbuf()+virtqueue_kick()) as an example, a complete transfer flow is as follows:
123456789101112131415161718192021 | 驱动(生产者) 设备(消费者)──────────────────────────── ────────────────────────────1. 构造 buffer,写入消息内容2. virtqueue_add_outbuf() - 取空闲描述符,填 addr/len - 链头写入 avail ring[avail.idx] - avail.idx++(伴随内存屏障)3. virtqueue_kick() - 写 VRING_USED_F_NO_NOTIFY 检查是否需要通知 ---------- 通知 ----------> 4. 收到通知(VMExit / IPCC 中断) - 触发 notify - 读 avail.idx,发现新链 5. 按描述符 addr 到共享内存 读取数据并处理 6. 把链头和写入长度放入 used ring[used.idx] used.idx++(伴随内存屏障) <---------- 中断 ------ 7. 触发回调(vring 回调 / vq callback)8. virtqueue_get_buf() - 读 used.idx,取出已处理链 - 回收描述符到空闲池 - 释放或复用 buffer |
A few key points:
- Notification (kick/notify) and data are two independent paths. VRing is responsible for ‘placing data’, and notification is responsible for ‘giving a shout’. In QEMU/KVM, notification is a VMExit plus a write to a device register (PCI BAR / MMIO); in OpenAMP, it is writing to IPCC/mailbox to trigger an interrupt on the peer.
idxYesMonotonically increasing then wrappingthe free-running counter; to determine ‘whether there is something new’, you only need to compare the locally cachedidxwith the ring’sidx, which is also the basis for batch processing (one kick handles multiple buffers).- After the driver callback, it is common to use
virtqueue_disable_cb()disable interrupt callbacks and batch reapvirtqueue_get_buf(), and after reaping,virtqueue_enable_cb()to avoid interrupt storms.
Notification suppression
Notification (VMExit / inter-core interrupt) is the most expensive operation in the entire process, so the spec provides suppression switches in two directions, set bythe receiversetting the bit,the senderchecks before notifying:
VRING_USED_F_NO_NOTIFY(flags on the avail ring, written by the device for the driver to see): the device setting the bit means ‘don’t keep kicking me, I’ll poll myself’. The driver, invirtqueue_kick()checks this flag to decide whether to actually send a notification.VRING_AVAIL_F_NO_INTERRUPT(flags on the used ring, written by the driver for the device to see): the driver setting the bit means ‘you don’t need to interrupt me immediately after processing’.- Event Idx(
VIRTIO_F_EVENT_IDXfeature): instead of using a switch value, the two sides useused_event/avail_eventprecisely inform ‘don’t notify me until idx reaches a certain value’, making the suppression granularity at the buffer level.
Packed Virtqueue
VirtIO 1.1 introduced packed ring, aiming to reduce cache line usage and eliminate the multiple memory accesses caused by the separation of avail/used and descriptor table in split ring:
- The descriptor table and the avail/used ringsare merged into a single ring descriptor array, with each descriptor carrying its own available/used flag;
- Each descriptor uses
VRING_DESC_F_AVAILandVRING_DESC_F_USEDof two flagsflipto indicate state changes (distinguishing equal/not equal, with wrap counter wrap-around), the driver and device each only write their own flags; - The number of descriptors can be less than the queue depth; the driver can publish them incrementally one by one, without needing to build the complete chain at once.
123456 | struct vring_packed_desc { __le64 addr; /* buffer address */ __le32 len; /* buffer length */ __le16 id; /* An identifier freely used by the driver, usually the chain head index. */ __le16 flags; /* NEXT/WRITE/AVAIL/USED */}; |
The device and driver, through negotiated feature bits,VIRTIO_F_RING_PACKEDdecide whether to use packed or split layout; the two are functionally equivalent. The Linux kernel’svirtio_ring.cimplements both sets (vring_create_virtqueue_split()/vring_create_virtqueue_packed()), while the OpenAMP library currently only implements the split ring.
virtqueue in the Linux kernel
The kernel-side VirtQueue isdrivers/virtio/virtio_ring.cimplemented, exposing a unifiedstruct virtqueueAPI (defined ininclude/linux/virtio.h):
1234567891011 | struct virtqueue { struct list_head list; /* attached to the vq linked list of virtio_device */ void (*callback)(struct virtqueue *vq); /* The device notifies "a buffer has been used up" */ const char *name; struct virtio_device *vdev; unsigned int index; /* vq number */ unsigned int num_free; /* number of free descriptors */ unsigned int num_max; void *priv; bool reset;}; |
Commonly used APIs (in Virtio Rpmsg Bus will be heavily used in the series):
| API | Function |
|---|---|
virtqueue_add_sgs() | Generic interface: insert multiple scatterlists into the avail ring with mixed in/out directions |
virtqueue_add_outbuf() | Give the device a “readable” buffer (driver -> device direction) |
virtqueue_add_inbuf() | Give the device a “writable” buffer (device -> driver direction) |
virtqueue_kick() | Notify the device: the queue has new buffers |
virtqueue_get_buf() | Retrieve the buffer processed by the device from the used ring |
virtqueue_disable_cb()/virtqueue_enable_cb() | Disable/restore the completion interrupt callback |
virtqueue_detach_unused_buf() | Retrieve buffers that have been queued but not yet used by the device (used on remove) |
virtqueue_get_vring_size() | Get queue depth |
An example of the Linux kernel using VRing
Minimal example: virtio-rng
drivers/char/hw_random/virtio-rng.cIt is the simplest virtqueue user in the kernel: the device has only one queue, the driver submits an 8-byte writable buffer, and the device fills it with random numbers. The entire driver core has only three steps, exactly corresponding to the transfer flow in the previous section:
1234567891011121314151617181920212223242526272829303132333435363738394041 | /* drivers/char/hw_random/virtio-rng.c (simplified, omitting hwrng registration and error handling) */struct virtio_rng { struct virtqueue *vq; unsigned char buf[64];};static int virtio_rng_probe(struct virtio_device *vdev){ struct virtio_rng *vi = ...; /* There is only one queue, and a callback is provided at registration (triggered by an interrupt when the device has finished filling the data) */ vi->vq = virtio_find_single_vq(vdev, virtio_rng_done, "input"); ...}/* Steps 1 and 2: submit an empty buffer + kick */static void register_buf(struct virtio_rng *vi){ struct scatterlist sg; sg_init_one(&sg, vi->buf, sizeof(vi->buf)); if (virtqueue_add_inbuf(vi->vq, &sg, 1, vi->buf, GFP_KERNEL) < 0) return; virtqueue_kick(vi->vq);}/* Step 8: reap in the callback */static void virtio_rng_done(struct virtqueue *vq){ struct virtio_rng *vi = vq->vdev->priv; unsigned int len; unsigned char *buf; /* There may be spurious callbacks (shared IRQ, etc.); get_buf returning NULL means there is no new data */ buf = virtqueue_get_buf(vq, &len); if (!buf) return; add_device_randomness(vi->buf, len); register_buf(vi); /* Immediately after reaping, submit a new buffer to keep the pipeline flowing */} |
Map each step onto the vring (assuming queue depth is 64, the driver obtains a free descriptordesc[5]):
| Step | API | Actual effect on the vring |
|---|---|---|
| Submit an empty buffer | virtqueue_add_inbuf() | Filldesc[5]:addr = vi->buf、len = 64、flags = VRING_DESC_F_WRITE; put5write intoavail.ring[avail.idx],avail.idx++ |
| notify the device | virtqueue_kick() | Readused.flagsCheckNO_NOTIFY, if not set, write the device notification register (under KVM this is a VMExit) |
| Device fills random numbers | (device-side action) | Tovi->bufWrite random numbers; put{id=5, len}write intoused.ring[used.idx],used.idx++, trigger an interrupt |
| Reap in callback | virtqueue_get_buf() | Readused.ringretrieve the chain head5and the write length,desc[5]goes back to the free pool, returning the originally passed-indata(i.e.,vi->buf) |
It can be seen that the driver, from start to finish,never touches any vring structure: fetching descriptors, updating avail, and checking used are all encapsulated invirtio_ring.cinside, the driver only deals with scatterlist anddataprivate pointers.
Mixed-direction example: virtio-blk
drivers/block/virtio_blk.cdemonstratesvirtqueue_add_sgs()the typical usage of – a request consists ofmultiple segments with different directionsforming a descriptor chain. Take a 512-byte disk read as an example, the request looks like this in the queue:
123456789 | sgs[0](out,设备读) sgs[1..n](in,设备写) sgs[末尾](in,设备写)┌────────────────────┐ ┌──────────────────────┐ ┌─────────────────┐│ virtio_blk_outhdr │ │ 512B 页缓冲 │ │ 1B status 字节 ││ type/sector/数据 │ │ (读盘的落盘位置) │ │ (设备回写 OK) │└─────────┬──────────┘ └──────────┬───────────┘ └─────────────────┘ │ NEXT │ NEXT ▼ ▼ desc[i] ─────────────► desc[i+1..] ─────────────► desc[j](VRING_DESC_F_WRITE) avail ring 里登记链头 i |
123456789101112131415161718192021222324 | /* drivers/block/virtio_blk.c (simplified) */static int virtblk_add_req(struct virtblk_req *vbr, struct virtqueue *vq){ struct scatterlist hdr, status, *sgs[...]; unsigned int num_out = 0, num_in = 0; /* Request header: block device request type + starting sector number, direction out */ sg_init_one(&hdr, &vbr->out_hdr, sizeof(vbr->out_hdr)); sgs[num_out++] = &hdr; /* Data segment: direction in for read requests, direction out for write requests */ blk_mq_rq_to_sgl(rq, data_sgs); /* Map the bio's physical pages into sg */ if (req_op(rq) == REQ_OP_WRITE) sgs[num_out++] = data_sgs; /* The out segment must be placed before the in segment */ else sgs[num_out + num_in++] = data_sgs; /* Status byte: after the device finishes processing, it writes back success or failure, direction in */ sg_init_one(&status, &vbr->status, sizeof(vbr->status)); sgs[num_out + num_in++] = &status; /* The data parameter is vbr; at interrupt harvest time, virtqueue_get_buf() returns it as-is */ return virtqueue_add_sgs(vq, sgs, num_out, num_in, vbr, GFP_ATOMIC);} |
virtqueue_add_sgs()Internally (the split implementation isvirtqueue_add_split()) translates this sg array into the descriptor chain from the previous section:out segment first, in segment afterstrung into a chain (directions cannot be mixed on the same chain; the spec requires out before in), andvbrThis private pointer is attached tovirtqueueIn the management structure. After the device finishes processing and writes back to the used ring,virtblk_done()In the callbackvirtqueue_get_buf()Retrievevbr, checkvbr->statusand terminate this request to the upper-layer block device stack.
Compared to virtio-rng, the difference of virtio-blk is:
- Attach multiple segments at once, mixed directions: The request header, data, and status byte share a descriptor chain, relying on
NEXT+WRITEThe sign distinguishes direction; this is precisely the purpose of the descriptor chain’s existence. - True asynchronous: The driver can attach N requests consecutively and then kick once (batching), and loop in the callback.
while ((vbr = virtqueue_get_buf(...)))Harvest, amortizing the notification overhead across each request – this is also Virtio Rpmsg Bus The source of the idea of pre-delivering 256 receive buffers in RPMsg.
virtqueue in OpenAMP
In the OpenAMP scenario of this series (STM32MP157, A7 running Linux, M4 bare-metal firmware), the M4 side is OpenAMP library(lib/virtio/virtqueue.c) Implement the same split ring protocol, with the division of labor between the two sides as follows:
- The Memory Source of VRing: The position and size of vring are determined by the M4 firmware.Resource Tablein (resource table)
struct fw_rsc_vdev_vringDescription (num、align、daetc.), on the A7 side, after remoteproc parses the resource table, it allocates actual memory in the shared memory (carveout) and establishes mapping according to the same vring ID. This ensures that the A7’svirtio_ring.cAnd what is seen with the M4’s OpenAMP isthe same ring on the same block of physical memory。 - Data path: The RPMsg TX/RX buffers themselves are placed in shared memory and are “handed over” to each other via vring descriptors. The specific call flow is Virtio Rpmsg Bus analyzed in detail.
- Notification path: OpenAMP’s
virtqueue_kick()ultimately reaches the virtio device’s notify callback. On STM32MP157, it writes a channel flag to IPCC to trigger an interrupt on the remote side (see IPCC)。
The OpenAMP-side functions corresponding to the Linux API (semantics are basically one-to-one, but aimed at bare-metal environments, more streamlined):
| OpenAMP API | Corresponding Linux API | Function |
|---|---|---|
virtqueue_create() | vring_create_virtqueue() | Initialize the queue on the given vring memory, register notify and callback. |
virtqueue_fill_avail_buffers() | virtqueue_add_inbuf() | Pre-deliver writable buffers in batches (RPMsg RX queue usage). |
virtqueue_add_buffer() | virtqueue_add_sgs() | Attach the buffer chain to the avail ring. |
virtqueue_get_buffer() | virtqueue_get_buf() | Retrieve buffers from the used ring (RX receives data / TX reclaims). |
virtqueue_kick() | virtqueue_kick() | Notify the peer (via IPCC). |
virtqueue_disable_cb()/virtqueue_enable_cb() | Same name | Notification suppression |
OpenAMP and the Linux kernel’sstruct virtqueuedefinitions are not the same (this also confirms what was said at the beginning: the specification only constrains the vring protocol and behavior, not the organization of software structures), but as long as both parties follow the same version of the virtio specification (feature negotiation is consistent), the interaction on the vring is compatible.
