Leveraging Native-Level NVIDIA GPU Performance in Virtual Machines with Virtio- Technology

A new technology has emerged that fundamentally resolves the severe performance degradation and latency issues that previously occurred when using NVIDIA graphics devices in virtual machine environments. Traditional virtualization methods often wasted enormous computational resources by translating and serializing API calls one by one within the virtual machine. However, the newly developed virtual device driver approach allows kernel-level I/O commands to be passed directly without any complex conversion processes. In actual tests involving gaming and real-time video streaming, the technology demonstrated remarkable processing speeds that were on par with running on a standard physical computer. For developers aiming to build headless streaming environments, this technology is a godsend. Going forward, if you want to fully extract graphics acceleration performance when building virtualization servers, you must examine the operating principles of this innovative technology.

=

Leveraging Native-Level NVIDIA GPU Performance in Virtual Machines with Virtio- Technology

Leveraging Native-Level NVIDIA GPU Performance in Virtual Machines with Virtio- Technology

1. Subtitle

1. Subtitle
1. Subtitle

Structural Limitations and Issues of Existing Virtualization Graphics Technologies

The biggest bottleneck when utilizing high-performance graphics cards inside a virtual machine is the unnecessary command serialization process. For example, typical virtualization devices intercept thousands of graphic commands sent by applications and subject them to translation work that crosses the virtual machine boundary. This process results in inefficiency, as a significant portion of the total computational budget is wasted merely on packaging, unpacking, or re-executing commands. In high-speed environments where the screen rendering cycle lasts only a few milliseconds, this serialization latency alone causes a noticeable drop in overall performance. Furthermore, there is a structural drawback where the complex compositor inside the virtual machine cannot directly access the host computer’s actual buffer memory. Consequently, the CPU load spikes due to the unnecessary memory copy process required to compress the displayed video and send it externally. While solutions were somewhat available for Intel or AMD graphics devices through Mesa drivers, there was no suitable alternative for NVIDIA environments.

The traditional passthrough method has a fatal weakness in terms of resource utilization because it allocates the entire graphics device to a single virtual machine. In multi-tenant environments where multiple users share a single hardware resource, this whole-device allocation significantly reduces business efficiency. However, using the existing virtualization graphics API translation method was too burdensome to handle the massive volume of commands required by the latest games or high-performance rendering applications. Ultimately, developers concluded that a new path was needed to communicate at the actual ABI level of the device driver, bypassing application-level APIs. Against this backdrop, an innovative structure was devised where a lightweight driver inside the virtual machine communicates directly with the host to transmit commands. By adopting this method, unnecessary translation stages are drastically reduced, achieving efficiency comparable to plugging a graphics card directly into a physical computer.

💡 Key Point
Existing virtualization graphics methods suffered from significant performance waste due to command serialization and memory copying, but this has been overcome with a new approach.

2. Subtitle

2. Subtitle
2. Subtitle

Driver-Level I/O Command Translation and Efficient Execution Structure

The new technology, virtio-, allows the NVIDIA user-mode library inside the virtual machine to directly generate graphics command buffers. The kernel driver within the virtual machine registers standard NVIDIA device files as-is and serializes user-sent I/O requests through a control queue. At this stage, individual graphics drawing calls are handled smoothly like local function calls, rather than being newly translated each time they cross the virtual machine boundary. The device management layer connects handles inside the virtual machine to the host device’s file descriptors and accurately translates pointers within the parameters. Shared memory regions are mapped with correct cache attributes, so only raw byte data is copied quickly, without interfering with the device’s own ABI decisions. By largely omitting these complex conversion processes, communication costs between the virtual machine and the host have been successfully minimized.

The real-time loop for rendering the screen has a unique structure where there are almost no I/O control commands that actually need to be transmitted. This is because the user-mode driver submits commands by directly writing work content to the already mapped host memory space. Therefore, graphic tasks can be processed at very high speeds without incurring any individual submission costs for each screen draw. The event handling queue, which operates in the opposite direction, has the host monitor the state of open descriptors and wake up the virtual machine waiting for graphic tasks at the appropriate time. If this wake-up path did not exist, the virtual machine would fall into polling, meaninglessly consuming resources by assuming the device is always ready. Thanks to this sophisticated event management, a foundation has been laid to minimize resource waste while pushing hardware performance to its limits.

💡 Key Point
The core innovation is that the user-mode driver writes commands directly to host memory, eliminating unnecessary communication costs and latency.

3. Subtitle

3. Subtitle
3. Subtitle

Performance Metrics and Detailed Analysis Measured in Real Hardware Environments

The development team meticulously measured performance by applying the same headless rendering load in a real system environment equipped with an NVIDIA RTX 3060 graphics card. When the host computer’s frame processing time was 39 milliseconds, the speed inside the virtual machine was actually measured to be 0.4 percent faster, demonstrating identical performance within the margin of error. In environments with a processing time of 9.9 milliseconds, it recorded a figure 0.7 percent faster, and in the 2.0-millisecond range, there was only a slight difference of 1.7 percent. In general load situations where frame processing time is approximately 2 milliseconds or more, it perfectly maintained performance levels of 98 to 100 percent of the physical computer. However, in very light load situations where frame processing time drops below 0.5 milliseconds, the cost of the graphics device waking up from standby became relatively prominent. In extreme light loads where tasks must be completed in very short moments, performance dropped to around 70 percent compared to the physical computer.

When run for 12 seconds at approximately 100 frames per second without any specific frame rate limit, CPU usage time was surprisingly low. While the host computer consumed 0.40 seconds, the inside of the virtual machine used only 0.37 seconds, proving it ran smoothly with even fewer resources. During the processing of over 800,000 frames, the backend exchanged only about 13,000 messages, demonstrating extremely high communication efficiency. On average, it crossed the virtual machine boundary only once every 59 frames, and even those were mostly initial operations for device configuration. In the actual rendering loop section where screens are drawn continuously, the number of boundary crossings per frame was only about 0.02, making the virtualization barrier virtually imperceptible. Compared to past virtualization methods that struggled while exchanging thousands of messages per frame, this can be seen as a transformative development.

💡 Key Point
In general load environments of about 2 milliseconds or more, it shows perfect performance retention indistinguishable from a physical computer.

4. Subtitle

4. Subtitle
4. Subtitle

Powerful Real-Time Encoding Capabilities in Headless Streaming Environments

The primary arena this technology ultimately targets is headless streaming, which involves rendering and compressing screens in server environments without connected monitors. In cloud gaming services or remote virtual desktop environments, the ability to quickly render and compress screens for real-time transmission to user terminals is vital. By utilizing the newly developed method, graphics can be rendered directly inside the virtual machine, windows composited, and connected directly to the hardware encoder. Since only the compressed bitstream generated inside the virtual machine needs to be sent externally, it provides smooth visuals while drastically saving network bandwidth. To achieve this, full driver-level access rights to buffer handles, fences, device pointers, and hardware encoding sessions must be perfectly supported. Existing virtualization technologies suffered from skyrocketing latency issues because they had to pass through the CPU every time while processing these complex pipelines.

Tests were also successfully completed where multiple virtual machines simultaneously performed rendering tasks and video encoding on a single high-performance graphics device. Even with four virtual machines simultaneously encoding high-quality video from a single physical device, the total throughput was almost identical to that of a single virtual machine. Since hardware resources are distributed equally and fairly to each virtual machine, the phenomenon of the entire system slowing down due to a specific virtual machine did not occur. The outgoing data is a compressed bitstream of only around 100 kilobytes per frame, so the burden on the company’s network equipment is also very low. For companies operating remote work environments or cloud-based real-time collaboration tools, this is an excellent alternative that can significantly reduce server operating costs. In the future, this virtualization technology will become an essential infrastructure choice for developers planning high-quality real-time streaming services.

💡 Key Point
It dramatically reduces network bandwidth while simultaneously performing rendering and real-time encoding inside the virtual machine.

5. Subtitle

5. Subtitle
5. Subtitle

Security Isolation Levels and Practical Application Scenarios in Multi-Tenant Environments

Although significant progress has been made in performance, there are important security characteristics that must be considered when deploying this technology in the field. The newly developed method does not provide hardware-level isolation based on IOMMU (Input-Output Memory Management Unit), which is commonly used in virtualization technologies. In other words, since it must fully trust the NVIDIA driver installed on the host computer, it is not suitable for untrusted multi-tenant environments. If multiple customers’ virtual machines share a single physical server in a public cloud environment, the traditional whole-device allocation method should be used. On the other hand, in places where security trust is established, such as building internal infrastructure or operating virtual servers within trusted development teams, it is undoubtedly the best choice. Since performance could be maximized by removing unnecessary hardware isolation layers, careful selection based on the intended use is necessary.

Corporate IT managers must accurately understand these technical characteristics when formulating server virtualization strategies and deploy them in appropriate environments. For example, when building high-performance remote work virtual machines for in-house designers, adopting this method can achieve both cost efficiency and speed. It flexibly distributes expensive graphics cards to multiple virtual machines while providing near-native screen smoothness, significantly improving user satisfaction. Additionally, when processing AI model training or large-scale data analysis tasks in virtualized environments, efficient computation without resource waste becomes possible. It is a clear trend that the future direction of virtualization technology is to break down hardware barriers and overcome performance limits. In the midst of this wave of change, it is worth paying attention to what changes this new driver translation technology will bring to practical fields.

💡 Key Point
Since it lacks hardware isolation features, it is most suitable for internal server builds of single enterprises rather than untrusted multi-user environments.

6. Subtitle

6. Subtitle
6. Subtitle

Future Prospects of Virtualization Graphics and Response Strategies for System Engineers

In the future, technology for handling graphics devices in virtual machine environments will evolve in a direction that further blurs the boundaries between hardware and software. In the past, performing high-performance graphics tasks in virtualized environments was either impossible or required accepting enormous performance sacrifices. However, with the popularization of direct driver-level communication and optimized memory sharing techniques, the concept of virtualization itself has entered a completely new phase. Now, system engineers must consider architectures that minimize communication costs between the host and guest, rather than simply launching virtual machines. How flexibly one can secure actual control rights of the driver while removing unnecessary translation stages will become the key criterion for the success of future infrastructure builds. The innovative technology that has emerged this time will serve as a catalyst that prompts more graphics manufacturers to consider virtualization-friendly driver structures.

We recommend that readers also review the graphics pipelines of their practical systems in line with the upcoming changes in the virtualization ecosystem. In particular, if you are operating remote streaming or cloud-based real-time rendering services, you must face the limitations of existing virtualization methods. Test new approaches that can handle encoding all at once inside the virtual machine and explore whether they can be applied to your company’s systems. It may seem like an unfamiliar structure at first, but once you directly measure the performance, you will be amazed at why this technology is attracting attention. In the increasingly fierce competition in cloud infrastructure, we hope you will complete a high-performance graphics virtualization system one step ahead of others. Only companies that accurately understand the essence of the technology and adopt it wisely can become winners in the upcoming digital innovation.

💡 Key Point
Driver-level optimization is changing the landscape of virtualization infrastructure, so the adoption of new architectures should be actively considered.

Frequently Asked Questions

How much performance difference is there compared to a regular computer when using this technology?
In general load environments where frame processing time is 2 milliseconds or more, it shows nearly perfect speed reaching 98 to 100 percent of physical computer performance.
In what environments is it most effective to adopt this technology?
It is most suitable for headless streaming environments where rendering is completed on servers without monitors and compressed video is transmitted, or for building high-performance virtual desktops within a company.
What is the biggest advantage compared to the existing passthrough method?
It can achieve near-native graphics performance inside the virtual machine without allocating the entire graphics device, resulting in very high resource utilization efficiency.
Are there any points to note regarding security or hardware isolation?
Since it does not provide IOMMU-based hardware isolation, it is suitable for trusted single-enterprise environments rather than untrusted multi-user public clouds.

=