vAiStrata

Run your AI pipeline
on any GPU

A cross-platform AI inference runtime on Vulkan.

Coming soonComing soon

Product

One code, Every GPU

The vAi and vStream pipelineCamera, LiDAR and radar input enters vStream, a zero-copy pipeline of sensor, preprocess, inference, postprocess and output stages. The inference stage is vAi, the Vulkan AI engine, built from NeuralNet, Tensor, Node and Loader modules. AI model input uses SafeTensors weights; ONNX import is planned. Output leaves as display, motor control and data stream. The same codebase runs on NVIDIA, AMD, Qualcomm, ARM and Intel GPUs.AI ModelStandard model fileSafeTensorsONNX (planned)Default onHugging Face HubCameraLiDARRadarvStreamSensorPreprocessvAiNeuralNetTensorNodeLoaderInferencePostprocessOutputDisplayMotor controlData streamSingle codebaseNVIDIAAMDQualcommARMIntel

vAi Inference engine

Runs your model on Vulkan compute, on any GPU with a Vulkan driver.

vStream Zero-copy pipeline

Keeps every stage of the pipeline in the same memory, end to end.

EVA Vulkan wrapper

Easy Vulkan API. A wrapper that makes Vulkan easy to use.

Landscape

vAi + vStream: one pipeline, from sensor to output, across GPU vendors.

For high-performance, multi-sensor Physical AI devices: robotics, vehicles, medical and industrial equipment.

Scroll horizontally to see both GPU categories.

Landscape by GPU vendor coverage and pipeline scope
Single-vendor chipsMultiple GPU vendors (cross-platform)
Sensor-to-output pipeline

Chip-vendor pipeline SDKs

NVIDIA DeepStream, Holoscan, Isaac ROS

Sensor to output on the vendor's GPUs

Cross-platform pipelines

Fluendo Raven, NNStreamer

Sensor to output across GPU vendors

Inference only

Chip-vendor inference SDKs

NVIDIA TensorRT, Intel OpenVINO, Qualcomm AI Engine Direct

Model execution on vendor chips

Cross-platform inference frameworks

ONNX Runtime, ExecuTorch, LiteRT, ncnn, MNN, llama.cpp

Model execution across chip vendors

Adjacent layers excluded from this comparison: compilers (TVM, IREE) and robotics middleware (ROS 2).

*vAi + vStream currently supports a single stream. Multi-sensor synchronization, branching and merging are in design. The discrete-GPU path avoids a CPU round-trip and includes one copy within GPU memory.

Why GPU-native

Keep data on the GPU

Every major vendor already ships a GPU, and real time depends on how often your data leaves it.

Host round-trip

GPU MEMORYDecodePreprocessInferencePostHost memory
COPY AT EVERY BOUNDARY

Each stage hands its result back through host memory, and those copies are where the time goes.

GPU-resident

GPU MEMORYDecodePreprocessInferencePost
ONE UPLOAD, ONE READBACK

Decode, preprocessing, inference and post-processing run on the same buffers, and nothing crosses back until the result does.

Open source

Get the code and be involved

vAi and vStream are Apache 2.0. EVA is MIT. No revenue threshold, no royalties.

vAi builds its model graphs through a C++ graph DSL. Adding or tuning a kernel is a contained change you can land without reading the whole engine.

  • Port a model in a few files; the demo ships in the tree for anyone to run
  • Add a node; a kernel that gets faster on yours gets faster for everyone
  • Cover a GPU we missed; cross-vendor support comes from people who share
Coming soonComing soon

Roadmap

What ships next

Dates are targets, not commitments. Anything without one is scoped but not scheduled.

EVAVulkan wrapper, MIT. Public and taking issues today.Live
vAiVulkan inference engine, Apache 2.0.Q4 2026
vStreamZero-copy GPU pipeline, Apache 2.0. Ships with vAi.Q4 2026
ONNX importLoad a model file directly. Today the path is SafeTensors weights with the graph in code.Planned
CDNA · XDNAAMD CDNA (ROCm) and Ryzen AI (XDNA) as backends beyond Vulkan.Planned
Safety trackDeterministic compute driver on Vulkan SC, targeting ISO 26262, IEC 61508 and ISO/PAS 8800.Planned
  • Live on GitHub
  • Dated target
  • Scoped, not scheduled