vAiStrata

VAISTRATA Co., Ltd.

Why now

AI inference is expanding from server-hosted single-shot services into the continuous control loops of Physical AI: sensor, inference, control. A late result is a wrong result.

Why a GPU

The GPU is the accelerator every major vendor already ships, and Physical AI asks it for more than inference: point clouds, localization, sensor fusion and rendering all run there, in the same memory. Raw model throughput is not always the reason to pick it. Where an NPU is the better target, the runtime offloads to it.

Why Vulkan

Vulkan compute runs on effectively every modern GPU across desktop, mobile and embedded, and the Khronos Group maintains it as an open standard rather than any single vendor. Building on it gives the runtime one implementation across every supported vendor.

Performance was the open question, and the gap has been closing. Phoronix measured llama.cpp cases where the Vulkan backend came out ahead of the CUDA one on NVIDIA hardware, and graph-level model execution is now arriving in the standard itself as vendor extensions from ARM and Qualcomm.

Contact