Hand-written GPU and NPU inference engines and open-source SDKs that run AI models on-device across iOS, Android, macOS, Windows, Linux, and web.
RunAnywhere writes custom inference kernels by hand for consumer silicon, including MetalRT for Apple GPUs and QHexRT for Qualcomm Hexagon NPUs supporting LLM, VLM, STT, TTS, and embeddings. A single C++ core handles inference, model management, routing, and telemetry, wrapped by six open-source SDKs (Swift, Kotlin, React Native, Flutter, TypeScript, C++). A hosted console adds OTA model updates and fleet management for enterprise deployments. RunAnywhere is a product of RunAnywhere.