The same capabilities run on different chip architectures. The usage above is identical; what differs is how fully each kind of hardware is used underneath.
Built on llama.cpp, ONNX Runtime and MLX, with the backend and quantization chosen per chip. No manual configuration needed.
We’re adapting model by model. Tell us your device and what you’d like to do with it.