LLM inference that fits your hardware. AI tools built around your work.
We build open-source software to make LLMs more useful on the hardware people already own — from the inference runtime to the everyday workspace.
Website · GitHub · X / Twitter
Imparo is a hardware- and workload-adaptive LLM inference engine built in Rust.
Apple Silicon with Metal is the primary development and validation path. Other backend/model combinations remain qualification-specific. Context reuse does not imply concurrent decoding or continuous batching.
Source & Quick Start · Performance · Contribute
Zeraix brings models, conversations, agents, files, and tools into an open-source desktop workspace, with local inference and optional cloud providers.
The desktop application and Imparo are separate projects. The current desktop release uses its documented llama.cpp-based runtimes; Imparo is not its default bundled engine.
Download Zeraix · Source & Docs
| Repository | Purpose |
|---|---|
| imparo-benchmarks | Published inference benchmark results and measurement notes. |
| llama-builds | Runtime and sandbox assets consumed by the Zeraix application. |
| llama-seeds | Version-matched prefix-cache seeds and auxiliary drafter assets for Zeraix. |
The two application asset repositories keep their existing download paths. They are not general-purpose standalone chat models or Imparo source mirrors.
We welcome contributions to model support, kernels, hardware validation, state management, benchmarks, integrations, and documentation. For performance changes, include a reproducible baseline and correctness checks.
Imparo issues · Zeraix issues · Contact
Project source licenses and third-party artifact licenses are documented in their respective repositories. Model weights and upstream runtime components retain their own terms.