One inference runtime. Different hardware.
EuLLM is the layer that makes local AI hardware-independent — the same runtime and API, from desktop-class systems to enterprise accelerators.
The runtime behind every Local AI Appliance
I3K builds the hardware. EuLLM is the software that runs on it: the inference engine, model catalog, and API layer that let you run LLMs and other AI workloads entirely on your own infrastructure.
- Fully local inference — no cloud round-trip, no data leaves your infrastructure
- OpenAI-compatible API — drop-in for existing tools and agent frameworks
- Model management with a built-in GGUF model catalog
- Hardware abstraction across CPU, unified memory, and discrete GPU
- Validated on AMD, NVIDIA, and other accelerators
- Multi-GPU support for larger models and higher concurrency
- Choose hardware for capacity (large models) or performance (low latency)
- Swap the underlying hardware later — same API, same application, zero rewrite
- Powers RAG Enterprise, agents, and external applications through one API
- No per-token cost from EuLLM when the model runs locally
Two ways to size a Local AI Appliance
Capacity architecture
Large unified memory pools for loading very large models whole, without offloading.
Ryzen AI Max+ 395 — 128 GB unified memory. A future platform extends this to up to 192 GB.
Performance architecture
Dedicated GPUs with higher memory bandwidth, built for interactive, low-latency inference.
32 GB accelerator · 64 GB dual accelerator · 96 GB enterprise accelerator
Appliance configurations
EuLLM Capacity 128
128 GB unified memory
EuLLM Performance 32
32 GB GPU memory
EuLLM Performance 64
64 GB aggregate GPU memory
EuLLM Enterprise 96
96 GB ECC GPU memory
EuLLM Capacity 192
Coming soon
Throughput and concurrency depend on the model, quantization, context length, and hardware. We publish benchmark numbers only after testing EuLLM directly on final configurations — not before.
How it fits together
“Applications integrate with EuLLM, not with a specific GPU vendor.”
That's what lets I3K choose the accelerator with the best price, memory, performance, and availability at any given time — without asking any application to be rewritten.
RAG Enterprise runs on EuLLM, not inside it
RAG Enterprise, agent frameworks, and other business applications talk to EuLLM through the OpenAI-compatible API. They are applications built on top of the runtime, not part of the core inference engine.
What's core, what's licensed separately
EuLLM core
Core updates and new model/backend support are part of the maintenance subscription.
Application modules
Advanced agents, voice platform, document intelligence, and other modules are licensed separately through I3K / RAG Enterprise.
Run local AI on your own hardware
EuLLM is open-source and runs anywhere. I3K packages it into appliances sized for your workload.
Powered by EuLLM. Available through I3K Local AI.
