Powered by EuLLM · Available through I3K

One inference runtime. Different hardware.

EuLLM is the layer that makes local AI hardware-independent — the same runtime and API, from desktop-class systems to enterprise accelerators.

Runtime

The runtime behind every Local AI Appliance

I3K builds the hardware. EuLLM is the software that runs on it: the inference engine, model catalog, and API layer that let you run LLMs and other AI workloads entirely on your own infrastructure.

  • Fully local inference — no cloud round-trip, no data leaves your infrastructure
  • OpenAI-compatible API — drop-in for existing tools and agent frameworks
  • Model management with a built-in GGUF model catalog
  • Hardware abstraction across CPU, unified memory, and discrete GPU
  • Validated on AMD, NVIDIA, and other accelerators
  • Multi-GPU support for larger models and higher concurrency
  • Choose hardware for capacity (large models) or performance (low latency)
  • Swap the underlying hardware later — same API, same application, zero rewrite
  • Powers RAG Enterprise, agents, and external applications through one API
  • No per-token cost from EuLLM when the model runs locally
Architectures

Two ways to size a Local AI Appliance

Capacity architecture

Large unified memory pools for loading very large models whole, without offloading.

Ryzen AI Max+ 395 — 128 GB unified memory. A future platform extends this to up to 192 GB.

Performance architecture

Dedicated GPUs with higher memory bandwidth, built for interactive, low-latency inference.

32 GB accelerator · 64 GB dual accelerator · 96 GB enterprise accelerator

Configurations

Appliance configurations

EuLLM Capacity 128

128 GB unified memory

EuLLM Performance 32

32 GB GPU memory

EuLLM Performance 64

64 GB aggregate GPU memory

EuLLM Enterprise 96

96 GB ECC GPU memory

EuLLM Capacity 192

Coming soon

Throughput and concurrency depend on the model, quantization, context length, and hardware. We publish benchmark numbers only after testing EuLLM directly on final configurations — not before.

Architecture

How it fits together

Business applications · Agents · RAG Enterprise
OpenAI-compatible API
EuLLM
Model runtime · scheduler · model catalog
CPU · AMD · NVIDIA · multi-accelerator hardware

Applications integrate with EuLLM, not with a specific GPU vendor.

That's what lets I3K choose the accelerator with the best price, memory, performance, and availability at any given time — without asking any application to be rewritten.

Applications

RAG Enterprise runs on EuLLM, not inside it

RAG Enterprise, agent frameworks, and other business applications talk to EuLLM through the OpenAI-compatible API. They are applications built on top of the runtime, not part of the core inference engine.

Licensing

What's core, what's licensed separately

EuLLM core

Core updates and new model/backend support are part of the maintenance subscription.

Application modules

Advanced agents, voice platform, document intelligence, and other modules are licensed separately through I3K / RAG Enterprise.

Run local AI on your own hardware

EuLLM is open-source and runs anywhere. I3K packages it into appliances sized for your workload.

Powered by EuLLM. Available through I3K Local AI.