Roadmap
Echtzeit-Status jeder EULLM-Komponente, die Meilensteine, die wir erreichen, und eine vollstÀndige Geschichte jeder veröffentlichten Version.
PlattformĂŒbersicht
EULLM Engine
v0.6.2Rust-Inferenz-Laufzeit. Multimodal Vision + Audio, direkter Ollama-Ersatz mit OpenAI-kompatibler API und integrierter Chat-OberflÀche auf localhost:11435.
259 tok/s
Throughput
Vision+Audio
Multimodal
â getestet
Windows
EULLM Forge
Modell-Vertikalisierungspipeline. Komponenten fertig, End-to-End-CLI-Integration in Arbeit.
30Bâ7B
GröĂenreduktion
GGUF
Export
Beta
Pipeline
EULLM Hub
In der EU gehostetes Modellregister mit AI Act-KonformitÀtskarten. Als Prototyp in Betrieb.
Prototyp
Modelle
3 geplant
Sektoren
Nur EU
Hosting
Engine-FĂ€higkeiten (v0.6.2)
Rust-Laufzeit · continuous batching · multimodal Vision + Audio · vollstÀndig lokal auf Consumer-GPUs
259 tok/s
Throughput
16 gleichzeitige Anfragen
Vision+Audio
Multimodal
OCR, Szenen, Transkription
~2-4Ă
Quantized KV
Kontext, Q4_0/Q5/Q8
--web
Web-Browsing
model-agnostic, jedes GGUF
Was wir aufbauen
Phase 01: Fundament
Q1 2026
Der Inferenz-Engine erreicht ProduktionsqualitÀt. Forge-Pipeline-Komponenten entwickelt. Hub als Prototyp in Betrieb.
- Engine: Standalone-BinÀrdateien (Linux x64, Windows x64)
- Multimodal Vision + Audio (Gemma 4)
- Continuous batching, 259 Tok/s
- Quantized KV cache, Q4_0/Q5/Q8 (~2-4Ă Kontext)
- OpenAI-compatible + Ollama Drop-in API
- GPU: CUDA (getestet), ROCm, Vulkan, Metal
- Integriertes EU AI Act-Audit-Logging
- Transparentes Web-Browsing (--web, model-agnostic)
- Interaktives REPL: /temp, /maxtokens, /system
- Integrierte Chat-OberflÀche, localhost:11435, ~29 KB im BinÀrformat
- Forge: structural pruning + knowledge distillation
- Forge: End-to-End-Pipeline CLI
- Demo-Modell: legal-it-7b
Phase 02: Ăkosystem
Q2 2026
Die ersten produktionsbereiten Hub-Modelle gehen live. Stabile Forge-CLI. Erweiterte PlattformunterstĂŒtzung.
- Hub: Modell fĂŒr den Rechtssektor (EU/italienisches Recht)
- Hub: Modell zur UnterstĂŒtzung der medizinischen Triage
- Hub: Finanz- und KYC-Compliance-Modell
- AI Act-KonformitĂ€tskarten fĂŒr alle Hub-Modelle
- Forge: stabile CLI + vollstÀndige Dokumentation
- Windows x64-UnterstĂŒtzung
- Multi-GPU-Inferenz
- Quantisierungsassistent fĂŒr Consumer-Hardware
Phase 03: Enterprise
H2 2026
Enterprise-HÀrtung: verteilte Inferenz, Zugriffskontrolle, Forge Studio visuelle OberflÀche.
- Multi-Node verteilte Inferenz
- Kubernetes-Operator
- SSO / RBAC access control
- Forge Studio: visuelle OberflĂ€che fĂŒr das fine-tuning
- Modellversionierung und Rollback im Hub
- Zertifizierte EU-Rechenzentrumspartnerschaften
- SLA-Support-Stufen
Versionshistorie
- Multimodal in the Chat UI: drop in an image or audio clip, fully local
- Vision + audio understanding stable (Gemma 4): OCR, scene description, transcription
- BOS token handling fix for multimodal prompts
- Multimodal vision launched: image OCR and scene description on consumer GPUs
- Audio understanding (experimental, CLI), transcription and in-content search
- Runs fully local, zero telemetry
- Math expression rendering in the Chat UI
- Quantized KV cache, Q4_0/Q5/Q8 for ~2-4Ă context on the same GPU
- Embedded chat UI on localhost:11435, ~29 KB in binary, zero CDN or external dependencies
- eullm -V now shows the active backend variant
- Standalone Windows binaries: CPU and CUDA
- Web tool calling: transparent URL fetching in conversation
- Legal-IT dataset preparation module
- GPU layer fitting improvements
- Drop-in Ollama replacement with continuous batching
- Quantized KV cache for larger context on 16 GB GPUs
- Transparent web browsing without function-call overhead
- EU AI Act audit logging built-in
- Interactive REPL: /temp, /maxtokens, /system commands
- Quantized KV cache quality/accuracy automatic recommendations
- Quantized KV cache math accuracy improvements
- 1% accuracy loss isolated to matrix operations only
- Default context window increased to 2 048 tokens
- Math accuracy benchmarking suite added
- Mixed KV cache type support
- Bug fixes
- Documentation updates
- Batch scheduler refinements
- Build pipeline stabilization
Die Roadmap mitgestalten
Ăffnen Sie ein Issue, stimmen Sie ĂŒber Funktionen ab oder tragen Sie Code bei. EULLM wird öffentlich entwickelt und jede Stimme zĂ€hlt.
