Skip to content
aiolmPublic

About

AioLM — All-In-One LM. Windows desktop llama.cpp runtime manager for models, streaming chat, tuning, and benchmarks.

Resources

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Latest commit

 

History

224 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AioLM

Language: English | 한국어 | 日本語 | 中文

Documentation index — installation, development, architecture, and policies.

AioLM (All-in-One LM) — Windows/macOS/Linux desktop runtime manager for llama.cpp. Uses llama-server with a Tauri v2 desktop UI for models, runtimes, chat, and benchmarks.

Features

  • GGUF model discovery and safe management
  • Managed runtimes (CPU/Vulkan/ROCm/CUDA/SYCL/OpenVINO) with portable ZIP support
  • PR builds by number/URL with provenance review
  • Streaming chat with local threads, document context, and embeddings
  • User AGENTS.md instructions, local skills, and an in-app personalization editor
  • Projects, Hugging Face Discover, and MCP servers
  • Independent model loading and one API server page for OpenAI- and Anthropic-compatible clients
  • Tuning for server and sampling parameters
  • New-version notification at startup and verified installer updates from Settings

In Tuning, Reset all tuning removes all overrides, including raw server arguments and chat JSON; each field also has Reset to default. Defaults are inherited from the selected llama.cpp runtime/model, not hard-coded recommendation presets. Set custom value opts back into an override, and profiles retain the default-mode selection. Model files, adapters, GPU assignments, runtime selection and saved profiles are preserved. Server changes require Apply & restart; default context size and memory use can vary by runtime/model.

Loading a model makes it available to internal chat. Start the API server separately from API server in the sidebar to use loaded models from other applications: copy the URL, key and model ID, and pick the OpenAI or Anthropic example on the same page. The API stays up when models are unloaded or replaced, and stopping it leaves internal chat and loaded models available. See API server and model lifecycle.

Use Settings → Personalization to edit your shared and AioLM-specific AGENTS.md instructions. Chat reads these files for new turns and can load local skills automatically or through the skill picker and $skill-name references. See chat personalization.

Platform support

Windows x64, macOS 13.3+ (Apple Silicon and Intel) and Linux x86_64 (Ubuntu 24.04+) builds are available. Windows uses NSIS/MSI, macOS uses DMG and Linux uses DEB/AppImage. Linux packages are included starting with v0.2.1; see the Linux installation guide. macOS 13.3+ DMGs for Apple Silicon (with Metal) and Intel (CPU) are included starting with v0.3.0, ad-hoc signed but not notarized, with a verified terminal installer; macOS validation describes the hosted checks and remaining device work. Linux ARM64/NVIDIA DGX acceptance remains pending. See cross-platform validation. Release assets are unsigned and include SHA-256 checksums.

Download

Linux: DEB / AppImage

Windows:

powershell -ExecutionPolicy Bypass -Command "irm https://github.com/aiolm/AioLM/releases/latest/download/install.ps1 | iex"

Or get installer files directly from GitHub Releases.

For advanced install options, verification, development setup, and CLI usage, see install.md, development.md, and cli.md.

For data migration and compatibility, see the migration guide.

Security and privacy

See Security and Privacy.

License

MIT — see LICENSE. llama.cpp binaries — see NOTICE.

About

AioLM — All-In-One LM. Windows desktop llama.cpp runtime manager for models, streaming chat, tuning, and benchmarks.

Resources

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages