Nexomia
v1.0.0 — Available now

Nexomia

Your AI. Your machine. No compromise.
Run powerful language models locally — or bring your own API key for cloud models.

Windows 10/11 64-bit · No installation required · SmartScreen: More info → Run anyway

Features

Everything you need. Nothing you don't.

Fully local by default

Models run directly on your hardware. Your conversations never leave your machine. No telemetry, no logs, no cloud sync.

Streaming chat interface

Tokens stream in real-time as the model generates. Conversation history is preserved across sessions automatically.

One-click model download

Browse curated open-source models, download in-app with real-time progress. No command line, no Python setup required.

Cloud API support

Bring your own OpenAI or Anthropic key to access GPT-4o, Claude, and more — or connect any OpenAI-compatible endpoint.

Configurable parameters

Tune temperature, max tokens, context length, and system prompt. Per-conversation or global defaults.

Single executable

Nexomia ships as a standalone EXE. No installer, no dependencies, no admin rights. Download and run.

Getting started

Up and running in four steps.

1

Download

Click the download button above. Nexomia is a single .exe file — no installer needed.

2

Launch

Run the file. If Windows SmartScreen appears, click More info → Run anyway. The app opens instantly.

3

Get a model

Open the Models tab. Download any model — we recommend Llama 3.2 3B to start (2 GB, balanced speed).

4

Start chatting

Hit Load on your downloaded model, switch to the Chat tab, and send your first message.

Local Models

Five curated models, ready to download.

All models run 100% on your hardware via llama.cpp — the fastest open-source local inference engine. No GPU required.

Llama 3.2 1B
0.8 GB ⚡ Fast Lightest model. Best for quick tasks on low-RAM machines.
Llama 3.2 3B Recommended
2.0 GB ⚖ Balanced Best balance of speed and quality for everyday use.
Phi-3.5 Mini
2.2 GB ⚖ Balanced Exceptional at reasoning and code. Punches above its size.
Gemma 2 2B
1.6 GB ⚡ Fast Natural, fluid conversation. Great all-rounder from Google.
Mistral 7B
4.4 GB 🔋 Slower Highest quality output. Requires 8 GB+ RAM.
Cloud Models

Bring your own API key.

Add your OpenAI or Anthropic key in Settings to use cloud models alongside local ones. Your key is stored locally and never shared with Vector Dynamics.

🤖
OpenAI
GPT-4o · GPT-4o mini · GPT-3.5 Turbo

Add your sk-… key from platform.openai.com. Usage billed directly to your OpenAI account.

🧠
Anthropic
Claude Sonnet 4.6 · Claude Haiku 4.5

Add your sk-ant-… key from console.anthropic.com. Billed to your Anthropic account.

⚙️
Custom Endpoint
Ollama · LM Studio · Any OpenAI-compatible API

Point Nexomia at any OpenAI-compatible server. Works with Ollama running locally or any self-hosted endpoint.

System Requirements

What you need.

Windows
OSWindows 10 / 11 (64-bit)
RAM8 GB minimum · 16 GB recommended
Storage500 MB app + model files (0.8–4.4 GB each)
GPUNot required — CPU inference
InternetOnly for initial model download
macOS (coming soon)
OSmacOS 12 Monterey or later
ChipApple Silicon (M1+) or Intel 64-bit
RAM8 GB minimum
StatusBuild in progress — sign up for early access
Documentation

Settings reference.

Temperature
Controls randomness. 0.1–0.4 for precise, focused answers. 0.7 (default) for balanced. 1.0+ for creative or exploratory outputs.
Max Tokens
Maximum response length. 512 for short answers. 2048 (default) for most tasks. 4096–8192 for long-form writing or code.
Context Length
How much of the conversation the model can "see" (local models only). Higher values use more RAM. Default 4096 tokens covers ~30 exchanges. Max depends on the model.
System Prompt
A persistent instruction prepended to every conversation. Use it to give the AI a persona, a focus area, or a response style. Leave blank for no system instruction.
API Keys
Stored in ~/Nexomia/settings.json on your local machine. Never transmitted to Vector Dynamics. Clear them any time from Settings → API Keys.
Model files
Downloaded to ~/Nexomia/models/. You can place any .gguf file there manually and Nexomia will detect it. Delete from the Models tab to free disk space.
FAQ

Common questions.

Is Nexomia really free?
Yes. Nexomia is completely free — no paid tier, no premium plan, no in-app purchases. Local models are free forever. Cloud model usage (OpenAI, Anthropic) is billed directly by those providers to your own API account.
Can I use my own API key?
Yes. Go to Settings → API Keys, enter your OpenAI or Anthropic key, and save. Then open the Models tab — cloud models will appear. Select one and start chatting. Your key is stored locally and never sent to Vector Dynamics.
Does it send data to the internet?
Only when you explicitly download a model (from Hugging Face) or use a cloud API (your messages go to OpenAI/Anthropic/your custom endpoint). Local conversations never touch the network. There's no telemetry, no crash reporting, no background connections.
Windows shows a SmartScreen warning. Is it safe?
Yes — the warning appears because Nexomia isn't yet code-signed. Click More info, then Run anyway. Code signing is planned for the next release. The source code and build process are fully controlled by Vector Dynamics.
Do I need a GPU?
No. Nexomia v1.0 runs entirely on CPU using llama.cpp. GPU acceleration is on the roadmap and will dramatically improve generation speed when added. For now, smaller models (1B–3B) feel snappy on modern CPUs.
Can I use Ollama or LM Studio with Nexomia?
Yes. In Settings → API Keys, add your endpoint URL (e.g. http://localhost:11434/v1 for Ollama) and the model name. Nexomia connects to any OpenAI-compatible HTTP API.

Run AI locally. Starting now.

Free, no account, no cloud. Just download and go.