Jan.ai is an open-source ChatGPT alternative that runs 100% offline on your computer, powered by llama.cpp. It runs an OpenAI-compatible API server at localhost:1337, so you can run models like Llama 3, Mistral, and Qwen locally without sending data to external servers.
I'll cover setting up Jan.ai, tuning its performance, and integrating it with your development workflow, plus a few lesser-known tricks that improve the experience.
Why Jan.ai Over Cloud Solutions?
Most people default to ChatGPT, but running AI locally offers distinct advantages:
- Zero latency for offline work - Perfect for air-gapped environments or when traveling
- Complete data sovereignty - Your prompts never leave your machine
- No rate limits or usage caps - Generate as much as your hardware allows
- Customizable censorship levels - Adjust alignment and moderation to your needs
- Free API endpoint - Build applications without worrying about OpenAI credits
Prerequisites
A jan.ai local AI setup runs the model on your own silicon, so the hardware you have decides which models you can actually use. Check these before you download anything, because the most common first-run disappointment is picking a model that is too big for the machine and watching it crawl at a word every few seconds.
Operating system
- Windows: Windows 10 or higher
- macOS: macOS 13.6 or higher, Intel and Apple Silicon both supported
- Linux: most modern distributions, via the .deb package or AppImage
Jan is released under Apache 2.0, which matters if you are rolling it out inside a company rather than on a personal laptop. The version line has moved to v0.8.x, so if you are following older screenshots elsewhere, expect some menus to have shifted.
CPU and RAM
- CPU: AVX2 support, which means Intel Haswell (2013 or later) or AMD Excavator (2015 or later)
- RAM: 8GB minimum, 16GB recommended
RAM is the real constraint on model choice. The official guidance maps to roughly this:
| System RAM | Model size that fits comfortably |
|---|---|
| 8GB | 3B |
| 16GB | 7B |
| 32GB | 13B |
These are quantized GGUF models, and the numbers assume you still want the rest of the machine to be usable. If you keep a browser with 40 tabs and an IDE open, drop a tier.
Pro tip: Run cat /proc/cpuinfo | grep avx2 on Linux or check with CPU-Z on Windows to verify AVX2 support.
GPU acceleration (optional but recommended)
- NVIDIA: 6GB+ VRAM with CUDA 12.0+
- AMD: 6GB+ VRAM with Vulkan support
- Intel Arc: supported on Windows
- Apple Silicon: Metal support, built in, no driver work needed
Without a GPU, Jan falls back to CPU inference. It works, and a 3B model on CPU is fine for short questions, though generation speed drops noticeably compared with the same model offloaded to VRAM.
One failure mode worth knowing about before you install: on some older NVIDIA cards, Jan tries to initialize CUDA, fails without printing anything, and hangs on launch with no error dialog and nothing useful in the logs. The process shows up as running, but the window never appears. If that happens, disable GPU in settings.json in the Jan data folder (~/.jan on Linux, ~/Library/Application Support/Jan on macOS) and start again in CPU mode to confirm the rest of the install is sound.
Storage
- Minimum: 10GB free for the app plus a first model
Budget more than the minimum if you plan to compare models. Jan keeps downloaded weights, conversation threads, settings, and logs in a local data folder, and a handful of 7B models will fill 10GB quickly. Nothing in that folder leaves your machine, which is the point of running locally, but it does mean disk usage grows every time you try something new from the Hub.
Step 1: Install Jan.ai and configure for maximum performance
A jan.ai local AI setup lives or dies on two decisions you make in the first ten minutes: whether your hardware can hold the model you want, and whether you fix the default settings before your first chat. Jan installs happily on machines that then generate text at a crawl, so check the numbers first.
Check what your hardware can run
Jan's official requirements tie model size to system RAM. The current v0.8.x line (v0.8.3 shipped 2026-06-24) runs on macOS 13.6 or newer and Windows 10 or newer, with these tiers:
| System RAM | Largest comfortable model |
|---|---|
| 8GB | 3B |
| 16GB | 7B |
| 32GB | 13B |
On Windows there are a few extra constraints worth confirming before you download: the CPU needs AVX2 support (Intel Haswell from 2013 onward, AMD Excavator from 2015 onward), you want at least 10GB of free storage for the app plus model files, and GPU acceleration expects 6GB of VRAM on an NVIDIA, AMD, or Intel Arc card. 8GB of system RAM is the stated minimum, 16GB the recommendation.
Jan is Apache 2.0 licensed and built by Menlo Research, which matters if you plan to roll it out inside a company rather than on a personal laptop.
Download and install
# macOS/Linux users can use Homebrew (unofficial)
brew install --cask jan
# Or download directly
wget https://github.com/menloresearch/jan/releases/latest/download/jan-mac-x64.dmg # macOS Intel
wget https://github.com/menloresearch/jan/releases/latest/download/jan-mac-arm64.dmg # macOS Silicon
wget https://github.com/menloresearch/jan/releases/latest/download/jan-linux-x86_64.AppImage # Linux
For Windows users, grab the .exe from jan.ai or GitHub releases. Run it, wait for the installer to finish, then launch the app. Linux users can take the .AppImage or the .deb package depending on the distribution.
On first launch Jan downloads a default foundation model on its own, so the app is chat-ready without any further setup. Everything it pulls down stays on disk locally.
If Jan hangs on first launch
One failure mode catches people out before they ever see the interface: the app starts, the process shows up in ps, and the window never appears. No crash dialog, no useful log output, even with verbose flags. The cause is GPU initialization. On some NVIDIA cards Jan tries to bring up CUDA, fails silently, and freezes instead of reporting the error.
The fix is to disable GPU in the config before the first successful run. Edit settings.json in Jan's data folder, set GPU off, then relaunch and confirm the app opens on CPU. Once you have a working window you can re-enable acceleration from the UI and see whether it holds.
Jan data folder locations:
Windows: ~\Users\<YourUsername>\AppData\Roaming\Jan\data
macOS: ~/Library/Application Support/Jan
Linux: ~/.jan
That folder holds your models, conversation threads, settings, and logs. Nothing in it is sent anywhere.
Initial performance optimization
Once installed, go to Settings > Hardware and:
- Enable GPU acceleration if you have a compatible GPU. This is the single biggest speed change available to you, and on a 6GB card it decides whether a 7B model responds in a conversational rhythm or in paragraphs you wait for.
- Set CPU threads to your physical core count minus 2 (leave some for the OS). Handing every core to inference makes the rest of the machine stutter while a response streams, and it rarely buys proportional speed.
- Adjust context size based on your RAM:Context is memory you pay for up front. Setting it far above what your conversations need eats the headroom the model itself wants, and on a tight machine it pushes you into swap.
- 8GB RAM: 2048 tokens
- 16GB RAM: 4096 tokens
- 32GB+ RAM: 8192+ tokens
Set these before you download anything from the Hub. Model choice and context size interact, and picking a model while the defaults are wrong leads to blaming the model for a settings problem.
Step 2: Download Models Strategically
Here's how to choose:
Quick Model Selection Guide
| RAM Available | Recommended Model Size | Example Models |
|---|---|---|
| 8GB | 3B-7B | Qwen2.5-3B, Mistral-7B-Q4 |
| 16GB | 7B-13B | Llama-3.1-8B, Mistral-Nemo-12B |
| 32GB+ | 13B-30B | Llama-3.1-70B-Q4, Qwen2.5-32B |
Download via Jan Hub
- Click the Hub icon (four squares)
- Filter by your hardware capabilities
- Look for models with these quantization levels:
- Q4_K_M: Best balance (recommended)
- Q5_K_M: Higher quality, more RAM
- Q3_K_S: Faster but lower quality
Advanced: Import Custom GGUF Models
# Download directly from Hugging Face
wget https://huggingface.co/TheBloke/Mistral-7B-Instruct-v0.2-GGUF/resolve/main/mistral-7b-instruct-v0.2.Q4_K_M.gguf
# Move to Jan's model directory
mv mistral-7b-instruct-v0.2.Q4_K_M.gguf ~/jan/models/
Then create a model.json in the same directory:
{
"id": "mistral-7b-custom",
"object": "model",
"name": "Mistral 7B Custom",
"version": "1.0",
"format": "gguf",
"settings": {
"ctx_len": 4096,
"ngl": 35,
"embedding": false,
"n_batch": 512,
"n_parallel": 4
}
}
Step 3: Optimize GPU Acceleration (The Big Speed Win)
This is where most tutorials stop, but proper GPU configuration can 10x your inference speed.
NVIDIA CUDA Setup
# Verify CUDA installation
nvidia-smi
# If you see your GPU, enable it in Jan
# Settings > Hardware > GPUs > Toggle ON
Critical setting: Adjust the ngl (GPU layers) parameter:
- Start with
ngl: 35for 7B models - Monitor VRAM usage with
nvidia-smi -l 1 - Increase until you hit 90% VRAM utilization
- Back off by 5 if you experience crashes
AMD Vulkan Configuration
For AMD GPUs, Jan uses Vulkan. Set the backend:
# In Jan's settings.json
"llama_cpp_backend": "vulkan"
Apple Silicon Optimization
M-series Macs use Metal acceleration automatically, but you can tune it:
{
"n_gpu_layers": -1, // Use all available
"metal_buffers": true,
"use_mlock": true
}
Step 4: Expose the Local API Server
Jan exposes an OpenAI-compatible API at localhost:1337. Here's how to put it to work:
Enable the API Server
- Click the <> button in Jan
- Navigate to Local API Server
- Configure:
- Host:
127.0.0.1(local only) or0.0.0.0(network access) - Port:
1337(or any available port) - API Key: Set any string (e.g., "jan-local-key")
- Enable CORS: Toggle on for web apps
- Host:
Test the API
import requests
import json
url = "http://localhost:1337/v1/chat/completions"
headers = {"Content-Type": "application/json"}
payload = {
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain quantum computing in one sentence."}
],
"model": "mistral-7b-q4", # Use your loaded model ID
"stream": False,
"temperature": 0.7
}
response = requests.post(url, headers=headers, json=payload)
print(response.json()['choices'][0]['message']['content'])
Expose via Reverse Proxy (Advanced)
Want to access Jan from other devices? Use ngrok or Cloudflare Tunnel:
# Using ngrok
ngrok http 1337
# Using Cloudflare Tunnel
cloudflared tunnel --url http://localhost:1337
⚠️ Security Warning: Only expose with proper authentication in production!
Step 5: Integrate with Development Tools
VS Code + Continue.dev Integration
This setup gives you GitHub Copilot-like features using local models:
- Install Continue extension in VS Code
- Configure
~/.continue/config.json:
{
"models": [
{
"title": "Jan Local",
"provider": "openai",
"model": "mistral-7b-q4",
"apiKey": "EMPTY",
"apiBase": "http://localhost:1337"
}
]
}
- Use
Ctrl+Lto chat,Ctrl+Ifor inline edits
Shell Integration with CLI
Create a bash function for quick AI queries:
# Add to ~/.bashrc or ~/.zshrc
ai() {
curl -s http://localhost:1337/v1/chat/completions \
-H "Content-Type: application/json" \
-d "{
\"model\": \"mistral-7b-q4\",
\"messages\": [{\"role\": \"user\", \"content\": \"$*\"}],
\"temperature\": 0.7
}" | jq -r '.choices[0].message.content'
}
# Usage
ai "Convert this to Python: ls -la | grep .txt"
Build Custom Applications
Since Jan provides an OpenAI-compatible API, you can use any OpenAI SDK:
// Node.js example
import OpenAI from 'openai';
const openai = new OpenAI({
baseURL: 'http://localhost:1337/v1',
apiKey: 'jan-local-key',
});
const completion = await openai.chat.completions.create({
model: 'mistral-7b-q4',
messages: [{ role: 'user', content: 'Hello!' }],
});
Step 6: Advanced Optimization Tricks
Memory Optimization for Large Models
If you're hitting RAM limits, use these techniques:
- Reduce context window: Settings > Model > Context Length
- Enable mmap: Allows the OS to page model data
- Use quantization: Q3 versions use ~30% less RAM than Q5
{
"use_mmap": true,
"mlock": false, // Don't lock in RAM
"n_batch": 256, // Smaller batches
"n_ctx": 2048 // Reduced context
}
Speed Hacks
- Disable unused features:
{
"use_flash_attn": false, // If not supported
"embedding": false, // If not needed
"low_vram": true // For limited VRAM
}
- Use CPU+GPU hybrid: Set
nglto partial layers (e.g., 20 out of 35) - Parallel processing: Increase
n_parallelfor batch inference
Model Mixing Strategy
Run multiple specialized models instead of one large general model:
- Coding: CodeQwen-7B for development tasks
- Writing: Mistral-7B for general text
- Analysis: Llama-3.1-8B for reasoning
- Small tasks: Qwen2.5-3B for quick responses
Switch between them based on task requirements - smaller models often outperform larger ones in specialized domains.
Troubleshooting Common Issues
"Model won't start" or "Failed to fetch"
- Reduce
ngl(GPU layers) in model settings - Clear browser cache if using web interface
- Check available RAM/VRAM
Slow inference on CPU
- Ensure AVX2 is enabled in BIOS
- Reduce context size to 2048
- Use Q3 quantization instead of Q4/Q5
GPU not detected
# For NVIDIA
sudo nvidia-modprobe -c 0 -u
# For AMD
sudo modprobe amdgpu
# Restart Jan after driver fixes
The Real Power Move: Building Your Own AI Stack
Jan is also infrastructure for building private AI applications. Combine it with:
- LangChain for complex chains
- Qdrant for vector search
- n8n for automation workflows
- Open WebUI as a team interface
Together these tools give you a local AI stack with no per-query cost.
Conclusion
Jan.ai transforms your computer into a private AI powerhouse. While cloud services have their place, the ability to run models locally - with complete privacy, no rate limits, and full customization - is invaluable for developers, researchers, and privacy-conscious users.
Start with a smaller model like Mistral-7B-Q4, get comfortable with the API, then scale up as needed. The ecosystem is evolving rapidly, and having local inference capability puts you ahead of the curve.
Remember: The best model is the one that runs reliably on your hardware. Don't chase parameter counts - chase actual utility.