Using Ollama Self-hosted AI
Using Ollama Self-hosted AI
What is Ollama?
Ollama is an open-source tool for running large language models (LLMs) locally on your own server. Kinexa integrates Ollama as a self-hosted AI provider, giving you:
- Zero per-request cost — No API fees, no token charges
- Data privacy — Your data never leaves your server
- Full control — Choose your own models and configurations
- No rate limits — Only limited by your hardware capacity
Available Ollama Models
| Model | Parameters | RAM Required | Best For |
|---|---|---|---|
| Qwen 2.5 | 7B | 8 GB | General purpose, multilingual (excellent Indonesian) |
| Llama 3.2 | 8B | 8 GB | General purpose, reasoning |
| Gemma 2 | 9B | 10 GB | Knowledge, summarization |
| Phi-3 | 3.8B | 4 GB | Lightweight, fast responses |
| Mistral | 7B | 8 GB | Instruction following, coding |
When to Use Ollama
Free/Starter Tier:
Ollama is the default AI provider for Free and Starter plans. You get capable AI responses without any per-request charges.
Enterprise Dedicated:
Enterprise customers can deploy dedicated Ollama instances with larger models (e.g., Qwen 2.5 72B, Llama 3.1 70B) for premium self-hosted AI.
How to Set Up
If you are on the Free/Starter tier, Ollama is already configured as your default. To change the model:
- Go to Settings > AI > Provider
- Select Ollama (Self-hosted)
- Choose your preferred model from the dropdown
- Click Save
The change takes effect immediately for all new AI conversations.
Tips
- Qwen 2.5 is recommended for Indonesian businesses — it has excellent multilingual support
- Phi-3 is the fastest option if response speed is your priority
- For customer support use cases, Qwen 2.5 or Llama 3.2 provide the best quality