KinexaSupport

Using Ollama Self-hosted AI

Using Ollama Self-hosted AI

What is Ollama?

Ollama is an open-source tool for running large language models (LLMs) locally on your own server. Kinexa integrates Ollama as a self-hosted AI provider, giving you:

  • Zero per-request cost No API fees, no token charges
  • Data privacy Your data never leaves your server
  • Full control Choose your own models and configurations
  • No rate limits Only limited by your hardware capacity

Available Ollama Models

ModelParametersRAM RequiredBest For
Qwen 2.57B8 GBGeneral purpose, multilingual (excellent Indonesian)
Llama 3.28B8 GBGeneral purpose, reasoning
Gemma 29B10 GBKnowledge, summarization
Phi-33.8B4 GBLightweight, fast responses
Mistral7B8 GBInstruction following, coding

When to Use Ollama

Free/Starter Tier:

Ollama is the default AI provider for Free and Starter plans. You get capable AI responses without any per-request charges.

Enterprise Dedicated:

Enterprise customers can deploy dedicated Ollama instances with larger models (e.g., Qwen 2.5 72B, Llama 3.1 70B) for premium self-hosted AI.

How to Set Up

If you are on the Free/Starter tier, Ollama is already configured as your default. To change the model:

  1. Go to Settings > AI > Provider
  2. Select Ollama (Self-hosted)
  3. Choose your preferred model from the dropdown
  4. Click Save

The change takes effect immediately for all new AI conversations.

Tips

  • Qwen 2.5 is recommended for Indonesian businesses — it has excellent multilingual support
  • Phi-3 is the fastest option if response speed is your priority
  • For customer support use cases, Qwen 2.5 or Llama 3.2 provide the best quality
Was this article helpful?

Need more help? Contact our support team