Nitin Kumar SinghSolutions Architect

Type to search. to move, Enter to open.

    move open esc close
    Deep DiveDocker

    Deploying Ollama with Open WebUI Locally: A Step-by-Step Guide

    Deploy Ollama with Open WebUI locally via Docker Compose — run open-source language models on your own hardware for privacy and cost savings.

    I run LLMs on my own machine for two reasons: my prompts never leave my hardware, and I can experiment all day without watching an API meter. The setup used to be the hard part. Wrangling Python environments, CUDA drivers, and model weights by hand was enough friction that most people quit before the first token.

    Ollama removes most of that friction, and Open WebUI puts a ChatGPT-style interface on top of it. The pair runs cleanly in Docker, works offline once the models are downloaded, and the whole stack comes up with a single compose file.

    What we’ll cover

    • Two ways to run Ollama with Open WebUI locally: Docker Compose (the one I use) and a manual install
    • Pulling and managing open-source models like Llama 3 and Mistral
    • Customizing model behavior with a Modelfile
    • The failures you’re most likely to hit, and their fixes
    • Hardware sizing, security notes, and calling the Ollama API from your own code

    Hardware Requirements

    Before starting, ensure your system meets these minimum requirements:

    • CPU: 4+ cores (8+ recommended for larger models)
    • RAM: 8GB minimum (16GB+ recommended, more better.)
    • Storage: 10GB+ free space (models can range from 4GB to 50GB depending on size)
    • GPU: Optional but recommended for faster inference (NVIDIA with CUDA support)

    Understanding Ollama and Open WebUI

    What is Ollama?

    Ollama is a lightweight tool designed to simplify the deployment of large language models on local machines. It provides a user-friendly way to run, manage, and interact with open-source models like LLaMA, Mistral, and others without dealing with complex configurations.

    Key Features:

    • A CLI and an HTTP API, so you can drive it from a terminal or from code
    • Support for multiple open-source models
    • Easy model installation and management
    • Optimized for running on consumer hardware
    • Built-in parameter customization (temperature, context length, etc.)

    Official Repository: Ollama GitHub

    What is Open WebUI?

    Open WebUI is an intuitive, browser-based interface for interacting with language models. It serves as the front-end to Ollama’s backend, providing a user-friendly experience similar to commercial AI platforms.

    Key Features:

    • Clean, ChatGPT-like user interface
    • Model management capabilities
    • Conversation history tracking
    • Customizable system prompts
    • Model parameter adjustments
    • Visual response streaming

    Official Repository: Open WebUI GitHub

    Step-by-Step Implementation

    Docker Compose offers the simplest way to deploy both Ollama and Open WebUI, especially for beginners. This method requires minimal configuration and works across different operating systems.

    Prerequisites

    Before starting, ensure you have:

    • Docker1 installed and running
    • Docker Compose installed (often bundled with Docker Desktop)
    • Terminal or command prompt access

    Step 1: Create a Project Directory

    First, create a dedicated directory for our project:

    Terminal window
    mkdir ollama-webui && cd ollama-webui

    Step 2: Create the Docker Compose File

    Create a file named docker-compose.yml with the following content:

    version: '3.8'
    services:
    ollama:
    image: ollama/ollama:latest
    container_name: ollama
    ports:
    - "11434:11434" # Ollama API port
    volumes:
    - ollama_data:/root/.ollama
    restart: unless-stopped
    open-webui:
    image: ghcr.io/open-webui/open-webui:main
    container_name: open-webui
    ports:
    - "3000:8080" # Open Web UI port
    environment:
    - OLLAMA_API_BASE_URL=http://ollama:11434
    depends_on:
    - ollama
    restart: unless-stopped
    volumes:
    ollama_data:

    Step 3: Start the Services

    From your project directory, run:

    Terminal window
    docker-compose up -d

    This command starts both services in detached mode (running in the background). The first time you run this, Docker will download the necessary images, which might take a few minutes depending on your internet connection.

    Step 4: Access Open WebUI

    Once the containers are running:

    1. Open your web browser
    2. Navigate to http://localhost:3000

    You’ll see the Open WebUI landing page and signup screen:

    Open WebUI Signup Screen
    Open WebUI Signup Screen

    Step 5: Create an Account and Pull a Model

    After creating an account, you’ll need to pull a language model. The interface will look like this:

    Open WebUI Model Selection
    Open WebUI Model Selection

    When you select a model to download, you’ll see a progress indicator:

    Model Download Progress
    Model Download Progress

    Once the model is downloaded, you can start chatting:

    Chat Interface with Model
    Chat Interface with Model

    Method 2: Manual Setup

    For users who prefer more control over the installation or cannot use Docker, this method provides step-by-step instructions for setting up Ollama and Open WebUI separately.

    Prerequisites

    Ensure you have:

    • Node.js and npm2 (for Open WebUI)
    • Python 3.7+3 and pip
    • Git4

    Step 1: Install Ollama

    1. Download Ollama for your operating system:

    2. Complete the installation by following the installer instructions

    3. Open a terminal or command prompt and verify the installation:

      Terminal window
      ollama --version
    4. Pull a model to test your installation:

      Terminal window
      ollama pull mistral

    Step 2: Install and Run Open WebUI

    1. Clone the Open WebUI repository:

      Terminal window
      git clone https://github.com/open-webui/open-webui.git
    2. Navigate to the project directory:

      Terminal window
      cd open-webui
    3. Install Open WebUI using pip:

      Terminal window
      pip install open-webui
    4. Start the server:

      Terminal window
      open-webui serve
    5. Access the interface in your browser at http://localhost:3000

    Working with Models

    Available Models

    Ollama supports a variety of open-source models. Here are some popular ones to try:

    ModelSizeBest ForSample Command
    Llama38BGeneral purpose, instruction followingollama pull llama3
    Mistral7BBalanced performance and sizeollama pull mistral
    Gemma7BGoogle’s lightweight modelollama pull gemma
    Phi-22.7BEfficient for basic tasksollama pull phi
    CodeLlama7B/13BProgramming and code generationollama pull codellama

    Pulling Models

    You can pull models either through the Open WebUI interface or using the Ollama CLI:

    Using Open WebUI:

    1. Navigate to the “Models” section
    2. Search for the model you want
    3. Click “Download” or “Pull”

    Using Ollama CLI:

    Terminal window
    ollama pull mistral

    Creating Custom Models

    You can customize existing models with specific instructions using a Modelfile. This is especially useful for creating assistants with specialized knowledge or behavior.

    1. Create a file named Modelfile (no extension):

      FROM llama3
      SYSTEM "You are an AI assistant specializing in JavaScript programming. Provide code examples when asked."
    2. Create your custom model:

      Terminal window
      ollama create js-assistant -f Modelfile
    3. Run your custom model:

      Terminal window
      ollama run js-assistant

    Troubleshooting Common Issues

    • Issue: Docker containers won’t start
      Solution: Ensure Docker Desktop is running and has sufficient resources allocated

    • Issue: “port is already allocated” error
      Solution: Change the port mappings in docker-compose.yml or stop services using ports 11434 or 3000

    • Issue: Model download fails
      Solution: Check your internet connection and try again; verify you have enough disk space

    • Issue: Out of memory errors
      Solution: Try a smaller model or increase Docker’s memory allocation (in Docker Desktop settings)

    • Issue: Slow model responses
      Solution: Consider using a GPU for acceleration or switch to a smaller model

    Interface Issues

    • Issue: Can’t connect to Open WebUI
      Solution: Verify both containers are running with docker container ls and check logs with docker logs open-webui

    • Issue: Authentication problems
      Solution: Reset your browser cache or try incognito mode; restart the container if needed

    Best Practices for Local LLM Deployment

    Performance Optimization

    1. Allocate Sufficient Resources:

      • Increase Docker memory limits for better performance
      • If using NVIDIA GPU, enable CUDA support
    2. Choose the Right Model Size:

      • Smaller models (7B or less) for basic tasks and limited hardware
      • Larger models (13B+) for more complex reasoning when hardware allows
    3. Manage System Resources:

      • Close resource-intensive applications when running models
      • Monitor CPU, RAM, and GPU usage with system tools

    Security Considerations

    1. Local Network Exposure:

      • By default, the services are only available on localhost
      • Be cautious when exposing to your network (e.g., changing to 0.0.0.0:3000:8080)
    2. Data Privacy:

      • While data stays local, be mindful of what information you input
      • No data is sent to external servers unless you configure external API usage

    Advanced Use Cases

    Integration with Other Applications

    Ollama provides an API (port 11434) that you can use to integrate with custom applications:

    import requests
    def query_ollama(prompt, model="llama3"):
    response = requests.post(
    "http://localhost:11434/api/generate",
    json={"model": model, "prompt": prompt}
    )
    return response.json()["response"]
    result = query_ollama("Explain quantum computing in simple terms")
    print(result)

    RAG (Retrieval-Augmented Generation)

    You can enhance your models with local knowledge by implementing RAG:

    1. Use Open WebUI’s document upload feature
    2. Create embeddings from your documents
    3. Enable the model to reference these documents when answering questions

    Where to go next

    Pull a small model first (ollama pull llama3.2), confirm it answers in the WebUI, then swap in something larger once you know your hardware handles it. If you want this reachable from other devices on your network, put it behind a reverse proxy with auth rather than exposing the port directly. The Ollama model library is the fastest way to see what fits your memory budget before you commit to a download.

    Further reading

    References

    1. Docker — docker.com

    2. Node.js and npm — nodejs.org

    3. Python 3.7+ — python.org

    4. Git — git-scm.com

    5. Windows — ollama.com

    Comments

    Comments are GitHub discussions. Sign in with GitHub to post; reactions need no account.