Running a Local LLM Agent with Pi + LM Studio: A Complete Setup Guide

·Toolin Editorial Team

Combine the Pi agent framework with LM Studio's inference server to run Gemma 4 series models on a local Mac for coding, proofreading, and agent tasks, reaching about 75% of frontier-model performance.

Running a Local LLM Agent with Pi + LM Studio: A Complete Setup Guide

Local LLMs have crossed the "barely usable" watershed. With Google's newly released Gemma 4 series, you can run agent coding, script refactoring, and unit test writing on a 2022 M2 Mac (64GB of memory), reaching roughly 75% of frontier models' loop accuracy and speed. This tutorial breaks the whole local agent workflow into reproducible steps, for developers who want an AI coding assistant in a privacy-sensitive or offline setting.

Preparation

  • Hardware: an Apple Silicon Mac recommended (M2 or newer), 16GB of unified memory minimum, 64GB better (the KV cache will eat your memory)
  • Software: Docker Desktop, LM Studio (local inference server), the Pi agent framework
  • Model: Gemma 4 series; gemma-4-12b-qat recommended (newer, smaller, faster, with little accuracy loss), or gemma-4-26b-a4b
  • Estimated time: about 30 minutes for first-time setup

Where Local Models Stand Now

Before OpenAI released GPT-OSS in August 2025, local models weren't accurate enough for most programming tasks. Since the Gemma 4 series shipped, the author has actually used it to complete these tasks:

  • Refactoring a Python notebook into a repository of 5-6 modules
  • Linting a module and fixing generic type hints
  • Proofreading blog posts and writing unit tests
  • Scaffolding a recommender-system repo based on a two-tower model, from scratch

All of these were tasks local models simply could not handle 6 months ago.

Step 1: Start a Local Inference Server with LM Studio

  1. Install and open LM Studio.
  2. Download gemma-4-12b-qat from the model marketplace.
  3. Start the Local Server; it listens on http://localhost:1234/v1 by default (OpenAI-compatible endpoint).

Once started, LM Studio provides an OpenAI-format inference endpoint, and Pi calls the local model through it.

Tip: you can also substitute Ollama, llama.cpp, or Open WebUI for LM Studio. Going straight to llama.cpp is faster and a worthwhile optimization to try later.

Step 2: Point Pi at LM Studio

Pi configures models through models.json. Write the following into ~/.pi/agent/models.json, pointing the endpoint at LM Studio on the Docker host:

{
  "lmstudio": {
    "baseUrl": "http://host.docker.internal:1234/v1",
    "api": "openai-completions",
    "apiKey": "not-needed",
    "models": [
      {
        "id": "google/gemma-4-12b-qat",
        "input": ["text", "image"]
      }
    ]
  }
}

Note that host.docker.internal is the alias Docker containers use to reach the host machine. If you're also using the real OpenAI API, specify a separate base for OPENAI_API_BASE to avoid a conflict.

Step 3: Run Pi in Docker Compose (a Safe Sandbox)

It's strongly recommended to run Pi inside a restricted Docker container with only bash granted, so it can't read and write your physical disk directly.

services:
  pi:
    build:
      context: .
      dockerfile: Dockerfile
    image: pi-agent:0.74.0
    init: true
    stdin_open: true
    tty: true
    extra_hosts:
      - "host.docker.internal:host-gateway"
    environment:
      ANTHROPIC_API_KEY: ${ANTHROPIC_API_KEY:-}
      OPENAI_API_KEY: ${OPENAI_API_KEY:-not-needed}
      GEMINI_API_KEY: ${GEMINI_API_KEY:-}
      OPENAI_API_BASE: ${OPENAI_API_BASE:-http://host.docker.internal:1234/v1}
    volumes:
      - ${HOME}/.pi/agent/models.json:/config/models.json
      - ${WORKSPACE:-.}:/workspace
      - pi-config:/config
      - pi-sessions:/sessions
    working_dir: /workspace

volumes:
  pi-config:
  pi-sessions:

The key point: mount the current working directory as /workspace, and Pi can operate on files inside your code repository without touching system directories outside the container.

Step 4: Write a Launch Script

Here is a recommended pi launch script that auto-generates the container name from the working directory and supports a --sandbox flag for a stricter sandbox:

#!/usr/bin/env bash
# Pi — Start the containerized Pi agent.

SCRIPT_DIR="$(cd -- "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
WORKSPACE_DIR="${WORKSPACE:-$(pwd)}"
case "$WORKSPACE_DIR" in
  /) ;;
  *) WORKSPACE_DIR="$(cd -- "$WORKSPACE_DIR" && pwd)" ;;
esac
export WORKSPACE="$WORKSPACE_DIR"

sandbox="${PI_SANDBOX:-0}"
pi_args=()
while (($#)); do
  case "$1" in
    --sandbox)    sandbox=1 ;;
    --no-sandbox) sandbox=0 ;;
    *)            pi_args+=("$1") ;;
  esac
  shift
done

compose_files=(-f "$SCRIPT_DIR/docker-compose.yml")
if [[ "$sandbox" == "1" ]]; then
  compose_files+=(-f "$SCRIPT_DIR/docker-compose.sandbox.yml")
fi

repo_slug="$(basename -- "$WORKSPACE_DIR" | tr -c 'a-zA-Z0-9_.-' '-' | sed 's/^-//')"
[[ -z "$repo_slug" ]] && repo_slug="workspace"
container_name="pi-${repo_slug}-$$"

api_key_args=(-e OPENAI_API_KEY -e DEEPSEEK_API_KEY -e ANTHROPIC_API_KEY -e GEMINI_API_KEY)

cmd=(docker compose --project-directory "$SCRIPT_DIR" "${compose_files[@]}" run --rm --name "$container_name" "${api_key_args[@]}" pi)
if ((${#pi_args[@]})); then cmd+=("${pi_args[@]}"); fi
exec "${cmd[@]}"

Verify the Results

Run ./pi in the directory of a repository you're editing, and Pi will start Docker and enter /workspace. Have it do these things to verify:

  • Ask Pi to refactor one module of the current repo and watch whether it splits files correctly
  • Ask it to check whether type hints are correct
  • Watch token inference, KV cache usage, and context window changes live in the LM Studio UI

On a successful run, you'll see Pi modifying files and calling bash inside Docker while the local model infers token by token in the background.

FAQ

  • Slow inference: local models are hardware-bound with small context windows. Switch to a smaller quantized model (such as 12B QAT), or shorten the context.
  • Prompt template mismatch: a common issue with early releases; LM Studio's and HuggingFace's "use this model" buttons usually fix it quickly — just keep your tools updated.
  • KV cache full: 64GB of memory can be exhausted by long contexts; monitor memory usage and restart the session when needed.
  • Can it be used for production software development? Not fully mature yet, but it's already practical as a fast, personalized local documentation lookup and coding aid — and well worth investing in for privacy-sensitive scenarios.