Running a Local LLM Agent with Pi + LM Studio: A Complete Setup Guide
Combine the Pi agent framework with LM Studio's inference server to run Gemma 4 series models on a local Mac for coding, proofreading, and agent tasks, reaching about 75% of frontier-model performance.


Running a Local LLM Agent with Pi + LM Studio: A Complete Setup Guide
Combine the Pi agent framework with LM Studio's inference server to run Gemma 4 series models on a local Mac for coding, proofreading, and agent tasks, reaching about 75% of frontier-model performance.
Local LLMs have crossed the "barely usable" watershed. With Google's newly released Gemma 4 series, you can run agent coding, script refactoring, and unit test writing on a 2022 M2 Mac (64GB of memory), reaching roughly 75% of frontier models' loop accuracy and speed. This tutorial breaks the whole local agent workflow into reproducible steps, for developers who want an AI coding assistant in a privacy-sensitive or offline setting.
Preparation
- Hardware: an Apple Silicon Mac recommended (M2 or newer), 16GB of unified memory minimum, 64GB better (the KV cache will eat your memory)
- Software: Docker Desktop, LM Studio (local inference server), the Pi agent framework
- Model: Gemma 4 series;
gemma-4-12b-qatrecommended (newer, smaller, faster, with little accuracy loss), orgemma-4-26b-a4b - Estimated time: about 30 minutes for first-time setup
Where Local Models Stand Now
Before OpenAI released GPT-OSS in August 2025, local models weren't accurate enough for most programming tasks. Since the Gemma 4 series shipped, the author has actually used it to complete these tasks:
- Refactoring a Python notebook into a repository of 5-6 modules
- Linting a module and fixing generic type hints
- Proofreading blog posts and writing unit tests
- Scaffolding a recommender-system repo based on a two-tower model, from scratch
All of these were tasks local models simply could not handle 6 months ago.
Step 1: Start a Local Inference Server with LM Studio
- Install and open LM Studio.
- Download
gemma-4-12b-qatfrom the model marketplace. - Start the Local Server; it listens on
http://localhost:1234/v1by default (OpenAI-compatible endpoint).
Once started, LM Studio provides an OpenAI-format inference endpoint, and Pi calls the local model through it.
Tip: you can also substitute Ollama, llama.cpp, or Open WebUI for LM Studio. Going straight to llama.cpp is faster and a worthwhile optimization to try later.
Step 2: Point Pi at LM Studio
Pi configures models through models.json. Write the following into ~/.pi/agent/models.json, pointing the endpoint at LM Studio on the Docker host:
{
"lmstudio": {
"baseUrl": "http://host.docker.internal:1234/v1",
"api": "openai-completions",
"apiKey": "not-needed",
"models": [
{
"id": "google/gemma-4-12b-qat",
"input": ["text", "image"]
}
]
}
}Note that host.docker.internal is the alias Docker containers use to reach the host machine. If you're also using the real OpenAI API, specify a separate base for OPENAI_API_BASE to avoid a conflict.
Step 3: Run Pi in Docker Compose (a Safe Sandbox)
It's strongly recommended to run Pi inside a restricted Docker container with only bash granted, so it can't read and write your physical disk directly.
services:
pi:
build:
context: .
dockerfile: Dockerfile
image: pi-agent:0.74.0
init: true
stdin_open: true
tty: true
extra_hosts:
- "host.docker.internal:host-gateway"
environment:
ANTHROPIC_API_KEY: ${ANTHROPIC_API_KEY:-}
OPENAI_API_KEY: ${OPENAI_API_KEY:-not-needed}
GEMINI_API_KEY: ${GEMINI_API_KEY:-}
OPENAI_API_BASE: ${OPENAI_API_BASE:-http://host.docker.internal:1234/v1}
volumes:
- ${HOME}/.pi/agent/models.json:/config/models.json
- ${WORKSPACE:-.}:/workspace
- pi-config:/config
- pi-sessions:/sessions
working_dir: /workspace
volumes:
pi-config:
pi-sessions:The key point: mount the current working directory as /workspace, and Pi can operate on files inside your code repository without touching system directories outside the container.
Step 4: Write a Launch Script
Here is a recommended pi launch script that auto-generates the container name from the working directory and supports a --sandbox flag for a stricter sandbox:
#!/usr/bin/env bash
# Pi — Start the containerized Pi agent.
SCRIPT_DIR="$(cd -- "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
WORKSPACE_DIR="${WORKSPACE:-$(pwd)}"
case "$WORKSPACE_DIR" in
/) ;;
*) WORKSPACE_DIR="$(cd -- "$WORKSPACE_DIR" && pwd)" ;;
esac
export WORKSPACE="$WORKSPACE_DIR"
sandbox="${PI_SANDBOX:-0}"
pi_args=()
while (($#)); do
case "$1" in
--sandbox) sandbox=1 ;;
--no-sandbox) sandbox=0 ;;
*) pi_args+=("$1") ;;
esac
shift
done
compose_files=(-f "$SCRIPT_DIR/docker-compose.yml")
if [[ "$sandbox" == "1" ]]; then
compose_files+=(-f "$SCRIPT_DIR/docker-compose.sandbox.yml")
fi
repo_slug="$(basename -- "$WORKSPACE_DIR" | tr -c 'a-zA-Z0-9_.-' '-' | sed 's/^-//')"
[[ -z "$repo_slug" ]] && repo_slug="workspace"
container_name="pi-${repo_slug}-$$"
api_key_args=(-e OPENAI_API_KEY -e DEEPSEEK_API_KEY -e ANTHROPIC_API_KEY -e GEMINI_API_KEY)
cmd=(docker compose --project-directory "$SCRIPT_DIR" "${compose_files[@]}" run --rm --name "$container_name" "${api_key_args[@]}" pi)
if ((${#pi_args[@]})); then cmd+=("${pi_args[@]}"); fi
exec "${cmd[@]}"Verify the Results
Run ./pi in the directory of a repository you're editing, and Pi will start Docker and enter /workspace. Have it do these things to verify:
- Ask Pi to refactor one module of the current repo and watch whether it splits files correctly
- Ask it to check whether type hints are correct
- Watch token inference, KV cache usage, and context window changes live in the LM Studio UI
On a successful run, you'll see Pi modifying files and calling bash inside Docker while the local model infers token by token in the background.
FAQ
- Slow inference: local models are hardware-bound with small context windows. Switch to a smaller quantized model (such as 12B QAT), or shorten the context.
- Prompt template mismatch: a common issue with early releases; LM Studio's and HuggingFace's "use this model" buttons usually fix it quickly — just keep your tools updated.
- KV cache full: 64GB of memory can be exhausted by long contexts; monitor memory usage and restart the session when needed.
- Can it be used for production software development? Not fully mature yet, but it's already practical as a fast, personalized local documentation lookup and coding aid — and well worth investing in for privacy-sensitive scenarios.
Related articles

Build a One-Person Company Brand System with Lovart
A hands-on guide to managing brand assets with Lovart's Brand Kit and unifying your visual style across platforms, from $19/month.

Aholo Viewer: A Browser for 3D Scenes with 1 Billion Gaussian Points
Manycore Tech open-sources Aholo Viewer, a 3D Gaussian browser with half the memory and 3x faster rendering, letting any device's browser smoothly load massive 3D scenes with 1 billion+ points.

Claude's Dual Memory System: The Permanent Brain Is Finally Here
Anthropic is testing a brand-new dual-mode memory system for Claude, with Memory Files for file-based memory and Dreams for memory consolidation, paired with the Conway Agent for 24/7 persistent memory.