OpenSquilla: A Token-Saving Middleware Layer for Agents

·Toolin Editorial Team

An open-source Agent Harness framework that cuts Agent token costs by more than 50% through intelligent routing, context management, and self-evolution.

OpenSquilla: A Token-Saving Middleware Layer for Agents

If you're building an AI Agent product, the token bill has probably already given you a headache. Agents burn far more tokens than chatbots — a single task may run a dozen-plus steps behind the scenes, each one spending money. And not every step needs the strongest model; for simple tasks like classification, summarization, and formatting cleanup, running a flagship model is pure waste.

OpenSquilla is an open-source Agent Harness framework that adds a "runtime hub" layer between the Agent application and the large model. It has one core goal: stop Agents from spending tokens they shouldn't, while making them understand the user better over time.

What Is OpenSquilla

OpenSquilla is a middleware framework with a four-layer architecture, deployed between the Agent application and the underlying large model. It doesn't replace the model itself; it manages the key decisions of "which model to call, how much context to feed in, how to orchestrate tasks, and how to keep evolving."

Image

OpenSquilla's four-layer architecture between the Agent application and the model: intelligent routing, context management, MetaSkill orchestration, and a self-evolution mechanism.

Four Core Mechanisms

Layer 1: Intelligent Routing — Call the Right Model, Save the Right Money

Most Agent teams bind to a single primary model. Results are stable and development is easy, but the bill quickly spirals out of control.

Before a task enters the large model, OpenSquilla uses a local routing model to gauge task complexity. Based on features like semantics, keywords, language, context length, and conversation turns, it grades tasks into tiers, then matches each tier to a different model.

The key difference: this routing is not a static rule table but a parameterized model that keeps optimizing from task feedback. Which tasks succeeded, where tokens got burned, which models offer better value for money — these signals flow back into the router and continuously train it.

Per the team's data, OpenSquilla's intelligent routing beats OpenRouter's routing by 4.4 percentage points in accuracy while costing 75% less.

Layer 2: Context Management — Don't Make the Model Read Filler

Many Agent systems stuff Skill descriptions, tool documentation, historical memory, and web content into the prompt all at once. The model has to re-read everything on every call, and the tokens it doesn't use still get billed.

OpenSquilla's approach:

  • Load Skills on demand: a single task only injects the Skills it might use, instead of cramming in the descriptions of dozens of Skills
  • Precise memory recall: retrieve relevant fragments from a local database rather than copying in entire passages
  • Preprocess tool results: strip tags, styles, navbars, ads, and other irrelevant content out of web page HTML

Per the team's data, context management delivers an additional 20% to 50% cost reduction on top.

Layer 3: MetaSkill — Let the Agent Organize Its Own Capabilities

The more Skills, the stronger the Agent — in theory. In real-world use, users start losing track of how to combine them. Writing an article means researching, fact-checking, studying style, drafting, and proofreading — there's a Skill for every step, but who sequences them?

OpenSquilla's MetaSkill mechanism lets the user simply state the goal, and the AI automatically breaks it into steps, picks Skill combinations, and arranges the dependencies. Each step gets its own dedicated context to avoid mutual interference.

Image

Layer 4: Self-Evolution — Train User Preferences into the Harness

The first time a user hands an Agent a task, it usually takes several rounds of correction. The problem is that once fixed, the Agent makes the same mistake again next time — the experience never accumulates.

OpenSquilla looks back over the whole interaction: which conditions the user added, which deviations they corrected, what final result they accepted — then distills all of it into a Skill or workflow. The next time a similar task comes up, the Agent doesn't start from zero.

One fewer correction from the user = one fewer run for the system = one fewer round of burned tokens.

Who It's For

  • Agent product teams: teams whose token gross margin is below 30% and who need to systematically cut call costs
  • Multi-model mixed calls: teams already using multiple models but lacking a unified routing solution
  • Complex workflow orchestration: scenarios where the Agent spans multiple steps and Skill combinations
  • User retention optimization: anyone who wants the Agent to remember user preferences and reduce repeated guidance

Tip: OpenSquilla's routing model is a locally running ensemble tree model. It requires no extra calls to a large-model API and adds no token overhead of its own.

Related articles

Finding Overseas Influencers with AI Employees: A Hands-On AhaCreator Guide
AI Tutorials

Finding Overseas Influencers with AI Employees: A Hands-On AhaCreator Guide

A step-by-step walkthrough of the full overseas influencer marketing workflow with AhaCreator — from creator sourcing and content review to cross-border payouts. Ideal for indie developers and teams going global.

Toolin Editorial Team
AI Long-Video Generation: A Head-to-Head Review of Two Open-Source Frameworks
AI Products

AI Long-Video Generation: A Head-to-Head Review of Two Open-Source Frameworks

VideoClaw and JoyAI-Echo, two open-source frameworks released the same day, tackle AI long-video consistency through multi-agent collaboration and cross-modal memory banks respectively — this article compares their technical approaches.

Toolin Editorial Team
ChatGPT's Memory System Gets a Full Overhaul: Dreaming V3 Is Live
AI Products

ChatGPT's Memory System Gets a Full Overhaul: Dreaming V3 Is Live

OpenAI ships the new Dreaming V3 memory architecture — ChatGPT now "dreams" in the background to organize what it knows about you. Free access for 1 billion users for the first time, with doubled memory capacity for Plus/Pro.

Toolin Editorial Team
Cloudflare Integrates Claude Managed Agents: A Developer's Practical Guide
AI Products

Cloudflare Integrates Claude Managed Agents: A Developer's Practical Guide

Cloudflare adds support for Claude Managed Agents, letting developers run Claude agents on Cloudflare's platform, connect to private systems, and deploy AI agents securely.

Toolin Editorial Team
Claude Code vs Codex: A Panoramic Timeline of 24 Features
AI Products

Claude Code vs Codex: A Panoramic Timeline of 24 Features

A timeline breakdown of the 24 features both AI coding agents share — Claude Code shipped 18 of them first, but the gap is now closing in days, not months.

Toolin Editorial Team
Higress: The AI-Assisted K8s Gateway Migration Tool
AI Products

Higress: The AI-Assisted K8s Gateway Migration Tool

An AI-assisted migration approach showcased by CNCF converts 60 ingress-nginx resources into Higress configuration automatically in about 30 minutes, slashing K8s gateway migration costs.

Toolin Editorial Team