OpenSquilla: A Token-Saving Middleware Layer for Agents
An open-source Agent Harness framework that cuts Agent token costs by more than 50% through intelligent routing, context management, and self-evolution.


OpenSquilla: A Token-Saving Middleware Layer for Agents
An open-source Agent Harness framework that cuts Agent token costs by more than 50% through intelligent routing, context management, and self-evolution.
If you're building an AI Agent product, the token bill has probably already given you a headache. Agents burn far more tokens than chatbots — a single task may run a dozen-plus steps behind the scenes, each one spending money. And not every step needs the strongest model; for simple tasks like classification, summarization, and formatting cleanup, running a flagship model is pure waste.
OpenSquilla is an open-source Agent Harness framework that adds a "runtime hub" layer between the Agent application and the large model. It has one core goal: stop Agents from spending tokens they shouldn't, while making them understand the user better over time.
What Is OpenSquilla
OpenSquilla is a middleware framework with a four-layer architecture, deployed between the Agent application and the underlying large model. It doesn't replace the model itself; it manages the key decisions of "which model to call, how much context to feed in, how to orchestrate tasks, and how to keep evolving."

OpenSquilla's four-layer architecture between the Agent application and the model: intelligent routing, context management, MetaSkill orchestration, and a self-evolution mechanism.
Four Core Mechanisms
Layer 1: Intelligent Routing — Call the Right Model, Save the Right Money
Most Agent teams bind to a single primary model. Results are stable and development is easy, but the bill quickly spirals out of control.
Before a task enters the large model, OpenSquilla uses a local routing model to gauge task complexity. Based on features like semantics, keywords, language, context length, and conversation turns, it grades tasks into tiers, then matches each tier to a different model.
The key difference: this routing is not a static rule table but a parameterized model that keeps optimizing from task feedback. Which tasks succeeded, where tokens got burned, which models offer better value for money — these signals flow back into the router and continuously train it.
Per the team's data, OpenSquilla's intelligent routing beats OpenRouter's routing by 4.4 percentage points in accuracy while costing 75% less.
Layer 2: Context Management — Don't Make the Model Read Filler
Many Agent systems stuff Skill descriptions, tool documentation, historical memory, and web content into the prompt all at once. The model has to re-read everything on every call, and the tokens it doesn't use still get billed.
OpenSquilla's approach:
- Load Skills on demand: a single task only injects the Skills it might use, instead of cramming in the descriptions of dozens of Skills
- Precise memory recall: retrieve relevant fragments from a local database rather than copying in entire passages
- Preprocess tool results: strip tags, styles, navbars, ads, and other irrelevant content out of web page HTML
Per the team's data, context management delivers an additional 20% to 50% cost reduction on top.
Layer 3: MetaSkill — Let the Agent Organize Its Own Capabilities
The more Skills, the stronger the Agent — in theory. In real-world use, users start losing track of how to combine them. Writing an article means researching, fact-checking, studying style, drafting, and proofreading — there's a Skill for every step, but who sequences them?
OpenSquilla's MetaSkill mechanism lets the user simply state the goal, and the AI automatically breaks it into steps, picks Skill combinations, and arranges the dependencies. Each step gets its own dedicated context to avoid mutual interference.

Layer 4: Self-Evolution — Train User Preferences into the Harness
The first time a user hands an Agent a task, it usually takes several rounds of correction. The problem is that once fixed, the Agent makes the same mistake again next time — the experience never accumulates.
OpenSquilla looks back over the whole interaction: which conditions the user added, which deviations they corrected, what final result they accepted — then distills all of it into a Skill or workflow. The next time a similar task comes up, the Agent doesn't start from zero.
One fewer correction from the user = one fewer run for the system = one fewer round of burned tokens.
Who It's For
- Agent product teams: teams whose token gross margin is below 30% and who need to systematically cut call costs
- Multi-model mixed calls: teams already using multiple models but lacking a unified routing solution
- Complex workflow orchestration: scenarios where the Agent spans multiple steps and Skill combinations
- User retention optimization: anyone who wants the Agent to remember user preferences and reduce repeated guidance
Tip: OpenSquilla's routing model is a locally running ensemble tree model. It requires no extra calls to a large-model API and adds no token overhead of its own.
Toolin Editorial Team
Categories
Related articles

Finding Overseas Influencers with AI Employees: A Hands-On AhaCreator Guide
A step-by-step walkthrough of the full overseas influencer marketing workflow with AhaCreator — from creator sourcing and content review to cross-border payouts. Ideal for indie developers and teams going global.

AI Long-Video Generation: A Head-to-Head Review of Two Open-Source Frameworks
VideoClaw and JoyAI-Echo, two open-source frameworks released the same day, tackle AI long-video consistency through multi-agent collaboration and cross-modal memory banks respectively — this article compares their technical approaches.

ChatGPT's Memory System Gets a Full Overhaul: Dreaming V3 Is Live
OpenAI ships the new Dreaming V3 memory architecture — ChatGPT now "dreams" in the background to organize what it knows about you. Free access for 1 billion users for the first time, with doubled memory capacity for Plus/Pro.

Cloudflare Integrates Claude Managed Agents: A Developer's Practical Guide
Cloudflare adds support for Claude Managed Agents, letting developers run Claude agents on Cloudflare's platform, connect to private systems, and deploy AI agents securely.

Claude Code vs Codex: A Panoramic Timeline of 24 Features
A timeline breakdown of the 24 features both AI coding agents share — Claude Code shipped 18 of them first, but the gap is now closing in days, not months.

Higress: The AI-Assisted K8s Gateway Migration Tool
An AI-assisted migration approach showcased by CNCF converts 60 ingress-nginx resources into Higress configuration automatically in about 30 minutes, slashing K8s gateway migration costs.